Question generation method, knowledge extraction method, question answering system construction method, question generation program, knowledge extraction program, question generation system, and knowledge extraction system
The method addresses inefficiencies in extracting tacit knowledge by using a structured approach with a large-scale language model to refine questions and store knowledge in both text and graph formats, ensuring relevance and reducing duplication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- JFE STEEL CORP
- Filing Date
- 2025-07-22
- Publication Date
- 2026-05-11
AI Technical Summary
Existing methods for extracting tacit knowledge, such as questionnaires and interviews, require skilled questioners and are inefficient in capturing relevant information without duplication or deviation from the theme.
A method involving theme information reception, local information extraction, supplemental information generation, text acquisition using a large-scale language model, and question refinement to ensure relevance and uniqueness, followed by knowledge extraction and storage in both text and graph formats.
Efficient and reliable extraction of tacit knowledge without requiring skilled questioners, reducing duplication and ensuring theme relevance, facilitating easy integration and retrieval.
Smart Images

Figure 0007856222000001 
Figure 0007856222000002 
Figure 0007856222000003
Abstract
Description
Technical Field
[0001] The present invention relates to a question generation method, a knowledge extraction method, a question-and-answer system construction method, a question generation program, a knowledge extraction program, a question generation system, and a knowledge extraction system.
Background Art
[0002] In recent years, the inheritance of skilled techniques in the manufacturing industry has become a problem. In particular, the knowledge of skilled workers, especially tacit knowledge that has not been documented, urgently needs to be formally stored in some form. In order to formalize human tacit knowledge, it is common to use the form of questionnaires or interviews. As a technique for formalizing tacit knowledge, for example, as in Non-Patent Document 1, a method of conducting an interview suitable for extracting tacit knowledge is known.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, one of the conventional methods, the questionnaire, has no time constraints on the questioner or the respondent, but can only obtain answers to the pre-determined questions, and it is difficult to obtain tacit knowledge. In addition, in the interview format, questions can be set according to the answers, so it is easy to extract tacit knowledge, but on the other hand, skills and techniques on the questioner side are required as shown in Non-Patent Document 1. For this reason, there has been a demand for a technique that does not require the skills and techniques of the questioner side and can surely extract tacit knowledge.
[0005] The present invention has been made in view of these circumstances, and its objective is to provide a question generation method, a knowledge extraction method, a question answering system construction method, a question generation program, a knowledge extraction program, a question generation system, and a knowledge extraction system that can efficiently and reliably extract tacit knowledge. [Means for solving the problem]
[0006] To solve the above-mentioned problems and achieve the objective, a question generation method according to one aspect of the present invention includes: a theme information reception step for receiving theme information to generate a question; an extraction step for extracting local information from the received theme information; a supplemental information generation step for generating supplemental information for the extracted local information; a text acquisition step for acquiring text generated using a large-scale language model that learns local information as features based on a prompt including the theme information, the local information and the supplemental information; and a question refinement step for determining whether or not to set the text as a question related to the theme information based on the degree of overlap in content and / or theme relevance.
[0007] A question generation method according to one aspect of the present invention, in the above invention, if it is determined in the question review step that the question should not be set as a question relating to the theme information, the text obtained in the text acquisition step is set as a supplementary prompt that includes the text as an inappropriate specific example and includes instructions to encourage improvement of the output, and the text acquisition step is executed again based on the prompt and the supplementary prompt. .
[0008] A question generation method according to one aspect of the present invention further includes a topic setting step of setting a topic for questions related to theme information as a guideline for questions related to theme information, and in the text acquisition step, information related to the set topic for questions is added to the prompt.
[0009] A question generation method according to one aspect of the present invention, in the above invention, the topic setting step includes theme information and sets the topic of the question using a large-scale language model that is generated by learning local information as features, based on prompts that prompt the generation of a question topic.
[0010] In one aspect of the present invention, the question generation method, in the above invention, involves comparing the similarity between the information obtained by vectorizing the text and the information obtained by vectorizing other texts generated before the text to be evaluated, in order to determine whether or not the content is redundant.
[0011] In one aspect of the present invention, the question generation method, in the above invention, determines whether the sentence is related to the set theme based on the relationship between the information obtained by vectorizing the sentence and the information obtained by vectorizing the theme information.
[0012] In one aspect of the present invention, the question generation method, in the above invention, involves the extraction step of extracting local information from the theme information by referring to a database that has been pre-stored local terms and sentences containing such local terms.
[0013] In one aspect of the present invention, the question generation method, in the above invention, involves the extraction step of referring to a database that stores information obtained by pre-vectorizing local terms and sentences containing said local terms, and extracting the local information by comparing the similarity between said information and information obtained by vectorizing the theme information.
[0014] A knowledge extraction method according to one aspect of the present invention includes: a theme information receiving step for receiving theme information to generate a question; a first extraction step for extracting first local information from the received theme information; a first supplementary information generation step for generating first supplementary information for the extracted first local information; a text acquisition step for acquiring text generated using a large-scale language model that learns local information as features based on a prompt including the theme information, the first local information, and the first supplementary information; a question refinement step for determining whether to set the text as a question related to the theme information based on the degree of overlap in content and / or theme relevance; a question output step for outputting a question that has been determined to be set as a question by the question refinement step; a response receiving step for receiving response information from a respondent; a second extraction step for extracting second local information from the response information; and a second supplementary information generation step for generating second supplementary information for the extracted second local information. The process includes a knowledge information generation step of generating knowledge information by converting information from a series of processes, including the aforementioned question, the aforementioned answer information, second local information, and second supplementary information, into a set format.
[0015] A knowledge extraction method according to one aspect of the present invention further includes an update step in which the prompt is updated by adding the answer information, the second local information, and the second supplementary information to the prompt, and the text acquisition step, the question review step, the quality output step, the answer reception step, the second extraction step, and the second supplementary information generation step are repeatedly executed based on the updated prompt.
[0016] In one aspect of the present invention, the knowledge extraction method, in the above invention, generates the knowledge information by converting it into a knowledge graph format and a text format in the knowledge information generation step.
[0017] A method for constructing a question-answering system according to one aspect of the present invention involves constructing a question-answering system that utilizes generated knowledge information using the knowledge extraction method according to the above invention.
[0018] A question generation program according to one aspect of the present invention causes a computer to perform the following steps: a theme information receiving step of receiving theme information for generating a question; an extraction step of extracting local information from the received theme information by referring to a storage unit; a supplemental information generation step of generating supplemental information for the extracted local information; a text acquisition step of acquiring text generated using a large-scale language model that learns the local information as features based on a prompt including the theme information, the local information and the supplemental information; and a question refinement step of determining whether or not to set the text as a question related to the theme information based on the degree of overlap in content and / or theme relevance.
[0019] A knowledge extraction program according to one aspect of the present invention includes: a theme information receiving step for receiving theme information to generate a question; a first extraction step for extracting first local information from the received theme information by referring to a memory unit; a first supplementary information generation step for generating first supplementary information for the extracted first local information; a text acquisition step for acquiring text generated using a large-scale language model that learns the local information as features based on a prompt including the theme information, the first local information, and the first supplementary information; and a method for analyzing the text based on the degree of content overlap and / or theme relevance. Accordingly, the computer is made to execute the following steps: a question review step to determine whether or not to set the above text as a question relating to the theme information; a question output step to output the question set by the question review step; a response reception step to receive response information from the respondent; a second extraction step to extract second local information from the response information; a second supplementary information generation step to generate second supplementary information for the extracted second local information; and a knowledge information generation step to generate knowledge information by converting the information of a series of processes including the question, the response information, the second local information, and the second supplementary information into a set format.
[0020] A question generation system according to one aspect of the present invention includes: a theme information receiving unit that receives theme information for generating a question; an extraction unit that extracts local information from the received theme information by referring to a storage unit; a supplemental information generating unit that generates supplemental information for the extracted local information; a text acquisition unit that acquires text generated using a large-scale language model that learns the local information as features based on prompts including the theme information, the local information and the supplemental information; and a question refinement unit that determines whether or not to set the text as a question related to the theme information based on the degree of overlap in content and / or theme relevance.
[0021] A knowledge extraction system according to an aspect of the present invention includes a theme information reception unit that receives an input of theme information for generating a question, a first extraction unit that extracts first local information by referring to a storage unit from the received theme information, a first supplementary information generation unit that generates first supplementary information for the extracted first local information, a sentence acquisition unit that acquires a sentence generated using a large language model generated by learning local information as a feature amount based on a prompt including the theme information, the first local information, and the first supplementary information, a question scrutiny unit that determines whether to set the sentence as a question regarding the theme information based on the degree of content duplication and / or theme relevance of the sentence, a question output unit that outputs a question determined by the question scrutiny unit to be set as the question, an answer reception unit that receives answer information from a responder, a second extraction unit that extracts second local information from the answer information, a second supplementary information generation unit that generates second supplementary information for the extracted second local information, and a knowledge information generation unit that generates knowledge information by converting information on a series of processes including the question, the answer information, the second local information, and the second supplementary information into a set format.
Advantages of the Invention
[0022] According to the question generation method, knowledge extraction method, question-and-answer system construction method, question generation program, knowledge extraction program, question generation system, and knowledge extraction system according to the present invention, there is an effect that tacit knowledge can be efficiently and reliably extracted.
Brief Description of the Drawings
[0023] [Figure 1] FIG. 1 is a block diagram showing a schematic configuration of a knowledge processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing a configuration of a knowledge extraction device included in the knowledge processing system according to an embodiment of the present invention. [Figure 3] FIG. 3 is a block diagram showing a configuration of a sentence generation device included in the knowledge processing system according to an embodiment of the present invention. [Figure 4] FIG. 4 is a sequence diagram for explaining the flow of knowledge extraction processing according to an embodiment of the present invention. [Figure 5] FIG. 5 is a diagram (part 1) for explaining an example of interaction with a respondent during knowledge extraction according to an embodiment of the present invention. [Figure 6] FIG. 6 is a diagram (part 2) for explaining an example of interaction with a respondent during knowledge extraction according to an embodiment of the present invention. [Figure 7] FIG. 7 is a sequence diagram for explaining the flow of knowledge extraction processing according to Modification 1 of the present invention. [Figure 8] FIG. 8 is a sequence diagram for explaining the flow of knowledge extraction processing according to Modification 2 of the present invention. [Figure 9] FIG. 9 is a diagram showing, in text summary information and knowledge graph format, knowledge information of tacit knowledge that has not been verbalized, extracted from skilled workers in charge of equipment maintenance at a steelworks, by the knowledge extraction processing of the present invention. [Figure 10A] FIG. 10A is a diagram showing an example of an answer by a large language model given knowledge information generated by the knowledge extraction processing of the present invention in a RAG environment, as a result of asking a question-and-answer system about causes and countermeasures when damage occurs to the main machines (main drives and main shafts) of the hot rolling line at a steelworks. [Figure 10B] FIG. 10B shows a comparison of answers by a large language model that was not given in a RAG environment, as a result of asking a question-and-answer system about causes and countermeasures when damage occurs to the main machines (main drives and main shafts) of the hot rolling line at a steelworks.
Mode for Carrying Out the Invention
[0024] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In all the drawings of the following embodiment, the same or corresponding parts are denoted by the same reference numerals. Also, the present invention is not limited to the embodiment described below. .
[0025] (Embodiment) Figure 1 is a block diagram illustrating the schematic configuration of a knowledge processing system according to one embodiment of the present invention. As shown in Figure 1, the knowledge processing system 1 comprises a knowledge extraction device 10, a text generation device 20, and a server 30. The knowledge processing system 1 is configured such that the knowledge extraction device 10 and the text generation device 20 can input and output data from each other, and the knowledge extraction device 10 and the text generation device 20 can each read information from the server 30. In this embodiment, the knowledge extraction device 10 is configured with a question generation system and a knowledge extraction system using at least some of its components.
[0026] Here, data input and output of the knowledge extraction device 10 and the text generation device 20 can be performed via network communication through a network or cloud, or via contactless communication such as Bluetooth®. It is also possible to move data via USB (Universal Serial Bus) memory or disk recording media such as CD (Compact Disc), DVD (Digital Versatile Disc), or BD (Blu-ray® Disc). The network is composed of a combination of wired and wireless communication as appropriate, and consists of communication networks such as the Internet network and mobile phone network. The network may consist of one or more combinations of, for example, dedicated lines, public communication networks such as the Internet, such as LAN (Local Area Network), WAN (Wide Area Network), telephone communication networks and public lines such as mobile phones, and VPN (Virtual Private Network).
[0027] Figure 2 is a block diagram showing the configuration of a knowledge extraction device in a knowledge processing system according to one embodiment of the present invention. The knowledge extraction device 10 includes a communication unit 11, an input / output unit 12, an extraction unit 13, a prompt generation unit 14, a question review unit 15, a knowledge information generation unit 16, a control unit 17, and a storage unit 18. The question generation device is comprised of a communication unit 11, an input / output unit 12, an extraction unit 13, a prompt generation unit 14, a question review unit 15, a control unit 17, and a storage unit 18.
[0028] The communication unit 11 is, for example, a LAN interface board, a wired communication circuit for wired communication, or a wireless communication circuit for wireless communication. The LAN interface board, wired communication circuit, or wireless communication circuit is connected to the network. The communication unit 11, acting as both a transmitter and receiver, is connected to the network and communicates with the document generation device 20 and the server 30.
[0029] The input / output unit 12 can be composed of, for example, a touch panel display or a speaker microphone. The input / output unit 12 as an input means may include an interface that receives various information transmitted from an external server via the communication unit and outputs it to the control unit 17. The input / output unit 12 also includes a user interface such as a keyboard, input buttons, levers, a touch panel for manual input superimposed on a display such as an LCD, or a microphone for voice recognition. The input / output unit 12 is configured to allow predetermined information to be input to the control unit 17 by an operator or other person operating it. The input / output unit 12 as an output means displays predetermined images on a display monitor, displays characters or figures on the screen of a touch panel display, or outputs sound from a speaker, according to the control unit 17's control. In other words, the input / output unit 12 is configured to be able to notify the outside of predetermined information. The input unit and output unit of the input / output unit 12 may be configured as separate units.
[0030] The extraction unit 13 decomposes the input language information into word units and extracts local terms. Local terms are specialized terms that are not commonly used, such as those used in a manufacturing environment. For example, the extraction unit 13 extracts local terms from language information input via the input / output unit 12. In this process, the extraction unit 13 reads a pre-configured database of local terms by referring to the server 30 and extracts local terms and sentences that use those local terms. The decomposition of language information into words can be done using methods such as morphological analysis or large language models (LLM).
[0031] The prompt generation unit 14 generates prompts for generating questions to obtain tacit knowledge from the respondent. Specifically, the prompt generation unit 14 generates prompts by referring to the server 30 and performing searches on extracted local terms to supplement the meaning of the terms. In this embodiment, for example, information is linked from a word list in which terms and explanations are paired. The prompt generated by the prompt generation unit 14 is sent to the text generation device 20, and the text generation device 20 generates a question (text) for acquiring tacit knowledge based on the prompt.
[0032] The question review unit 15 functions as a judgment unit that determines whether the content of a question (text) obtained from the text generation device 20 is duplicated and whether it is related to the set theme. If it is determined that the generated question does not duplicate the content of a previously adopted question and is related to the theme, the prompt is confirmed as a question.
[0033] The knowledge information generation unit 16 generates knowledge information as a record of the interaction with the respondent. The knowledge information generation unit 16 generates knowledge information in a set format. Examples of formats include knowledge graphs, text data, and JSON files intended for use in other applications.
[0034] The extraction unit 13, prompt generation unit 14, question refinement unit 15, and knowledge information generation unit 16 are configured using processors such as a CPU (Central Processing Unit), DSP (Digital Signal Processor), and FPGA (Field-Programmable Gate Array).
[0035] Specifically, the control unit 17 includes a processor having hardware such as a CPU, DSP, and FPGA, and a main memory unit such as RAM (Random Access Memory) and ROM (Read Only Memory) (none of which are shown). In the following description, the knowledge extraction device 10 is assumed to have a function that performs speech recognition on audio data input via the input / output unit 12, etc., and converts it into text.
[0036] The storage unit 18 is composed of a storage medium selected from volatile memory such as RAM, non-volatile memory such as ROM, EPROM (Erasable Programmable ROM), hard disk drive (HDD), and removable media. The removable media can be, for example, a USB memory stick, or a disk recording medium such as a CD, DVD, or BD. Alternatively, the storage unit 18 may be configured using a computer-readable recording medium such as an externally insertable memory card.
[0037] The memory unit 18 can store various programs such as an operating system (OS), knowledge extraction applications, various tables, and various databases for executing the operation of the knowledge extraction device 10. These various programs can also be recorded on computer-readable recording media such as hard disks, flash memory, CD-ROMs, DVD-ROMs, and flexible disks and widely distributed.
[0038] Figure 3 is a block diagram showing the configuration of a text generation device included in a knowledge processing system according to one embodiment of the present invention. The text generation device 20 comprises a communication unit 21, an input / output unit 22, a text generation unit 23, a vectorization unit 24, a control unit 25, and a storage unit 26.
[0039] The communication unit 21 is, for example, a LAN interface board, a wired communication circuit for wired communication, or a wireless communication circuit for wireless communication. The LAN interface board, wired communication circuit, or wireless communication circuit is connected to the network. The communication unit 21, acting as both a transmitter and receiver, is connected to the network and communicates with the knowledge extraction device 10 and the server 30.
[0040] The input / output unit 22 can be composed of, for example, a display, a touch panel display, or a speaker microphone. The input / output unit 22 as an input means may include an interface that receives various information transmitted from an external server via the communication unit and outputs it to the control unit 25. The input / output unit 12 also includes a user interface such as a keyboard, input buttons, a lever, a touch panel for manual input superimposed on a display such as an LCD, or a microphone for voice recognition. The input / output unit 22 is configured to allow predetermined information to be input to the control unit 25 by an operator or other person operating it. The input / output unit 22 as an output means displays predetermined images on a display monitor, displays characters or figures on the screen of a touch panel display, or outputs sound from a speaker, according to the control unit 25. In other words, the input / output unit 22 is configured to be able to notify the outside of predetermined information. The input unit and output unit of the input / output unit 22 may be configured as separate units.
[0041] The text generation unit 23 generates sentences to be used as questions using prompts obtained from the knowledge extraction device 10. The text generation unit 23 generates sentences using a trained model. The trained model in this embodiment is a large-scale language model (LLM) generated by machine learning using local terms as features with a multilayer neural network comprising an input layer, an intermediate layer, and an output layer. Known methods can be used for machine learning.
[0042] The vectorization unit 24 vectorizes the text generated by the text generation unit 23. Through this vectorization, the text is converted into numerical values such as strings or coordinates in a multidimensional space. Known methods such as Bag of Words and distributed representations can be used for vectorization.
[0043] The text generation unit 23 and the vectorization unit 24 are configured using processors such as CPUs, DSPs, and FPGAs.
[0044] The control unit 25 specifically includes a processor having hardware such as a CPU, DSP, and FPGA, and a main memory unit such as RAM and ROM (none of which are shown).
[0045] The storage unit 26 is composed of a storage medium selected from volatile memory such as RAM, non-volatile memory such as ROM, EPROM, HDD, and removable media. The removable media can be, for example, a USB memory stick, or a disk recording medium such as a CD, DVD, or BD. Alternatively, the storage unit 26 may be configured using a computer-readable recording medium such as an externally insertable memory card.
[0046] The memory unit 26 can store various programs such as the OS and text generation applications, various tables, and various databases necessary for the operation of the text generation device 20. These various programs can also be recorded on computer-readable recording media such as hard disks, flash memory, CD-ROMs, DVD-ROMs, and flexible disks and widely distributed.
[0047] Server 30 is implemented using a computer equipped with a general-purpose processor such as a CPU and a storage medium selected from volatile memory such as RAM, non-volatile memory such as ROM, EPROM, HDD, and removable media. Server 30 includes a storage unit in which local terminology and documents using local terminology (local documents) are pre-stored.
[0048] (Knowledge extraction process) Figure 4 is a sequence diagram illustrating the flow of knowledge extraction processing according to one embodiment of the present invention. When the knowledge extraction process is started in the knowledge extraction device 10, the control unit 17 determines whether or not it has received the input for setting theme information (step S101: theme information reception step). At this time, the theme information is input, for example, by voice or character input on a keyboard via the input / output unit 12 which functions as a theme information reception unit. If there is no input for setting theme information (step S101: No), the control unit 17 repeats the input confirmation. If the control unit 17 determines that theme information has been input (step S101: Yes), it proceeds to step S102.
[0049] In step S102, the extraction unit 13 (first extraction unit) extracts local terms (first local terms) from the input theme information ((first) extraction step). At this time, the extraction unit 13 extracts local information (first local information) from the theme information by referring to a database (e.g., server 30) that has been pre-stored with local terms and sentences containing those local terms. Alternatively, the extraction unit 13 may extract local information by referring to a database (e.g., server 30) that has been pre-stored with information obtained by vectorizing local terms and sentences containing those local terms, and comparing the similarity between that information and the information obtained by vectorizing the theme information.
[0050] Then, the prompt generation unit 14, which functions as a supplementary information generation unit (first supplementary information generation unit), searches for local information by referring to the server 30 for the extracted local term, and generates supplementary information (first supplementary information) that supplements the meaning of the local term based on the search results (step S103).
[0051] Subsequently, the prompt generation unit 14 generates a prompt by supplementing the meaning of local terms based on the search results (step S104). The prompt generation unit 14 transmits the generated prompt to the text generation device 20 via the communication unit 11. This prompt includes, for example, theme information, local information (first local information), and supplementary information that supplements the meaning of local terms (first supplementary information). Here, steps S103 and S104 correspond to the supplementary information generation steps.
[0052] When the text generation device 20 receives a prompt from the knowledge extraction device 10 via the communication unit 21, it uses the received prompt to generate a text that will serve as a question for acquiring tacit knowledge (step S105). In this embodiment, the text generation unit 23 generates the text using LLM. The knowledge extraction device 10 acquires the text generated by the text generation device 20 via the communication unit 11, which functions as a text acquisition unit (step S106: text acquisition step).
[0053] Furthermore, in the text generation device 20, the vectorization unit 24 vectorizes the text generated in step S105 (step S107). The text generation device 20 transmits the generated text and its vector information to the knowledge extraction device 10 via the communication unit 21. At this time, the text generation device 20 stores the generated text and its vector information in the storage unit 26 and the server 30. The knowledge extraction device 10 acquires the vector information generated by the text generation device 20 via the communication unit 21 (step S108).
[0054] In the knowledge extraction device 10, after acquiring text and vectors from the text generation device 20, the question refinement unit 15 determines whether the acquired questions have duplicate content (step S109). The question refinement unit 15 calculates the similarity between the questions used in the current exchange and the current text from the vector information, and compares the calculated similarity with a pre-set threshold. If the similarity is higher than the threshold, the question refinement unit 15 determines that the question content is duplicated (step S109: Yes), and proceeds to step S104 to generate text with different content. If the similarity is below the threshold, the question refinement unit 15 determines that the question content is not duplicated (step S109: No), and proceeds to step S110. In step S109, we explained an example of determining whether there is any duplication of questions used in the current exchange, but it is also possible to determine similarity with questions used in past exchanges prior to the current one.
[0055] In step S110, the question refinement unit 15 determines whether or not the question is related to the current theme. The question refinement unit 15 calculates the distance between the theme and the text (local terminology) from the vector and compares the calculated distance with a pre-set threshold. If the distance is greater than the threshold, the question refinement unit 15 determines that the text and the theme are not very related (step S110: No) and proceeds to step S104. If the distance is less than or equal to the threshold, the question refinement unit 15 determines that the text and the theme are highly related (step S110: Yes) and proceeds to step S111. The question review unit 15 may execute step S111 before step S109, or it may execute steps S109 and S110 simultaneously. Furthermore, when proceeding to step S104 to generate text again, a supplementary prompt is set that includes text deemed to be duplicates or unrelated to the theme as inappropriate examples, and provides instructions to encourage improvement of the output by creating text different from this text. In addition, a threshold for the number of repetitions of the question may be set, and the interaction may be terminated when the number of repetitions exceeds the threshold.
[0056] In step S111, the question review unit 15 determines that, in steps S109 and S110, the content is not redundant and is related to the theme, and decides to set that sentence as a question. These steps S109 to S111 represent a series of processes that constitute a decision step.
[0057] In step S112, the control unit 17 causes the input / output unit 12, which functions as a question output unit, to output the text generated in S105 as a question (question output step). As a result, the question is output as a display or as audio.
[0058] Subsequently, the control unit 17 determines whether or not there is input of an answer (answer information) to the outputted question from the input / output unit 12, which functions as an answer receiving unit (step S113: answer receiving step). If there is no answer input (step S113: No), the control unit 17 repeats the input confirmation. If the input of an answer is confirmed via the input / output unit 12 (step S113: Yes), the control unit 17 proceeds to step 114. Furthermore, if the control unit 17 does not receive an answer after a predetermined time has elapsed since outputting the question, it may proceed to step S114.
[0059] In step S114, the extraction unit 13 (second extraction unit) extracts local terms (second local terms) from the input language information, which is the response (response information), as related terms related to the theme information (second extraction step).
[0060] Then, the prompt generation unit 14, which functions as a supplementary information generation unit (second supplementary information generation unit), generates supplementary information (second supplementary information) for the answer (step S115). At this time, the prompt generation unit 14 searches for local information (second local information) by referring to the server 30 for the local terms extracted in step S114, and generates supplementary information that supplements the meaning of the local terms in the answer based on the search results.
[0061] After generating supplemental information, the control unit 17 determines whether or not termination information has been entered (step S116). The control unit 17 determines whether or not a pre-set input, such as a voice message containing a keyword related to termination, such as "I'm ending," has been entered. If the control unit 17 determines that no termination information has been entered (step S116: No), it proceeds to step S104 to generate a new question. Conversely, if the control unit 17 determines that termination information has been entered (step S116: Yes), it proceeds to step S117. Furthermore, if the system proceeds to step S104 to continue generating questions, the prompt generation unit 14 updates the prompt by adding the respondent's answer (answer information), the local information (second local information) and supplementary information (second supplementary information) generated in step S115 (update step).
[0062] In step S117, the knowledge information generation unit 16 generates knowledge information based on the information from a series of processes, including questions, answers (answer information), local information (second local information), and supplementary information (second supplementary information) from the series of processes (knowledge information generation step). The knowledge information generation unit 16 generates knowledge information in an output format specified by the user via voice input or the like, or in a pre-set output format.
[0063] Here, the output format should organize the knowledge information obtained through the interaction between Knowledge System 1 and the respondent, and make it easily usable by other users and systems. Specific output formats for knowledge information include text format and knowledge graph format.
[0064] Text format has the advantage of allowing for detailed recording of specific information. Examples include the FAQ (Frequently Asked Questions) format, which compiles questions and answers related to the topic, and the summary information (document summary) format, which uses natural language processing (NLP) to summarize knowledge gained through question-and-answer exchanges and concisely summarizes important points and concepts related to the topic. The FAQ format allows for easy integration with other systems and applications because it stores data as structured information, and information can be retrieved via APIs. The organized information is also easier for other users to understand. The summary format allows other users to quickly grasp important information and enables efficient referencing. Because the summarized information can be quickly referenced by other systems, it improves the efficiency of information retrieval.
[0065] The knowledge graph format visually represents the relationships and structure of information using nodes (entities) and edges (relationships), based on the key entities extracted from questions and answers related to thematic information, and the relationships between those entities. It has the advantage of clarifying the relationships between data and improving search performance.
[0066] It is advisable to store tacit knowledge information in databases in both knowledge graph and text formats. Storing tacit knowledge information in both formats allows for technological advancements in the system. The knowledge graph format significantly improves the efficiency of information retrieval in other systems by explicitly showing the relationships between entities. This enables the provision of fast and accurate search results even for complex queries, and allows for advanced analysis utilizing the interrelationships of the data. Furthermore, saving in text format allows for the use of natural language processing technology, providing the flexibility to search for information using natural language. Combining these two formats ensures data integrity and consistency, improving overall system performance. As a result, it facilitates the use of tacit knowledge and facilitates integration with other systems, contributing to the overall evolution of data retrieval techniques.
[0067] Furthermore, the knowledge graph format visually represents the relationships between entities, making it easier for users to intuitively understand complex information, while the text format utilizes natural language processing technology, allowing users to input questions in natural language, thus improving search flexibility. By combining these two formats, users can gain deeper insights, significantly improving the overall efficiency of data retrieval and the user experience.
[0068] In this embodiment, the input / output unit 12 is described as combining the functions of the theme information receiving unit, the question output unit, and the answer receiving unit, and the extraction unit 13 is described as combining the functions of the first and second extraction units. However, these functions may be handled by separate blocks.
[0069] (Example of processing) Next, an example of processing according to this embodiment will be described with reference to Figures 5 and 6. Figures 5 and 6 are diagrams illustrating an example of interaction with a respondent during knowledge extraction according to one embodiment of the present invention. Figure 5 shows an example of interaction at the start of processing. Figure 6 shows an example of interaction at the end of processing. Figures 5 and 6 show the content of the interaction between the knowledge extraction device 10 and the respondent. In the figures, the nozzles with the outlet on the left show the output content of the knowledge extraction device 10, and the nozzles with the outlet on the right show the response content from the respondent.
[0070] When the knowledge extraction process is started, the knowledge extraction device 10 executes the process. The following describes an example in which the following themes and prompts are set and the process is executed. Note that when setting prompts, local terminology is supplemented and supplementary prompts are generated (see steps S103 and S104). (Topic) About CC's BO (Local term) CC: Abbreviation for continuous casting machine BO: Abbreviation for breakout (Here, we'll refer to this as a "molten steel leak" problem.) (Basic prompt) Based on conversations between humans and large-scale language models, interviews were conducted. We extract knowledge from people. Therefore, large-scale language models should generate the necessary questions. Please continue asking questions that delve deeper into the reasons behind the answers given. You can only ask one question at a time. (Theme settings prompt: Theme information) The topic is "About CC's BO (Build-Oriented Agreement)." (Theme settings prompt: Local information + supplementary information) Here, CC stands for continuous casting machine, and BO stands for breakout. Breakout is a typical operational problem in continuous casting machines. This is a problem where molten steel is leaking.
[0071] First, the knowledge extraction device 10 outputs the following question (see Figure 5). "Let me begin with the question: What causes a blackout (BO)?" In response, the following answers were obtained from the respondents. "In cases of restrictive over-engineering, it's said that the cause is overheating."
[0072] In response to the above answer, the knowledge extraction device 10 outputs the following question. "I understand. So, what caused the seizure?" In response, the following answers were obtained from the respondents. "Powder and VC are involved."
[0073] In response to the above answer, the knowledge extraction device 10 outputs the following question. "So the type of powder and the pouring speed are factors, right?" How does powder come into play? In response, the following answers were obtained from the respondents. "I don't know." In this case, the knowledge extraction device 10 adds that "Vc" is the "casting speed" (see step S114).
[0074] Then, in response to the above answer, the knowledge extraction device 10 outputs the following question. "Understood. I have another question." What are the main measures to prevent blackouts? In this case, the knowledge extraction device 10 internally generated a candidate question, for example, "What are some possible reasons why powder might be involved?", to reconfirm the relationship with the answer. However, based on vector comparison, it was deemed to have a high degree of similarity and replaced with another question (What are the main measures to prevent BO?).
[0075] In this way, the knowledge extraction device 10 continues to output questions based on whether there is any overlap and the relevance of the themes.
[0076] Then, when the respondent enters the answer "I'm done," the knowledge extraction device 10 outputs the following question (see Figure 6). "Understood. What output format would you like?" In response, the following answers were obtained from the respondents. "Please output this as a knowledge graph."
[0077] Based on the above response, the knowledge extraction device 10 outputs the current exchange in a knowledge graph and stores it in the memory unit 18. The results of the exchange may be displayed, or they may be stored in the memory unit 18 without being displayed. Also, if the output format is predetermined, the question about the output format may be omitted, or the output format may be explicitly specified before proceeding with the processing.
[0078] According to the embodiment described above, a large-scale language model (LLM) is used to generate questions, and the questions are refined by evaluating question overlap and theme relevance through question vectorization. This suppresses the output of duplicate questions and questions that deviate from the theme, thereby reducing the duplication of answers related to the extraction of tacit knowledge and preventing answers that are off-topic. As a result, tacit knowledge can be extracted efficiently and reliably.
[0079] In conventional interview formats, it is necessary to coordinate schedules between the interviewer and the respondent, and the aggregation of tacit knowledge takes time. Furthermore, the subsequent work of summarizing the interview content is required, resulting in a time-consuming information processing problem per person. In contrast, with this embodiment, the extraction work can be performed solely at the respondent's convenience, and the aggregation is also performed by the device, thus enabling efficient extraction and storage of tacit knowledge.
[0080] (Variation 1) Next, a modification 1 of this embodiment will be described with reference to Figure 7. The configuration of the knowledge processing system according to this modification 1 is the same as that of the knowledge processing system 1 described above, so its description will be omitted. Figure 7 is a sequence diagram illustrating the flow of the knowledge extraction process according to modification 1 of the present invention.
[0081] When the knowledge extraction process is started in the knowledge extraction device 10, it accepts the setting input of theme information in the same manner as in steps S101 to S103, extracts local terms (first local terms) from the input theme information, and generates supplementary information (first supplementary information) that supplements the meaning of the extracted local terms (steps S201 to S203).
[0082] Subsequently, in this modified example 1, the prompt generation unit 14 sets the topic of the question related to the theme information as a preprocessing process for generating guidelines for questions related to the theme information (steps S204-S207).
[0083] Question topics refer to specific aspects or important elements related to the theme information that deepen the information and discussion. For example, if the theme information is breakout in continuous casting, examples of question topics include the causes of breakout, detection methods, preventive measures, impacts, and post-processing methods. By identifying one or more question topics related to the theme information as guidelines and repeatedly generating questions along these identified topics, deviations from the theme information can be prevented, unnecessary questions can be reduced, and a wide range of tacit knowledge related to the theme information can be efficiently acquired from various perspectives. Furthermore, systematically organizing the collected tacit knowledge related to the theme information makes it easier to refer to and use later, promoting knowledge accumulation. Thus, before generating prompts for the text generation device 20 to generate sentences that will serve as questions for acquiring tacit knowledge related to the theme information, guidelines for questions regarding the theme information are generated.
[0084] In step S204, the prompt generation unit 14 generates a prompt that encourages the generation of one or more topics related to the theme information, serving as a guideline for questions about the theme information. A specific example of this prompt that encourages the generation of question topics is: "Please generate one or more topics related to the theme information. The topics to be generated should reflect important aspects related to the theme information. Specifically, consider aspects such as cause, detection method, prevention measures, impact, and post-processing method. However, it is not limited to these. Please limit the number of topics to about five." The prompt generation unit 14 transmits the generated prompts to the document generation device 20 via the input / output unit 12.
[0085] When the text generation device 20 receives a prompt from the knowledge extraction device 10 via the input / output unit 22 to prompt the generation of a question topic, it generates a text to be used as the question topic as a guideline for a question regarding the theme information (step S205). In this embodiment, the text generation unit 23 generates a text to be used as the question topic using LLM.
[0086] Then, the knowledge extraction device 10 acquires the text generated by the text generation device 20 (one or more topics of the question) as guidelines for questions regarding the theme information (step S206).
[0087] After obtaining the question guidelines, the prompt generation unit 14 sets the question topic (step S207: topic setting step). Then, the prompt generation unit 14 generates a prompt that adds information related to the updated question topic to the theme information, local information, and supplementary information that supplements the meaning of local terms (step S208). The prompt generation unit 14 sends the generated prompt to the document generation device 20. After that, in the same manner as in steps S105 to S116, the device generates text, obtains vector information, confirms the question, extracts local terms and generates supplementary information, and determines whether or not termination information has been entered (steps S209 to S220).
[0088] If the control unit 17 determines that no termination information has been entered (step S220: No), it proceeds to step S207 and sets the topic for the next question. Then, it performs steps S208 to S219 again for the set topic. Conversely, if the control unit 17 determines that termination information has been entered (step S220: Yes), it proceeds to step S221.
[0089] In step S221, the knowledge information generation unit 16 generates knowledge information based on the information from a series of processes, including questions, answers (answer information), local information (second local information), and supplementary information (second supplementary information) from the series of processes (knowledge information generation step). The knowledge information generation unit 16 generates knowledge information in an output format specified by the user via voice input or the like, or in a pre-set output format.
[0090] According to the modified example 1 described above, a large-scale language model (LLM) is used to generate questions, and the questions are refined by evaluating question overlap and theme relevance through question vectorization. This suppresses the output of duplicate questions and questions that deviate from the theme, thus reducing the duplication of answers related to the extraction of tacit knowledge and preventing answers that are off-topic. As a result, tacit knowledge can be extracted efficiently and reliably.
[0091] Furthermore, according to Modification 1, one or more topics related to the theme information are pre-set as guidelines when generating questions, which prevents deviations from the theme information, reduces unnecessary questions, and enables the efficient acquisition of a wide range of tacit knowledge related to the theme information from various perspectives. Moreover, according to Modification 1, the tacit knowledge related to the collected theme information is systematically organized, making it easier to refer to and use after it is stored in the database, and enabling the acquisition of tacit knowledge even more efficiently.
[0092] (Modification 2) Next, a modified example 2 of this embodiment will be described with reference to Figure 8. Figure 8 is a sequence diagram illustrating the flow of the knowledge extraction process according to modified example 2 of the present invention. Figure 8 is a configuration diagram illustrating modified example 2, in which the sequences shown in Figures 4 and 7 are executed using multiple agents. In the following description, the processing flow will be explained using, for example, the processing flow shown in Figure 7.
[0093] An agent is a program that automatically performs specific tasks, and it collaborates with a generative AI via an API to utilize a large-scale language model to collect information and generate questions. The agent shown in Figure 8 consists of a question structuring agent 41, a local term reference agent 42, a question adjustment agent 43, a question duplicate confirmation agent 44, a theme relevance confirmation agent 45, and a knowledge organization agent 46, and it realizes the functions of each part included in the knowledge extraction device 10. Each agent communicates with each other via an API and exchanges information. The agent also collaborates with an input / output unit including a user interface, a server 51, and a database 50 via an API.
[0094] The question configuration agent 41 interacts with the user interface (input / output unit) via an API and is responsible for determining whether or not the theme information setting input has been accepted (see step S201). The question configuration agent 41 is responsible for identifying the topic of the question related to the accepted theme information (see steps S204 and S206). The question configuration agent 41 shares the theme information and the identified question topic information with other agents via the API. The question configuration agent 41 performs some of the functions of the input / output unit 12 and the prompt generation unit 14, which function as a theme information receiving unit.
[0095] The local term reference agent 42 accesses servers and databases via APIs, collaborates with the generation AI to utilize a large-scale language model, extracts local terms (first local terms or second local terms) from theme information shared by the question structuring agent 41 (step S202 or step S218), and generates supplementary information ((first) supplementary information or second supplementary information) for the extracted local information. The local term reference agent 42 performs some of the functions of the extraction unit 13 and the prompt generation unit 14.
[0096] The question adjustment agent 43 is responsible for generating prompts that include theme information, local information, and supplementary information, and for obtaining candidate sentences for the generated questions using a large-scale language model that is generated by learning the local information as features via an API (see steps S208 and S210). The question adjustment agent 43 performs some of the functions of the prompt generation unit 14 and is responsible for some of the functions of the prompt generation unit 14.
[0097] The Question Coordination Agent 43 also plays a role in coordinating other agents and assigning tasks to them via API. The Question Coordination Agent 43 instructs the Question Duplicate Confirmation Agent 44 and the Theme Relevance Confirmation Agent 45 via API to perform the evaluations in steps S213 and S214. Based on the evaluations from each agent, the Question Coordination Agent 43 determines whether the question content is duplicated and whether the text is related to the theme, and decides whether to re-obtain candidate texts for questions (see steps S208 and S210) or to confirm the candidate texts for questions as questions (see step S215). The Question Coordination Agent 43 performs some of the functions of the Question Review Unit 15.
[0098] The question adjustment agent 43 is responsible for determining, based on the question topics generated by the question configuration agent 41, whether to repeatedly generate questions that delve deeper into the respondent's answer information on the set question topic, or to move on to the next question topic (see step S220). The question adjustment agent 43 performs some of the functions of the control unit 17.
[0099] The Question Duplicate Confirmation Agent 44, based on instructions from the Question Adjustment Agent 43, evaluates whether the content of the candidate sentences for questions obtained by the Question Adjustment Agent 43 overlaps with previous questions (step S213), and is responsible for providing the evaluation to the Question Adjustment Agent 43. The Question Duplicate Confirmation Agent 44 performs a part of the function of the Question Review Unit 15. The theme-related confirmation agent 45, based on instructions from the question adjustment agent 43, evaluates the relevance of candidate question sentences obtained by the question adjustment agent 43 to theme-related information (step S214), and is responsible for providing the evaluation to the question adjustment agent 43. The theme-related confirmation agent 45 performs a part of the function of the question review unit 15.
[0100] The knowledge organization agent 46 is responsible for generating knowledge information based on questions, answers (answer information), local information, and supplementary information processed by a series of agents, and for saving the obtained knowledge information to the database via an API in knowledge graph format or text format (see step S221). The knowledge organization agent 46 performs the functions of the knowledge information generation unit 16.
[0101] (Method for building a question-answering system) A question-answering system can be constructed that utilizes the knowledge information generated by the knowledge extraction process described above. Specifically, a RAG (Retrieval Augmented Generation) environment can be constructed as a question-answering system by providing the generated knowledge information to a large-scale language model. In a RAG environment, the large-scale language model retrieves knowledge information from an external, implicit knowledge base to generate responses, enabling responses to questions based on the respondent's experience, insights, and expertise.
[0102] An example of the processing of this embodiment will be described with reference to Figures 9 and 10, showing its application to a question-answering system for troubleshooting manufacturing equipment in a steel mill.
[0103] Figure 9 shows (a) undocumented tacit knowledge information extracted from skilled technicians responsible for equipment maintenance at a steel mill, as a text-format summary, and (b) knowledge graph format, respectively, using the knowledge extraction process of the present invention. As shown in Figure 9, the knowledge extraction process of the present invention extracts undocumented tacit knowledge information from skilled technicians responsible for equipment maintenance at a steel mill, and stores it in a database as text-format summary information and as a knowledge graph.
[0104] Figures 10A and 10B show the results of a question-answering system being asked to provide answers regarding the causes and countermeasures in the case of damage to the main machinery (main drive and spindle) of a hot rolling line in a steel mill. Figure 10A shows an example of a response from a large-scale language model provided with knowledge information generated by the knowledge extraction process of the present invention in a RAG environment. Figure 10B shows an example of a response from a large-scale language model that was not provided with knowledge information in a RAG environment. Comparing Figures 10A and 10B, the answers shown in Figure 10B illustrate generally expected factors and countermeasures, while the answers shown in Figure 10A present specific factors and countermeasures, as well as the implementation status of countermeasures, based on the deeper insights and specialized knowledge of skilled equipment maintenance technicians.
[0105] By constructing a question-answering system that utilizes knowledge information generated through knowledge extraction processing, users can quickly acquire specialized knowledge by obtaining information based on the experience and knowledge of skilled technicians. Because tacit knowledge is based on practical experience, past failures and errors can be avoided, enabling safer and more efficient work. Furthermore, users can respond in a way that is appropriate to the actual situation, improving the speed of troubleshooting. In addition, by incorporating the tacit knowledge of skilled technicians into the question-answering system, knowledge can be passed on, allowing the next generation of users to inherit and utilize valuable knowledge.
[0106] (Recording medium) In one embodiment described above, Knowledge extraction device 10 or Sentence generator 20A program that executes a processing method can be recorded on a recording medium readable by a computer or other machine or wearable device (hereinafter referred to as "computer, etc."). By having a computer, etc. read and execute this program on the recording medium, the computer, etc. functions as a mobile device control device. Here, a recording medium readable by a computer, etc. refers to a non-temporary recording medium that stores information such as data and programs through electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer, etc. Examples of such recording media that can be removed from a computer, etc. include flexible disks, magneto-optical disks, CD-ROMs, CD-R / Ws, DVDs, BDs, DATs, magnetic tapes, and memory cards such as flash memory. In addition, recording media fixed to a computer, etc. include hard disks and ROMs. Furthermore, SSDs can be used as both a recording medium that can be removed from a computer, etc. and a recording medium that is fixed to a computer, etc.
[0107] Although one embodiment of the present invention has been described in detail above, the present invention is not limited to the above-described embodiment, and various modifications are possible based on the technical idea of the present invention. For example, the model given in the above-described embodiment is merely an example, and a different model may be used as needed, and the present invention is not limited by the description and drawings that constitute part of the disclosure of the present invention in this embodiment. For example, in the above-described embodiment, an example was described in which text is generated using a trained model in which features have been trained on a large-scale language model, but a configuration may also be used in which a pre-created (untrained) large-scale language model or a model based on a large-scale language model such as a model (framework) that uses RAG (Retrieval-augmented Generation) may be used.
[0108] Furthermore, while the above-described embodiment uses deep learning with a neural network as an example of machine learning, machine learning based on other methods may also be performed. For example, other supervised learning methods such as support vector machines, decision trees, Naive Bayes, and k-nearest neighbors may be used. Also, semi-supervised learning may be used instead of supervised learning.
[0109] Furthermore, in one embodiment, an example was described in which the extraction unit 13 extracts by separating into individual words, but other methods are also available, such as extracting by separating into words using morphological analysis, or extracting only nouns. In this embodiment, information was linked from a word list in which terms and explanations are paired, but information may also be supplemented by performing a RAG search on local documents using a large-scale language model. Finally, a prompt should be created that is provided to the large-scale language model in advance by combining the basic prompt, the theme setting prompt, and the supplementary prompt for the theme setting content.
[0110] Furthermore, in one embodiment, an example was described in which the knowledge extraction device 10 and the text generation device 20 read information from a common server 30 as a storage unit. However, they may each read information from a different server, or they may read information from the storage unit provided by each device.
[0111] Furthermore, in one embodiment, the terms "part" as described above can be replaced with "circuit" or the like. For example, the control unit can be replaced with a control circuit.
[0112] Furthermore, in one embodiment, the "parts" of each device can be configured such that at least some of the "parts" are implemented in different devices. In this case, for example, the question generation system and the knowledge extraction system are configured by combining different devices.
[0113] Further effects and modifications can be readily derived by those skilled in the art. Broader aspects of this disclosure are not limited to the specific details and representative embodiments expressed and described above. Therefore, various modifications are possible without departing from the spirit or scope of the overall concept of the invention as defined by the appended claims and their equivalents. [Industrial applicability]
[0114] As described above, the question generation method, knowledge extraction method, question answering system construction method, question generation program, knowledge extraction program, question generation system, and knowledge extraction system according to the present invention are suitable for efficiently and reliably extracting tacit knowledge. [Explanation of Symbols]
[0115] 1. Knowledge Processing System 10 Knowledge extraction device 11, 21 Communications Department 12, 22 Input / output section 13 Extraction part 14. Prompt generation unit 15. Question Review Department 16 Knowledge information generation section 17, 25 Control Unit 18, 26 Memory section 20 Sentence generator 23 Sentence generation section 24 Vectorization section
Claims
1. A method for generating questions performed by a computer, A theme information reception step that accepts theme information input for generating questions, From the received theme information, there is an extraction step to extract local information, A supplementary information generation step for generating supplementary information for the extracted local information, A text acquisition step involves acquiring text generated using a large-scale language model that learns the local information as a feature, based on a prompt including the theme information, the local information, and the supplementary information; A question refinement step in which, with respect to the aforementioned text, a question is refined to determine whether or not to set the text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information; A method for generating questions that includes this.
2. If, in the question review step, it is determined that the question should not be set as a question relating to the theme information, the text obtained in the text acquisition step is set as a supplementary prompt that includes the text as an inappropriate example and provides instructions to encourage improvement of the output. Based on the above prompt and the supplementary prompt, the text acquisition step is executed again. The question generation method according to claim 1.
3. As a guideline for questions regarding thematic information, it further includes a topic setting step to define topics for questions related to thematic information. In the text acquisition step, information related to the topic of the question that was set is added to the prompt. The question generation method according to claim 1 or 2.
4. The aforementioned topic setting step sets the topic of the question using a large-scale language model that is generated by learning local information as features, based on prompts that include theme information and prompt the generation of a question topic. The question generation method according to claim 3.
5. The aforementioned question review step compares the similarity between the information obtained by vectorizing the text and the information obtained by vectorizing other texts generated before the text being evaluated, in order to determine whether or not there is any overlap in content. The question generation method according to claim 1.
6. The aforementioned question review step determines whether the text is related to the set theme based on the relationship between the information obtained by vectorizing the text and the information obtained by vectorizing the theme information. The question generation method according to claim 1.
7. The extraction step involves extracting local information from the theme information by referring to a database that has been pre-stored with local terms and sentences containing those local terms. The question generation method according to claim 1.
8. The extraction step involves extracting the local information by referring to a database that stores information obtained by pre-vectorizing local terms and sentences containing those local terms, and comparing the similarity between that information and the information obtained by vectorizing the theme information. The question generation method according to claim 1.
9. A method of knowledge extraction performed by a computer, A theme information reception step that accepts theme information input for generating questions, The first extraction step involves extracting local information from the received theme information, A first supplementary information generation step in which first supplementary information is generated with respect to the extracted first local information, A text acquisition step in which text is acquired using a large-scale language model that is generated by learning the local information as a feature, based on a prompt including the theme information, the first local information, and the first supplementary information, A question refinement step in which, with respect to the aforementioned text, a question is refined to determine whether or not to set the text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information; A question output step that outputs a question that has been determined to be set as the question by the question review step, The response reception step involves receiving response information from respondents, A second extraction step is performed to extract second local information from the aforementioned response information, A second supplemental information generation step is performed to generate second supplemental information with respect to the extracted second local information, A knowledge information generation step that generates knowledge information by converting the information of a series of processes, including the aforementioned question, the aforementioned answer information, second local information, and second supplementary information, into a set format, A knowledge extraction method that includes this.
10. The method further includes an update step to update the prompt by adding the response information, the second local information, and the second supplementary information to the prompt, The document acquisition step, the question review step, the question output step, the answer reception step, the second extraction step, and the second supplemental information generation step are repeatedly executed based on the updated prompt. The knowledge extraction method according to claim 9.
11. In the knowledge information generation step, the knowledge information is converted into a knowledge graph format and a text format to generate the knowledge information. The knowledge extraction method according to claim 9 or 10.
12. A step of generating knowledge information converted into knowledge graph format and text format using the knowledge extraction method described in claim 11, The steps include saving the generated knowledge information to a database, A Retrieval Augmented Generation (RAG) environment is constructed as a question answering system, in which the knowledge information stored in the database is provided to a large-scale language model, enabling the acquisition of the knowledge information from the database and the generation of responses. Method for building a question-answering system.
13. A theme information reception step that accepts theme information input for generating questions, The extraction step involves referencing the memory unit to extract local information from the received theme information, A supplementary information generation step for generating supplementary information for the extracted local information, A text acquisition step involves acquiring text generated using a large-scale language model that learns the local information as a feature, based on a prompt including the theme information, the local information, and the supplementary information; A question refinement step in which, with respect to the aforementioned text, a question is refined to determine whether or not to set the text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information; A question generation program that causes a computer to execute a question.
14. A theme information reception step that accepts theme information input for generating questions, A first extraction step involves extracting first local information by referring to the memory unit from the received theme information, A first supplementary information generation step in which first supplementary information is generated with respect to the extracted first local information, A text acquisition step in which text is acquired using a large-scale language model that is generated by learning the local information as a feature, based on a prompt including the theme information, the first local information, and the first supplementary information, A question refinement step in which, with respect to the aforementioned text, a question is refined to determine whether or not to set the text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information; A question output step that outputs the questions set by the question review step, The response reception step involves receiving response information from respondents, A second extraction step is performed to extract second local information from the aforementioned response information, A second supplemental information generation step is performed to generate second supplemental information with respect to the extracted second local information, A knowledge information generation step that generates knowledge information by converting the information of a series of processes, including the aforementioned question, the aforementioned answer information, second local information, and second supplementary information, into a set format, A knowledge extraction program that causes a computer to execute a command.
15. A theme information receiving unit that accepts theme information input for generating questions, An extraction unit extracts local information by referencing the memory unit from the received theme information, A supplementary information generation unit generates supplementary information for the extracted local information, A text acquisition unit that acquires text generated using a large-scale language model that learns the local information as a feature based on a prompt including the theme information, the local information and the supplementary information, A question review unit determines whether or not to set the aforementioned text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information. A question generation system equipped with the following features.
16. A theme information receiving unit that accepts theme information input for generating questions, A first extraction unit extracts first local information by referring to the memory unit from the received theme information, A first supplementary information generation unit generates first supplementary information with respect to the extracted first local information, A text acquisition unit acquires text generated using a large-scale language model that learns the local information as features based on a prompt including the theme information, the first local information, and the first supplementary information. A question review unit determines whether or not to set the aforementioned text as a question relating to the theme information, based on the degree of overlap in content with other texts generated prior to the text being evaluated, and / or the degree of theme relevance to the theme information. A question output unit that outputs a question that the question review unit has determined to set as the question, The response reception department receives response information from respondents, A second extraction unit extracts second local information from the aforementioned response information, A second supplemental information generation unit generates second supplemental information for the extracted second local information, A knowledge information generation unit generates knowledge information by converting information from a series of processes, including the aforementioned question, the aforementioned answer information, second local information, and second supplementary information, into a set format. A knowledge extraction system equipped with the following features.