System
A system using AI to analyze and generate construction-related queries and improve through feedback addresses the shortage of skilled craftsmen by providing efficient and accurate construction skills transfer.
Patent Information
- Application Number
- JP2024126240
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-13
AI Technical Summary
The construction industry faces rising material costs and a shortage of skilled craftsmen, leading to a decline in experienced workers and difficulties in passing on construction skills effectively.
A system that accepts user questions and requests in natural language, analyzes them, generates relevant answers, and improves through user feedback, utilizing AI to provide construction skills efficiently and systematically.
Enables quick and accurate provision of construction knowledge, addressing the skills shortage and enhancing the transfer of architectural techniques.
Smart Images

Figure 2026023919000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Rising material costs and a shortage of craftsmen have become problems in the construction industry in recent years. As a result, the number of experienced carpenters is decreasing, making it difficult to pass on skills. The skills shortage is particularly pronounced in a wide range of fields, including carpentry, scaffolding, exterior construction, electrical work, equipment installation, interior construction, plastering, painting, demolition work, sheet metal work, and roof siding (tilt repair). To address this, there is a need for a means to efficiently and accurately provide construction skills and promote the passing on of skills. [Means for solving the problem]
[0005] The present invention solves the above problems with a system that includes an input means for accepting specific questions and requests about architecture from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intent, an information search and generation means for searching for related information based on the extracted keywords and generating answers, a display means for displaying the generated answers to the user, and a feedback processing means for accepting feedback from users and using it to improve the AI model. In other words, the system of the present invention solves the problems of a shortage of craftsmen and the difficulty of passing on skills by covering a wide range of architectural techniques using AI and providing those techniques systematically and efficiently.
[0006] "Input means" refers to a means for accepting specific questions and requests about architecture from users in natural language.
[0007] The "analysis means" is a means for analyzing questions and requests entered by users and extracting important keywords and intentions.
[0008] "Information search and generation means" refers to a means of searching for related information based on extracted keywords and generating appropriate answers.
[0009] The "display means" is a means for visually displaying the generated answer to the user.
[0010] A "feedback processing means" is a means of accepting feedback from users, analyzing its contents, and using it to improve the AI model.
[0011] "Tokenization" is a method of dividing a sentence into words and identifying each word.
[0012] "Morphological analysis" is a method of analyzing the morphemes (smallest semantic units) of words and phrases that make up a sentence and identifying parts of speech and other characteristics.
[0013] "Grammar structure analysis" is a means of analyzing the grammatical structure of a sentence and understanding the meaning of the entire sentence.
[0014] The "internal database" is a database that accumulates information on construction techniques and knowledge stored within the system.
[0015] An "external knowledge base" is a database or knowledge source that provides information about construction techniques and knowledge that exist outside the system.
[0016] An "AI model" is a model trained using artificial intelligence technology and is a means of generating appropriate answers to user questions and requests. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] MODE FOR CARRYING OUT THE INVENTION
[0039] This section describes an embodiment of the present invention. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system.
[0040] System Configuration
[0041] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[0042] 1. User Input
[0043] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[0044] 2. Sending Input
[0045] The terminal sends questions and requests entered by the user to the server.
[0046] 3. Natural Language Processing (NLP)
[0047] The server analyzes the received user question, specifically:
[0048] Tokenization: Divide the question into words.
[0049] Morphological analysis: Identifying the part of speech of each word.
[0050] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0051] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0052] 4. Information Retrieval and Answer Generation
[0053] The server uses the keywords extracted from the analysis results to search for related information from its internal database and external knowledge bases, and generates appropriate answers based on the information obtained as search results.
[0054] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0055] 5. View Answers
[0056] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0057] 6. Get feedback
[0058] The user provides feedback on the displayed answers, including a rating of whether the answer was helpful or if more detail is needed.
[0059] 7. Processing Feedback
[0060] The device sends user feedback to the server, which analyzes it and uses it to improve the AI model, allowing the system to continuously learn and provide better answers in the future.
[0061] Specific examples
[0062] For example, the user inputs "Please tell me the basics of electrical wiring."
[0063] Terminal: Sends user input to the server.
[0064] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[0065] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0066] Terminal: Displays the generated answer to the user.
[0067] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[0068] Server: Receives feedback and helps improve the AI model.
[0069] In this way, the system of the present invention can solve the problems of a shortage of craftsmen and the transfer of skills by providing information on construction technology quickly and appropriately, and contribute to improving technology throughout the construction industry.
[0070] The processing flow will be explained below.
[0071] Step 1: User enters question
[0072] User: Enters specific questions or requests about construction into the terminal.
[0073] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[0074] Step 2: Sending Input
[0075] Terminal: Sends questions and requests entered by the user to the server.
[0076] Specifically, the user input data is sent to the server as an HTTP request.
[0077] Step 3: Tokenize the Question
[0078] Server: Tokenizes the received question.
[0079] Specifically, the question is divided into words and phrases.
[0080] For example, it is divided into "wooden house," "foundation work," and "procedure."
[0081] Step 4: Morphological analysis
[0082] Server: Performs morphological analysis of the question.
[0083] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[0084] Step 5: Grammatical structure analysis
[0085] Server: Analyzes the grammatical structure of the question.
[0086] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[0087] Step 6: Finding information and generating answers
[0088] Server: Searches for relevant information based on the analysis results and generates answers.
[0089] Specifically, it retrieves relevant data from internal databases and external knowledge bases.
[0090] Based on the search results, answers that are easy for users to understand are generated.
[0091] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0092] Step 7: Submit your response
[0093] Server: Sends the generated answer to the device.
[0094] Step 8: View your answers
[0095] Terminal: Displays the received answer to the user.
[0096] As a specific operation, the answer content is displayed on the screen.
[0097] Step 9: Get feedback
[0098] User: Provide feedback on the displayed answers.
[0099] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[0100] Step 10: Submit your feedback
[0101] Terminal: Sends user feedback to the server.
[0102] Step 11: Processing feedback
[0103] Server: Analyzes the feedback received and helps improve the AI model.
[0104] Specifically, the content of the feedback is analyzed and fed back to the AI model as adaptable data.
[0105] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[0106] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[0107] Example 1
[0108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0109] In recent years, while the demand for technical information on architecture has increased, the shortage of engineers with specialized knowledge has become a problem. The difficulty of transferring skills and the need for efficient information provision are also increasing. In response to these issues, there is a demand for systems that can provide information on architectural technology quickly and accurately, and that can continuously improve the system's performance based on user feedback.
[0110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0111] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, a means for transmitting the input questions and requests to the server, an analysis means for extracting important keywords and intent by tokenizing, morphologically analyzing, and grammatically analyzing the received questions, an information search and generation means for searching for related information from an internal database and an external knowledge base based on the extracted keywords and generating an answer, a means for transmitting and displaying the generated answer to the user, and a feedback processing means for accepting and analyzing feedback from users and using it to improve the AI model. This allows users to quickly and accurately obtain information about architecture technology, and enables the system to continuously learn and improve its performance.
[0112] "Input means" refers to a device or interface for accepting specific questions or requests about architecture from users in natural language.
[0113] The "server" is a central information processing device that analyzes information sent by users and generates appropriate responses.
[0114] "Tokenization" is a natural language processing process that divides an input question or request into words.
[0115] "Morphological analysis" is a natural language processing technique for identifying the part of speech of each word.
[0116] "Grammar structure analysis" is a process in natural language processing that analyzes the grammatical structure of a sentence and extracts important keywords and intent.
[0117] The "analysis means" is a device or software that analyzes the received question through tokenization, morphological analysis, and grammatical structure analysis to extract important keywords and intent.
[0118] "Information search and generation means" refers to a device or software that searches for related information from internal databases and external knowledge bases based on extracted keywords and generates answers.
[0119] A "display means" is a device or interface for visually displaying the generated answers to the user.
[0120] "Feedback processing means" refers to a device or software that accepts and analyzes feedback from users and uses it to improve the AI model.
[0121] MODE FOR CARRYING OUT THE INVENTION
[0122] The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. An embodiment of the present invention is described below.
[0123] System Configuration
[0124] This system consists of a device operated by the user (e.g., PC, smartphone, tablet) and a server for processing information sent from these devices.
[0125] User Input
[0126] The user inputs specific questions or requests about construction in natural language using the input means of the terminal. For example, a question such as "Please tell me the procedure for laying the foundations of a wooden house" is input using the keyboard or touch screen of the terminal.
[0127] Sending Input
[0128] The device sends the questions and requests entered by the user to the server. In this process, the device packages the entered text data in JSON format or similar and sends it using an HTTP request to send it to the server over the network.
[0129] Natural Language Processing (NLP)
[0130] The server analyzes the received user question. Specifically, it performs the following analyses: tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent, such as "wooden house," "foundation work," and "procedure."
[0131] Information retrieval and answer generation
[0132] The server uses keywords extracted from the analysis results to search for related information from its internal database and external knowledge base. Based on the information obtained as a search result, it generates an appropriate answer. For example, it generates an answer such as, "The steps for laying the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0133] Show Answers
[0134] The server sends the generated answer to the terminal, which then visually displays the received answer to the user, allowing the user to check the answer on the terminal screen.
[0135] Get feedback
[0136] The user provides feedback on the displayed answer, for example, with options such as "helpful" or "more information," and the user enters a prompt such as:
[0137] "It was helpful"
[0138] "I want more details."
[0139] Processing Feedback
[0140] The device sends user feedback to a server, which analyzes it and uses it to improve the AI model. The feedback data can be saved and used as training data for the machine learning model to improve the accuracy of answers in future searches.
[0141] Specific examples
[0142] For example, if a user types "What are the basics of electrical wiring?", the following steps are taken:
[0143] Terminal: The user types in a question and sends it to the server.
[0144] Server: Receives the question, performs tokenization, morphological analysis, and grammatical structure analysis, and extracts important keywords (e.g., "electrical wiring" and "basic knowledge").
[0145] Server: Searches internal databases and external knowledge bases to generate appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0146] Terminal: Receives the generated answers and displays them on the screen.
[0147] User: Review the answer and provide feedback such as "this was helpful" or "I'd like more information."
[0148] Device: Sends feedback to the server.
[0149] Server: Analyzes feedback and improves the AI model.
[0150] By implementing the system in this way, information on building technology can be provided quickly and appropriately, and user feedback can be utilized to continuously improve the system's performance.
[0151] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0152] Step 1: User Input
[0153] The user inputs specific questions or requests about construction into a terminal (PC, smartphone, tablet). A question such as "Please tell me the steps for laying the foundations for a wooden house" is input using a keyboard or touch screen. The input text data is received by the input means of the terminal. (Input) The user's question text. (Output) The input question as text data.
[0154] Step 2: Sending Input
[0155] The device sends questions and requests entered by the user to the server. Specifically, the entered text data is packaged in JSON format or similar over the network and sent to the server using an HTTP request. (Input) Question as text data. (Output) HTTP request sent to the server.
[0156] Step 3: Natural Language Processing (NLP)
[0157] The server analyzes the received user question and performs the following specific processing:
[0158] Tokenization: The server divides the question into words. For example, the sentence "Please tell me the procedure for laying the foundation for a wooden house" is divided into "wooden," "house," "foundation," "construction," "procedure," "tell me," and "please." (Input) The question as text data. (Output) A tokenized word list.
[0159] Morphological analysis: The server identifies the part of speech of each word. For example, "wooden structure (noun)", "house (noun)", "foundation (noun)", "construction (noun)", "procedure (noun)", "teaching (verb)". (Input) A tokenized word list. (Output) A word list with parts of speech assigned.
[0160] Grammatical structure analysis: The server analyzes the grammatical structure of the sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords. (Input) A list of words with parts of speech assigned. (Output) A list of extracted keywords.
[0161] Step 4: Information retrieval and answer generation
[0162] Based on the analysis results, the server uses the extracted keywords to search for relevant information from its internal database and external knowledge bases. The specific processes are as follows:
[0163] Internal database search: The server searches the internal database (technical manuals and industry standard information) using keywords such as "wooden house," "foundation work," and "procedure." (Input) Extracted keyword list. (Output) Information obtained from the internal database.
[0164] External knowledge base search: The server searches an external knowledge base (such as public documents on the Internet). (Input) Extracted keyword list. (Output) Information obtained from the external knowledge base.
[0165] Answer generation: Generate an answer in a format appropriate for the user from the search results. For example, generate an answer in the format "The steps for laying the foundation for a wooden house are as follows: 1. Level the site 2. Pour concrete for the foundation 3. Install rebar 4. Install formwork 5. Re-pour concrete." (Input) Information from the search results. (Output) Generated answer text.
[0166] Step 5: View your answers
[0167] The server sends the generated answer to the device. The device visually displays the received answer to the user. Specifically, the server packages the answer in JSON format and sends it back to the device as an HTTP response. The device displays the answer on the screen so that the user can confirm it. (Input) Generated answer. (Output) Answer displayed on the device.
[0168] Step 6: Getting feedback
[0169] The user provides feedback on the displayed answer. For example, options such as "Helpful" or "Need more details" are displayed and the user selects one. The device receives this feedback and sends it to the server. (Input) User feedback. (Output) Feedback sent to the server.
[0170] Step 7: Processing feedback
[0171] The server analyzes the received feedback and uses it to improve the AI model. Specifically, it stores the feedback data and uses it as training data for the machine learning model. This data can be used to improve the accuracy of the AI model. (Input) User feedback. (Output) Improved AI model, improving the accuracy of answers from next time onwards.
[0172] (Application example 1)
[0173] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0174] When field workers perform tasks that require specialized knowledge, they need to be provided with information quickly and accurately. However, in the past, they often relied on specialized books and manuals, which lacked immediacy and reduced work efficiency. Furthermore, there was a lack of mechanisms for utilizing feedback to continuously improve the system. This has led to concerns that this increases the burden on workers and reduces the accuracy and efficiency of work, especially in sites with a wide range of complex tasks, such as factories.
[0175] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0176] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intent, an information search and generation means for searching for related information based on the extracted keywords and generating answers, a display means for displaying the generated answers to the user, a feedback processing means for receiving feedback from the user and using it to improve the AI model, and a smart gadget equipped with a voice input device and a display device as a terminal used by the user, wherein the input means accepts voice input and the display means has the function of displaying answers on the display of the smart gadget. This allows on-site workers to input questions hands-free and receive answers quickly, improving work efficiency and enabling continuous improvement of the system.
[0177] 1. "Specific questions or requests regarding construction"
[0178] "Specific questions or requests regarding construction" refers to specific information or instructions that a user requests regarding the design, construction, maintenance, etc. of a building.
[0179] 2. "Input methods that accept natural language"
[0180] "Input means that accepts natural language" refers to devices or software that accept questions or requests entered by the user using everyday language or technical terms.
[0181] 3. “Analysis means”
[0182] "Analysis means" refers to the technology and functions used to process input data and extract important keywords and the intent of the question.
[0183] 4. "Extract important keywords and intent"
[0184] "Extracting key keywords and intent" means identifying specific words and phrases, as well as the purpose and meaning behind them, from a user's question or request.
[0185] 5. "Information retrieval and generation methods"
[0186] "Information search and generation means" refers to the technology and functions that search for related information from databases and external information sources based on extracted keywords and generate appropriate answers.
[0187] 6. "A means of displaying the generated answer to the user"
[0188] "Display means for displaying the generated answer to the user" refers to technology or devices for visually presenting the answer generated by the system on a display of a computer, mobile device, etc.
[0189] 7. "Feedback Processing Means"
[0190] "Feedback processing means" refers to the technology and functions that receive evaluations and opinions on responses from users, analyze them, and use them to improve the system.
[0191] 8. "Smart Gadgets"
[0192] "Smart gadgets" are portable devices with internet connectivity and various functions, such as smart glasses and smartphones.
[0193] 9. "Voice input device"
[0194] A "voice input device" is a microphone and associated software that allows a user to speak questions or commands.
[0195] 10. “Display device”
[0196] "Display device" means a display or screen for visually displaying information or data.
[0197] An embodiment of the present invention will be described. The present invention is a system that enables factory workers to use smart gadgets to ask questions about building and equipment maintenance in real time and quickly receive appropriate answers.
[0198] System Configuration
[0199] The system consists of the following main components:
[0200] Smart gadgets (e.g., smart glasses)
[0201] server
[0202] Database
[0203] display device
[0204] 1. User Input
[0205] The user wears a smart gadget and inputs questions or requests in natural language using voice, such as "Please tell me the basic procedures for equipment inspection."
[0206] 2. Sending Input
[0207] A voice input device in the smart gadget converts the user's voice into text data (using voice recognition software) and sends the text data to a server.
[0208] 3. Natural Language Processing (NLP)
[0209] The server analyzes the received user question using the following techniques:
[0210] Tokenization: Divide the question into words.
[0211] Morphological analysis: Identifying the part of speech of each word.
[0212] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0213] For example, "equipment," "inspection," and "basic procedures" are extracted as main keywords.
[0214] 4. Information Retrieval and Answer Generation
[0215] The server uses keywords extracted from the analysis results to search for relevant information from its internal database and external knowledge bases, and generates appropriate answers using a generative AI model based on the information obtained as search results.
[0216] Example: "The basic steps for equipment inspection are as follows: 1. Visual inspection 2. Operation check 3. Connection check 4. Lubrication check"
[0217] 5. View Answers
[0218] The server transmits the generated answer to the smart gadget, which visually displays the answer on its display device.
[0219] 6. Get feedback
[0220] Users provide feedback on the displayed answers, including whether they found the answer helpful or if they need more detail.
[0221] 7. Processing Feedback
[0222] Through the feedback collection function of the smart gadget, user feedback is sent to the server, which analyzes this feedback and helps improve the generative AI model.
[0223] Usage example
[0224] For example, the user inputs, "Please tell me the basic procedures for equipment inspection."
[0225] Smart gadget: Converts user's voice input into text and sends it to the server.
[0226] Server: Performs natural language processing and analyzes the intent of the question (e.g., "equipment," "inspection," "basic procedures").
[0227] Server: Searches internal databases and external knowledge bases and generates appropriate answers using generative AI models (e.g., "The basic steps for equipment inspection are as follows...").
[0228] Smart Gadget: Displays the generated answers to the user.
[0229] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[0230] Server: Receives feedback and helps improve the generative AI model.
[0231] In this way, the system of the present invention allows field workers to input questions hands-free and receive quick answers, thereby improving work efficiency and enabling continuous improvement of the system.
[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0233] Step 1:
[0234] The user uses a smart gadget to provide voice input. For voice input, the user speaks, "Please tell me the basic procedures for equipment inspection." The smart gadget converts this voice into text data. The input is voice data, and the output is text data. Specifically, the gadget converts the voice data into text using voice recognition software (e.g., Google Speech-to-Text API).
[0235] Step 2:
[0236] The terminal transmits text data to the server. This is done using a network. In this case, the input is text data from the user, and the output is data sent to the server. In concrete terms, the smart gadget transmits text data to the server using a wireless network.
[0237] Step 3:
[0238] The server applies natural language processing (NLP) to the received text data. This processing includes tokenization, morphological analysis, and grammatical structure analysis. The input is the text data, and the output is the key keywords and the intent of the question. Specifically, the server uses the spaCy library to analyze the text data and extract key keywords.
[0239] Step 4:
[0240] The server searches for related information from the internal database and external knowledge base based on the analysis results. The input is the extracted keywords, and the output is related information as a search result. Specifically, it searches the internal database using SQL queries and retrieves information from the external knowledge base using APIs.
[0241] Step 5:
[0242] The server uses a generative AI model to generate an appropriate answer. The input is relevant information, and the output is a response to the user. Specifically, the server uses Hugging Face's transformers library to input relevant information into the generative AI model as a prompt, and generates an answer. Take the prompt "Please tell me the basic procedures for equipment inspection" as an example.
[0243] Step 6:
[0244] The server sends the generated answer to the smart gadget. The input is the answer, and the output is the display data for the smart gadget. Specifically, the server sends text data to the smart gadget via the network.
[0245] Step 7:
[0246] The smart gadget displays the received answer on a display device. The input is the text data of the answer, and the output is visual information displayed on the display. Specifically, the smart gadget renders the received text data on the display.
[0247] Step 8:
[0248] The user provides feedback on the displayed answer by speaking a rating such as "helpful" or "I'd like more information." The input is voice data, and the output is text feedback. Specifically, as with the initial voice input, speech recognition software is used to convert the speech into text.
[0249] Step 9:
[0250] The terminal transmits the feedback text data to the server. The input is the feedback text data from the user, and the output is the data to be transmitted to the server. In concrete terms, the smart gadget transmits the feedback data to the server using a wireless network.
[0251] Step 10:
[0252] The server analyzes the feedback and uses it to improve the AI model. The input is the text feedback, and the output is an improved AI model. Specifically, the server uses the feedback data to train and update the AI model.
[0253] This allows field workers to input questions hands-free and receive quick answers, improving work efficiency and enabling continuous improvement of the system.
[0254] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0255] MODE FOR CARRYING OUT THE INVENTION
[0256] An embodiment of the present invention will be described. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. In addition, it combines this with an emotion engine that recognizes the user's emotions.
[0257] System Configuration
[0258] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[0259] 1. User Input
[0260] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[0261] 2. Sending Input
[0262] The terminal sends questions and requests entered by the user to the server.
[0263] 3. Natural Language Processing (NLP)
[0264] The server analyzes the received user question, specifically:
[0265] Tokenization: Divide the question into words.
[0266] Morphological analysis: Identifying the part of speech of each word.
[0267] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0268] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0269] 4. Emotion analysis
[0270] The server uses an emotion engine to recognize the user's emotion in the question. The following process is performed:
[0271] Emotion recognition: Extracting user emotions (e.g., confusion, excitement, anger, etc.) from input text.
[0272] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[0273] For example, if a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[0274] 5. Information Retrieval and Answer Generation
[0275] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[0276] Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[0277] Generate answers that are interesting and considerate of the user's emotions.
[0278] For example, an answer might be generated along the lines of, "Please understand that foundation work is difficult, and refer to the following steps."
[0279] 6. View Answers
[0280] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0281] 7. Get feedback
[0282] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[0283] 8. Submitting Feedback
[0284] The device sends the feedback from the user to the server.
[0285] 9. Processing Feedback
[0286] The server analyzes this feedback and uses it to improve the AI model. Specifically, it analyzes the feedback, extracts evaluation information, and reflects it in the AI model for re-learning. This allows the system to continuously improve.
[0287] Specific examples
[0288] For example, if a user types "What are the basics of electrical wiring?":
[0289] Terminal: Sends user input to the server.
[0290] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[0291] Server: Uses an emotion engine to recognize the user's emotions (e.g., confusion, interest).
[0292] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0293] Terminal: Displays the generated answer to the user.
[0294] User: Provide feedback on the answer.
[0295] Server: Receives feedback and helps improve the AI model.
[0296] In this way, the system of the present invention not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take into consideration the user's feelings. This effectively solves the problems of passing on construction techniques and the shortage of craftsmen, and contributes to the improvement of technology in the construction industry as a whole.
[0297] The processing flow will be explained below.
[0298] Step 1: User enters question
[0299] User: Enters specific questions or requests about construction into the terminal.
[0300] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[0301] Step 2: Sending Input
[0302] Terminal: Sends questions and requests entered by the user to the server.
[0303] Specifically, the user input data is sent to the server as an HTTP request.
[0304] Step 3: Tokenize the Question
[0305] Server: Tokenizes the received question.
[0306] Specifically, the question is divided into words and phrases.
[0307] For example, it is divided into "wooden house," "foundation work," and "procedure."
[0308] Step 4: Morphological analysis
[0309] Server: Performs morphological analysis of the question.
[0310] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[0311] Step 5: Grammatical structure analysis
[0312] Server: Analyzes the grammatical structure of the question.
[0313] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[0314] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0315] Step 6: Sentiment Analysis
[0316] Server: Recognizes the user's emotions contained in the question using an emotion engine.
[0317] Specifically, it extracts the user's emotions (e.g., confusion, excitement, anger, etc.) from the input text and classifies them into predefined emotion categories.
[0318] Example: If a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[0319] Step 7: Information retrieval and answer generation
[0320] Server: Searches for relevant information and generates answers based on the analysis results and emotion recognition results.
[0321] Specifically, based on the extracted keywords and recognized emotions, the system retrieves relevant data from an internal database and an external knowledge base.
[0322] Generate answers that are interesting and considerate of the user's emotions.
[0323] For example, an answer like "Please understand that foundation work is difficult, and refer to the following steps." is generated.
[0324] Step 8: Submit your response
[0325] Server: Sends the generated answer to the device.
[0326] Step 9: View your answers
[0327] Terminal: Displays the received answer to the user.
[0328] As a specific operation, the answer content is displayed on the screen.
[0329] Step 10: Get feedback
[0330] User: Provide feedback on the displayed answers.
[0331] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[0332] Step 11: Submit your feedback
[0333] Terminal: Sends user feedback to the server.
[0334] Step 12: Processing feedback
[0335] Server: Analyzes the feedback received and helps improve the AI model.
[0336] Specifically, the feedback content is analyzed and fed back to the AI model as adaptable data.
[0337] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[0338] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[0339] Example 2
[0340] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0341] In the modern construction industry, there are limited means of instantly obtaining technical information and procedures related to construction, creating a demand for fast, accurate answers to technical questions and problems. It is also important to increase user satisfaction by providing answers that take the user's emotions into consideration. However, conventional systems struggle to respond in a way that takes user emotions into account, and they lack mechanisms for appropriately utilizing user feedback to improve the system. Therefore, there is a need for a system that can analyze user questions and requests, recognize emotions, generate and display appropriate answers, and further utilize feedback.
[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0343] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an information search and generation means for searching for related information based on the extracted keywords and the user's emotions and generating answers, a display means for displaying the generated answers to the user, and a feedback processing means for receiving feedback from the user and using it to improve the AI model. This makes it possible to quickly and appropriately answer user questions, provide answers that take user emotions into consideration, and utilize feedback to continuously improve the system.
[0344] "Input means" refers to devices or functions that accept specific questions or requests about architecture from users in natural language.
[0345] "Analysis means" refers to a device or function that analyzes questions or requests entered by users and extracts important keywords and intentions.
[0346] "Information search and generation means" refers to devices or functions that search for related information and generate answers based on the keywords extracted by the analysis means and the user's emotions.
[0347] "Display means" refers to a device or function that visually displays the generated answer to the user.
[0348] A "feedback processing means" is a device or function that accepts feedback provided by users and uses that feedback to help improve the AI model.
[0349] "Tokenization" is the process of dividing an input natural language sentence into words.
[0350] "Morphological analysis" is the process of identifying the part of speech of each tokenized word.
[0351] "Grammar structure analysis" is the process of analyzing the grammatical structure of an input sentence and extracting important keywords and intent.
[0352] "Emotion recognition" is the process of extracting a user's emotions from input text.
[0353] "Emotion classification" refers to the process of classifying extracted emotions into predefined emotion categories.
[0354] An "internal database" is a database that manages information and data stored within the system.
[0355] An "external knowledge base" is an information source or database that is externally accessible, such as the Internet.
[0356] An "AI model" is a model that uses artificial intelligence to learn and perform processes such as answering questions and recognizing emotions.
[0357] "Feedback" refers to the ratings and opinions that users provide in response to the answers displayed.
[0358] The system of the present invention allows users to input specific questions or requests about architecture in natural language, analyzes them, and generates and displays answers. Furthermore, it can collect feedback from users and use it to improve the system. The system consists of the following main components:
[0359] System Configuration
[0360] 1. Hardware configuration:
[0361] Terminal: A device such as a PC, smartphone, or tablet that allows users to input questions or requests and display answers.
[0362] Server: A high-performance cloud server or on-premise server with high data processing capacity and fast response is recommended.
[0363] 2. Software configuration:
[0364] Natural language processing engine: An engine that tokenizes text, analyzes morphology, and analyzes grammatical structures (e.g., Google NLP).
[0365] Emotion engine: An engine for recognizing and classifying user emotions (e.g., IBM Watson Tone Analyzer).
[0366] Database management system: Databases such as MySQL and MongoDB are used as internal databases.
[0367] AI model: An artificial intelligence model that analyzes a user's question or request and generates an answer. Use a generative AI model.
[0368] Operation procedure and processing contents
[0369] 1. User input:
[0370] Users operate their own terminals to input specific questions or requests about construction in natural language, such as "Please tell me the steps for laying the foundations for a wooden house."
[0371] 2. Sending input:
[0372] The device sends the user-entered questions or requests to a server over the Internet. The data is sent as an HTTP request.
[0373] 3. Natural Language Processing (NLP):
[0374] The server analyzes the received user question, specifically by performing the following steps:
[0375] Tokenization: Divide the input question into words.
[0376] Morphological analysis: Identifying the parts of speech of segmented words.
[0377] Grammatical structure analysis: Analyze grammatical structures and extract important keywords and intent.
[0378] 4. Emotion analysis:
[0379] The server uses an emotion engine to recognize the user's emotion in the question, which includes the following steps:
[0380] Emotion Recognition: Extracting user emotions from input text.
[0381] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[0382] 5. Information retrieval and answer generation:
[0383] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[0384] Information retrieval: Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[0385] Answer generation: Based on the searched information, an answer is generated that takes into consideration the user's feelings.
[0386] 6. Show Answer:
[0387] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0388] 7. Getting feedback:
[0389] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[0390] 8. Submitting Feedback:
[0391] The device sends the user's feedback to the server, again as an HTTP request.
[0392] 9. Handling feedback (retraining the AI model):
[0393] The server analyzes the received feedback and extracts evaluation information, which is then used to retrain the AI model and improve response accuracy from the next time onwards.
[0394] Specific examples
[0395] For example, if a user types "Please tell me the basics of electrical wiring," the following will be processed:
[0396] User: Enters a question into the terminal and sends it.
[0397] Server: Receives input and performs natural language processing and sentiment analysis.
[0398] Server: Searches for relevant information from internal databases and external knowledge bases and generates appropriate answers.
[0399] Terminal: Display the generated answer.
[0400] User: Provide feedback on the answer.
[0401] Server: Receives feedback and helps retrain the AI model.
[0402] This system not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take the user's emotions into consideration. This will effectively solve the problems of passing on construction techniques and the shortage of craftsmen, and contribute to improving the technology of the entire construction industry.
[0403] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0404] Step 1:
[0405] The user uses the input means to input specific questions or requests about construction in natural language. For example, they may input a question such as, "Please tell me the steps for laying the foundations for a wooden house." This input is saved as text data on the terminal.
[0406] Step 2:
[0407] The terminal sends the text data entered by the user to the server. Specifically, the data is sent to the server as an HTTP request. At this time, the input data is temporarily stored in the terminal's buffer memory. The input is the user's question text, and the output is the sent HTTP request.
[0408] Step 3:
[0409] The server analyzes the received text data. First, it uses a natural language processing engine to tokenize the question. As a result of tokenization, the input sentence can be divided into words, and the output is a word list. For example, it may be divided into words such as "wooden house," "foundation work," "procedure," "tell me," and "please."
[0410] Step 4:
[0411] The server performs morphological analysis. As a result of the morphological analysis, it is possible to identify the part of speech for each word. For example, "wooden house" is recognized as a noun and "teach me" as a verb. The input is a list of tokenized words, and the output is a list of words with identified parts of speech.
[0412] Step 5:
[0413] The server performs grammatical analysis. This analyzes the grammatical structure of the entire sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as key keywords. The results of the grammatical analysis are the extracted keywords and a structural analysis tree.
[0414] Step 6:
[0415] The server uses an emotion recognition engine to recognize the user's emotions. During the emotion recognition process, the server extracts the user's emotional characteristics (e.g., confusion, excitement, anger) from the text data and classifies them into categories. For example, the emotion "confusion" is extracted from the sentence "I'm having trouble because the foundation work on a wooden house is difficult." The input is the analyzed text data, and the output is the recognized emotion and its category.
[0416] Step 7:
[0417] The server performs an information search based on the analysis results and emotion recognition results. It searches its internal database and external knowledge base to obtain relevant information. For example, it searches the internal database for technical information using keywords such as "wooden house," "foundation work," and "procedure." The information obtained at this stage is output in the form of text or images.
[0418] Step 8:
[0419] The server generates an answer based on the search results. Using a generative AI model, it creates an answer that takes into account the acquired information and the user's emotions. For example, an answer such as "Please understand that foundation construction is difficult, and refer to the following steps" may be generated. The input is the search results and emotion data, and the output is the answer text for the user.
[0420] Step 9:
[0421] The server sends the generated answer to the terminal. The terminal displays the received answer to the user. Specifically, the answer is displayed in text format on the terminal screen. The input is the answer data from the server, and the output is the display for the user.
[0422] Step 10:
[0423] The user provides feedback on the displayed answer, including an evaluation of whether the answer was helpful or whether further explanation is needed. The feedback is again input into the terminal as text data.
[0424] Step 11:
[0425] The device sends the user's feedback to the server, again as an HTTP request. The input is the user's feedback text, and the output is the sent HTTP request.
[0426] Step 12:
[0427] The server receives and analyzes the feedback data. As a result of the analysis, specific evaluation information is extracted and used to retrain the AI model. This improves response accuracy from the next time onwards. The input is the feedback data, and the output is an updated AI model.
[0428] (Application example 2)
[0429] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0430] Lack of technical knowledge and information on construction sites, especially complex work procedures and troubleshooting, makes it difficult to quickly implement them. Furthermore, the inability to provide appropriate advice that takes into account the feelings of workers makes it difficult to work efficiently. Furthermore, there is a lack of means to continuously improve the system based on user feedback.
[0431] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0432] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an emotion analysis means for analyzing emotions using an emotion engine that recognizes the user's emotions, an information search and generation means for searching for related information and generating answers based on the extracted keywords and recognized emotions, a display means for displaying the generated answers to the user, and a feedback processing means for accepting feedback from users and using it to improve the AI model. This enables workers at construction sites to respond quickly to questions and requests, provides appropriate advice that takes workers' emotions into consideration, and continuously improves the system, enabling efficient and effective work.
[0433] "Input means" refers to a device or system that accepts specific questions or requests about architecture from users in natural language.
[0434] The "analysis means" is a device or program that analyzes the input question or request and extracts important keywords and intentions.
[0435] "Emotion analysis means" refers to a device or system for analyzing emotions from input text using an emotion engine that recognizes the emotions of a user.
[0436] "Information search and generation means" refers to a device or program that searches for related information based on the extracted keywords and recognized emotions and generates answers.
[0437] A "display means" is a device or system that visually displays the generated answers to the user.
[0438] A "feedback processing means" is a device or program that accepts feedback from users and uses it to improve the AI model.
[0439] The present invention is a system in which a factory robot responds to questions and requests from workers at a construction site and provides advice that takes into consideration their emotions. An embodiment of the system is described below.
[0440] System Configuration
[0441] The system includes the following components:
[0442] 1. Input Method
[0443] The factory robot can receive specific questions and requests about construction from workers in natural language. For example, the factory robot can receive a question from a worker such as, "Please tell me the key points of foundation work."
[0444] 2. Analysis method
[0445] The factory robot analyzes the questions and requests it receives. This analysis involves tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent. For example, keywords such as "foundation work" and "key points" are extracted.
[0446] 3. Emotion analysis means
[0447] The factory robot uses an emotion engine to analyze the emotions of workers from input text. Specifically, the input text is input into the emotion engine, and emotions such as "confusion" or "interest" are recognized.
[0448] 4. Information retrieval and generation methods
[0449] Based on the extracted keywords and the recognized emotions, the factory robot searches for relevant information from its internal database and external knowledge base, and generates an appropriate answer taking into account the search results and the emotions of the worker. For example, it generates an answer in the form of "The main points of foundation work are as follows..."
[0450] 5. Display means
[0451] The generated answers are displayed and presented visually or audibly by the factory robot to the worker, allowing the worker to obtain the appropriate information.
[0452] 6. Feedback Processing Methods
[0453] Workers can provide feedback on the displayed answers, which helps improve the robot's AI model, such as "This answer was helpful" or "Please give me more details."
[0454] Hardware and software used
[0455] Hardware
[0456] Factory robots are equipped with standard input devices (microphones, cameras, etc.) to receive user input, and display devices (displays, speakers, etc.).
[0457] software
[0458] The factory robots are installed with Transformers, a library for natural language processing (NLP). They also use the Sentiment-Analysis model, which is known for its stability and performance, for sentiment analysis. They use the Requests library to communicate with the server.
[0459] Specific examples
[0460] For example, if a factory robot receives a question from a worker saying, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?", it will process it as follows:
[0461] 1. Input Method
[0462] The worker types in a question.
[0463] 2. Analysis method
[0464] The content of the question is analyzed and keywords and intent such as "wooden house," "foundation work," and "having trouble" are extracted.
[0465] 3. Emotion analysis means
[0466] The emotion of "being in trouble" is recognized using an emotion engine.
[0467] 4. Information retrieval and generation methods
[0468] It searches for relevant information based on keywords and sentiment and generates answers such as, "For the foundation construction procedure, please try the following steps."
[0469] 5. Display means
[0470] The generated answer is displayed on the screen and communicated to the worker by voice.
[0471] 6. Feedback Processing Methods
[0472] The worker provides feedback, saying, "This information was useful," and the AI model is improved based on that feedback.
[0473] An example of a prompt sentence is, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?"
[0474] This allows factory robots to respond quickly and appropriately to questions from workers and take workers' feelings into consideration, making on-site work more efficient.
[0475] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0476] Step 1:
[0477] The user inputs specific questions and requests about construction to the factory robot in natural language.
[0478] Input: Questions or requests in natural language (e.g., "I'm having trouble with the foundation work on my wooden house. What should I do specifically?")
[0479] Output: The input text is sent to the system
[0480] Step 2:
[0481] The terminal sends the input text to the server.
[0482] Input: Text entered by the user
[0483] Output: The input text is passed to the server
[0484] Step 3:
[0485] The server processes the received text using analytical means, specifically tokenizing, morphological analysis, and grammatical structure analysis to extract important keywords and intent from the question or request.
[0486] Input: Input text (e.g., "Foundation work for wooden houses")
[0487] Data processing / data operations: tokenization, morphological analysis, grammatical analysis
[0488] Output: Keywords and intent (e.g., "foundation work," "difficult," etc.)
[0489] Step 4:
[0490] Recognize user emotions through emotion analysis. Using an emotion engine, emotions (e.g., "confused") are extracted from the input text and classified into emotion categories.
[0491] Input: Parsed text and extracted keywords
[0492] Data processing / data calculation: emotion extraction and classification
[0493] Output: Recognized emotion (e.g., "confused")
[0494] Step 5:
[0495] The information search and generation means searches for related information from the internal database and external knowledge base based on the extracted keywords and emotions, and generates answers based on the search results.
[0496] Input: Keywords, Recognized Sentiments
[0497] Data processing / data calculation: database search, answer generation
[0498] Output: Generated answer (e.g., "Please proceed as follows for the foundation construction procedure.")
[0499] Step 6:
[0500] The server transmits the generated answer to the terminal, which then presents the answer to the user visually or audibly through a display means.
[0501] Input: Generated Answer
[0502] Output: The answer that is presented to the user
[0503] Step 7:
[0504] The user provides feedback on the displayed answer, including whether it was helpful or if more details are needed.
[0505] Input: User feedback (e.g., satisfied, dissatisfied, further questions, etc.)
[0506] Output: Feedback is sent to the system
[0507] Step 8:
[0508] The device sends feedback to the server, which then receives and analyzes the feedback to help improve the AI model. Specifically, the feedback content is analyzed, evaluation information is extracted, and the feedback is reflected in the AI model for re-learning.
[0509] Input: User feedback
[0510] Data processing / data calculation: Feedback analysis, evaluation information extraction, AI model retraining
[0511] Output: An improved AI model
[0512] As described above, we will explain how the system operates at each processing step and what inputs and outputs there are.
[0513] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0514] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0515] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0516] [Second embodiment]
[0517] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0518] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0519] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0520] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0521] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0522] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0523] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0524] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0525] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0526] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0527] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0528] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0529] MODE FOR CARRYING OUT THE INVENTION
[0530] This section describes an embodiment of the present invention. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system.
[0531] System Configuration
[0532] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[0533] 1. User Input
[0534] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[0535] 2. Sending Input
[0536] The terminal sends questions and requests entered by the user to the server.
[0537] 3. Natural Language Processing (NLP)
[0538] The server analyzes the received user question, specifically:
[0539] Tokenization: Divide the question into words.
[0540] Morphological analysis: Identifying the part of speech of each word.
[0541] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0542] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0543] 4. Information Retrieval and Answer Generation
[0544] The server uses the keywords extracted from the analysis results to search for related information from its internal database and external knowledge bases, and generates appropriate answers based on the information obtained as search results.
[0545] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0546] 5. View Answers
[0547] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0548] 6. Get feedback
[0549] The user provides feedback on the displayed answers, including a rating of whether the answer was helpful or if more detail is needed.
[0550] 7. Processing Feedback
[0551] The device sends user feedback to the server, which analyzes it and uses it to improve the AI model, allowing the system to continuously learn and provide better answers in the future.
[0552] Specific examples
[0553] For example, the user inputs "Please tell me the basics of electrical wiring."
[0554] Terminal: Sends user input to the server.
[0555] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[0556] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0557] Terminal: Displays the generated answer to the user.
[0558] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[0559] Server: Receives feedback and helps improve the AI model.
[0560] In this way, the system of the present invention can solve the problems of a shortage of craftsmen and the transfer of skills by providing information on construction technology quickly and appropriately, and contribute to improving technology throughout the construction industry.
[0561] The processing flow will be explained below.
[0562] Step 1: User enters question
[0563] User: Enters specific questions or requests about construction into the terminal.
[0564] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[0565] Step 2: Sending Input
[0566] Terminal: Sends questions and requests entered by the user to the server.
[0567] Specifically, the user input data is sent to the server as an HTTP request.
[0568] Step 3: Tokenize the Question
[0569] Server: Tokenizes the received question.
[0570] Specifically, the question is divided into words and phrases.
[0571] For example, it is divided into "wooden house," "foundation work," and "procedure."
[0572] Step 4: Morphological analysis
[0573] Server: Performs morphological analysis of the question.
[0574] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[0575] Step 5: Grammatical structure analysis
[0576] Server: Analyzes the grammatical structure of the question.
[0577] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[0578] Step 6: Finding information and generating answers
[0579] Server: Searches for relevant information based on the analysis results and generates answers.
[0580] Specifically, it retrieves relevant data from internal databases and external knowledge bases.
[0581] Based on the search results, answers that are easy for users to understand are generated.
[0582] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0583] Step 7: Submit your response
[0584] Server: Sends the generated answer to the device.
[0585] Step 8: View your answers
[0586] Terminal: Displays the received answer to the user.
[0587] As a specific operation, the answer content is displayed on the screen.
[0588] Step 9: Get feedback
[0589] User: Provide feedback on the displayed answers.
[0590] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[0591] Step 10: Submit your feedback
[0592] Terminal: Sends user feedback to the server.
[0593] Step 11: Processing feedback
[0594] Server: Analyzes the feedback received and helps improve the AI model.
[0595] Specifically, the content of the feedback is analyzed and fed back to the AI model as adaptable data.
[0596] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[0597] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[0598] Example 1
[0599] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0600] In recent years, while the demand for technical information on architecture has increased, the shortage of engineers with specialized knowledge has become a problem. The difficulty of transferring skills and the need for efficient information provision are also increasing. In response to these issues, there is a demand for systems that can provide information on architectural technology quickly and accurately, and that can continuously improve the system's performance based on user feedback.
[0601] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0602] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, a means for transmitting the input questions and requests to the server, an analysis means for extracting important keywords and intent by tokenizing, morphologically analyzing, and grammatically analyzing the received questions, an information search and generation means for searching for related information from an internal database and an external knowledge base based on the extracted keywords and generating an answer, a means for transmitting and displaying the generated answer to the user, and a feedback processing means for accepting and analyzing feedback from users and using it to improve the AI model. This allows users to quickly and accurately obtain information about architecture technology, and enables the system to continuously learn and improve its performance.
[0603] "Input means" refers to a device or interface for accepting specific questions or requests about architecture from users in natural language.
[0604] The "server" is a central information processing device that analyzes information sent by users and generates appropriate responses.
[0605] "Tokenization" is a natural language processing process that divides an input question or request into words.
[0606] "Morphological analysis" is a natural language processing technique for identifying the part of speech of each word.
[0607] "Grammar structure analysis" is a process in natural language processing that analyzes the grammatical structure of a sentence and extracts important keywords and intent.
[0608] The "analysis means" is a device or software that analyzes the received question through tokenization, morphological analysis, and grammatical structure analysis to extract important keywords and intent.
[0609] "Information search and generation means" refers to a device or software that searches for related information from internal databases and external knowledge bases based on extracted keywords and generates answers.
[0610] A "display means" is a device or interface for visually displaying the generated answers to the user.
[0611] "Feedback processing means" refers to a device or software that accepts and analyzes feedback from users and uses it to improve the AI model.
[0612] MODE FOR CARRYING OUT THE INVENTION
[0613] The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. An embodiment of the present invention is described below.
[0614] System Configuration
[0615] This system consists of a device operated by the user (e.g., PC, smartphone, tablet) and a server for processing information sent from these devices.
[0616] User Input
[0617] The user inputs specific questions or requests about construction in natural language using the input means of the terminal. For example, a question such as "Please tell me the procedure for laying the foundations of a wooden house" is input using the keyboard or touch screen of the terminal.
[0618] Sending Input
[0619] The device sends the questions and requests entered by the user to the server. In this process, the device packages the entered text data in JSON format or similar and sends it using an HTTP request to send it to the server over the network.
[0620] Natural Language Processing (NLP)
[0621] The server analyzes the received user question. Specifically, it performs the following analyses: tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent, such as "wooden house," "foundation work," and "procedure."
[0622] Information retrieval and answer generation
[0623] The server uses keywords extracted from the analysis results to search for related information from its internal database and external knowledge base. Based on the information obtained as a search result, it generates an appropriate answer. For example, it generates an answer such as, "The steps for laying the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[0624] Show Answers
[0625] The server sends the generated answer to the terminal, which then visually displays the received answer to the user, allowing the user to check the answer on the terminal screen.
[0626] Get feedback
[0627] The user provides feedback on the displayed answer, for example, with options such as "helpful" or "more information," and the user enters a prompt such as:
[0628] "It was helpful"
[0629] "I want more details."
[0630] Processing Feedback
[0631] The device sends user feedback to a server, which analyzes it and uses it to improve the AI model. The feedback data can be saved and used as training data for the machine learning model to improve the accuracy of answers in future searches.
[0632] Specific examples
[0633] For example, if a user types "What are the basics of electrical wiring?", the following steps are taken:
[0634] Terminal: The user types in a question and sends it to the server.
[0635] Server: Receives the question, performs tokenization, morphological analysis, and grammatical structure analysis, and extracts important keywords (e.g., "electrical wiring" and "basic knowledge").
[0636] Server: Searches internal databases and external knowledge bases to generate appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0637] Terminal: Receives the generated answers and displays them on the screen.
[0638] User: Review the answer and provide feedback such as "this was helpful" or "I'd like more information."
[0639] Device: Sends feedback to the server.
[0640] Server: Analyzes feedback and improves the AI model.
[0641] By implementing the system in this way, information on building technology can be provided quickly and appropriately, and user feedback can be utilized to continuously improve the system's performance.
[0642] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0643] Step 1: User Input
[0644] The user inputs specific questions or requests about construction into a terminal (PC, smartphone, tablet). A question such as "Please tell me the steps for laying the foundations for a wooden house" is input using a keyboard or touch screen. The input text data is received by the input means of the terminal. (Input) The user's question text. (Output) The input question as text data.
[0645] Step 2: Sending Input
[0646] The device sends questions and requests entered by the user to the server. Specifically, the entered text data is packaged in JSON format or similar over the network and sent to the server using an HTTP request. (Input) Question as text data. (Output) HTTP request sent to the server.
[0647] Step 3: Natural Language Processing (NLP)
[0648] The server analyzes the received user question and performs the following specific processing:
[0649] Tokenization: The server divides the question into words. For example, the sentence "Please tell me the procedure for laying the foundation for a wooden house" is divided into "wooden," "house," "foundation," "construction," "procedure," "tell me," and "please." (Input) The question as text data. (Output) A tokenized word list.
[0650] Morphological analysis: The server identifies the part of speech of each word. For example, "wooden structure (noun)", "house (noun)", "foundation (noun)", "construction (noun)", "procedure (noun)", "teaching (verb)". (Input) A tokenized word list. (Output) A word list with parts of speech assigned.
[0651] Grammatical structure analysis: The server analyzes the grammatical structure of the sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords. (Input) A list of words with parts of speech assigned. (Output) A list of extracted keywords.
[0652] Step 4: Information retrieval and answer generation
[0653] Based on the analysis results, the server uses the extracted keywords to search for relevant information from its internal database and external knowledge bases. The specific processes are as follows:
[0654] Internal database search: The server searches the internal database (technical manuals and industry standard information) using keywords such as "wooden house," "foundation work," and "procedure." (Input) Extracted keyword list. (Output) Information obtained from the internal database.
[0655] External knowledge base search: The server searches an external knowledge base (such as public documents on the Internet). (Input) Extracted keyword list. (Output) Information obtained from the external knowledge base.
[0656] Answer generation: Generate an answer in a format appropriate for the user from the search results. For example, generate an answer in the format "The steps for laying the foundation for a wooden house are as follows: 1. Level the site 2. Pour concrete for the foundation 3. Install rebar 4. Install formwork 5. Re-pour concrete." (Input) Information from the search results. (Output) Generated answer text.
[0657] Step 5: View your answers
[0658] The server sends the generated answer to the device. The device visually displays the received answer to the user. Specifically, the server packages the answer in JSON format and sends it back to the device as an HTTP response. The device displays the answer on the screen so that the user can confirm it. (Input) Generated answer. (Output) Answer displayed on the device.
[0659] Step 6: Getting feedback
[0660] The user provides feedback on the displayed answer. For example, options such as "Helpful" or "Need more details" are displayed and the user selects one. The device receives this feedback and sends it to the server. (Input) User feedback. (Output) Feedback sent to the server.
[0661] Step 7: Processing feedback
[0662] The server analyzes the received feedback and uses it to improve the AI model. Specifically, it stores the feedback data and uses it as training data for the machine learning model. This data can be used to improve the accuracy of the AI model. (Input) User feedback. (Output) Improved AI model, improving the accuracy of answers from next time onwards.
[0663] (Application example 1)
[0664] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0665] When field workers perform tasks that require specialized knowledge, they need to be provided with information quickly and accurately. However, in the past, they often relied on specialized books and manuals, which lacked immediacy and reduced work efficiency. Furthermore, there was a lack of mechanisms for utilizing feedback to continuously improve the system. This has led to concerns that this increases the burden on workers and reduces the accuracy and efficiency of work, especially in sites with a wide range of complex tasks, such as factories.
[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0667] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intent, an information search and generation means for searching for related information based on the extracted keywords and generating answers, a display means for displaying the generated answers to the user, a feedback processing means for receiving feedback from the user and using it to improve the AI model, and a smart gadget equipped with a voice input device and a display device as a terminal used by the user, wherein the input means accepts voice input and the display means has the function of displaying answers on the display of the smart gadget. This allows on-site workers to input questions hands-free and receive answers quickly, improving work efficiency and enabling continuous improvement of the system.
[0668] 1. "Specific questions or requests regarding construction"
[0669] "Specific questions or requests regarding construction" refers to specific information or instructions that a user requests regarding the design, construction, maintenance, etc. of a building.
[0670] 2. "Input methods that accept natural language"
[0671] "Input means that accepts natural language" refers to devices or software that accept questions or requests entered by the user using everyday language or technical terms.
[0672] 3. “Analysis means”
[0673] "Analysis means" refers to the technology and functions used to process input data and extract important keywords and the intent of the question.
[0674] 4. "Extract important keywords and intent"
[0675] "Extracting key keywords and intent" means identifying specific words and phrases, as well as the purpose and meaning behind them, from a user's question or request.
[0676] 5. "Information retrieval and generation methods"
[0677] "Information search and generation means" refers to the technology and functions that search for related information from databases and external information sources based on extracted keywords and generate appropriate answers.
[0678] 6. "A means of displaying the generated answer to the user"
[0679] "Display means for displaying the generated answer to the user" refers to technology or devices for visually presenting the answer generated by the system on a display of a computer, mobile device, etc.
[0680] 7. "Feedback Processing Means"
[0681] "Feedback processing means" refers to the technology and functions that receive evaluations and opinions on responses from users, analyze them, and use them to improve the system.
[0682] 8. "Smart Gadgets"
[0683] "Smart gadgets" are portable devices with internet connectivity and various functions, such as smart glasses and smartphones.
[0684] 9. "Voice input device"
[0685] A "voice input device" is a microphone and associated software that allows a user to speak questions or commands.
[0686] 10. “Display device”
[0687] "Display device" means a display or screen for visually displaying information or data.
[0688] An embodiment of the present invention will be described. The present invention is a system that enables factory workers to use smart gadgets to ask questions about building and equipment maintenance in real time and quickly receive appropriate answers.
[0689] System Configuration
[0690] The system consists of the following main components:
[0691] Smart gadgets (e.g., smart glasses)
[0692] server
[0693] Database
[0694] display device
[0695] 1. User Input
[0696] The user wears a smart gadget and inputs questions or requests in natural language using voice, such as "Please tell me the basic procedures for equipment inspection."
[0697] 2. Sending Input
[0698] A voice input device in the smart gadget converts the user's voice into text data (using voice recognition software) and sends the text data to a server.
[0699] 3. Natural Language Processing (NLP)
[0700] The server analyzes the received user question using the following techniques:
[0701] Tokenization: Divide the question into words.
[0702] Morphological analysis: Identifying the part of speech of each word.
[0703] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0704] For example, "equipment," "inspection," and "basic procedures" are extracted as main keywords.
[0705] 4. Information Retrieval and Answer Generation
[0706] The server uses keywords extracted from the analysis results to search for relevant information from its internal database and external knowledge bases, and generates appropriate answers using a generative AI model based on the information obtained as search results.
[0707] Example: "The basic steps for equipment inspection are as follows: 1. Visual inspection 2. Operation check 3. Connection check 4. Lubrication check"
[0708] 5. View Answers
[0709] The server transmits the generated answer to the smart gadget, which visually displays the answer on its display device.
[0710] 6. Get feedback
[0711] Users provide feedback on the displayed answers, including whether they found the answer helpful or if they need more detail.
[0712] 7. Processing Feedback
[0713] Through the feedback collection function of the smart gadget, user feedback is sent to the server, which analyzes this feedback and helps improve the generative AI model.
[0714] Usage example
[0715] For example, the user inputs, "Please tell me the basic procedures for equipment inspection."
[0716] Smart gadget: Converts user's voice input into text and sends it to the server.
[0717] Server: Performs natural language processing and analyzes the intent of the question (e.g., "equipment," "inspection," "basic procedures").
[0718] Server: Searches internal databases and external knowledge bases and generates appropriate answers using generative AI models (e.g., "The basic steps for equipment inspection are as follows...").
[0719] Smart Gadget: Displays the generated answers to the user.
[0720] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[0721] Server: Receives feedback and helps improve the generative AI model.
[0722] In this way, the system of the present invention allows field workers to input questions hands-free and receive quick answers, thereby improving work efficiency and enabling continuous improvement of the system.
[0723] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0724] Step 1:
[0725] The user uses a smart gadget to provide voice input. For voice input, the user speaks, "Please tell me the basic procedures for equipment inspection." The smart gadget converts this voice into text data. The input is voice data, and the output is text data. Specifically, the gadget converts the voice data into text using voice recognition software (e.g., Google Speech-to-Text API).
[0726] Step 2:
[0727] The terminal transmits text data to the server. This is done using a network. In this case, the input is text data from the user, and the output is data sent to the server. In concrete terms, the smart gadget transmits text data to the server using a wireless network.
[0728] Step 3:
[0729] The server applies natural language processing (NLP) to the received text data. This processing includes tokenization, morphological analysis, and grammatical structure analysis. The input is the text data, and the output is the key keywords and the intent of the question. Specifically, the server uses the spaCy library to analyze the text data and extract key keywords.
[0730] Step 4:
[0731] The server searches for related information from the internal database and external knowledge base based on the analysis results. The input is the extracted keywords, and the output is related information as a search result. Specifically, it searches the internal database using SQL queries and retrieves information from the external knowledge base using APIs.
[0732] Step 5:
[0733] The server uses a generative AI model to generate an appropriate answer. The input is relevant information, and the output is a response to the user. Specifically, the server uses Hugging Face's transformers library to input relevant information into the generative AI model as a prompt, and generates an answer. Take the prompt "Please tell me the basic procedures for equipment inspection" as an example.
[0734] Step 6:
[0735] The server sends the generated answer to the smart gadget. The input is the answer, and the output is the display data for the smart gadget. Specifically, the server sends text data to the smart gadget via the network.
[0736] Step 7:
[0737] The smart gadget displays the received answer on a display device. The input is the text data of the answer, and the output is visual information displayed on the display. Specifically, the smart gadget renders the received text data on the display.
[0738] Step 8:
[0739] The user provides feedback on the displayed answer by speaking a rating such as "helpful" or "I'd like more information." The input is voice data, and the output is text feedback. Specifically, as with the initial voice input, speech recognition software is used to convert the speech into text.
[0740] Step 9:
[0741] The terminal transmits the feedback text data to the server. The input is the feedback text data from the user, and the output is the data to be transmitted to the server. In concrete terms, the smart gadget transmits the feedback data to the server using a wireless network.
[0742] Step 10:
[0743] The server analyzes the feedback and uses it to improve the AI model. The input is the text feedback, and the output is an improved AI model. Specifically, the server uses the feedback data to train and update the AI model.
[0744] This allows field workers to input questions hands-free and receive quick answers, improving work efficiency and enabling continuous improvement of the system.
[0745] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0746] MODE FOR CARRYING OUT THE INVENTION
[0747] An embodiment of the present invention will be described. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. In addition, it combines this with an emotion engine that recognizes the user's emotions.
[0748] System Configuration
[0749] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[0750] 1. User Input
[0751] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[0752] 2. Sending Input
[0753] The terminal sends questions and requests entered by the user to the server.
[0754] 3. Natural Language Processing (NLP)
[0755] The server analyzes the received user question, specifically:
[0756] Tokenization: Divide the question into words.
[0757] Morphological analysis: Identifying the part of speech of each word.
[0758] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[0759] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0760] 4. Emotion analysis
[0761] The server uses an emotion engine to recognize the user's emotion in the question. The following process is performed:
[0762] Emotion recognition: Extracting user emotions (e.g., confusion, excitement, anger, etc.) from input text.
[0763] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[0764] For example, if a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[0765] 5. Information Retrieval and Answer Generation
[0766] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[0767] Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[0768] Generate answers that are interesting and considerate of the user's emotions.
[0769] For example, an answer might be generated along the lines of, "Please understand that foundation work is difficult, and refer to the following steps."
[0770] 6. View Answers
[0771] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0772] 7. Get feedback
[0773] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[0774] 8. Submitting Feedback
[0775] The device sends the feedback from the user to the server.
[0776] 9. Processing Feedback
[0777] The server analyzes this feedback and uses it to improve the AI model. Specifically, it analyzes the feedback, extracts evaluation information, and reflects it in the AI model for re-learning. This allows the system to continuously improve.
[0778] Specific examples
[0779] For example, if a user types "What are the basics of electrical wiring?":
[0780] Terminal: Sends user input to the server.
[0781] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[0782] Server: Uses an emotion engine to recognize the user's emotions (e.g., confusion, interest).
[0783] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[0784] Terminal: Displays the generated answer to the user.
[0785] User: Provide feedback on the answer.
[0786] Server: Receives feedback and helps improve the AI model.
[0787] In this way, the system of the present invention not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take into consideration the user's feelings. This effectively solves the problems of passing on construction techniques and the shortage of craftsmen, and contributes to the improvement of technology in the construction industry as a whole.
[0788] The processing flow will be explained below.
[0789] Step 1: User enters question
[0790] User: Enters specific questions or requests about construction into the terminal.
[0791] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[0792] Step 2: Sending Input
[0793] Terminal: Sends questions and requests entered by the user to the server.
[0794] Specifically, the user input data is sent to the server as an HTTP request.
[0795] Step 3: Tokenize the Question
[0796] Server: Tokenizes the received question.
[0797] Specifically, the question is divided into words and phrases.
[0798] For example, it is divided into "wooden house," "foundation work," and "procedure."
[0799] Step 4: Morphological analysis
[0800] Server: Performs morphological analysis of the question.
[0801] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[0802] Step 5: Grammatical structure analysis
[0803] Server: Analyzes the grammatical structure of the question.
[0804] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[0805] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[0806] Step 6: Sentiment Analysis
[0807] Server: Recognizes the user's emotions contained in the question using an emotion engine.
[0808] Specifically, it extracts the user's emotions (e.g., confusion, excitement, anger, etc.) from the input text and classifies them into predefined emotion categories.
[0809] Example: If a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[0810] Step 7: Information retrieval and answer generation
[0811] Server: Searches for relevant information and generates answers based on the analysis results and emotion recognition results.
[0812] Specifically, based on the extracted keywords and recognized emotions, the system retrieves relevant data from an internal database and an external knowledge base.
[0813] Generate answers that are interesting and considerate of the user's emotions.
[0814] For example, an answer like "Please understand that foundation work is difficult, and refer to the following steps." is generated.
[0815] Step 8: Submit your response
[0816] Server: Sends the generated answer to the device.
[0817] Step 9: View your answers
[0818] Terminal: Displays the received answer to the user.
[0819] As a specific operation, the answer content is displayed on the screen.
[0820] Step 10: Get feedback
[0821] User: Provide feedback on the displayed answers.
[0822] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[0823] Step 11: Submit your feedback
[0824] Terminal: Sends user feedback to the server.
[0825] Step 12: Processing feedback
[0826] Server: Analyzes the feedback received and helps improve the AI model.
[0827] Specifically, the feedback content is analyzed and fed back to the AI model as adaptable data.
[0828] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[0829] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[0830] Example 2
[0831] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0832] In the modern construction industry, there are limited means of instantly obtaining technical information and procedures related to construction, creating a demand for fast, accurate answers to technical questions and problems. It is also important to increase user satisfaction by providing answers that take the user's emotions into consideration. However, conventional systems struggle to respond in a way that takes user emotions into account, and they lack mechanisms for appropriately utilizing user feedback to improve the system. Therefore, there is a need for a system that can analyze user questions and requests, recognize emotions, generate and display appropriate answers, and further utilize feedback.
[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0834] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an information search and generation means for searching for related information based on the extracted keywords and the user's emotions and generating answers, a display means for displaying the generated answers to the user, and a feedback processing means for receiving feedback from the user and using it to improve the AI model. This makes it possible to quickly and appropriately answer user questions, provide answers that take user emotions into consideration, and utilize feedback to continuously improve the system.
[0835] "Input means" refers to devices or functions that accept specific questions or requests about architecture from users in natural language.
[0836] "Analysis means" refers to a device or function that analyzes questions or requests entered by users and extracts important keywords and intentions.
[0837] "Information search and generation means" refers to devices or functions that search for related information and generate answers based on the keywords extracted by the analysis means and the user's emotions.
[0838] "Display means" refers to a device or function that visually displays the generated answer to the user.
[0839] A "feedback processing means" is a device or function that accepts feedback provided by users and uses that feedback to help improve the AI model.
[0840] "Tokenization" is the process of dividing an input natural language sentence into words.
[0841] "Morphological analysis" is the process of identifying the part of speech of each tokenized word.
[0842] "Grammar structure analysis" is the process of analyzing the grammatical structure of an input sentence and extracting important keywords and intent.
[0843] "Emotion recognition" is the process of extracting a user's emotions from input text.
[0844] "Emotion classification" refers to the process of classifying extracted emotions into predefined emotion categories.
[0845] An "internal database" is a database that manages information and data stored within the system.
[0846] An "external knowledge base" is an information source or database that is externally accessible, such as the Internet.
[0847] An "AI model" is a model that uses artificial intelligence to learn and perform processes such as answering questions and recognizing emotions.
[0848] "Feedback" refers to the ratings and opinions that users provide in response to the answers displayed.
[0849] The system of the present invention allows users to input specific questions or requests about architecture in natural language, analyzes them, and generates and displays answers. Furthermore, it can collect feedback from users and use it to improve the system. The system consists of the following main components:
[0850] System Configuration
[0851] 1. Hardware configuration:
[0852] Terminal: A device such as a PC, smartphone, or tablet that allows users to input questions or requests and display answers.
[0853] Server: A high-performance cloud server or on-premise server with high data processing capacity and fast response is recommended.
[0854] 2. Software configuration:
[0855] Natural language processing engine: An engine that tokenizes text, analyzes morphology, and analyzes grammatical structures (e.g., Google NLP).
[0856] Emotion engine: An engine for recognizing and classifying user emotions (e.g., IBM Watson Tone Analyzer).
[0857] Database management system: Databases such as MySQL and MongoDB are used as internal databases.
[0858] AI model: An artificial intelligence model that analyzes a user's question or request and generates an answer. Use a generative AI model.
[0859] Operation procedure and processing contents
[0860] 1. User input:
[0861] Users operate their own terminals to input specific questions or requests about construction in natural language, such as "Please tell me the steps for laying the foundations for a wooden house."
[0862] 2. Sending input:
[0863] The device sends the user-entered questions or requests to a server over the Internet. The data is sent as an HTTP request.
[0864] 3. Natural Language Processing (NLP):
[0865] The server analyzes the received user question, specifically by performing the following steps:
[0866] Tokenization: Divide the input question into words.
[0867] Morphological analysis: Identifying the parts of speech of segmented words.
[0868] Grammatical structure analysis: Analyze grammatical structures and extract important keywords and intent.
[0869] 4. Emotion analysis:
[0870] The server uses an emotion engine to recognize the user's emotion in the question, which includes the following steps:
[0871] Emotion Recognition: Extracting user emotions from input text.
[0872] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[0873] 5. Information retrieval and answer generation:
[0874] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[0875] Information retrieval: Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[0876] Answer generation: Based on the searched information, an answer is generated that takes into consideration the user's feelings.
[0877] 6. Show Answer:
[0878] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[0879] 7. Getting feedback:
[0880] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[0881] 8. Submitting Feedback:
[0882] The device sends the user's feedback to the server, again as an HTTP request.
[0883] 9. Handling feedback (retraining the AI model):
[0884] The server analyzes the received feedback and extracts evaluation information, which is then used to retrain the AI model and improve response accuracy from the next time onwards.
[0885] Specific examples
[0886] For example, if a user types "Please tell me the basics of electrical wiring," the following will be processed:
[0887] User: Enters a question into the terminal and sends it.
[0888] Server: Receives input and performs natural language processing and sentiment analysis.
[0889] Server: Searches for relevant information from internal databases and external knowledge bases and generates appropriate answers.
[0890] Terminal: Display the generated answer.
[0891] User: Provide feedback on the answer.
[0892] Server: Receives feedback and helps retrain the AI model.
[0893] This system not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take the user's emotions into consideration. This will effectively solve the problems of passing on construction techniques and the shortage of craftsmen, and contribute to improving the technology of the entire construction industry.
[0894] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0895] Step 1:
[0896] The user uses the input means to input specific questions or requests about construction in natural language. For example, they may input a question such as, "Please tell me the steps for laying the foundations for a wooden house." This input is saved as text data on the terminal.
[0897] Step 2:
[0898] The terminal sends the text data entered by the user to the server. Specifically, the data is sent to the server as an HTTP request. At this time, the input data is temporarily stored in the terminal's buffer memory. The input is the user's question text, and the output is the sent HTTP request.
[0899] Step 3:
[0900] The server analyzes the received text data. First, it uses a natural language processing engine to tokenize the question. As a result of tokenization, the input sentence can be divided into words, and the output is a word list. For example, it may be divided into words such as "wooden house," "foundation work," "procedure," "tell me," and "please."
[0901] Step 4:
[0902] The server performs morphological analysis. As a result of the morphological analysis, it is possible to identify the part of speech for each word. For example, "wooden house" is recognized as a noun and "teach me" as a verb. The input is a list of tokenized words, and the output is a list of words with identified parts of speech.
[0903] Step 5:
[0904] The server performs grammatical analysis. This analyzes the grammatical structure of the entire sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as key keywords. The results of the grammatical analysis are the extracted keywords and a structural analysis tree.
[0905] Step 6:
[0906] The server uses an emotion recognition engine to recognize the user's emotions. During the emotion recognition process, the server extracts the user's emotional characteristics (e.g., confusion, excitement, anger) from the text data and classifies them into categories. For example, the emotion "confusion" is extracted from the sentence "I'm having trouble because the foundation work on a wooden house is difficult." The input is the analyzed text data, and the output is the recognized emotion and its category.
[0907] Step 7:
[0908] The server performs an information search based on the analysis results and emotion recognition results. It searches its internal database and external knowledge base to obtain relevant information. For example, it searches the internal database for technical information using keywords such as "wooden house," "foundation work," and "procedure." The information obtained at this stage is output in the form of text or images.
[0909] Step 8:
[0910] The server generates an answer based on the search results. Using a generative AI model, it creates an answer that takes into account the acquired information and the user's emotions. For example, an answer such as "Please understand that foundation construction is difficult, and refer to the following steps" may be generated. The input is the search results and emotion data, and the output is the answer text for the user.
[0911] Step 9:
[0912] The server sends the generated answer to the terminal. The terminal displays the received answer to the user. Specifically, the answer is displayed in text format on the terminal screen. The input is the answer data from the server, and the output is the display for the user.
[0913] Step 10:
[0914] The user provides feedback on the displayed answer, including an evaluation of whether the answer was helpful or whether further explanation is needed. The feedback is again input into the terminal as text data.
[0915] Step 11:
[0916] The device sends the user's feedback to the server, again as an HTTP request. The input is the user's feedback text, and the output is the sent HTTP request.
[0917] Step 12:
[0918] The server receives and analyzes the feedback data. As a result of the analysis, specific evaluation information is extracted and used to retrain the AI model. This improves response accuracy from the next time onwards. The input is the feedback data, and the output is an updated AI model.
[0919] (Application example 2)
[0920] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0921] Lack of technical knowledge and information on construction sites, especially complex work procedures and troubleshooting, makes it difficult to quickly implement them. Furthermore, the inability to provide appropriate advice that takes into account the feelings of workers makes it difficult to work efficiently. Furthermore, there is a lack of means to continuously improve the system based on user feedback.
[0922] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0923] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an emotion analysis means for analyzing emotions using an emotion engine that recognizes the user's emotions, an information search and generation means for searching for related information and generating answers based on the extracted keywords and recognized emotions, a display means for displaying the generated answers to the user, and a feedback processing means for accepting feedback from users and using it to improve the AI model. This enables workers at construction sites to respond quickly to questions and requests, provides appropriate advice that takes workers' emotions into consideration, and continuously improves the system, enabling efficient and effective work.
[0924] "Input means" refers to a device or system that accepts specific questions or requests about architecture from users in natural language.
[0925] The "analysis means" is a device or program that analyzes the input question or request and extracts important keywords and intentions.
[0926] "Emotion analysis means" refers to a device or system for analyzing emotions from input text using an emotion engine that recognizes the emotions of a user.
[0927] "Information search and generation means" refers to a device or program that searches for related information based on the extracted keywords and recognized emotions and generates answers.
[0928] A "display means" is a device or system that visually displays the generated answers to the user.
[0929] A "feedback processing means" is a device or program that accepts feedback from users and uses it to improve the AI model.
[0930] The present invention is a system in which a factory robot responds to questions and requests from workers at a construction site and provides advice that takes into consideration their emotions. An embodiment of the system is described below.
[0931] System Configuration
[0932] The system includes the following components:
[0933] 1. Input Method
[0934] The factory robot can receive specific questions and requests about construction from workers in natural language. For example, the factory robot can receive a question from a worker such as, "Please tell me the key points of foundation work."
[0935] 2. Analysis method
[0936] The factory robot analyzes the questions and requests it receives. This analysis involves tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent. For example, keywords such as "foundation work" and "key points" are extracted.
[0937] 3. Emotion analysis means
[0938] The factory robot uses an emotion engine to analyze the emotions of workers from input text. Specifically, the input text is input into the emotion engine, and emotions such as "confusion" or "interest" are recognized.
[0939] 4. Information retrieval and generation methods
[0940] Based on the extracted keywords and the recognized emotions, the factory robot searches for relevant information from its internal database and external knowledge base, and generates an appropriate answer taking into account the search results and the emotions of the worker. For example, it generates an answer in the form of "The main points of foundation work are as follows..."
[0941] 5. Display means
[0942] The generated answers are displayed and presented visually or audibly by the factory robot to the worker, allowing the worker to obtain the appropriate information.
[0943] 6. Feedback Processing Methods
[0944] Workers can provide feedback on the displayed answers, which helps improve the robot's AI model, such as "This answer was helpful" or "Please give me more details."
[0945] Hardware and software used
[0946] Hardware
[0947] Factory robots are equipped with standard input devices (microphones, cameras, etc.) to receive user input, and display devices (displays, speakers, etc.).
[0948] software
[0949] The factory robots are installed with Transformers, a library for natural language processing (NLP). They also use the Sentiment-Analysis model, which is known for its stability and performance, for sentiment analysis. They use the Requests library to communicate with the server.
[0950] Specific examples
[0951] For example, if a factory robot receives a question from a worker saying, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?", it will process it as follows:
[0952] 1. Input Method
[0953] The worker types in a question.
[0954] 2. Analysis method
[0955] The content of the question is analyzed and keywords and intent such as "wooden house," "foundation work," and "having trouble" are extracted.
[0956] 3. Emotion analysis means
[0957] The emotion of "being in trouble" is recognized using an emotion engine.
[0958] 4. Information retrieval and generation methods
[0959] It searches for relevant information based on keywords and sentiment and generates answers such as, "For the foundation construction procedure, please try the following steps."
[0960] 5. Display means
[0961] The generated answer is displayed on the screen and communicated to the worker by voice.
[0962] 6. Feedback Processing Methods
[0963] The worker provides feedback, saying, "This information was useful," and the AI model is improved based on that feedback.
[0964] An example of a prompt sentence is, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?"
[0965] This allows factory robots to respond quickly and appropriately to questions from workers and take workers' feelings into consideration, making on-site work more efficient.
[0966] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0967] Step 1:
[0968] The user inputs specific questions and requests about construction to the factory robot in natural language.
[0969] Input: Questions or requests in natural language (e.g., "I'm having trouble with the foundation work on my wooden house. What should I do specifically?")
[0970] Output: The input text is sent to the system
[0971] Step 2:
[0972] The terminal sends the input text to the server.
[0973] Input: Text entered by the user
[0974] Output: The input text is passed to the server
[0975] Step 3:
[0976] The server processes the received text using analytical means, specifically tokenizing, morphological analysis, and grammatical structure analysis to extract important keywords and intent from the question or request.
[0977] Input: Input text (e.g., "Foundation work for wooden houses")
[0978] Data processing / data operations: tokenization, morphological analysis, grammatical analysis
[0979] Output: Keywords and intent (e.g., "foundation work," "difficult," etc.)
[0980] Step 4:
[0981] Recognize user emotions through emotion analysis. Using an emotion engine, emotions (e.g., "confused") are extracted from the input text and classified into emotion categories.
[0982] Input: Parsed text and extracted keywords
[0983] Data processing / data calculation: emotion extraction and classification
[0984] Output: Recognized emotion (e.g., "confused")
[0985] Step 5:
[0986] The information search and generation means searches for related information from the internal database and external knowledge base based on the extracted keywords and emotions, and generates answers based on the search results.
[0987] Input: Keywords, Recognized Sentiments
[0988] Data processing / data calculation: database search, answer generation
[0989] Output: Generated answer (e.g., "Please proceed as follows for the foundation construction procedure.")
[0990] Step 6:
[0991] The server transmits the generated answer to the terminal, which then presents the answer to the user visually or audibly through a display means.
[0992] Input: Generated Answer
[0993] Output: The answer that is presented to the user
[0994] Step 7:
[0995] The user provides feedback on the displayed answer, including whether it was helpful or if more details are needed.
[0996] Input: User feedback (e.g., satisfied, dissatisfied, further questions, etc.)
[0997] Output: Feedback is sent to the system
[0998] Step 8:
[0999] The device sends feedback to the server, which then receives and analyzes the feedback to help improve the AI model. Specifically, the feedback content is analyzed, evaluation information is extracted, and the feedback is reflected in the AI model for re-learning.
[1000] Input: User feedback
[1001] Data processing / data calculation: Feedback analysis, evaluation information extraction, AI model retraining
[1002] Output: An improved AI model
[1003] As described above, we will explain how the system operates at each processing step and what inputs and outputs there are.
[1004] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1005] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1006] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1007] [Third embodiment]
[1008] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1009] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1010] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1011] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1012] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1013] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1014] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1015] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1016] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1017] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1018] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1019] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1020] MODE FOR CARRYING OUT THE INVENTION
[1021] This section describes an embodiment of the present invention. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system.
[1022] System Configuration
[1023] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[1024] 1. User Input
[1025] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[1026] 2. Sending Input
[1027] The terminal sends questions and requests entered by the user to the server.
[1028] 3. Natural Language Processing (NLP)
[1029] The server analyzes the received user question, specifically:
[1030] Tokenization: Divide the question into words.
[1031] Morphological analysis: Identifying the part of speech of each word.
[1032] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1033] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1034] 4. Information Retrieval and Answer Generation
[1035] The server uses the keywords extracted from the analysis results to search for related information from its internal database and external knowledge bases, and generates appropriate answers based on the information obtained as search results.
[1036] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1037] 5. View Answers
[1038] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1039] 6. Get feedback
[1040] The user provides feedback on the displayed answers, including a rating of whether the answer was helpful or if more detail is needed.
[1041] 7. Processing Feedback
[1042] The device sends user feedback to the server, which analyzes it and uses it to improve the AI model, allowing the system to continuously learn and provide better answers in the future.
[1043] Specific examples
[1044] For example, the user inputs "Please tell me the basics of electrical wiring."
[1045] Terminal: Sends user input to the server.
[1046] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[1047] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1048] Terminal: Displays the generated answer to the user.
[1049] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[1050] Server: Receives feedback and helps improve the AI model.
[1051] In this way, the system of the present invention can solve the problems of a shortage of craftsmen and the transfer of skills by providing information on construction technology quickly and appropriately, and contribute to improving technology throughout the construction industry.
[1052] The processing flow will be explained below.
[1053] Step 1: User enters question
[1054] User: Enters specific questions or requests about construction into the terminal.
[1055] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[1056] Step 2: Sending Input
[1057] Terminal: Sends questions and requests entered by the user to the server.
[1058] Specifically, the user input data is sent to the server as an HTTP request.
[1059] Step 3: Tokenize the Question
[1060] Server: Tokenizes the received question.
[1061] Specifically, the question is divided into words and phrases.
[1062] For example, it is divided into "wooden house," "foundation work," and "procedure."
[1063] Step 4: Morphological analysis
[1064] Server: Performs morphological analysis of the question.
[1065] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[1066] Step 5: Grammatical structure analysis
[1067] Server: Analyzes the grammatical structure of the question.
[1068] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[1069] Step 6: Finding information and generating answers
[1070] Server: Searches for relevant information based on the analysis results and generates answers.
[1071] Specifically, it retrieves relevant data from internal databases and external knowledge bases.
[1072] Based on the search results, answers that are easy for users to understand are generated.
[1073] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1074] Step 7: Submit your response
[1075] Server: Sends the generated answer to the device.
[1076] Step 8: View your answers
[1077] Terminal: Displays the received answer to the user.
[1078] As a specific operation, the answer content is displayed on the screen.
[1079] Step 9: Get feedback
[1080] User: Provide feedback on the displayed answers.
[1081] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[1082] Step 10: Submit your feedback
[1083] Terminal: Sends user feedback to the server.
[1084] Step 11: Processing feedback
[1085] Server: Analyzes the feedback received and helps improve the AI model.
[1086] Specifically, the content of the feedback is analyzed and fed back to the AI model as adaptable data.
[1087] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[1088] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[1089] Example 1
[1090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1091] In recent years, while the demand for technical information on architecture has increased, the shortage of engineers with specialized knowledge has become a problem. The difficulty of transferring skills and the need for efficient information provision are also increasing. In response to these issues, there is a demand for systems that can provide information on architectural technology quickly and accurately, and that can continuously improve the system's performance based on user feedback.
[1092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1093] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, a means for transmitting the input questions and requests to the server, an analysis means for extracting important keywords and intent by tokenizing, morphologically analyzing, and grammatically analyzing the received questions, an information search and generation means for searching for related information from an internal database and an external knowledge base based on the extracted keywords and generating an answer, a means for transmitting and displaying the generated answer to the user, and a feedback processing means for accepting and analyzing feedback from users and using it to improve the AI model. This allows users to quickly and accurately obtain information about architecture technology, and enables the system to continuously learn and improve its performance.
[1094] "Input means" refers to a device or interface for accepting specific questions or requests about architecture from users in natural language.
[1095] The "server" is a central information processing device that analyzes information sent by users and generates appropriate responses.
[1096] "Tokenization" is a natural language processing process that divides an input question or request into words.
[1097] "Morphological analysis" is a natural language processing technique for identifying the part of speech of each word.
[1098] "Grammar structure analysis" is a process in natural language processing that analyzes the grammatical structure of a sentence and extracts important keywords and intent.
[1099] The "analysis means" is a device or software that analyzes the received question through tokenization, morphological analysis, and grammatical structure analysis to extract important keywords and intent.
[1100] "Information search and generation means" refers to a device or software that searches for related information from internal databases and external knowledge bases based on extracted keywords and generates answers.
[1101] A "display means" is a device or interface for visually displaying the generated answers to the user.
[1102] "Feedback processing means" refers to a device or software that accepts and analyzes feedback from users and uses it to improve the AI model.
[1103] MODE FOR CARRYING OUT THE INVENTION
[1104] The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. An embodiment of the present invention is described below.
[1105] System Configuration
[1106] This system consists of a device operated by the user (e.g., PC, smartphone, tablet) and a server for processing information sent from these devices.
[1107] User Input
[1108] The user inputs specific questions or requests about construction in natural language using the input means of the terminal. For example, a question such as "Please tell me the procedure for laying the foundations of a wooden house" is input using the keyboard or touch screen of the terminal.
[1109] Sending Input
[1110] The device sends the questions and requests entered by the user to the server. In this process, the device packages the entered text data in JSON format or similar and sends it using an HTTP request to send it to the server over the network.
[1111] Natural Language Processing (NLP)
[1112] The server analyzes the received user question. Specifically, it performs the following analyses: tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent, such as "wooden house," "foundation work," and "procedure."
[1113] Information retrieval and answer generation
[1114] The server uses keywords extracted from the analysis results to search for related information from its internal database and external knowledge base. Based on the information obtained as a search result, it generates an appropriate answer. For example, it generates an answer such as, "The steps for laying the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1115] Show Answers
[1116] The server sends the generated answer to the terminal, which then visually displays the received answer to the user, allowing the user to check the answer on the terminal screen.
[1117] Get feedback
[1118] The user provides feedback on the displayed answer, for example, with options such as "helpful" or "more information," and the user enters a prompt such as:
[1119] "It was helpful"
[1120] "I want more details."
[1121] Processing Feedback
[1122] The device sends user feedback to a server, which analyzes it and uses it to improve the AI model. The feedback data can be saved and used as training data for the machine learning model to improve the accuracy of answers in future searches.
[1123] Specific examples
[1124] For example, if a user types "What are the basics of electrical wiring?", the following steps are taken:
[1125] Terminal: The user types in a question and sends it to the server.
[1126] Server: Receives the question, performs tokenization, morphological analysis, and grammatical structure analysis, and extracts important keywords (e.g., "electrical wiring" and "basic knowledge").
[1127] Server: Searches internal databases and external knowledge bases to generate appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1128] Terminal: Receives the generated answers and displays them on the screen.
[1129] User: Review the answer and provide feedback such as "this was helpful" or "I'd like more information."
[1130] Device: Sends feedback to the server.
[1131] Server: Analyzes feedback and improves the AI model.
[1132] By implementing the system in this way, information on building technology can be provided quickly and appropriately, and user feedback can be utilized to continuously improve the system's performance.
[1133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1134] Step 1: User Input
[1135] The user inputs specific questions or requests about construction into a terminal (PC, smartphone, tablet). A question such as "Please tell me the steps for laying the foundations for a wooden house" is input using a keyboard or touch screen. The input text data is received by the input means of the terminal. (Input) The user's question text. (Output) The input question as text data.
[1136] Step 2: Sending Input
[1137] The device sends questions and requests entered by the user to the server. Specifically, the entered text data is packaged in JSON format or similar over the network and sent to the server using an HTTP request. (Input) Question as text data. (Output) HTTP request sent to the server.
[1138] Step 3: Natural Language Processing (NLP)
[1139] The server analyzes the received user question and performs the following specific processing:
[1140] Tokenization: The server divides the question into words. For example, the sentence "Please tell me the procedure for laying the foundation for a wooden house" is divided into "wooden," "house," "foundation," "construction," "procedure," "tell me," and "please." (Input) The question as text data. (Output) A tokenized word list.
[1141] Morphological analysis: The server identifies the part of speech of each word. For example, "wooden structure (noun)", "house (noun)", "foundation (noun)", "construction (noun)", "procedure (noun)", "teaching (verb)". (Input) A tokenized word list. (Output) A word list with parts of speech assigned.
[1142] Grammatical structure analysis: The server analyzes the grammatical structure of the sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords. (Input) A list of words with parts of speech assigned. (Output) A list of extracted keywords.
[1143] Step 4: Information retrieval and answer generation
[1144] Based on the analysis results, the server uses the extracted keywords to search for relevant information from its internal database and external knowledge bases. The specific processes are as follows:
[1145] Internal database search: The server searches the internal database (technical manuals and industry standard information) using keywords such as "wooden house," "foundation work," and "procedure." (Input) Extracted keyword list. (Output) Information obtained from the internal database.
[1146] External knowledge base search: The server searches an external knowledge base (such as public documents on the Internet). (Input) Extracted keyword list. (Output) Information obtained from the external knowledge base.
[1147] Answer generation: Generate an answer in a format appropriate for the user from the search results. For example, generate an answer in the format "The steps for laying the foundation for a wooden house are as follows: 1. Level the site 2. Pour concrete for the foundation 3. Install rebar 4. Install formwork 5. Re-pour concrete." (Input) Information from the search results. (Output) Generated answer text.
[1148] Step 5: View your answers
[1149] The server sends the generated answer to the device. The device visually displays the received answer to the user. Specifically, the server packages the answer in JSON format and sends it back to the device as an HTTP response. The device displays the answer on the screen so that the user can confirm it. (Input) Generated answer. (Output) Answer displayed on the device.
[1150] Step 6: Getting feedback
[1151] The user provides feedback on the displayed answer. For example, options such as "Helpful" or "Need more details" are displayed and the user selects one. The device receives this feedback and sends it to the server. (Input) User feedback. (Output) Feedback sent to the server.
[1152] Step 7: Processing feedback
[1153] The server analyzes the received feedback and uses it to improve the AI model. Specifically, it stores the feedback data and uses it as training data for the machine learning model. This data can be used to improve the accuracy of the AI model. (Input) User feedback. (Output) Improved AI model, improving the accuracy of answers from next time onwards.
[1154] (Application example 1)
[1155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1156] When field workers perform tasks that require specialized knowledge, they need to be provided with information quickly and accurately. However, in the past, they often relied on specialized books and manuals, which lacked immediacy and reduced work efficiency. Furthermore, there was a lack of mechanisms for utilizing feedback to continuously improve the system. This has led to concerns that this increases the burden on workers and reduces the accuracy and efficiency of work, especially in sites with a wide range of complex tasks, such as factories.
[1157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1158] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intent, an information search and generation means for searching for related information based on the extracted keywords and generating answers, a display means for displaying the generated answers to the user, a feedback processing means for receiving feedback from the user and using it to improve the AI model, and a smart gadget equipped with a voice input device and a display device as a terminal used by the user, wherein the input means accepts voice input and the display means has the function of displaying answers on the display of the smart gadget. This allows on-site workers to input questions hands-free and receive answers quickly, improving work efficiency and enabling continuous improvement of the system.
[1159] 1. "Specific questions or requests regarding construction"
[1160] "Specific questions or requests regarding construction" refers to specific information or instructions that a user requests regarding the design, construction, maintenance, etc. of a building.
[1161] 2. "Input methods that accept natural language"
[1162] "Input means that accepts natural language" refers to devices or software that accept questions or requests entered by the user using everyday language or technical terms.
[1163] 3. “Analysis means”
[1164] "Analysis means" refers to the technology and functions used to process input data and extract important keywords and the intent of the question.
[1165] 4. "Extract important keywords and intent"
[1166] "Extracting key keywords and intent" means identifying specific words and phrases, as well as the purpose and meaning behind them, from a user's question or request.
[1167] 5. "Information retrieval and generation methods"
[1168] "Information search and generation means" refers to the technology and functions that search for related information from databases and external information sources based on extracted keywords and generate appropriate answers.
[1169] 6. "A means of displaying the generated answer to the user"
[1170] "Display means for displaying the generated answer to the user" refers to technology or devices for visually presenting the answer generated by the system on a display of a computer, mobile device, etc.
[1171] 7. "Feedback Processing Means"
[1172] "Feedback processing means" refers to the technology and functions that receive evaluations and opinions on responses from users, analyze them, and use them to improve the system.
[1173] 8. "Smart Gadgets"
[1174] "Smart gadgets" are portable devices with internet connectivity and various functions, such as smart glasses and smartphones.
[1175] 9. "Voice input device"
[1176] A "voice input device" is a microphone and associated software that allows a user to speak questions or commands.
[1177] 10. “Display device”
[1178] "Display device" means a display or screen for visually displaying information or data.
[1179] An embodiment of the present invention will be described. The present invention is a system that enables factory workers to use smart gadgets to ask questions about building and equipment maintenance in real time and quickly receive appropriate answers.
[1180] System Configuration
[1181] The system consists of the following main components:
[1182] Smart gadgets (e.g., smart glasses)
[1183] server
[1184] Database
[1185] display device
[1186] 1. User Input
[1187] The user wears a smart gadget and inputs questions or requests in natural language using voice, such as "Please tell me the basic procedures for equipment inspection."
[1188] 2. Sending Input
[1189] A voice input device in the smart gadget converts the user's voice into text data (using voice recognition software) and sends the text data to a server.
[1190] 3. Natural Language Processing (NLP)
[1191] The server analyzes the received user question using the following techniques:
[1192] Tokenization: Divide the question into words.
[1193] Morphological analysis: Identifying the part of speech of each word.
[1194] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1195] For example, "equipment," "inspection," and "basic procedures" are extracted as main keywords.
[1196] 4. Information Retrieval and Answer Generation
[1197] The server uses keywords extracted from the analysis results to search for relevant information from its internal database and external knowledge bases, and generates appropriate answers using a generative AI model based on the information obtained as search results.
[1198] Example: "The basic steps for equipment inspection are as follows: 1. Visual inspection 2. Operation check 3. Connection check 4. Lubrication check"
[1199] 5. View Answers
[1200] The server transmits the generated answer to the smart gadget, which visually displays the answer on its display device.
[1201] 6. Get feedback
[1202] Users provide feedback on the displayed answers, including whether they found the answer helpful or if they need more detail.
[1203] 7. Processing Feedback
[1204] Through the feedback collection function of the smart gadget, user feedback is sent to the server, which analyzes this feedback and helps improve the generative AI model.
[1205] Usage example
[1206] For example, the user inputs, "Please tell me the basic procedures for equipment inspection."
[1207] Smart gadget: Converts user's voice input into text and sends it to the server.
[1208] Server: Performs natural language processing and analyzes the intent of the question (e.g., "equipment," "inspection," "basic procedures").
[1209] Server: Searches internal databases and external knowledge bases and generates appropriate answers using generative AI models (e.g., "The basic steps for equipment inspection are as follows...").
[1210] Smart Gadget: Displays the generated answers to the user.
[1211] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[1212] Server: Receives feedback and helps improve the generative AI model.
[1213] In this way, the system of the present invention allows field workers to input questions hands-free and receive quick answers, thereby improving work efficiency and enabling continuous improvement of the system.
[1214] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1215] Step 1:
[1216] The user uses a smart gadget to provide voice input. For voice input, the user speaks, "Please tell me the basic procedures for equipment inspection." The smart gadget converts this voice into text data. The input is voice data, and the output is text data. Specifically, the gadget converts the voice data into text using voice recognition software (e.g., Google Speech-to-Text API).
[1217] Step 2:
[1218] The terminal transmits text data to the server. This is done using a network. In this case, the input is text data from the user, and the output is data sent to the server. In concrete terms, the smart gadget transmits text data to the server using a wireless network.
[1219] Step 3:
[1220] The server applies natural language processing (NLP) to the received text data. This processing includes tokenization, morphological analysis, and grammatical structure analysis. The input is the text data, and the output is the key keywords and the intent of the question. Specifically, the server uses the spaCy library to analyze the text data and extract key keywords.
[1221] Step 4:
[1222] The server searches for related information from the internal database and external knowledge base based on the analysis results. The input is the extracted keywords, and the output is related information as a search result. Specifically, it searches the internal database using SQL queries and retrieves information from the external knowledge base using APIs.
[1223] Step 5:
[1224] The server uses a generative AI model to generate an appropriate answer. The input is relevant information, and the output is a response to the user. Specifically, the server uses Hugging Face's transformers library to input relevant information into the generative AI model as a prompt, and generates an answer. Take the prompt "Please tell me the basic procedures for equipment inspection" as an example.
[1225] Step 6:
[1226] The server sends the generated answer to the smart gadget. The input is the answer, and the output is the display data for the smart gadget. Specifically, the server sends text data to the smart gadget via the network.
[1227] Step 7:
[1228] The smart gadget displays the received answer on a display device. The input is the text data of the answer, and the output is visual information displayed on the display. Specifically, the smart gadget renders the received text data on the display.
[1229] Step 8:
[1230] The user provides feedback on the displayed answer by speaking a rating such as "helpful" or "I'd like more information." The input is voice data, and the output is text feedback. Specifically, as with the initial voice input, speech recognition software is used to convert the speech into text.
[1231] Step 9:
[1232] The terminal transmits the feedback text data to the server. The input is the feedback text data from the user, and the output is the data to be transmitted to the server. In concrete terms, the smart gadget transmits the feedback data to the server using a wireless network.
[1233] Step 10:
[1234] The server analyzes the feedback and uses it to improve the AI model. The input is the text feedback, and the output is an improved AI model. Specifically, the server uses the feedback data to train and update the AI model.
[1235] This allows field workers to input questions hands-free and receive quick answers, improving work efficiency and enabling continuous improvement of the system.
[1236] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1237] MODE FOR CARRYING OUT THE INVENTION
[1238] An embodiment of the present invention will be described. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. In addition, it combines this with an emotion engine that recognizes the user's emotions.
[1239] System Configuration
[1240] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[1241] 1. User Input
[1242] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[1243] 2. Sending Input
[1244] The terminal sends questions and requests entered by the user to the server.
[1245] 3. Natural Language Processing (NLP)
[1246] The server analyzes the received user question, specifically:
[1247] Tokenization: Divide the question into words.
[1248] Morphological analysis: Identifying the part of speech of each word.
[1249] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1250] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1251] 4. Emotion analysis
[1252] The server uses an emotion engine to recognize the user's emotion in the question. The following process is performed:
[1253] Emotion recognition: Extracting user emotions (e.g., confusion, excitement, anger, etc.) from input text.
[1254] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[1255] For example, if a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[1256] 5. Information Retrieval and Answer Generation
[1257] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[1258] Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[1259] Generate answers that are interesting and considerate of the user's emotions.
[1260] For example, an answer might be generated along the lines of, "Please understand that foundation work is difficult, and refer to the following steps."
[1261] 6. View Answers
[1262] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1263] 7. Get feedback
[1264] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[1265] 8. Submitting Feedback
[1266] The device sends the feedback from the user to the server.
[1267] 9. Processing Feedback
[1268] The server analyzes this feedback and uses it to improve the AI model. Specifically, it analyzes the feedback, extracts evaluation information, and reflects it in the AI model for re-learning. This allows the system to continuously improve.
[1269] Specific examples
[1270] For example, if a user types "What are the basics of electrical wiring?":
[1271] Terminal: Sends user input to the server.
[1272] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[1273] Server: Uses an emotion engine to recognize the user's emotions (e.g., confusion, interest).
[1274] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1275] Terminal: Displays the generated answer to the user.
[1276] User: Provide feedback on the answer.
[1277] Server: Receives feedback and helps improve the AI model.
[1278] In this way, the system of the present invention not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take into consideration the user's feelings. This effectively solves the problems of passing on construction techniques and the shortage of craftsmen, and contributes to the improvement of technology in the construction industry as a whole.
[1279] The processing flow will be explained below.
[1280] Step 1: User enters question
[1281] User: Enters specific questions or requests about construction into the terminal.
[1282] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[1283] Step 2: Sending Input
[1284] Terminal: Sends questions and requests entered by the user to the server.
[1285] Specifically, the user input data is sent to the server as an HTTP request.
[1286] Step 3: Tokenize the Question
[1287] Server: Tokenizes the received question.
[1288] Specifically, the question is divided into words and phrases.
[1289] For example, it is divided into "wooden house," "foundation work," and "procedure."
[1290] Step 4: Morphological analysis
[1291] Server: Performs morphological analysis of the question.
[1292] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[1293] Step 5: Grammatical structure analysis
[1294] Server: Analyzes the grammatical structure of the question.
[1295] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[1296] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1297] Step 6: Sentiment Analysis
[1298] Server: Recognizes the user's emotions contained in the question using an emotion engine.
[1299] Specifically, it extracts the user's emotions (e.g., confusion, excitement, anger, etc.) from the input text and classifies them into predefined emotion categories.
[1300] Example: If a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[1301] Step 7: Information retrieval and answer generation
[1302] Server: Searches for relevant information and generates answers based on the analysis results and emotion recognition results.
[1303] Specifically, based on the extracted keywords and recognized emotions, the system retrieves relevant data from an internal database and an external knowledge base.
[1304] Generate answers that are interesting and considerate of the user's emotions.
[1305] For example, an answer like "Please understand that foundation work is difficult, and refer to the following steps." is generated.
[1306] Step 8: Submit your response
[1307] Server: Sends the generated answer to the device.
[1308] Step 9: View your answers
[1309] Terminal: Displays the received answer to the user.
[1310] As a specific operation, the answer content is displayed on the screen.
[1311] Step 10: Get feedback
[1312] User: Provide feedback on the displayed answers.
[1313] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[1314] Step 11: Submit your feedback
[1315] Terminal: Sends user feedback to the server.
[1316] Step 12: Processing feedback
[1317] Server: Analyzes the feedback received and helps improve the AI model.
[1318] Specifically, the feedback content is analyzed and fed back to the AI model as adaptable data.
[1319] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[1320] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[1321] Example 2
[1322] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1323] In the modern construction industry, there are limited means of instantly obtaining technical information and procedures related to construction, creating a demand for fast, accurate answers to technical questions and problems. It is also important to increase user satisfaction by providing answers that take the user's emotions into consideration. However, conventional systems struggle to respond in a way that takes user emotions into account, and they lack mechanisms for appropriately utilizing user feedback to improve the system. Therefore, there is a need for a system that can analyze user questions and requests, recognize emotions, generate and display appropriate answers, and further utilize feedback.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1325] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an information search and generation means for searching for related information based on the extracted keywords and the user's emotions and generating answers, a display means for displaying the generated answers to the user, and a feedback processing means for receiving feedback from the user and using it to improve the AI model. This makes it possible to quickly and appropriately answer user questions, provide answers that take user emotions into consideration, and utilize feedback to continuously improve the system.
[1326] "Input means" refers to devices or functions that accept specific questions or requests about architecture from users in natural language.
[1327] "Analysis means" refers to a device or function that analyzes questions or requests entered by users and extracts important keywords and intentions.
[1328] "Information search and generation means" refers to devices or functions that search for related information and generate answers based on the keywords extracted by the analysis means and the user's emotions.
[1329] "Display means" refers to a device or function that visually displays the generated answer to the user.
[1330] A "feedback processing means" is a device or function that accepts feedback provided by users and uses that feedback to help improve the AI model.
[1331] "Tokenization" is the process of dividing an input natural language sentence into words.
[1332] "Morphological analysis" is the process of identifying the part of speech of each tokenized word.
[1333] "Grammar structure analysis" is the process of analyzing the grammatical structure of an input sentence and extracting important keywords and intent.
[1334] "Emotion recognition" is the process of extracting a user's emotions from input text.
[1335] "Emotion classification" refers to the process of classifying extracted emotions into predefined emotion categories.
[1336] An "internal database" is a database that manages information and data stored within the system.
[1337] An "external knowledge base" is an information source or database that is externally accessible, such as the Internet.
[1338] An "AI model" is a model that uses artificial intelligence to learn and perform processes such as answering questions and recognizing emotions.
[1339] "Feedback" refers to the ratings and opinions that users provide in response to the answers displayed.
[1340] The system of the present invention allows users to input specific questions or requests about architecture in natural language, analyzes them, and generates and displays answers. Furthermore, it can collect feedback from users and use it to improve the system. The system consists of the following main components:
[1341] System Configuration
[1342] 1. Hardware configuration:
[1343] Terminal: A device such as a PC, smartphone, or tablet that allows users to input questions or requests and display answers.
[1344] Server: A high-performance cloud server or on-premise server with high data processing capacity and fast response is recommended.
[1345] 2. Software configuration:
[1346] Natural language processing engine: An engine that tokenizes text, analyzes morphology, and analyzes grammatical structures (e.g., Google NLP).
[1347] Emotion engine: An engine for recognizing and classifying user emotions (e.g., IBM Watson Tone Analyzer).
[1348] Database management system: Databases such as MySQL and MongoDB are used as internal databases.
[1349] AI model: An artificial intelligence model that analyzes a user's question or request and generates an answer. Use a generative AI model.
[1350] Operation procedure and processing contents
[1351] 1. User input:
[1352] Users operate their own terminals to input specific questions or requests about construction in natural language, such as "Please tell me the steps for laying the foundations for a wooden house."
[1353] 2. Sending input:
[1354] The device sends the user-entered questions or requests to a server over the Internet. The data is sent as an HTTP request.
[1355] 3. Natural Language Processing (NLP):
[1356] The server analyzes the received user question, specifically by performing the following steps:
[1357] Tokenization: Divide the input question into words.
[1358] Morphological analysis: Identifying the parts of speech of segmented words.
[1359] Grammatical structure analysis: Analyze grammatical structures and extract important keywords and intent.
[1360] 4. Emotion analysis:
[1361] The server uses an emotion engine to recognize the user's emotion in the question, which includes the following steps:
[1362] Emotion Recognition: Extracting user emotions from input text.
[1363] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[1364] 5. Information retrieval and answer generation:
[1365] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[1366] Information retrieval: Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[1367] Answer generation: Based on the searched information, an answer is generated that takes into consideration the user's feelings.
[1368] 6. Show Answer:
[1369] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1370] 7. Getting feedback:
[1371] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[1372] 8. Submitting Feedback:
[1373] The device sends the user's feedback to the server, again as an HTTP request.
[1374] 9. Handling feedback (retraining the AI model):
[1375] The server analyzes the received feedback and extracts evaluation information, which is then used to retrain the AI model and improve response accuracy from the next time onwards.
[1376] Specific examples
[1377] For example, if a user types "Please tell me the basics of electrical wiring," the following will be processed:
[1378] User: Enters a question into the terminal and sends it.
[1379] Server: Receives input and performs natural language processing and sentiment analysis.
[1380] Server: Searches for relevant information from internal databases and external knowledge bases and generates appropriate answers.
[1381] Terminal: Display the generated answer.
[1382] User: Provide feedback on the answer.
[1383] Server: Receives feedback and helps retrain the AI model.
[1384] This system not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take the user's emotions into consideration. This will effectively solve the problems of passing on construction techniques and the shortage of craftsmen, and contribute to improving the technology of the entire construction industry.
[1385] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1386] Step 1:
[1387] The user uses the input means to input specific questions or requests about construction in natural language. For example, they may input a question such as, "Please tell me the steps for laying the foundations for a wooden house." This input is saved as text data on the terminal.
[1388] Step 2:
[1389] The terminal sends the text data entered by the user to the server. Specifically, the data is sent to the server as an HTTP request. At this time, the input data is temporarily stored in the terminal's buffer memory. The input is the user's question text, and the output is the sent HTTP request.
[1390] Step 3:
[1391] The server analyzes the received text data. First, it uses a natural language processing engine to tokenize the question. As a result of tokenization, the input sentence can be divided into words, and the output is a word list. For example, it may be divided into words such as "wooden house," "foundation work," "procedure," "tell me," and "please."
[1392] Step 4:
[1393] The server performs morphological analysis. As a result of the morphological analysis, it is possible to identify the part of speech for each word. For example, "wooden house" is recognized as a noun and "teach me" as a verb. The input is a list of tokenized words, and the output is a list of words with identified parts of speech.
[1394] Step 5:
[1395] The server performs grammatical analysis. This analyzes the grammatical structure of the entire sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as key keywords. The results of the grammatical analysis are the extracted keywords and a structural analysis tree.
[1396] Step 6:
[1397] The server uses an emotion recognition engine to recognize the user's emotions. During the emotion recognition process, the server extracts the user's emotional characteristics (e.g., confusion, excitement, anger) from the text data and classifies them into categories. For example, the emotion "confusion" is extracted from the sentence "I'm having trouble because the foundation work on a wooden house is difficult." The input is the analyzed text data, and the output is the recognized emotion and its category.
[1398] Step 7:
[1399] The server performs an information search based on the analysis results and emotion recognition results. It searches its internal database and external knowledge base to obtain relevant information. For example, it searches the internal database for technical information using keywords such as "wooden house," "foundation work," and "procedure." The information obtained at this stage is output in the form of text or images.
[1400] Step 8:
[1401] The server generates an answer based on the search results. Using a generative AI model, it creates an answer that takes into account the acquired information and the user's emotions. For example, an answer such as "Please understand that foundation construction is difficult, and refer to the following steps" may be generated. The input is the search results and emotion data, and the output is the answer text for the user.
[1402] Step 9:
[1403] The server sends the generated answer to the terminal. The terminal displays the received answer to the user. Specifically, the answer is displayed in text format on the terminal screen. The input is the answer data from the server, and the output is the display for the user.
[1404] Step 10:
[1405] The user provides feedback on the displayed answer, including an evaluation of whether the answer was helpful or whether further explanation is needed. The feedback is again input into the terminal as text data.
[1406] Step 11:
[1407] The device sends the user's feedback to the server, again as an HTTP request. The input is the user's feedback text, and the output is the sent HTTP request.
[1408] Step 12:
[1409] The server receives and analyzes the feedback data. As a result of the analysis, specific evaluation information is extracted and used to retrain the AI model. This improves response accuracy from the next time onwards. The input is the feedback data, and the output is an updated AI model.
[1410] (Application example 2)
[1411] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1412] Lack of technical knowledge and information on construction sites, especially complex work procedures and troubleshooting, makes it difficult to quickly implement them. Furthermore, the inability to provide appropriate advice that takes into account the feelings of workers makes it difficult to work efficiently. Furthermore, there is a lack of means to continuously improve the system based on user feedback.
[1413] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1414] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an emotion analysis means for analyzing emotions using an emotion engine that recognizes the user's emotions, an information search and generation means for searching for related information and generating answers based on the extracted keywords and recognized emotions, a display means for displaying the generated answers to the user, and a feedback processing means for accepting feedback from users and using it to improve the AI model. This enables workers at construction sites to respond quickly to questions and requests, provides appropriate advice that takes workers' emotions into consideration, and continuously improves the system, enabling efficient and effective work.
[1415] "Input means" refers to a device or system that accepts specific questions or requests about architecture from users in natural language.
[1416] The "analysis means" is a device or program that analyzes the input question or request and extracts important keywords and intentions.
[1417] "Emotion analysis means" refers to a device or system for analyzing emotions from input text using an emotion engine that recognizes the emotions of a user.
[1418] "Information search and generation means" refers to a device or program that searches for related information based on the extracted keywords and recognized emotions and generates answers.
[1419] A "display means" is a device or system that visually displays the generated answers to the user.
[1420] A "feedback processing means" is a device or program that accepts feedback from users and uses it to improve the AI model.
[1421] The present invention is a system in which a factory robot responds to questions and requests from workers at a construction site and provides advice that takes into consideration their emotions. An embodiment of the system is described below.
[1422] System Configuration
[1423] The system includes the following components:
[1424] 1. Input Method
[1425] The factory robot can receive specific questions and requests about construction from workers in natural language. For example, the factory robot can receive a question from a worker such as, "Please tell me the key points of foundation work."
[1426] 2. Analysis method
[1427] The factory robot analyzes the questions and requests it receives. This analysis involves tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent. For example, keywords such as "foundation work" and "key points" are extracted.
[1428] 3. Emotion analysis means
[1429] The factory robot uses an emotion engine to analyze the emotions of workers from input text. Specifically, the input text is input into the emotion engine, and emotions such as "confusion" or "interest" are recognized.
[1430] 4. Information retrieval and generation methods
[1431] Based on the extracted keywords and the recognized emotions, the factory robot searches for relevant information from its internal database and external knowledge base, and generates an appropriate answer taking into account the search results and the emotions of the worker. For example, it generates an answer in the form of "The main points of foundation work are as follows..."
[1432] 5. Display means
[1433] The generated answers are displayed and presented visually or audibly by the factory robot to the worker, allowing the worker to obtain the appropriate information.
[1434] 6. Feedback Processing Methods
[1435] Workers can provide feedback on the displayed answers, which helps improve the robot's AI model, such as "This answer was helpful" or "Please give me more details."
[1436] Hardware and software used
[1437] Hardware
[1438] Factory robots are equipped with standard input devices (microphones, cameras, etc.) to receive user input, and display devices (displays, speakers, etc.).
[1439] software
[1440] The factory robots are installed with Transformers, a library for natural language processing (NLP). They also use the Sentiment-Analysis model, which is known for its stability and performance, for sentiment analysis. They use the Requests library to communicate with the server.
[1441] Specific examples
[1442] For example, if a factory robot receives a question from a worker saying, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?", it will process it as follows:
[1443] 1. Input Method
[1444] The worker types in a question.
[1445] 2. Analysis method
[1446] The content of the question is analyzed and keywords and intent such as "wooden house," "foundation work," and "having trouble" are extracted.
[1447] 3. Emotion analysis means
[1448] The emotion of "being in trouble" is recognized using an emotion engine.
[1449] 4. Information retrieval and generation methods
[1450] It searches for relevant information based on keywords and sentiment and generates answers such as, "For the foundation construction procedure, please try the following steps."
[1451] 5. Display means
[1452] The generated answer is displayed on the screen and communicated to the worker by voice.
[1453] 6. Feedback Processing Methods
[1454] The worker provides feedback, saying, "This information was useful," and the AI model is improved based on that feedback.
[1455] An example of a prompt sentence is, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?"
[1456] This allows factory robots to respond quickly and appropriately to questions from workers and take workers' feelings into consideration, making on-site work more efficient.
[1457] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1458] Step 1:
[1459] The user inputs specific questions and requests about construction to the factory robot in natural language.
[1460] Input: Questions or requests in natural language (e.g., "I'm having trouble with the foundation work on my wooden house. What should I do specifically?")
[1461] Output: The input text is sent to the system
[1462] Step 2:
[1463] The terminal sends the input text to the server.
[1464] Input: Text entered by the user
[1465] Output: The input text is passed to the server
[1466] Step 3:
[1467] The server processes the received text using analytical means, specifically tokenizing, morphological analysis, and grammatical structure analysis to extract important keywords and intent from the question or request.
[1468] Input: Input text (e.g., "Foundation work for wooden houses")
[1469] Data processing / data operations: tokenization, morphological analysis, grammatical analysis
[1470] Output: Keywords and intent (e.g., "foundation work," "difficult," etc.)
[1471] Step 4:
[1472] Recognize user emotions through emotion analysis. Using an emotion engine, emotions (e.g., "confused") are extracted from the input text and classified into emotion categories.
[1473] Input: Parsed text and extracted keywords
[1474] Data processing / data calculation: emotion extraction and classification
[1475] Output: Recognized emotion (e.g., "confused")
[1476] Step 5:
[1477] The information search and generation means searches for related information from the internal database and external knowledge base based on the extracted keywords and emotions, and generates answers based on the search results.
[1478] Input: Keywords, Recognized Sentiments
[1479] Data processing / data calculation: database search, answer generation
[1480] Output: Generated answer (e.g., "Please proceed as follows for the foundation construction procedure.")
[1481] Step 6:
[1482] The server transmits the generated answer to the terminal, which then presents the answer to the user visually or audibly through a display means.
[1483] Input: Generated Answer
[1484] Output: The answer that is presented to the user
[1485] Step 7:
[1486] The user provides feedback on the displayed answer, including whether it was helpful or if more details are needed.
[1487] Input: User feedback (e.g., satisfied, dissatisfied, further questions, etc.)
[1488] Output: Feedback is sent to the system
[1489] Step 8:
[1490] The device sends feedback to the server, which then receives and analyzes the feedback to help improve the AI model. Specifically, the feedback content is analyzed, evaluation information is extracted, and the feedback is reflected in the AI model for re-learning.
[1491] Input: User feedback
[1492] Data processing / data calculation: Feedback analysis, evaluation information extraction, AI model retraining
[1493] Output: An improved AI model
[1494] As described above, we will explain how the system operates at each processing step and what inputs and outputs there are.
[1495] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1496] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1497] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1498] [Fourth embodiment]
[1499] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1500] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1501] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1502] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1503] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1504] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1505] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1506] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1507] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1510] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1511] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] MODE FOR CARRYING OUT THE INVENTION
[1513] This section describes an embodiment of the present invention. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system.
[1514] System Configuration
[1515] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[1516] 1. User Input
[1517] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[1518] 2. Sending Input
[1519] The terminal sends questions and requests entered by the user to the server.
[1520] 3. Natural Language Processing (NLP)
[1521] The server analyzes the received user question, specifically:
[1522] Tokenization: Divide the question into words.
[1523] Morphological analysis: Identifying the part of speech of each word.
[1524] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1525] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1526] 4. Information Retrieval and Answer Generation
[1527] The server uses the keywords extracted from the analysis results to search for related information from its internal database and external knowledge bases, and generates appropriate answers based on the information obtained as search results.
[1528] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1529] 5. View Answers
[1530] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1531] 6. Get feedback
[1532] The user provides feedback on the displayed answers, including a rating of whether the answer was helpful or if more detail is needed.
[1533] 7. Processing Feedback
[1534] The device sends user feedback to the server, which analyzes it and uses it to improve the AI model, allowing the system to continuously learn and provide better answers in the future.
[1535] Specific examples
[1536] For example, the user inputs "Please tell me the basics of electrical wiring."
[1537] Terminal: Sends user input to the server.
[1538] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[1539] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1540] Terminal: Displays the generated answer to the user.
[1541] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[1542] Server: Receives feedback and helps improve the AI model.
[1543] In this way, the system of the present invention can solve the problems of a shortage of craftsmen and the transfer of skills by providing information on construction technology quickly and appropriately, and contribute to improving technology throughout the construction industry.
[1544] The processing flow will be explained below.
[1545] Step 1: User enters question
[1546] User: Enters specific questions or requests about construction into the terminal.
[1547] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[1548] Step 2: Sending Input
[1549] Terminal: Sends questions and requests entered by the user to the server.
[1550] Specifically, the user input data is sent to the server as an HTTP request.
[1551] Step 3: Tokenize the Question
[1552] Server: Tokenizes the received question.
[1553] Specifically, the question is divided into words and phrases.
[1554] For example, it is divided into "wooden house," "foundation work," and "procedure."
[1555] Step 4: Morphological analysis
[1556] Server: Performs morphological analysis of the question.
[1557] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[1558] Step 5: Grammatical structure analysis
[1559] Server: Analyzes the grammatical structure of the question.
[1560] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[1561] Step 6: Finding information and generating answers
[1562] Server: Searches for relevant information based on the analysis results and generates answers.
[1563] Specifically, it retrieves relevant data from internal databases and external knowledge bases.
[1564] Based on the search results, answers that are easy for users to understand are generated.
[1565] Example: "The steps for constructing the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1566] Step 7: Submit your response
[1567] Server: Sends the generated answer to the device.
[1568] Step 8: View your answers
[1569] Terminal: Displays the received answer to the user.
[1570] As a specific operation, the answer content is displayed on the screen.
[1571] Step 9: Get feedback
[1572] User: Provide feedback on the displayed answers.
[1573] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[1574] Step 10: Submit your feedback
[1575] Terminal: Sends user feedback to the server.
[1576] Step 11: Processing feedback
[1577] Server: Analyzes the feedback received and helps improve the AI model.
[1578] Specifically, the content of the feedback is analyzed and fed back to the AI model as adaptable data.
[1579] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[1580] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[1581] Example 1
[1582] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1583] In recent years, while the demand for technical information on architecture has increased, the shortage of engineers with specialized knowledge has become a problem. The difficulty of transferring skills and the need for efficient information provision are also increasing. In response to these issues, there is a demand for systems that can provide information on architectural technology quickly and accurately, and that can continuously improve the system's performance based on user feedback.
[1584] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1585] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, a means for transmitting the input questions and requests to the server, an analysis means for extracting important keywords and intent by tokenizing, morphologically analyzing, and grammatically analyzing the received questions, an information search and generation means for searching for related information from an internal database and an external knowledge base based on the extracted keywords and generating an answer, a means for transmitting and displaying the generated answer to the user, and a feedback processing means for accepting and analyzing feedback from users and using it to improve the AI model. This allows users to quickly and accurately obtain information about architecture technology, and enables the system to continuously learn and improve its performance.
[1586] "Input means" refers to a device or interface for accepting specific questions or requests about architecture from users in natural language.
[1587] The "server" is a central information processing device that analyzes information sent by users and generates appropriate responses.
[1588] "Tokenization" is a natural language processing process that divides an input question or request into words.
[1589] "Morphological analysis" is a natural language processing technique for identifying the part of speech of each word.
[1590] "Grammar structure analysis" is a process in natural language processing that analyzes the grammatical structure of a sentence and extracts important keywords and intent.
[1591] The "analysis means" is a device or software that analyzes the received question through tokenization, morphological analysis, and grammatical structure analysis to extract important keywords and intent.
[1592] "Information search and generation means" refers to a device or software that searches for related information from internal databases and external knowledge bases based on extracted keywords and generates answers.
[1593] A "display means" is a device or interface for visually displaying the generated answers to the user.
[1594] "Feedback processing means" refers to a device or software that accepts and analyzes feedback from users and uses it to improve the AI model.
[1595] MODE FOR CARRYING OUT THE INVENTION
[1596] The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. An embodiment of the present invention is described below.
[1597] System Configuration
[1598] This system consists of a device operated by the user (e.g., PC, smartphone, tablet) and a server for processing information sent from these devices.
[1599] User Input
[1600] The user inputs specific questions or requests about construction in natural language using the input means of the terminal. For example, a question such as "Please tell me the procedure for laying the foundations of a wooden house" is input using the keyboard or touch screen of the terminal.
[1601] Sending Input
[1602] The device sends the questions and requests entered by the user to the server. In this process, the device packages the entered text data in JSON format or similar and sends it using an HTTP request to send it to the server over the network.
[1603] Natural Language Processing (NLP)
[1604] The server analyzes the received user question. Specifically, it performs the following analyses: tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent, such as "wooden house," "foundation work," and "procedure."
[1605] Information retrieval and answer generation
[1606] The server uses keywords extracted from the analysis results to search for related information from its internal database and external knowledge base. Based on the information obtained as a search result, it generates an appropriate answer. For example, it generates an answer such as, "The steps for laying the foundation of a wooden house are as follows: 1. Level the site. 2. Pour concrete for the foundation. 3. Install rebar. 4. Install formwork. 5. Re-pour concrete."
[1607] Show Answers
[1608] The server sends the generated answer to the terminal, which then visually displays the received answer to the user, allowing the user to check the answer on the terminal screen.
[1609] Get feedback
[1610] The user provides feedback on the displayed answer, for example, with options such as "helpful" or "more information," and the user enters a prompt such as:
[1611] "It was helpful"
[1612] "I want more details."
[1613] Processing Feedback
[1614] The device sends user feedback to a server, which analyzes it and uses it to improve the AI model. The feedback data can be saved and used as training data for the machine learning model to improve the accuracy of answers in future searches.
[1615] Specific examples
[1616] For example, if a user types "What are the basics of electrical wiring?", the following steps are taken:
[1617] Terminal: The user types in a question and sends it to the server.
[1618] Server: Receives the question, performs tokenization, morphological analysis, and grammatical structure analysis, and extracts important keywords (e.g., "electrical wiring" and "basic knowledge").
[1619] Server: Searches internal databases and external knowledge bases to generate appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1620] Terminal: Receives the generated answers and displays them on the screen.
[1621] User: Review the answer and provide feedback such as "this was helpful" or "I'd like more information."
[1622] Device: Sends feedback to the server.
[1623] Server: Analyzes feedback and improves the AI model.
[1624] By implementing the system in this way, information on building technology can be provided quickly and appropriately, and user feedback can be utilized to continuously improve the system's performance.
[1625] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1626] Step 1: User Input
[1627] The user inputs specific questions or requests about construction into a terminal (PC, smartphone, tablet). A question such as "Please tell me the steps for laying the foundations for a wooden house" is input using a keyboard or touch screen. The input text data is received by the input means of the terminal. (Input) The user's question text. (Output) The input question as text data.
[1628] Step 2: Sending Input
[1629] The device sends questions and requests entered by the user to the server. Specifically, the entered text data is packaged in JSON format or similar over the network and sent to the server using an HTTP request. (Input) Question as text data. (Output) HTTP request sent to the server.
[1630] Step 3: Natural Language Processing (NLP)
[1631] The server analyzes the received user question and performs the following specific processing:
[1632] Tokenization: The server divides the question into words. For example, the sentence "Please tell me the procedure for laying the foundation for a wooden house" is divided into "wooden," "house," "foundation," "construction," "procedure," "tell me," and "please." (Input) The question as text data. (Output) A tokenized word list.
[1633] Morphological analysis: The server identifies the part of speech of each word. For example, "wooden structure (noun)", "house (noun)", "foundation (noun)", "construction (noun)", "procedure (noun)", "teaching (verb)". (Input) A tokenized word list. (Output) A word list with parts of speech assigned.
[1634] Grammatical structure analysis: The server analyzes the grammatical structure of the sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords. (Input) A list of words with parts of speech assigned. (Output) A list of extracted keywords.
[1635] Step 4: Information retrieval and answer generation
[1636] Based on the analysis results, the server uses the extracted keywords to search for relevant information from its internal database and external knowledge bases. The specific processes are as follows:
[1637] Internal database search: The server searches the internal database (technical manuals and industry standard information) using keywords such as "wooden house," "foundation work," and "procedure." (Input) Extracted keyword list. (Output) Information obtained from the internal database.
[1638] External knowledge base search: The server searches an external knowledge base (such as public documents on the Internet). (Input) Extracted keyword list. (Output) Information obtained from the external knowledge base.
[1639] Answer generation: Generate an answer in a format appropriate for the user from the search results. For example, generate an answer in the format "The steps for laying the foundation for a wooden house are as follows: 1. Level the site 2. Pour concrete for the foundation 3. Install rebar 4. Install formwork 5. Re-pour concrete." (Input) Information from the search results. (Output) Generated answer text.
[1640] Step 5: View your answers
[1641] The server sends the generated answer to the device. The device visually displays the received answer to the user. Specifically, the server packages the answer in JSON format and sends it back to the device as an HTTP response. The device displays the answer on the screen so that the user can confirm it. (Input) Generated answer. (Output) Answer displayed on the device.
[1642] Step 6: Getting feedback
[1643] The user provides feedback on the displayed answer. For example, options such as "Helpful" or "Need more details" are displayed and the user selects one. The device receives this feedback and sends it to the server. (Input) User feedback. (Output) Feedback sent to the server.
[1644] Step 7: Processing feedback
[1645] The server analyzes the received feedback and uses it to improve the AI model. Specifically, it stores the feedback data and uses it as training data for the machine learning model. This data can be used to improve the accuracy of the AI model. (Input) User feedback. (Output) Improved AI model, improving the accuracy of answers from next time onwards.
[1646] (Application example 1)
[1647] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1648] When field workers perform tasks that require specialized knowledge, they need to be provided with information quickly and accurately. However, in the past, they often relied on specialized books and manuals, which lacked immediacy and reduced work efficiency. Furthermore, there was a lack of mechanisms for utilizing feedback to continuously improve the system. This has led to concerns that this increases the burden on workers and reduces the accuracy and efficiency of work, especially in sites with a wide range of complex tasks, such as factories.
[1649] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1650] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intent, an information search and generation means for searching for related information based on the extracted keywords and generating answers, a display means for displaying the generated answers to the user, a feedback processing means for receiving feedback from the user and using it to improve the AI model, and a smart gadget equipped with a voice input device and a display device as a terminal used by the user, wherein the input means accepts voice input and the display means has the function of displaying answers on the display of the smart gadget. This allows on-site workers to input questions hands-free and receive answers quickly, improving work efficiency and enabling continuous improvement of the system.
[1651] 1. "Specific questions or requests regarding construction"
[1652] "Specific questions or requests regarding construction" refers to specific information or instructions that a user requests regarding the design, construction, maintenance, etc. of a building.
[1653] 2. "Input methods that accept natural language"
[1654] "Input means that accepts natural language" refers to devices or software that accept questions or requests entered by the user using everyday language or technical terms.
[1655] 3. “Analysis means”
[1656] "Analysis means" refers to the technology and functions used to process input data and extract important keywords and the intent of the question.
[1657] 4. "Extract important keywords and intent"
[1658] "Extracting key keywords and intent" means identifying specific words and phrases, as well as the purpose and meaning behind them, from a user's question or request.
[1659] 5. "Information retrieval and generation methods"
[1660] "Information search and generation means" refers to the technology and functions that search for related information from databases and external information sources based on extracted keywords and generate appropriate answers.
[1661] 6. "A means of displaying the generated answer to the user"
[1662] "Display means for displaying the generated answer to the user" refers to technology or devices for visually presenting the answer generated by the system on a display of a computer, mobile device, etc.
[1663] 7. "Feedback Processing Means"
[1664] "Feedback processing means" refers to the technology and functions that receive evaluations and opinions on responses from users, analyze them, and use them to improve the system.
[1665] 8. "Smart Gadgets"
[1666] "Smart gadgets" are portable devices with internet connectivity and various functions, such as smart glasses and smartphones.
[1667] 9. "Voice input device"
[1668] A "voice input device" is a microphone and associated software that allows a user to speak questions or commands.
[1669] 10. “Display device”
[1670] "Display device" means a display or screen for visually displaying information or data.
[1671] An embodiment of the present invention will be described. The present invention is a system that enables factory workers to use smart gadgets to ask questions about building and equipment maintenance in real time and quickly receive appropriate answers.
[1672] System Configuration
[1673] The system consists of the following main components:
[1674] Smart gadgets (e.g., smart glasses)
[1675] server
[1676] Database
[1677] display device
[1678] 1. User Input
[1679] The user wears a smart gadget and inputs questions or requests in natural language using voice, such as "Please tell me the basic procedures for equipment inspection."
[1680] 2. Sending Input
[1681] A voice input device in the smart gadget converts the user's voice into text data (using voice recognition software) and sends the text data to a server.
[1682] 3. Natural Language Processing (NLP)
[1683] The server analyzes the received user question using the following techniques:
[1684] Tokenization: Divide the question into words.
[1685] Morphological analysis: Identifying the part of speech of each word.
[1686] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1687] For example, "equipment," "inspection," and "basic procedures" are extracted as main keywords.
[1688] 4. Information Retrieval and Answer Generation
[1689] The server uses keywords extracted from the analysis results to search for relevant information from its internal database and external knowledge bases, and generates appropriate answers using a generative AI model based on the information obtained as search results.
[1690] Example: "The basic steps for equipment inspection are as follows: 1. Visual inspection 2. Operation check 3. Connection check 4. Lubrication check"
[1691] 5. View Answers
[1692] The server transmits the generated answer to the smart gadget, which visually displays the answer on its display device.
[1693] 6. Get feedback
[1694] Users provide feedback on the displayed answers, including whether they found the answer helpful or if they need more detail.
[1695] 7. Processing Feedback
[1696] Through the feedback collection function of the smart gadget, user feedback is sent to the server, which analyzes this feedback and helps improve the generative AI model.
[1697] Usage example
[1698] For example, the user inputs, "Please tell me the basic procedures for equipment inspection."
[1699] Smart gadget: Converts user's voice input into text and sends it to the server.
[1700] Server: Performs natural language processing and analyzes the intent of the question (e.g., "equipment," "inspection," "basic procedures").
[1701] Server: Searches internal databases and external knowledge bases and generates appropriate answers using generative AI models (e.g., "The basic steps for equipment inspection are as follows...").
[1702] Smart Gadget: Displays the generated answers to the user.
[1703] Users: Provide feedback on the answer, such as "this was helpful" or "I'd like more details."
[1704] Server: Receives feedback and helps improve the generative AI model.
[1705] In this way, the system of the present invention allows field workers to input questions hands-free and receive quick answers, thereby improving work efficiency and enabling continuous improvement of the system.
[1706] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1707] Step 1:
[1708] The user uses a smart gadget to provide voice input. For voice input, the user speaks, "Please tell me the basic procedures for equipment inspection." The smart gadget converts this voice into text data. The input is voice data, and the output is text data. Specifically, the gadget converts the voice data into text using voice recognition software (e.g., Google Speech-to-Text API).
[1709] Step 2:
[1710] The terminal transmits text data to the server. This is done using a network. In this case, the input is text data from the user, and the output is data sent to the server. In concrete terms, the smart gadget transmits text data to the server using a wireless network.
[1711] Step 3:
[1712] The server applies natural language processing (NLP) to the received text data. This processing includes tokenization, morphological analysis, and grammatical structure analysis. The input is the text data, and the output is the key keywords and the intent of the question. Specifically, the server uses the spaCy library to analyze the text data and extract key keywords.
[1713] Step 4:
[1714] The server searches for related information from the internal database and external knowledge base based on the analysis results. The input is the extracted keywords, and the output is related information as a search result. Specifically, it searches the internal database using SQL queries and retrieves information from the external knowledge base using APIs.
[1715] Step 5:
[1716] The server uses a generative AI model to generate an appropriate answer. The input is relevant information, and the output is a response to the user. Specifically, the server uses Hugging Face's transformers library to input relevant information into the generative AI model as a prompt, and generates an answer. Take the prompt "Please tell me the basic procedures for equipment inspection" as an example.
[1717] Step 6:
[1718] The server sends the generated answer to the smart gadget. The input is the answer, and the output is the display data for the smart gadget. Specifically, the server sends text data to the smart gadget via the network.
[1719] Step 7:
[1720] The smart gadget displays the received answer on a display device. The input is the text data of the answer, and the output is visual information displayed on the display. Specifically, the smart gadget renders the received text data on the display.
[1721] Step 8:
[1722] The user provides feedback on the displayed answer by speaking a rating such as "helpful" or "I'd like more information." The input is voice data, and the output is text feedback. Specifically, as with the initial voice input, speech recognition software is used to convert the speech into text.
[1723] Step 9:
[1724] The terminal transmits the feedback text data to the server. The input is the feedback text data from the user, and the output is the data to be transmitted to the server. In concrete terms, the smart gadget transmits the feedback data to the server using a wireless network.
[1725] Step 10:
[1726] The server analyzes the feedback and uses it to improve the AI model. The input is the text feedback, and the output is an improved AI model. Specifically, the server uses the feedback data to train and update the AI model.
[1727] This allows field workers to input questions hands-free and receive quick answers, improving work efficiency and enabling continuous improvement of the system.
[1728] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1729] MODE FOR CARRYING OUT THE INVENTION
[1730] An embodiment of the present invention will be described. The present invention accepts specific questions and requests about architecture from users in natural language, analyzes them, generates and displays appropriate answers, and further accepts feedback from users to improve the system. In addition, it combines this with an emotion engine that recognizes the user's emotions.
[1731] System Configuration
[1732] This system consists of devices operated by users (e.g., PCs, smartphones, tablets) and a server for processing information sent from these devices. The role of each component is explained below.
[1733] 1. User Input
[1734] The user inputs specific questions or requests about construction in natural language using the input means of the terminal, for example, "Please tell me the steps for laying the foundations of a wooden house."
[1735] 2. Sending Input
[1736] The terminal sends questions and requests entered by the user to the server.
[1737] 3. Natural Language Processing (NLP)
[1738] The server analyzes the received user question, specifically:
[1739] Tokenization: Divide the question into words.
[1740] Morphological analysis: Identifying the part of speech of each word.
[1741] Grammatical structure analysis: Analyze the grammatical structure of the sentence and extract important keywords and intent.
[1742] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1743] 4. Emotion analysis
[1744] The server uses an emotion engine to recognize the user's emotion in the question. The following process is performed:
[1745] Emotion recognition: Extracting user emotions (e.g., confusion, excitement, anger, etc.) from input text.
[1746] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[1747] For example, if a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[1748] 5. Information Retrieval and Answer Generation
[1749] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[1750] Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[1751] Generate answers that are interesting and considerate of the user's emotions.
[1752] For example, an answer might be generated along the lines of, "Please understand that foundation work is difficult, and refer to the following steps."
[1753] 6. View Answers
[1754] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1755] 7. Get feedback
[1756] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[1757] 8. Submitting Feedback
[1758] The device sends the feedback from the user to the server.
[1759] 9. Processing Feedback
[1760] The server analyzes this feedback and uses it to improve the AI model. Specifically, it analyzes the feedback, extracts evaluation information, and reflects it in the AI model for re-learning. This allows the system to continuously improve.
[1761] Specific examples
[1762] For example, if a user types "What are the basics of electrical wiring?":
[1763] Terminal: Sends user input to the server.
[1764] Server: Performs natural language processing and analyzes the intent of the question (e.g., "electrical wiring" or "basic knowledge").
[1765] Server: Uses an emotion engine to recognize the user's emotions (e.g., confusion, interest).
[1766] Server: Searches internal databases and external knowledge bases and generates appropriate answers (e.g., "Basic knowledge about electrical wiring is as follows...").
[1767] Terminal: Displays the generated answer to the user.
[1768] User: Provide feedback on the answer.
[1769] Server: Receives feedback and helps improve the AI model.
[1770] In this way, the system of the present invention not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take into consideration the user's feelings. This effectively solves the problems of passing on construction techniques and the shortage of craftsmen, and contributes to the improvement of technology in the construction industry as a whole.
[1771] The processing flow will be explained below.
[1772] Step 1: User enters question
[1773] User: Enters specific questions or requests about construction into the terminal.
[1774] Example: Enter "Please tell me the steps for laying the foundation for a wooden house."
[1775] Step 2: Sending Input
[1776] Terminal: Sends questions and requests entered by the user to the server.
[1777] Specifically, the user input data is sent to the server as an HTTP request.
[1778] Step 3: Tokenize the Question
[1779] Server: Tokenizes the received question.
[1780] Specifically, the question is divided into words and phrases.
[1781] For example, it is divided into "wooden house," "foundation work," and "procedure."
[1782] Step 4: Morphological analysis
[1783] Server: Performs morphological analysis of the question.
[1784] Specifically, the part of speech (noun, verb, adjective, etc.) of each token is identified.
[1785] Step 5: Grammatical structure analysis
[1786] Server: Analyzes the grammatical structure of the question.
[1787] Specifically, it analyzes the structure of the entire sentence (subject, verb, object, etc.) and extracts important keywords and intent.
[1788] For example, "wooden house," "foundation work," and "procedure" are extracted as main keywords.
[1789] Step 6: Sentiment Analysis
[1790] Server: Recognizes the user's emotions contained in the question using an emotion engine.
[1791] Specifically, it extracts the user's emotions (e.g., confusion, excitement, anger, etc.) from the input text and classifies them into predefined emotion categories.
[1792] Example: If a user types, "I'm having trouble with the foundation work on a wooden house," the emotion "confused" is recognized.
[1793] Step 7: Information retrieval and answer generation
[1794] Server: Searches for relevant information and generates answers based on the analysis results and emotion recognition results.
[1795] Specifically, based on the extracted keywords and recognized emotions, the system retrieves relevant data from an internal database and an external knowledge base.
[1796] Generate answers that are interesting and considerate of the user's emotions.
[1797] For example, an answer like "Please understand that foundation work is difficult, and refer to the following steps." is generated.
[1798] Step 8: Submit your response
[1799] Server: Sends the generated answer to the device.
[1800] Step 9: View your answers
[1801] Terminal: Displays the received answer to the user.
[1802] As a specific operation, the answer content is displayed on the screen.
[1803] Step 10: Get feedback
[1804] User: Provide feedback on the displayed answers.
[1805] For example, enter feedback such as "This answer was helpful" or "I'd like more specific explanation."
[1806] Step 11: Submit your feedback
[1807] Terminal: Sends user feedback to the server.
[1808] Step 12: Processing feedback
[1809] Server: Analyzes the feedback received and helps improve the AI model.
[1810] Specifically, the feedback content is analyzed and fed back to the AI model as adaptable data.
[1811] The model is retrained based on this feedback to improve the accuracy of answers from next time onwards.
[1812] This process allows users to quickly and accurately obtain the necessary information about construction, and the system is continuously improved, effectively resolving the issues of passing on construction skills and the shortage of craftsmen.
[1813] Example 2
[1814] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1815] In the modern construction industry, there are limited means of instantly obtaining technical information and procedures related to construction, creating a demand for fast, accurate answers to technical questions and problems. It is also important to increase user satisfaction by providing answers that take the user's emotions into consideration. However, conventional systems struggle to respond in a way that takes user emotions into account, and they lack mechanisms for appropriately utilizing user feedback to improve the system. Therefore, there is a need for a system that can analyze user questions and requests, recognize emotions, generate and display appropriate answers, and further utilize feedback.
[1816] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1817] In this invention, the server includes an input means for accepting specific questions and requests about architecture from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an information search and generation means for searching for related information based on the extracted keywords and the user's emotions and generating answers, a display means for displaying the generated answers to the user, and a feedback processing means for receiving feedback from the user and using it to improve the AI model. This makes it possible to quickly and appropriately answer user questions, provide answers that take user emotions into consideration, and utilize feedback to continuously improve the system.
[1818] "Input means" refers to devices or functions that accept specific questions or requests about architecture from users in natural language.
[1819] "Analysis means" refers to a device or function that analyzes questions or requests entered by users and extracts important keywords and intentions.
[1820] "Information search and generation means" refers to devices or functions that search for related information and generate answers based on the keywords extracted by the analysis means and the user's emotions.
[1821] "Display means" refers to a device or function that visually displays the generated answer to the user.
[1822] A "feedback processing means" is a device or function that accepts feedback provided by users and uses that feedback to help improve the AI model.
[1823] "Tokenization" is the process of dividing an input natural language sentence into words.
[1824] "Morphological analysis" is the process of identifying the part of speech of each tokenized word.
[1825] "Grammar structure analysis" is the process of analyzing the grammatical structure of an input sentence and extracting important keywords and intent.
[1826] "Emotion recognition" is the process of extracting a user's emotions from input text.
[1827] "Emotion classification" refers to the process of classifying extracted emotions into predefined emotion categories.
[1828] An "internal database" is a database that manages information and data stored within the system.
[1829] An "external knowledge base" is an information source or database that is externally accessible, such as the Internet.
[1830] An "AI model" is a model that uses artificial intelligence to learn and perform processes such as answering questions and recognizing emotions.
[1831] "Feedback" refers to the ratings and opinions that users provide in response to the answers displayed.
[1832] The system of the present invention allows users to input specific questions or requests about architecture in natural language, analyzes them, and generates and displays answers. Furthermore, it can collect feedback from users and use it to improve the system. The system consists of the following main components:
[1833] System Configuration
[1834] 1. Hardware configuration:
[1835] Terminal: A device such as a PC, smartphone, or tablet that allows users to input questions or requests and display answers.
[1836] Server: A high-performance cloud server or on-premise server with high data processing capacity and fast response is recommended.
[1837] 2. Software configuration:
[1838] Natural language processing engine: An engine that tokenizes text, analyzes morphology, and analyzes grammatical structures (e.g., Google NLP).
[1839] Emotion engine: An engine for recognizing and classifying user emotions (e.g., IBM Watson Tone Analyzer).
[1840] Database management system: Databases such as MySQL and MongoDB are used as internal databases.
[1841] AI model: An artificial intelligence model that analyzes a user's question or request and generates an answer. Use a generative AI model.
[1842] Operation procedure and processing contents
[1843] 1. User input:
[1844] Users operate their own terminals to input specific questions or requests about construction in natural language, such as "Please tell me the steps for laying the foundations for a wooden house."
[1845] 2. Sending input:
[1846] The device sends the user-entered questions or requests to a server over the Internet. The data is sent as an HTTP request.
[1847] 3. Natural Language Processing (NLP):
[1848] The server analyzes the received user question, specifically by performing the following steps:
[1849] Tokenization: Divide the input question into words.
[1850] Morphological analysis: Identifying the parts of speech of segmented words.
[1851] Grammatical structure analysis: Analyze grammatical structures and extract important keywords and intent.
[1852] 4. Emotion analysis:
[1853] The server uses an emotion engine to recognize the user's emotion in the question, which includes the following steps:
[1854] Emotion Recognition: Extracting user emotions from input text.
[1855] Sentiment Classification: Classifying the extracted emotions into predefined emotion categories.
[1856] 5. Information retrieval and answer generation:
[1857] The server searches for relevant information and generates answers based on the analysis and emotion recognition results. Specifically:
[1858] Information retrieval: Based on the extracted keywords and recognized sentiment, relevant information is retrieved from internal databases and external knowledge bases.
[1859] Answer generation: Based on the searched information, an answer is generated that takes into consideration the user's feelings.
[1860] 6. Show Answer:
[1861] The server then sends the generated answer to the terminal, which visually displays the received answer to the user.
[1862] 7. Getting feedback:
[1863] Users provide feedback on the displayed answers, including whether the answer was helpful or if it needs further clarification.
[1864] 8. Submitting Feedback:
[1865] The device sends the user's feedback to the server, again as an HTTP request.
[1866] 9. Handling feedback (retraining the AI model):
[1867] The server analyzes the received feedback and extracts evaluation information, which is then used to retrain the AI model and improve response accuracy from the next time onwards.
[1868] Specific examples
[1869] For example, if a user types "Please tell me the basics of electrical wiring," the following will be processed:
[1870] User: Enters a question into the terminal and sends it.
[1871] Server: Receives input and performs natural language processing and sentiment analysis.
[1872] Server: Searches for relevant information from internal databases and external knowledge bases and generates appropriate answers.
[1873] Terminal: Display the generated answer.
[1874] User: Provide feedback on the answer.
[1875] Server: Receives feedback and helps retrain the AI model.
[1876] This system not only provides information on construction techniques quickly and appropriately, but also increases user satisfaction by providing responses that take the user's emotions into consideration. This will effectively solve the problems of passing on construction techniques and the shortage of craftsmen, and contribute to improving the technology of the entire construction industry.
[1877] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1878] Step 1:
[1879] The user uses the input means to input specific questions or requests about construction in natural language. For example, they may input a question such as, "Please tell me the steps for laying the foundations for a wooden house." This input is saved as text data on the terminal.
[1880] Step 2:
[1881] The terminal sends the text data entered by the user to the server. Specifically, the data is sent to the server as an HTTP request. At this time, the input data is temporarily stored in the terminal's buffer memory. The input is the user's question text, and the output is the sent HTTP request.
[1882] Step 3:
[1883] The server analyzes the received text data. First, it uses a natural language processing engine to tokenize the question. As a result of tokenization, the input sentence can be divided into words, and the output is a word list. For example, it may be divided into words such as "wooden house," "foundation work," "procedure," "tell me," and "please."
[1884] Step 4:
[1885] The server performs morphological analysis. As a result of the morphological analysis, it is possible to identify the part of speech for each word. For example, "wooden house" is recognized as a noun and "teach me" as a verb. The input is a list of tokenized words, and the output is a list of words with identified parts of speech.
[1886] Step 5:
[1887] The server performs grammatical analysis. This analyzes the grammatical structure of the entire sentence and extracts important keywords and intent. For example, "wooden house," "foundation work," and "procedure" are extracted as key keywords. The results of the grammatical analysis are the extracted keywords and a structural analysis tree.
[1888] Step 6:
[1889] The server uses an emotion recognition engine to recognize the user's emotions. During the emotion recognition process, the server extracts the user's emotional characteristics (e.g., confusion, excitement, anger) from the text data and classifies them into categories. For example, the emotion "confusion" is extracted from the sentence "I'm having trouble because the foundation work on a wooden house is difficult." The input is the analyzed text data, and the output is the recognized emotion and its category.
[1890] Step 7:
[1891] The server performs an information search based on the analysis results and emotion recognition results. It searches its internal database and external knowledge base to obtain relevant information. For example, it searches the internal database for technical information using keywords such as "wooden house," "foundation work," and "procedure." The information obtained at this stage is output in the form of text or images.
[1892] Step 8:
[1893] The server generates an answer based on the search results. Using a generative AI model, it creates an answer that takes into account the acquired information and the user's emotions. For example, an answer such as "Please understand that foundation construction is difficult, and refer to the following steps" may be generated. The input is the search results and emotion data, and the output is the answer text for the user.
[1894] Step 9:
[1895] The server sends the generated answer to the terminal. The terminal displays the received answer to the user. Specifically, the answer is displayed in text format on the terminal screen. The input is the answer data from the server, and the output is the display for the user.
[1896] Step 10:
[1897] The user provides feedback on the displayed answer, including an evaluation of whether the answer was helpful or whether further explanation is needed. The feedback is again input into the terminal as text data.
[1898] Step 11:
[1899] The device sends the user's feedback to the server, again as an HTTP request. The input is the user's feedback text, and the output is the sent HTTP request.
[1900] Step 12:
[1901] The server receives and analyzes the feedback data. As a result of the analysis, specific evaluation information is extracted and used to retrain the AI model. This improves response accuracy from the next time onwards. The input is the feedback data, and the output is an updated AI model.
[1902] (Application example 2)
[1903] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1904] Lack of technical knowledge and information on construction sites, especially complex work procedures and troubleshooting, makes it difficult to quickly implement them. Furthermore, the inability to provide appropriate advice that takes into account the feelings of workers makes it difficult to work efficiently. Furthermore, there is a lack of means to continuously improve the system based on user feedback.
[1905] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1906] In this invention, the server includes an input means for accepting specific questions and requests about construction from users in natural language, an analysis means for analyzing the input questions and requests and extracting important keywords and intentions, an emotion analysis means for analyzing emotions using an emotion engine that recognizes the user's emotions, an information search and generation means for searching for related information and generating answers based on the extracted keywords and recognized emotions, a display means for displaying the generated answers to the user, and a feedback processing means for accepting feedback from users and using it to improve the AI model. This enables workers at construction sites to respond quickly to questions and requests, provides appropriate advice that takes workers' emotions into consideration, and continuously improves the system, enabling efficient and effective work.
[1907] "Input means" refers to a device or system that accepts specific questions or requests about architecture from users in natural language.
[1908] The "analysis means" is a device or program that analyzes the input question or request and extracts important keywords and intentions.
[1909] "Emotion analysis means" refers to a device or system for analyzing emotions from input text using an emotion engine that recognizes the emotions of a user.
[1910] "Information search and generation means" refers to a device or program that searches for related information based on the extracted keywords and recognized emotions and generates answers.
[1911] A "display means" is a device or system that visually displays the generated answers to the user.
[1912] A "feedback processing means" is a device or program that accepts feedback from users and uses it to improve the AI model.
[1913] The present invention is a system in which a factory robot responds to questions and requests from workers at a construction site and provides advice that takes into consideration their emotions. An embodiment of the system is described below.
[1914] System Configuration
[1915] The system includes the following components:
[1916] 1. Input Method
[1917] The factory robot can receive specific questions and requests about construction from workers in natural language. For example, the factory robot can receive a question from a worker such as, "Please tell me the key points of foundation work."
[1918] 2. Analysis method
[1919] The factory robot analyzes the questions and requests it receives. This analysis involves tokenization, morphological analysis, and grammatical structure analysis. This allows it to extract important keywords and intent. For example, keywords such as "foundation work" and "key points" are extracted.
[1920] 3. Emotion analysis means
[1921] The factory robot uses an emotion engine to analyze the emotions of workers from input text. Specifically, the input text is input into the emotion engine, and emotions such as "confusion" or "interest" are recognized.
[1922] 4. Information retrieval and generation methods
[1923] Based on the extracted keywords and the recognized emotions, the factory robot searches for relevant information from its internal database and external knowledge base, and generates an appropriate answer taking into account the search results and the emotions of the worker. For example, it generates an answer in the form of "The main points of foundation work are as follows..."
[1924] 5. Display means
[1925] The generated answers are displayed and presented visually or audibly by the factory robot to the worker, allowing the worker to obtain the appropriate information.
[1926] 6. Feedback Processing Methods
[1927] Workers can provide feedback on the displayed answers, which helps improve the robot's AI model, such as "This answer was helpful" or "Please give me more details."
[1928] Hardware and software used
[1929] Hardware
[1930] Factory robots are equipped with standard input devices (microphones, cameras, etc.) to receive user input, and display devices (displays, speakers, etc.).
[1931] software
[1932] The factory robots are installed with Transformers, a library for natural language processing (NLP). They also use the Sentiment-Analysis model, which is known for its stability and performance, for sentiment analysis. They use the Requests library to communicate with the server.
[1933] Specific examples
[1934] For example, if a factory robot receives a question from a worker saying, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?", it will process it as follows:
[1935] 1. Input Method
[1936] The worker types in a question.
[1937] 2. Analysis method
[1938] The content of the question is analyzed and keywords and intent such as "wooden house," "foundation work," and "having trouble" are extracted.
[1939] 3. Emotion analysis means
[1940] The emotion of "being in trouble" is recognized using an emotion engine.
[1941] 4. Information retrieval and generation methods
[1942] It searches for relevant information based on keywords and sentiment and generates answers such as, "For the foundation construction procedure, please try the following steps."
[1943] 5. Display means
[1944] The generated answer is displayed on the screen and communicated to the worker by voice.
[1945] 6. Feedback Processing Methods
[1946] The worker provides feedback, saying, "This information was useful," and the AI model is improved based on that feedback.
[1947] An example of a prompt sentence is, "I'm having trouble with the foundation work on a wooden house. How should I go about it specifically?"
[1948] This allows factory robots to respond quickly and appropriately to questions from workers and take workers' feelings into consideration, making on-site work more efficient.
[1949] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1950] Step 1:
[1951] The user inputs specific questions and requests about construction to the factory robot in natural language.
[1952] Input: Questions or requests in natural language (e.g., "I'm having trouble with the foundation work on my wooden house. What should I do specifically?")
[1953] Output: The input text is sent to the system
[1954] Step 2:
[1955] The terminal sends the input text to the server.
[1956] Input: Text entered by the user
[1957] Output: The input text is passed to the server
[1958] Step 3:
[1959] The server processes the received text using analytical means, specifically tokenizing, morphological analysis, and grammatical structure analysis to extract important keywords and intent from the question or request.
[1960] Input: Input text (e.g., "Foundation work for wooden houses")
[1961] Data processing / data operations: tokenization, morphological analysis, grammatical analysis
[1962] Output: Keywords and intent (e.g., "foundation work," "difficult," etc.)
[1963] Step 4:
[1964] Recognize user emotions through emotion analysis. Using an emotion engine, emotions (e.g., "confused") are extracted from the input text and classified into emotion categories.
[1965] Input: Parsed text and extracted keywords
[1966] Data processing / data calculation: emotion extraction and classification
[1967] Output: Recognized emotion (e.g., "confused")
[1968] Step 5:
[1969] The information search and generation means searches for related information from the internal database and external knowledge base based on the extracted keywords and emotions, and generates answers based on the search results.
[1970] Input: Keywords, Recognized Sentiments
[1971] Data processing / data calculation: database search, answer generation
[1972] Output: Generated answer (e.g., "Please proceed as follows for the foundation construction procedure.")
[1973] Step 6:
[1974] The server transmits the generated answer to the terminal, which then presents the answer to the user visually or audibly through a display means.
[1975] Input: Generated Answer
[1976] Output: The answer that is presented to the user
[1977] Step 7:
[1978] The user provides feedback on the displayed answer, including whether it was helpful or if more details are needed.
[1979] Input: User feedback (e.g., satisfied, dissatisfied, further questions, etc.)
[1980] Output: Feedback is sent to the system
[1981] Step 8:
[1982] The device sends feedback to the server, which then receives and analyzes the feedback to help improve the AI model. Specifically, the feedback content is analyzed, evaluation information is extracted, and the feedback is reflected in the AI model for re-learning.
[1983] Input: User feedback
[1984] Data processing / data calculation: Feedback analysis, evaluation information extraction, AI model retraining
[1985] Output: An improved AI model
[1986] As described above, we will explain how the system operates at each processing step and what inputs and outputs there are.
[1987] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1988] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1989] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1990] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1991] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1992] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1993] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1994] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1995] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1996] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1997] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1998] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1999] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2000] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2001] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2002] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2003] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2004] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2005] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2006] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2007] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2008] The following is further disclosed regarding the above embodiment.
[2009] (Claim 1)
[2010] An input method for accepting specific questions and requests about architecture from users in natural language;
[2011] An analytical method for analyzing input questions and requests and extracting important keywords and intent;
[2012] An information search and generation means for searching related information based on the extracted keywords and generating answers;
[2013] a display means for displaying the generated answer to the user;
[2014] A feedback processing method to receive feedback from users and help improve the AI model;
[2015] A system including:
[2016] (Claim 2)
[2017] 2. The system according to claim 1, wherein the analyzing means performs tokenization, morphological analysis, and grammatical structure analysis of the question sentence.
[2018] (Claim 3)
[2019] 2. The system according to claim 1, wherein the information search and generation means searches for related information from an internal database and an external knowledge base based on the extracted keywords.
[2020] "Example 1"
[2021] (Claim 1)
[2022] An input method that accepts specific questions and requests about architecture from users in natural language, and
[2023] means for transmitting the input question or request to a server;
[2024] An analysis method for extracting important keywords and intent from received questions by tokenizing, morphologically analyzing, and analyzing grammatical structures;
[2025] An information search and generation means for searching for related information from an internal database and an external knowledge base based on the extracted keywords and generating answers;
[2026] a means for transmitting and displaying the generated answers to the user;
[2027] A feedback processing means for receiving and analyzing user feedback to help improve the AI model;
[2028] A system including:
[2029] (Claim 2)
[2030] 2. The system according to claim 1, wherein the analyzing means performs tokenization, morphological analysis, and grammatical structure analysis of the question sentence.
[2031] (Claim 3)
[2032] 2. The system according to claim 1, wherein the information search and generation means searches for related information from an internal database and an external knowledge base based on the extracted keywords.
[2033] "Application Example 1"
[2034] (Claim 1)
[2035] An input method for accepting specific questions and requests about architecture from users in natural language;
[2036] An analytical method for analyzing input questions and requests and extracting important keywords and intent;
[2037] An information search and generation means for searching related information based on the extracted keywords and generating answers;
[2038] a display means for displaying the generated answer to the user;
[2039] A feedback processing method to receive feedback from users and help improve the AI model;
[2040] Furthermore, the terminal used by the user includes a smart gadget equipped with a voice input device and a display device, the input means having a function of receiving voice input and the display means having a function of displaying an answer on the display of the smart gadget;
[2041] A system including:
[2042] (Claim 2)
[2043] 2. The system according to claim 1, wherein the analyzing means performs tokenization, morphological analysis, and grammatical structure analysis of the question sentence.
[2044] (Claim 3)
[2045] 2. The system according to claim 1, wherein the information search and generation means searches for related information from an internal database and an external knowledge base based on the extracted keywords.
[2046] "Example 2: Combining Emotion Engines"
[2047] (Claim 1)
[2048] An input method for accepting specific questions and requests about architecture from users in natural language;
[2049] An analytical method for analyzing input questions and requests and extracting important keywords and intent;
[2050] An information search and generation means for searching for related information based on the extracted keywords and user sentiment and generating answers;
[2051] a display means for displaying the generated answer to the user;
[2052] A feedback processing method to receive feedback from users and help improve the AI model;
[2053] A system including:
[2054] (Claim 2)
[2055] 2. The system according to claim 1, wherein the analyzing means performs tokenization of the question sentence, morphological analysis, grammatical structure analysis, and emotion recognition.
[2056] (Claim 3)
[2057] The system according to claim 1, wherein the information retrieval and generation means retrieves related information from an internal database and an external knowledge base based on the extracted keywords and recognized emotions.
[2058] "Application example 2 when combining emotion engines"
[2059] (Claim 1)
[2060] An input method for accepting specific questions and requests about architecture from users in natural language;
[2061] An analytical method for analyzing input questions and requests and extracting important keywords and intent;
[2062] emotion analysis means for analyzing emotions using an emotion engine that recognizes ...
Claims
1. An input method for accepting specific questions and requests about architecture from users in natural language; An analytical method for analyzing input questions and requests and extracting important keywords and intent; an information search and generation means for searching for related information based on the extracted keywords and generating answers; a display means for displaying the generated answer to the user; A feedback processing method to receive feedback from users and help improve the AI model; A system including:
2. 2. The system according to claim 1, wherein the analyzing means performs tokenization, morphological analysis, and grammatical structure analysis of the question sentence.
3. 2. The system according to claim 1, wherein said information search and generation means searches for related information from an internal database and an external knowledge base based on said extracted keywords.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A