system
A server-based system addresses the challenge of accessing manuals by verifying file formats, structuring text data, and building a generative chatbot to provide quick and personalized information access.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing systems face challenges with excessive information and difficulty in accessing manuals, leading to inefficiencies and reliance on experienced employees, with a need for a system that allows quick access to necessary information.
A server-based system that receives and verifies file formats, extracts and hierarchically structures text data, constructs a generative model, and builds a chatbot to provide answers to user questions, reducing the need to read manuals and improving efficiency.
Enables rapid and accurate access to information, eliminating the need to read manuals and enhancing user experience through personalized and emotionally sensitive responses.
Smart Images

Figure 2026074873000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, in many companies, the excess of information and the difficulty of access have been problems particularly related to a large number of manuals. Especially when introducing new equipment or systems, it is necessary to carefully read detailed manuals, which has become a factor hindering the efficiency of business. In addition, dependence on specific employees with certain experiences may occur, and there may be a bias in knowledge and insufficient sharing. There is a demand for a system that solves such problems and enables anyone to quickly access necessary information.
Means for Solving the Problems
[0005] This invention first provides a means for a server to receive a file uploaded by a user and verify its format. Next, it employs a means for the server to extract text data from the received file and structure it hierarchically. Then, it constructs a generative model by supplying the hierarchically structured data to a generative model and training it. Furthermore, it provides a means for constructing a chatbot that generates answers to questions using the generative model. Finally, it provides a system that includes a means for the server to present appropriate answers to questions from the user. This system makes it possible to reduce the effort required to read manuals and improve the efficiency of work.
[0006] A "file" is a collection of digital data that a user uploads and a server receives and processes.
[0007] "Format verification" is the process of verifying whether the format of the received file is supported by the system.
[0008] "Text data" refers to character information extracted from a file, and is the subject of analysis and structuring.
[0009] "Hierarchical structuring" is the process of organizing extracted text data by item or heading, making it easy to access and search.
[0010] A "generative model" is an AI algorithm that learns from supplied data and generates appropriate answers to user inquiries.
[0011] "Training" is the process by which a generative model learns from data and improves its ability to generate answers.
[0012] A "chatbot" is a system that uses generative models to automatically generate answers to questions through dialogue with users.
[0013] A "user" is an entity that uses a system to search for information and ask questions.
[0014] A "server" is a computer system that receives, analyzes, supplies data to, and generates responses from uploaded files. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] The following systems are conceivable as embodiments for carrying out the present invention.
[0037] First, the user uploads a file from their device to the system. This file contains digital data of the manual, and the server checks the format of the received file before processing it. If it is in a supported format, it proceeds to the next step.
[0038] Next, the server extracts text data from the received file. Once the text data is extracted, it is structured hierarchically. At this stage, the data is organized by category and headings so that users can easily search for it.
[0039] The server then feeds this structured data to a generative model. This generative model uses AI to learn from the supplied data. This prepares the model to generate appropriate answers to queries.
[0040] Next, the server uses the trained generative model to build an FAQ chatbot. This chatbot is programmed to receive questions from users and generate answers in real time.
[0041] Through a conversational interface accessible from the terminal, users input questions in natural language. The server receives these questions, uses a generative model to generate appropriate answers, and sends them to the user's terminal, enabling rapid information delivery.
[0042] As a concrete example, consider a question about how to operate a new device. When a user asks the chatbot on their device, "How do I set up this device?", the server uses a generative model to find the relevant data and provides a specific answer such as, "Please follow these steps for the initial setup."
[0043] This system allows users to quickly access important information, eliminating the need to read the entire manual. This is an embodiment of the present invention.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user selects a manual file from their terminal and uploads it. The uploaded file is sent to the server.
[0047] Step 2:
[0048] The server checks the received file and verifies that it is in a supported format. If the format is not appropriate, it returns an error message to the user.
[0049] Step 3:
[0050] The server uses OCR (Optical Character Recognition) or other parsing techniques to extract text data from the appropriate files.
[0051] Step 4:
[0052] The server organizes and hierarchically structures the extracted text data. This process classifies and structures the data by headings and paragraphs.
[0053] Step 5:
[0054] The server supplies the generated hierarchical structured data to the generative model and performs training. This training builds the model's ability to answer future questions.
[0055] Step 6:
[0056] The server uses a trained generative model to build an FAQ chatbot, enabling automated responses to user questions.
[0057] Step 7:
[0058] Users access the chatbot through their device and input the information they want to obtain in natural language.
[0059] Step 8:
[0060] The server sends the question received from the user to a generative model, which generates the optimal answer. The generated answer is then returned to the user's device and displayed.
[0061] Through this series of steps, users can easily and quickly access the information in the manual.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] This invention aims to improve a system that quickly and accurately retrieves useful information from files containing large amounts of data and provides appropriate answers to user inquiries. Furthermore, it seeks to solve the problem of improving the accuracy of answers by utilizing user feedback and enhancing the comprehensiveness of information provision by integrating data from diverse information sources.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for receiving information and confirming its structure, means for extracting and hierarchizing data from the received information, and means for providing the hierarchical data to a learning model for training. This makes it possible to provide quick and accurate answers to user inquiries. Furthermore, by utilizing feedback, it is possible to improve the accuracy of the system and provide comprehensive information through information integration.
[0067] "Means of receiving information and verifying its structure" refers to the process of taking in information provided from external sources and checking whether that information conforms to a specific format or standard.
[0068] "Methods for extracting and hierarchizing data" refers to the process of extracting necessary items from received information and organizing that data based on a logical order or classification.
[0069] "Means of providing and training a learning model" refers to the process of inputting prepared data into an artificial intelligence model, allowing that model to learn patterns and knowledge from the data.
[0070] "Methods for constructing dialogue programs" refers to the process of developing a system that uses a trained artificial intelligence model to automatically communicate with users and generate appropriate answers to questions.
[0071] "A means of providing appropriate answers to user inquiries" refers to a function that analyzes questions received from users, searches for relevant information, and provides responses.
[0072] "Methods for collecting evaluations, using that information to improve the learning model, and enhancing the accuracy of responses" refers to methods of collecting feedback from users, retraining the model based on that feedback, and improving the system's performance.
[0073] "Means of connecting diverse information sources to provide comprehensive information" refers to methods of aggregating data from multiple different databases and information sets to provide users with consistent and detailed information.
[0074] To implement this invention, the system is primarily operated through the collaboration of three parties: a server, a terminal, and a user. The specific method is described below.
[0075] First, users upload digital information to the system using their own devices. These devices are expected to be computers or smartphones. Examples of uploaded information include product manuals and technical documents.
[0076] The server automatically analyzes the format and structure of the received digital information. This analysis can utilize software that recognizes specific file formats, such as PDF analysis tools or OCR (optical character recognition) technology. Based on the analysis results, the server extracts the information as text and organizes the data structurally. This organization process involves completing a hierarchical format using a database management system or similar method.
[0077] The organized data is fed into a generative AI model, and the learning process begins. An example of a generative AI model is a natural language processing model using open-source libraries. This model learns from the supplied data and prepares answers to subsequent user inquiries.
[0078] The completed model is implemented as an interactive program by the server. Users can use the terminal's interactive interface to input questions in natural language. For example, they can ask, "Please tell me the initial setup procedure for this product." This question is processed by the server, and the optimal answer is generated based on the generative model. The answer is then immediately sent to the user's terminal.
[0079] This system allows users to obtain information instantly and accurately. Furthermore, user questions and feedback are recorded on the server, and the model is regularly updated based on this information to improve its accuracy.
[0080] An example of a prompt would be, "Please tell me how to troubleshoot this." This would prepare the model to provide relevant solutions.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] Users upload digital files from their devices. These files include, for example, product manuals and technical documents. They click the "Select File" button on their device, choose the desired file, and submit it. The input is a digital file, and the output is the transfer of the file to the server.
[0084] Step 2:
[0085] The server verifies the format of the received file. Specifically, the server checks the file extension and internal metadata to determine if it is a supported format (e.g., PDF, DOCX). The input for this step is the uploaded file, and the output is the result of the verification that the file is in the correct format.
[0086] Step 3:
[0087] The server extracts text data from the file. If the format is supported, it extracts the text using appropriate software (e.g., PDF analysis tools or OCR technology). The input for this step is a verified digital file, and the output is the extracted text data.
[0088] Step 4:
[0089] The server hierarchically structures the extracted text data. Using a database management system, it structures the data by section and topic. This process makes it easy for users to search for information. The input for this step is text data, and the output is hierarchically structured data.
[0090] Step 5:
[0091] The server supplies hierarchical data to a generative AI model for training. The model analyzes the supplied data and is trained to prepare appropriate answers to queries. The input for this step is hierarchical data, and the output is the trained generative model.
[0092] Step 6:
[0093] The server uses a trained generative model to build a conversational program (chatbot). This chatbot is designed to receive questions from users and generate answers in real time using the model. The input to this step is the trained generative model, and the output is a working chatbot.
[0094] Step 7:
[0095] The user enters a question through the terminal's interactive interface. For example, they might ask, "Please tell me how to set up this product." In this step, the input is the user's question, and the output is the sending of the question to the server.
[0096] Step 8:
[0097] The server analyzes the received question using a generative model and generates an appropriate answer. Based on the analysis results, it searches for relevant information, generates an answer, and sends it to the user's terminal. The input for this step is the user's question, and the output is the generated answer.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] Currently, many physical stores are required to provide customers with prompt information regarding product details and service information in response to their inquiries. However, there is a lack of efficient methods to achieve this, which can lead to decreased customer satisfaction. The objective of this invention is to solve this problem by providing a system that provides customers with the information they need quickly and accurately in physical stores.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for constructing an interactive program that generates answers to questions using a generative model, means for users to access product information using identification information and for presenting appropriate answers to questions, and means for acquiring identification information through the user terminal and redirecting to the relevant information based on that information. This enables rapid information provision to customers in physical stores.
[0103] "Means for receiving files and verifying their format" refers to a function that receives digital data sent from a user to the system and determines whether the data format is supported by the system.
[0104] "A means of extracting text data from received files and structuring it hierarchically" refers to a function that extracts textual information from transmitted digital data and organizes and structures that information by item or heading.
[0105] "A means of building an interactive program that generates answers to questions using generative models" refers to a function that utilizes machine learning techniques to create a program that generates appropriate responses to user questions.
[0106] "A means for users to access product information using identification information and to be presented with appropriate answers to questions" refers to a function that allows users to access a digital system using a specific ID or code and automatically answers questions based on detailed information of the selected product.
[0107] "A means of acquiring identification information through the user's terminal and redirecting them to relevant information based on that information" refers to a function that collects identification data supplied via the user's device and uses that data to guide the user to related content or information.
[0108] To implement this invention, the following system is necessary. Specifically, a system is constructed that utilizes the user's terminal and a server installed in the store to provide product information and generate quick responses to user inquiries.
[0109] The server receives manual-format files uploaded by users. Next, it verifies the file format and extracts and hierarchically structures the text data. Natural language processing libraries (e.g., TENSORFLOW® or PyTorch) are used for this text data processing. The server then feeds the organized data to a machine learning model, training a generative AI model to respond to user questions. This model forms the basis for real-time generated interactive programs.
[0110] Users can scan QR codes (registered trademarks) attached to products in physical stores using their smartphones or other devices. This sends identification information about the relevant product or service to a server. The server uses a generative model based on this information to instantly provide appropriate answers to user questions. This entire process enhances the user's in-store shopping experience and allows them to receive quick and detailed information.
[0111] As a concrete example, consider a case where a user scans a QR code for a specific product—for example, a camera—and asks about its initial setup. In this case, the server inputs the prompt "How do I set up the camera?" into the AI model and provides detailed setup instructions to the user's device.
[0112] This system allows customers to instantly obtain the information they need without any hassle, effectively complementing in-store customer service.
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] The user scans the QR code included with the product using their device.
[0116] Input: QR code data
[0117] Output: Identification information
[0118] Specific operation: The device uses its camera function to read the QR code, analyzes its contents, and obtains specific identification information.
[0119] Step 2:
[0120] The terminal sends identification information to the server.
[0121] Input: Identification Information
[0122] Output: Formal request to the server
[0123] Specific operation: The terminal sends identification information to the server via the network and requests information related to the product.
[0124] Step 3:
[0125] Based on the identification information received by the server, it searches for and retrieves the relevant manual files.
[0126] Input: Identification Information
[0127] Output: Manual file
[0128] Specific operation: The server refers to an internal database and searches for and retrieves the digital manual corresponding to the specified identification information.
[0129] Step 4:
[0130] The server extracts text data from the acquired manuals and structures it hierarchically.
[0131] Input: Manual file
[0132] Output: Structured text data
[0133] Specific operation: The server extracts necessary text from the scanned manual using natural language processing and organizes it by headings and items.
[0134] Step 5:
[0135] The server supplies structured data to the generative model and prepares it for response generation.
[0136] Input: Structured text data
[0137] Output: Feed to the generative model
[0138] Specific operation: The server inputs data into the AI model and performs the necessary training to generate appropriate answers to user questions.
[0139] Step 6:
[0140] The user sends the question to the server via their device.
[0141] Input: User Question
[0142] Output: Query to the server
[0143] Specific operation: The user uses the app's interface to input questions about the product in natural language and sends them to the server.
[0144] Step 7:
[0145] The server uses a generative model to generate an answer to the question and sends it to the terminal.
[0146] Input: User Question Query
[0147] Output: Answer to the question
[0148] Specific operation: The server analyzes the received question using a generative AI model, generates an answer based on the trained data, and sends it to the user's terminal.
[0149] Step 8:
[0150] The terminal displays the response received from the server to the user.
[0151] Input: Response from the server
[0152] Output: Answer displayed on the user interface
[0153] Specific operation: The terminal displays the response provided by the server on the screen, providing the user with detailed information.
[0154] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0155] In embodiments of the present invention, the process begins with a user uploading a file from their terminal. This file contains digital data of a manual, and the server receives the file and verifies its format. If it is determined to be in a supported format, the server extracts text data from the file and structures it hierarchically. This process organizes the data by category, making it easy to search.
[0156] The hierarchically structured data is fed to a generative model by the server, where it is trained. This training enables the model to generate appropriate answers to user inquiries. Next, the server utilizes the trained model to build a chatbot. This chatbot has the functionality to respond to user questions and provides information through dialogue.
[0157] Furthermore, this invention incorporates an emotion engine. This engine has the function of recognizing emotions from the user's input and conversation tone. The chatbot determines the user's emotions using the emotion engine and adjusts the tone and content of its response based on that information.
[0158] As a concrete example, consider a scenario where a user asks for instructions on how to use a product for the first time and feels anxious. They might consult a chatbot via their device, saying, "This is my first time using this device, so I'm worried about whether I've set it up correctly." In this case, the server utilizes an emotion engine to recognize the user's anxiety and provides a reassuring response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence." This approach improves the user experience and enhances the quality of information provided.
[0159] The above describes an embodiment of the present invention that combines an emotion engine, enabling the construction of a system that provides users with information quickly and with consideration for their emotions.
[0160] The following describes the processing flow.
[0161] Step 1:
[0162] The user selects a manual file from their terminal and uploads it to the system. This file is a digitized version of the manual.
[0163] Step 2:
[0164] The server receives the uploaded file and checks its format. If the file is in a supported format, it proceeds to the next step.
[0165] Step 3:
[0166] The server extracts text data from the received file. During this process, OCR and file parsing technologies are used to accurately extract the necessary text information.
[0167] Step 4:
[0168] The server organizes the extracted text data into a hierarchical structure. It classifies the data by item and heading, making it easy for users to quickly access the information they need.
[0169] Step 5:
[0170] The server supplies hierarchically structured data to a generative model, which is then trained. This training enables the model to generate appropriate answers to questions.
[0171] Step 6:
[0172] The server builds an FAQ chatbot based on a pre-trained generative model. This chatbot provides answers to questions in real time.
[0173] Step 7:
[0174] Users access the chatbot through their device and input questions in natural language. They can ask questions in a conversational format.
[0175] Step 8:
[0176] The server analyzes user input and uses an emotion engine to determine the user's emotions. By extracting emotion data from the input, it understands the user's feelings.
[0177] Step 9:
[0178] The server uses a generative model to generate responses with a tone adjusted according to the user's emotions. It then sends the generated responses to the user's device.
[0179] Step 10:
[0180] Users can review the answers on their device and ask additional questions to the chatbot as needed. Through this process, users can obtain information smoothly. This system enables the delivery of more personalized and emotionally sensitive information to users.
[0181] (Example 2)
[0182] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0183] In modern information technology, there is a need for systems that can respond quickly and accurately to user inquiries. However, conventional systems struggle to generate responses that take user emotions into account, and there is a demand for technologies that can achieve more human-like interaction. Furthermore, there is a lack of mechanisms to improve response accuracy by utilizing feedback, and data integration to provide comprehensive information is insufficient.
[0184] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0185] In this invention, the server includes means for receiving a file and verifying its format, means for extracting information from the received file and structuring it hierarchically, means for supplying the hierarchically structured information to a generative model for training, and means for providing an appropriate response to a user inquiry. This enables the generation of responses that take into account the user's emotional perception, improvement of response accuracy through feedback, and comprehensive information provision through the integration of multiple information sources.
[0186] "Receiving a file" refers to the act of a server retrieving digital data provided by a user from their device.
[0187] "Verifying the format" is the process of verifying whether the received digital data is in a specific, compatible data format.
[0188] "Information extraction" is the process of extracting necessary text data and related information from received digital data.
[0189] "Hierarchical structuring" refers to classifying extracted information into chapters or items and organizing them hierarchically.
[0190] A "generative model" is an algorithm that is trained using large amounts of digital data and generates responses to queries through natural language processing.
[0191] "Training" refers to the process of improving the accuracy of a generative model's responses by using input data to learn from it.
[0192] An "automated response system" is software or a system that provides a mechanism for the system to automatically generate a response to a user's inquiry.
[0193] "Identifying emotions" is the process of analyzing user input and context to detect the emotions contained within it.
[0194] "Adjusting responses" refers to changing the content and tone of responses provided to the user based on identified emotional information.
[0195] "Integrating information sources" is the process of aggregating data from multiple related databases and information streams and structuring it into a single, comprehensive piece of information.
[0196] In implementing this invention, the process begins with a user uploading a file containing digital data using their own device. This file contains information such as product manuals and instructions. Once the user uploads the file, the server receives the data. The server first checks the format of the received file, verifying, for example, whether it is in PDF or DOCX format. This verification process can be automated using open-source libraries or proprietary programs.
[0197] If the server determines that the file is in a supported format, it begins data extraction. For example, it uses tools such as Apache® Tika or Python's PDFMiner to extract the necessary text data from the file. Since the extracted data is difficult to handle as is, the server performs a step to structure the data hierarchically based on chapters and headings. This organizes the information by category, making subsequent searches easier.
[0198] The hierarchically structured information is then fed to the generative AI model. The server trains the model using open-source machine learning libraries such as TensorFlow and PyTorch. This training enables the model to generate fast and appropriate responses to real-time user inquiries.
[0199] Using a pre-trained generative AI model, the server builds a chatbot that automatically generates responses to user questions. Furthermore, the invention integrates an emotion engine, allowing the server to identify emotions from user input. This enables the adjustment of response tone based on emotional information obtained through data analysis.
[0200] As a concrete example, consider a scenario where a user is unsure about the initial setup of a product. When the user enters a prompt message through their device such as, "This is my first time using this device, so I'm worried about whether I've set it up correctly," the server uses an emotion engine to detect the user's anxiety. As a result, the chatbot generates a response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence," and provides it to the user. This system allows users to receive thoughtful and emotionally sensitive information.
[0201] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0202] Step 1:
[0203] The user uploads a digital data file to the server using a terminal. The input is a digital data file, and this process saves the file to the server. The server receives the uploaded file and temporarily stores it for processing in the next step.
[0204] Step 2:
[0205] The server checks the format of the uploaded file. Specifically, it verifies whether the file is in a supported format such as PDF or DOCX. The input is the uploaded file, and the output is the result of the file format verification. If the format is not appropriate, the server sends an error message to the user and prompts them to re-upload the file.
[0206] Step 3:
[0207] The server extracts information from files in the correct format. This extraction process uses tools such as Apache Tika or PDFMiner. The input is a formatted file, and the output is extracted text data. The server stores the extracted text data in memory and proceeds to the next structuring step.
[0208] Step 4:
[0209] The server hierarchically structures the extracted text data. The data is primarily organized based on document chapters and headings. The input is extracted text data, and the output is hierarchically structured information. This hierarchical structure organizes the information, making it easier to search and use later.
[0210] Step 5:
[0211] The server supplies hierarchically structured information to a generative AI model for training. This process involves training the generative model with new data patterns to improve the accuracy of response generation. The input is hierarchically structured information, and the output is the trained model. Using this model enables accurate and appropriate responses.
[0212] Step 6:
[0213] The server uses a trained generative AI model to build a chatbot. This automated response system has the ability to instantly respond to user prompts. The input is the user's question, and the output is an automatically generated response based on the model.
[0214] Step 7:
[0215] The server integrates an emotion engine, adding the ability to identify the user's input emotions. This system analyzes the emotions the user expresses and adjusts the response accordingly. The input is the user's response, including the prompt, and the output is the response adjusted based on the emotion identification.
[0216] Through these steps, the server enables the provision of advanced information to improve the user experience.
[0217] (Application Example 2)
[0218] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0219] In systems that provide advanced information, users often find it difficult to easily understand operation manuals and setup procedures, which can be particularly psychologically stressful and anxiety-inducing. This invention aims to reduce user anxiety during the information provision process and enable smooth information access by adjusting responses to take into account the user's emotional state.
[0220] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0221] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for supplying the hierarchically structured data to a generative model and training it, means for constructing a dialogue system that generates answers to questions using the generative model, and means for analyzing the emotions of the user in response to a question, adjusting the answer, and presenting an appropriate answer. This enables effective information provision that takes into account the user's psychological state.
[0222] A "file" is a collection of data used to store digital information and process it within a computer system.
[0223] "Means of checking format" refers to a function that determines whether the received data conforms to the specifications that the system can process.
[0224] "Text data" refers to digital information that includes string information, and is data recorded in the form of sentences or words.
[0225] "Means of hierarchical structuring" refers to the function of organizing information structurally and clarifying the relationships between each element.
[0226] A "generative model" is an algorithm that uses machine learning techniques to generate new information or results based on input data.
[0227] "Training methods" refer to the process of using data to train a generative model and improve its accuracy and capabilities.
[0228] A "dialogue system" is a computer-based conversation management system that provides information and generates responses based on input from users.
[0229] A "user" is an individual or organization that interacts with the system and receives or provides information.
[0230] "Means of analyzing emotions" refers to a function that identifies and evaluates a user's psychological state based on input data.
[0231] "Methods for adjusting responses" refer to the process of optimizing the content and expression of generated responses based on analyzed emotional information.
[0232] Modes for carrying out the invention
[0233] To implement this invention, a server plays a central role. The server receives files uploaded by users using their terminals. After receiving the files, the server uses means to verify that their format conforms to predetermined specifications. Subsequently, text data is extracted from the received files, and the information is organized through hierarchical structuring. This hierarchical structuring clarifies the relationships between data, facilitating subsequent processing.
[0234] Next, the server supplies the hierarchically structured data to the generative model for training. The machine learning framework "TensorFlow" is used to train the generative model. As a result, the trained generative model has the ability to function as a dialogue system. This dialogue system provides information and generates responses based on user input.
[0235] Furthermore, the server uses an emotion analysis engine to analyze the input data received from the user and evaluate the user's emotional state. The "Hugging Face Transformers" library is utilized for emotion analysis. Based on this emotional information, the server optimizes the generated response by adjusting its content and tone.
[0236] As a concrete application example, consider a scenario where a first-time user of an autonomous vehicle asks, "I'm worried about operating the autonomous driving mode." In this case, the server uses sentiment analysis to identify the user's anxiety and provides a reassuring response such as, "The autonomous driving mode has been confirmed to be safe. Please follow the steps below to use it with confidence." Through this process, the user can access accurate information to safely operate the vehicle system.
[0237] An example of a prompt message is, "Generate a method for providing vehicle information to the server that takes user sentiment into consideration."
[0238] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0239] Step 1:
[0240] The server receives files sent from the terminal. It parses the received file to determine if it is in a supported format. The input is a digital file uploaded by the user via the terminal, and the output is a boolean value indicating whether the format is valid.
[0241] Step 2:
[0242] The server extracts text data from verified digital files. The extracted text data is then organized into higher and lower levels to create a hierarchical structure. The input is raw data extracted according to a verified file format, and the output is hierarchically structured data.
[0243] Step 3:
[0244] The server supplies hierarchically structured data to a generative model using the machine learning framework "TensorFlow" and trains the model. The input is hierarchically structured data, and the output is a trained generative model.
[0245] Step 4:
[0246] When a user enters a question using a terminal, the server receives the question. Using a trained generative model, it generates an appropriate answer to the question. The input is a question from the user in natural language, and the output is the generated answer.
[0247] Step 5:
[0248] The server uses the emotion analysis engine "Hugging Face Transformers" to evaluate the user's emotional state from the context of the question. The input is the user's question text, and the output is an identified emotional state (e.g., anxiety, joy).
[0249] Step 6:
[0250] The server adjusts the content and expression of the generated response based on the emotion assessment results and provides it to the user in the most appropriate form. The input is the emotional state analyzed and the generated response, and the output is the adjusted response.
[0251] Step 7:
[0252] The server responds to the user's question by sending a pre-configured answer to the terminal and presenting it to the user. The input is the optimized answer, and the output is the final information provided to the user.
[0253] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0254] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0255] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0256] [Second Embodiment]
[0257] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0258] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0259] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0260] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0261] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0262] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0263] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0264] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0265] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0266] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0267] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0268] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0269] The following systems are conceivable as embodiments for carrying out the present invention.
[0270] First, the user uploads a file from their device to the system. This file contains digital data of the manual, and the server checks the format of the received file before processing it. If it is in a supported format, it proceeds to the next step.
[0271] Next, the server extracts text data from the received file. Once the text data is extracted, it is structured hierarchically. At this stage, the data is organized by category and headings so that users can easily search for it.
[0272] The server then feeds this structured data to a generative model. This generative model uses AI to learn from the supplied data. This prepares the model to generate appropriate answers to queries.
[0273] Next, the server uses the trained generative model to build an FAQ chatbot. This chatbot is programmed to receive questions from users and generate answers in real time.
[0274] Through a conversational interface accessible from the terminal, users input questions in natural language. The server receives these questions, uses a generative model to generate appropriate answers, and sends them to the user's terminal, enabling rapid information delivery.
[0275] As a concrete example, consider a question about how to operate a new device. When a user asks the chatbot on their device, "How do I set up this device?", the server uses a generative model to find the relevant data and provides a specific answer such as, "Please follow these steps for the initial setup."
[0276] This system allows users to quickly access important information, eliminating the need to read the entire manual. This is an embodiment of the present invention.
[0277] The following describes the processing flow.
[0278] Step 1:
[0279] The user selects a manual file from their terminal and uploads it. The uploaded file is sent to the server.
[0280] Step 2:
[0281] The server checks the received file and verifies whether it is in a supported format. If the format is inappropriate, an error message is returned to the user.
[0282] Step 3:
[0283] The server uses OCR (Optical Character Recognition) and other parsing techniques to extract text data from the appropriate file.
[0284] Step 4:
[0285] The server organizes and hierarchically structures the extracted text data. Through this operation, the data is classified and structured by headings and paragraphs.
[0286] Step 5:
[0287] The server supplies the generated hierarchically structured data to the generation model for training. Through this training, the model constructs the ability to answer future questions.
[0288] Step 6:
[0289] The server constructs a FAQ chatbot using the trained generation model, enabling automatic responses to questions from users.
[0290] Step 7:
[0291] The user accesses the chatbot through the terminal and enters the content for which information is desired in natural language.
[0292] Step 8:
[0293] The server sends the question received from the user to a generative model, which generates the optimal answer. The generated answer is then returned to the user's device and displayed.
[0294] Through this series of steps, users can easily and quickly access the information in the manual.
[0295] (Example 1)
[0296] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0297] This invention aims to improve a system that quickly and accurately retrieves useful information from files containing large amounts of data and provides appropriate answers to user inquiries. Furthermore, it seeks to solve the problem of improving the accuracy of answers by utilizing user feedback and enhancing the comprehensiveness of information provision by integrating data from diverse information sources.
[0298] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0299] In this invention, the server includes means for receiving information and confirming its structure, means for extracting and hierarchizing data from the received information, and means for providing the hierarchical data to a learning model for training. This makes it possible to provide quick and accurate answers to user inquiries. Furthermore, by utilizing feedback, it is possible to improve the accuracy of the system and provide comprehensive information through information integration.
[0300] "Means of receiving information and verifying its structure" refers to the process of taking in information provided from external sources and checking whether that information conforms to a specific format or standard.
[0301] The "means for extracting and hierarchizing data" is a process of extracting required items from the received information and organizing the data based on logical order or classification.
[0302] The "means for providing to a learning model and performing training" is a process of inputting the prepared data into an artificial intelligence model so that the model can learn patterns and knowledge from the data.
[0303] The "means for constructing an interactive program" is a process of developing a system that automatically communicates with users using the learned artificial intelligence model and generates appropriate answers to questions.
[0304] The "means for presenting an appropriate answer to an inquiry from a user" is a function of analyzing the question received from the user, searching for relevant information, and providing a response.
[0305] The "means for collecting evaluations, using the information to improve the learning model, and improving the accuracy of answers" is a method of collecting feedback obtained from users, retraining the model based on it, and improving the performance of the system.
[0306] The "means for connecting to various information sources to achieve comprehensive information provision" is a method of aggregating data from multiple different databases or information sets and providing consistent and detailed information to users.
[0307] To implement this invention, mainly three parties, namely the server, the terminal, and the user, cooperate to operate the system. The specific method will be described below.
[0308] First, the user uploads digital-formatted information into the system using their own device. The terminals assumed to be used are computers and smartphones. Examples of the information to be uploaded include product manuals and technical documents.
[0309] The server automatically analyzes the format and structure of the received digital information. This analysis can utilize software that recognizes specific file formats, such as PDF analysis tools or OCR (optical character recognition) technology. Based on the analysis results, the server extracts the information as text and organizes the data structurally. This organization process involves completing a hierarchical format using a database management system or similar method.
[0310] The organized data is fed into a generative AI model, and the learning process begins. An example of a generative AI model is a natural language processing model using open-source libraries. This model learns from the supplied data and prepares answers to subsequent user inquiries.
[0311] The completed model is implemented as an interactive program by the server. Users can use the terminal's interactive interface to input questions in natural language. For example, they can ask, "Please tell me the initial setup procedure for this product." This question is processed by the server, and the optimal answer is generated based on the generative model. The answer is then immediately sent to the user's terminal.
[0312] This system allows users to obtain information instantly and accurately. Furthermore, user questions and feedback are recorded on the server, and the model is regularly updated based on this information to improve its accuracy.
[0313] An example of a prompt would be, "Please tell me how to troubleshoot this." This would prepare the model to provide relevant solutions.
[0314] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0315] Step 1:
[0316] Users upload digital files from their devices. These files include, for example, product manuals and technical documents. They click the "Select File" button on their device, choose the desired file, and submit it. The input is a digital file, and the output is the transfer of the file to the server.
[0317] Step 2:
[0318] The server verifies the format of the received file. Specifically, the server checks the file extension and internal metadata to determine if it is a supported format (e.g., PDF, DOCX). The input for this step is the uploaded file, and the output is the result of the verification that the file is in the correct format.
[0319] Step 3:
[0320] The server extracts text data from the file. If the format is supported, it extracts the text using appropriate software (e.g., PDF analysis tools or OCR technology). The input for this step is a verified digital file, and the output is the extracted text data.
[0321] Step 4:
[0322] The server hierarchically structures the extracted text data. Using a database management system, it structures the data by section and topic. This process makes it easy for users to search for information. The input for this step is text data, and the output is hierarchically structured data.
[0323] Step 5:
[0324] The server supplies hierarchical data to a generative AI model for training. The model analyzes the supplied data and is trained to prepare appropriate answers to queries. The input for this step is hierarchical data, and the output is the trained generative model.
[0325] Step 6:
[0326] The server uses a trained generative model to build a conversational program (chatbot). This chatbot is designed to receive questions from users and generate answers in real time using the model. The input to this step is the trained generative model, and the output is a working chatbot.
[0327] Step 7:
[0328] The user enters a question through the terminal's interactive interface. For example, they might ask, "Please tell me how to set up this product." In this step, the input is the user's question, and the output is the sending of the question to the server.
[0329] Step 8:
[0330] The server analyzes the received question using a generative model and generates an appropriate answer. Based on the analysis results, it searches for relevant information, generates an answer, and sends it to the user's terminal. The input for this step is the user's question, and the output is the generated answer.
[0331] (Application Example 1)
[0332] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0333] Currently, many physical stores are required to provide customers with prompt information regarding product details and service information in response to their inquiries. However, there is a lack of efficient methods to achieve this, which can lead to decreased customer satisfaction. The objective of this invention is to solve this problem by providing a system that provides customers with the information they need quickly and accurately in physical stores.
[0334] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0335] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for constructing an interactive program that generates answers to questions using a generative model, means for users to access product information using identification information and for presenting appropriate answers to questions, and means for acquiring identification information through the user terminal and redirecting to the relevant information based on that information. This enables rapid information provision to customers in physical stores.
[0336] "Means for receiving files and verifying their format" refers to a function that receives digital data sent from a user to the system and determines whether the data format is supported by the system.
[0337] "A means of extracting text data from received files and structuring it hierarchically" refers to a function that extracts textual information from transmitted digital data and organizes and structures that information by item or heading.
[0338] "A means of building an interactive program that generates answers to questions using generative models" refers to a function that utilizes machine learning techniques to create a program that generates appropriate responses to user questions.
[0339] "A means for users to access product information using identification information and to be presented with appropriate answers to questions" refers to a function that allows users to access a digital system using a specific ID or code and automatically answers questions based on detailed information of the selected product.
[0340] "A means of acquiring identification information through the user's terminal and redirecting them to relevant information based on that information" refers to a function that collects identification data supplied via the user's device and uses that data to guide the user to related content or information.
[0341] To implement this invention, the following system is necessary. Specifically, a system is constructed that utilizes the user's terminal and a server installed in the store to provide product information and generate quick responses to user inquiries.
[0342] The server receives manual files uploaded by users. Next, it verifies the file format and extracts and hierarchically structures the text data. Natural language processing libraries (e.g., TensorFlow or PyTorch) are used for this text data processing. The server then feeds the organized data to a machine learning model, training a generative AI model to respond to user questions. This model forms the basis for real-time generated interactive programs.
[0343] Users can use their smartphones or other devices to scan QR codes attached to products in physical stores. This sends identifying information about the relevant product or service to a server. The server then uses a generative model based on this information to instantly provide appropriate answers to user questions. This entire process enhances the user's in-store shopping experience and allows them to receive quick and detailed information.
[0344] As a concrete example, consider a case where a user scans a QR code for a specific product—for example, a camera—and asks about its initial setup. In this case, the server inputs the prompt "How do I set up the camera?" into the AI model and provides detailed setup instructions to the user's device.
[0345] This system allows customers to instantly obtain the information they need without any hassle, effectively complementing in-store customer service.
[0346] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0347] Step 1:
[0348] The user scans the QR code included with the product using their device.
[0349] Input: QR code data
[0350] Output: Identification information
[0351] Specific operation: The device uses its camera function to read the QR code, analyzes its contents, and obtains specific identification information.
[0352] Step 2:
[0353] The terminal sends identification information to the server.
[0354] Input: Identification Information
[0355] Output: Formal request to the server
[0356] Specific operation: The terminal sends identification information to the server via the network and requests information related to the product.
[0357] Step 3:
[0358] Based on the identification information received by the server, it searches for and retrieves the relevant manual files.
[0359] Input: Identification Information
[0360] Output: Manual file
[0361] Specific operation: The server refers to an internal database and searches for and retrieves the digital manual corresponding to the specified identification information.
[0362] Step 4:
[0363] The server extracts text data from the acquired manuals and structures it hierarchically.
[0364] Input: Manual file
[0365] Output: Structured text data
[0366] Specific operation: The server extracts necessary text from the scanned manual using natural language processing and organizes it by headings and items.
[0367] Step 5:
[0368] The server supplies structured data to the generative model and prepares it for response generation.
[0369] Input: Structured text data
[0370] Output: Feed to the generative model
[0371] Specific operation: The server inputs data into the AI model and performs the necessary training to generate appropriate answers to user questions.
[0372] Step 6:
[0373] The user sends the question to the server via their device.
[0374] Input: User Question
[0375] Output: Query to the server
[0376] Specific operation: The user uses the app's interface to input questions about the product in natural language and sends them to the server.
[0377] Step 7:
[0378] The server uses a generative model to generate an answer to the question and sends it to the terminal.
[0379] Input: User Question Query
[0380] Output: Answer to the question
[0381] Specific operation: The server analyzes the received question using a generative AI model, generates an answer based on the trained data, and sends it to the user's terminal.
[0382] Step 8:
[0383] The terminal displays the response received from the server to the user.
[0384] Input: Response from the server
[0385] Output: Answer displayed on the user interface
[0386] Specific operation: The terminal displays the response provided by the server on the screen, providing the user with detailed information.
[0387] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0388] In embodiments of the present invention, the process begins with a user uploading a file from their terminal. This file contains digital data of a manual, and the server receives the file and verifies its format. If it is determined to be in a supported format, the server extracts text data from the file and structures it hierarchically. This process organizes the data by category, making it easy to search.
[0389] The hierarchically structured data is fed to a generative model by the server, where it is trained. This training enables the model to generate appropriate answers to user inquiries. Next, the server utilizes the trained model to build a chatbot. This chatbot has the functionality to respond to user questions and provides information through dialogue.
[0390] Furthermore, this invention incorporates an emotion engine. This engine has the function of recognizing emotions from the user's input and conversation tone. The chatbot determines the user's emotions using the emotion engine and adjusts the tone and content of its response based on that information.
[0391] As a concrete example, consider a scenario where a user asks for instructions on how to use a product for the first time and feels anxious. They might consult a chatbot via their device, saying, "This is my first time using this device, so I'm worried about whether I've set it up correctly." In this case, the server utilizes an emotion engine to recognize the user's anxiety and provides a reassuring response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence." This approach improves the user experience and enhances the quality of information provided.
[0392] The above describes an embodiment of the present invention that combines an emotion engine, enabling the construction of a system that provides users with information quickly and with consideration for their emotions.
[0393] The following describes the processing flow.
[0394] Step 1:
[0395] The user selects a manual file from their terminal and uploads it to the system. This file is a digitized version of the manual.
[0396] Step 2:
[0397] The server receives the uploaded file and checks its format. If the file is in a supported format, it proceeds to the next step.
[0398] Step 3:
[0399] The server extracts text data from the received file. During this process, OCR and file parsing technologies are used to accurately extract the necessary text information.
[0400] Step 4:
[0401] The server organizes the extracted text data into a hierarchical structure. It classifies the data by item and heading, making it easy for users to quickly access the information they need.
[0402] Step 5:
[0403] The server supplies hierarchically structured data to a generative model, which is then trained. This training enables the model to generate appropriate answers to questions.
[0404] Step 6:
[0405] The server builds an FAQ chatbot based on a pre-trained generative model. This chatbot provides answers to questions in real time.
[0406] Step 7:
[0407] Users access the chatbot through their device and input questions in natural language. They can ask questions in a conversational format.
[0408] Step 8:
[0409] The server analyzes user input and uses an emotion engine to determine the user's emotions. By extracting emotion data from the input, it understands the user's feelings.
[0410] Step 9:
[0411] The server uses a generative model to generate responses with a tone adjusted according to the user's emotions. It then sends the generated responses to the user's device.
[0412] Step 10:
[0413] Users can review the answers on their device and ask additional questions to the chatbot as needed. Through this process, users can obtain information smoothly. This system enables the delivery of more personalized and emotionally sensitive information to users.
[0414] (Example 2)
[0415] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0416] In modern information technology, there is a need for systems that can respond quickly and accurately to user inquiries. However, conventional systems struggle to generate responses that take user emotions into account, and there is a demand for technologies that can achieve more human-like interaction. Furthermore, there is a lack of mechanisms to improve response accuracy by utilizing feedback, and data integration to provide comprehensive information is insufficient.
[0417] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0418] In this invention, the server includes means for receiving a file and verifying its format, means for extracting information from the received file and structuring it hierarchically, means for supplying the hierarchically structured information to a generative model for training, and means for providing an appropriate response to a user inquiry. This enables the generation of responses that take into account the user's emotional perception, improvement of response accuracy through feedback, and comprehensive information provision through the integration of multiple information sources.
[0419] "Receiving a file" refers to the act of a server retrieving digital data provided by a user from their device.
[0420] "Verifying the format" is the process of verifying whether the received digital data is in a specific, compatible data format.
[0421] "Information extraction" is the process of extracting necessary text data and related information from received digital data.
[0422] "Hierarchical structuring" refers to classifying extracted information into chapters or items and organizing them hierarchically.
[0423] A "generative model" is an algorithm that is trained using large amounts of digital data and generates responses to queries through natural language processing.
[0424] "Training" refers to the process of improving the accuracy of a generative model's responses by using input data to learn from it.
[0425] An "automated response system" is software or a system that provides a mechanism for the system to automatically generate a response to a user's inquiry.
[0426] "Identifying emotions" is the process of analyzing user input and context to detect the emotions contained within it.
[0427] "Adjusting responses" refers to changing the content and tone of responses provided to the user based on identified emotional information.
[0428] "Integrating information sources" is the process of aggregating data from multiple related databases and information streams and structuring it into a single, comprehensive piece of information.
[0429] In implementing this invention, the process begins with a user uploading a file containing digital data using their own device. This file contains information such as product manuals and instructions. Once the user uploads the file, the server receives the data. The server first checks the format of the received file, verifying, for example, whether it is in PDF or DOCX format. This verification process can be automated using open-source libraries or proprietary programs.
[0430] If the server determines that the file is in a supported format, it begins data extraction. For example, it uses tools such as Apache Tika or Python's PDFMiner to extract the necessary text data from the file. Because the extracted data is difficult to handle directly, the server then performs a step to structure the data hierarchically based on chapters and headings. This organizes the information by category, making subsequent searches easier.
[0431] The hierarchically structured information is then fed to the generative AI model. The server trains the model using open-source machine learning libraries such as TensorFlow and PyTorch. This training enables the model to generate fast and appropriate responses to real-time user inquiries.
[0432] Using a pre-trained generative AI model, the server builds a chatbot that automatically generates responses to user questions. Furthermore, the invention integrates an emotion engine, allowing the server to identify emotions from user input. This enables the adjustment of response tone based on emotional information obtained through data analysis.
[0433] As a concrete example, consider a scenario where a user is unsure about the initial setup of a product. When the user enters a prompt message through their device such as, "This is my first time using this device, so I'm worried about whether I've set it up correctly," the server uses an emotion engine to detect the user's anxiety. As a result, the chatbot generates a response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence," and provides it to the user. This system allows users to receive thoughtful and emotionally sensitive information.
[0434] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0435] Step 1:
[0436] The user uploads a digital data file to the server using a terminal. The input is a digital data file, and this process saves the file to the server. The server receives the uploaded file and temporarily stores it for processing in the next step.
[0437] Step 2:
[0438] The server checks the format of the uploaded file. Specifically, it verifies whether the file is in a supported format such as PDF or DOCX. The input is the uploaded file, and the output is the result of the file format verification. If the format is not appropriate, the server sends an error message to the user and prompts them to re-upload the file.
[0439] Step 3:
[0440] The server extracts information from files in the correct format. This extraction process uses tools such as Apache Tika or PDFMiner. The input is a formatted file, and the output is extracted text data. The server stores the extracted text data in memory and proceeds to the next structuring step.
[0441] Step 4:
[0442] The server hierarchically structures the extracted text data. The data is primarily organized based on document chapters and headings. The input is extracted text data, and the output is hierarchically structured information. This hierarchical structure organizes the information, making it easier to search and use later.
[0443] Step 5:
[0444] The server supplies hierarchically structured information to a generative AI model for training. This process involves training the generative model with new data patterns to improve the accuracy of response generation. The input is hierarchically structured information, and the output is the trained model. Using this model enables accurate and appropriate responses.
[0445] Step 6:
[0446] The server uses a trained generative AI model to build a chatbot. This automated response system has the ability to instantly respond to user prompts. The input is the user's question, and the output is an automatically generated response based on the model.
[0447] Step 7:
[0448] The server integrates an emotion engine, adding the ability to identify the user's input emotions. This system analyzes the emotions the user expresses and adjusts the response accordingly. The input is the user's response, including the prompt, and the output is the response adjusted based on the emotion identification.
[0449] Through these steps, the server enables the provision of advanced information to improve the user experience.
[0450] (Application Example 2)
[0451] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0452] In systems that provide advanced information, users often find it difficult to easily understand operation manuals and setup procedures, which can be particularly psychologically stressful and anxiety-inducing. This invention aims to reduce user anxiety during the information provision process and enable smooth information access by adjusting responses to take into account the user's emotional state.
[0453] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0454] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for supplying the hierarchically structured data to a generative model and training it, means for constructing a dialogue system that generates answers to questions using the generative model, and means for analyzing the emotions of the user in response to a question, adjusting the answer, and presenting an appropriate answer. This enables effective information provision that takes into account the user's psychological state.
[0455] A "file" is a collection of data used to store digital information and process it within a computer system.
[0456] "Means of checking format" refers to a function that determines whether the received data conforms to the specifications that the system can process.
[0457] "Text data" refers to digital information that includes string information, and is data recorded in the form of sentences or words.
[0458] "Means of hierarchical structuring" refers to the function of organizing information structurally and clarifying the relationships between each element.
[0459] A "generative model" is an algorithm that uses machine learning techniques to generate new information or results based on input data.
[0460] "Training methods" refer to the process of using data to train a generative model and improve its accuracy and capabilities.
[0461] A "dialogue system" is a computer-based conversation management system that provides information and generates responses based on input from users.
[0462] A "user" is an individual or organization that interacts with the system and receives or provides information.
[0463] "Means of analyzing emotions" refers to a function that identifies and evaluates a user's psychological state based on input data.
[0464] "Methods for adjusting responses" refer to the process of optimizing the content and expression of generated responses based on analyzed emotional information.
[0465] Modes for carrying out the invention
[0466] To implement this invention, a server plays a central role. The server receives files uploaded by users using their terminals. After receiving the files, the server uses means to verify that their format conforms to predetermined specifications. Subsequently, text data is extracted from the received files, and the information is organized through hierarchical structuring. This hierarchical structuring clarifies the relationships between data, facilitating subsequent processing.
[0467] Next, the server supplies the hierarchically structured data to the generative model for training. The machine learning framework "TensorFlow" is used to train the generative model. As a result, the trained generative model has the ability to function as a dialogue system. This dialogue system provides information and generates responses based on user input.
[0468] Furthermore, the server uses an emotion analysis engine to analyze the input data received from the user and evaluate the user's emotional state. The "Hugging Face Transformers" library is utilized for emotion analysis. Based on this emotional information, the server optimizes the generated response by adjusting its content and tone.
[0469] As a concrete application example, consider a scenario where a first-time user of an autonomous vehicle asks, "I'm worried about operating the autonomous driving mode." In this case, the server uses sentiment analysis to identify the user's anxiety and provides a reassuring response such as, "The autonomous driving mode has been confirmed to be safe. Please follow the steps below to use it with confidence." Through this process, the user can access accurate information to safely operate the vehicle system.
[0470] An example of a prompt message is, "Generate a method for providing vehicle information to the server that takes user sentiment into consideration."
[0471] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0472] Step 1:
[0473] The server receives files sent from the terminal. It parses the received file to determine if it is in a supported format. The input is a digital file uploaded by the user via the terminal, and the output is a boolean value indicating whether the format is valid.
[0474] Step 2:
[0475] The server extracts text data from verified digital files. The extracted text data is then organized into higher and lower levels to create a hierarchical structure. The input is raw data extracted according to a verified file format, and the output is hierarchically structured data.
[0476] Step 3:
[0477] The server supplies hierarchically structured data to a generative model using the machine learning framework "TensorFlow" and trains the model. The input is hierarchically structured data, and the output is a trained generative model.
[0478] Step 4:
[0479] When a user enters a question using a terminal, the server receives the question. Using a trained generative model, it generates an appropriate answer to the question. The input is a question from the user in natural language, and the output is the generated answer.
[0480] Step 5:
[0481] The server uses the emotion analysis engine "Hugging Face Transformers" to evaluate the user's emotional state from the context of the question. The input is the user's question text, and the output is an identified emotional state (e.g., anxiety, joy).
[0482] Step 6:
[0483] The server adjusts the content and expression of the generated response based on the emotion assessment results and provides it to the user in the most appropriate form. The input is the emotional state analyzed and the generated response, and the output is the adjusted response.
[0484] Step 7:
[0485] The server responds to the user's question by sending a pre-configured answer to the terminal and presenting it to the user. The input is the optimized answer, and the output is the final information provided to the user.
[0486] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0487] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0488] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0489] [Third Embodiment]
[0490] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0491] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0492] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0493] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0494] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0495] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0496] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0497] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0498] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0499] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0500] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0501] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0502] The following systems are conceivable as embodiments for carrying out the present invention.
[0503] First, the user uploads a file from their device to the system. This file contains digital data of the manual, and the server checks the format of the received file before processing it. If it is in a supported format, it proceeds to the next step.
[0504] Next, the server extracts text data from the received file. Once the text data is extracted, it is structured hierarchically. At this stage, the data is organized by category and headings so that users can easily search for it.
[0505] The server then feeds this structured data to a generative model. This generative model uses AI to learn from the supplied data. This prepares the model to generate appropriate answers to queries.
[0506] Next, the server uses the trained generative model to build an FAQ chatbot. This chatbot is programmed to receive questions from users and generate answers in real time.
[0507] Through a conversational interface accessible from the terminal, users input questions in natural language. The server receives these questions, uses a generative model to generate appropriate answers, and sends them to the user's terminal, enabling rapid information delivery.
[0508] As a concrete example, consider a question about how to operate a new device. When a user asks the chatbot on their device, "How do I set up this device?", the server uses a generative model to find the relevant data and provides a specific answer such as, "Please follow these steps for the initial setup."
[0509] This system allows users to quickly access important information, eliminating the need to read the entire manual. This is an embodiment of the present invention.
[0510] The following describes the processing flow.
[0511] Step 1:
[0512] The user selects a manual file from their terminal and uploads it. The uploaded file is sent to the server.
[0513] Step 2:
[0514] The server checks the received file and verifies that it is in a supported format. If the format is not appropriate, it returns an error message to the user.
[0515] Step 3:
[0516] The server uses OCR (Optical Character Recognition) or other parsing techniques to extract text data from the appropriate files.
[0517] Step 4:
[0518] The server organizes and hierarchically structures the extracted text data. This process classifies and structures the data by headings and paragraphs.
[0519] Step 5:
[0520] The server supplies the generated hierarchical structured data to the generative model and performs training. This training builds the model's ability to answer future questions.
[0521] Step 6:
[0522] The server uses a trained generative model to build an FAQ chatbot, enabling automated responses to user questions.
[0523] Step 7:
[0524] Users access the chatbot through their device and input the information they want to obtain in natural language.
[0525] Step 8:
[0526] The server sends the question received from the user to a generative model, which generates the optimal answer. The generated answer is then returned to the user's device and displayed.
[0527] Through this series of steps, users can easily and quickly access the information in the manual.
[0528] (Example 1)
[0529] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] This invention aims to improve a system that quickly and accurately retrieves useful information from files containing large amounts of data and provides appropriate answers to user inquiries. Furthermore, it seeks to solve the problem of improving the accuracy of answers by utilizing user feedback and enhancing the comprehensiveness of information provision by integrating data from diverse information sources.
[0531] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0532] In this invention, the server includes means for receiving information and confirming its structure, means for extracting and hierarchizing data from the received information, and means for providing the hierarchical data to a learning model for training. This makes it possible to provide quick and accurate answers to user inquiries. Furthermore, by utilizing feedback, it is possible to improve the accuracy of the system and provide comprehensive information through information integration.
[0533] "Means of receiving information and verifying its structure" refers to the process of taking in information provided from external sources and checking whether that information conforms to a specific format or standard.
[0534] "Methods for extracting and hierarchizing data" refers to the process of extracting necessary items from received information and organizing that data based on a logical order or classification.
[0535] "Means of providing and training a learning model" refers to the process of inputting prepared data into an artificial intelligence model, allowing that model to learn patterns and knowledge from the data.
[0536] "Methods for constructing dialogue programs" refers to the process of developing a system that uses a trained artificial intelligence model to automatically communicate with users and generate appropriate answers to questions.
[0537] "A means of providing appropriate answers to user inquiries" refers to a function that analyzes questions received from users, searches for relevant information, and provides responses.
[0538] "Methods for collecting evaluations, using that information to improve the learning model, and enhancing the accuracy of responses" refers to methods of collecting feedback from users, retraining the model based on that feedback, and improving the system's performance.
[0539] "Means of connecting diverse information sources to provide comprehensive information" refers to methods of aggregating data from multiple different databases and information sets to provide users with consistent and detailed information.
[0540] To implement this invention, the system is primarily operated through the collaboration of three parties: a server, a terminal, and a user. The specific method is described below.
[0541] First, users upload digital information to the system using their own devices. These devices are expected to be computers or smartphones. Examples of uploaded information include product manuals and technical documents.
[0542] The server automatically analyzes the format and structure of the received digital information. This analysis can utilize software that recognizes specific file formats, such as PDF analysis tools or OCR (optical character recognition) technology. Based on the analysis results, the server extracts the information as text and organizes the data structurally. This organization process involves completing a hierarchical format using a database management system or similar method.
[0543] The organized data is fed into a generative AI model, and the learning process begins. An example of a generative AI model is a natural language processing model using open-source libraries. This model learns from the supplied data and prepares answers to subsequent user inquiries.
[0544] The completed model is implemented as an interactive program by the server. Users can use the terminal's interactive interface to input questions in natural language. For example, they can ask, "Please tell me the initial setup procedure for this product." This question is processed by the server, and the optimal answer is generated based on the generative model. The answer is then immediately sent to the user's terminal.
[0545] This system allows users to obtain information instantly and accurately. Furthermore, user questions and feedback are recorded on the server, and the model is regularly updated based on this information to improve its accuracy.
[0546] An example of a prompt would be, "Please tell me how to troubleshoot this." This would prepare the model to provide relevant solutions.
[0547] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0548] Step 1:
[0549] Users upload digital files from their devices. These files include, for example, product manuals and technical documents. They click the "Select File" button on their device, choose the desired file, and submit it. The input is a digital file, and the output is the transfer of the file to the server.
[0550] Step 2:
[0551] The server verifies the format of the received file. Specifically, the server checks the file extension and internal metadata to determine if it is a supported format (e.g., PDF, DOCX). The input for this step is the uploaded file, and the output is the result of the verification that the file is in the correct format.
[0552] Step 3:
[0553] The server extracts text data from the file. If the format is supported, it extracts the text using appropriate software (e.g., PDF analysis tools or OCR technology). The input for this step is a verified digital file, and the output is the extracted text data.
[0554] Step 4:
[0555] The server hierarchically structures the extracted text data. Using a database management system, it structures the data by section and topic. This process makes it easy for users to search for information. The input for this step is text data, and the output is hierarchically structured data.
[0556] Step 5:
[0557] The server supplies hierarchical data to a generative AI model for training. The model analyzes the supplied data and is trained to prepare appropriate answers to queries. The input for this step is hierarchical data, and the output is the trained generative model.
[0558] Step 6:
[0559] The server uses a trained generative model to build a conversational program (chatbot). This chatbot is designed to receive questions from users and generate answers in real time using the model. The input to this step is the trained generative model, and the output is a working chatbot.
[0560] Step 7:
[0561] The user enters a question through the terminal's interactive interface. For example, they might ask, "Please tell me how to set up this product." In this step, the input is the user's question, and the output is the sending of the question to the server.
[0562] Step 8:
[0563] The server analyzes the received question using a generative model and generates an appropriate answer. Based on the analysis results, it searches for relevant information, generates an answer, and sends it to the user's terminal. The input for this step is the user's question, and the output is the generated answer.
[0564] (Application Example 1)
[0565] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0566] Currently, many physical stores are required to provide customers with prompt information regarding product details and service information in response to their inquiries. However, there is a lack of efficient methods to achieve this, which can lead to decreased customer satisfaction. The objective of this invention is to solve this problem by providing a system that provides customers with the information they need quickly and accurately in physical stores.
[0567] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0568] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for constructing an interactive program that generates answers to questions using a generative model, means for users to access product information using identification information and for presenting appropriate answers to questions, and means for acquiring identification information through the user terminal and redirecting to the relevant information based on that information. This enables rapid information provision to customers in physical stores.
[0569] "Means for receiving files and verifying their format" refers to a function that receives digital data sent from a user to the system and determines whether the data format is supported by the system.
[0570] "A means of extracting text data from received files and structuring it hierarchically" refers to a function that extracts textual information from transmitted digital data and organizes and structures that information by item or heading.
[0571] "A means of building an interactive program that generates answers to questions using generative models" refers to a function that utilizes machine learning techniques to create a program that generates appropriate responses to user questions.
[0572] "A means for users to access product information using identification information and to be presented with appropriate answers to questions" refers to a function that allows users to access a digital system using a specific ID or code and automatically answers questions based on detailed information of the selected product.
[0573] "A means of acquiring identification information through the user's terminal and redirecting them to relevant information based on that information" refers to a function that collects identification data supplied via the user's device and uses that data to guide the user to related content or information.
[0574] To implement this invention, the following system is necessary. Specifically, a system is constructed that utilizes the user's terminal and a server installed in the store to provide product information and generate quick responses to user inquiries.
[0575] The server receives manual files uploaded by users. Next, it verifies the file format and extracts and hierarchically structures the text data. Natural language processing libraries (e.g., TensorFlow or PyTorch) are used for this text data processing. The server then feeds the organized data to a machine learning model, training a generative AI model to respond to user questions. This model forms the basis for real-time generated interactive programs.
[0576] Users can use their smartphones or other devices to scan QR codes attached to products in physical stores. This sends identifying information about the relevant product or service to a server. The server then uses a generative model based on this information to instantly provide appropriate answers to user questions. This entire process enhances the user's in-store shopping experience and allows them to receive quick and detailed information.
[0577] As a concrete example, consider a case where a user scans a QR code for a specific product—for example, a camera—and asks about its initial setup. In this case, the server inputs the prompt "How do I set up the camera?" into the AI model and provides detailed setup instructions to the user's device.
[0578] This system allows customers to instantly obtain the information they need without any hassle, effectively complementing in-store customer service.
[0579] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0580] Step 1:
[0581] The user scans the QR code included with the product using their device.
[0582] Input: QR code data
[0583] Output: Identification information
[0584] Specific operation: The device uses its camera function to read the QR code, analyzes its contents, and obtains specific identification information.
[0585] Step 2:
[0586] The terminal sends identification information to the server.
[0587] Input: Identification Information
[0588] Output: Formal request to the server
[0589] Specific operation: The terminal sends identification information to the server via the network and requests information related to the product.
[0590] Step 3:
[0591] Based on the identification information received by the server, it searches for and retrieves the relevant manual files.
[0592] Input: Identification Information
[0593] Output: Manual file
[0594] Specific operation: The server refers to an internal database and searches for and retrieves the digital manual corresponding to the specified identification information.
[0595] Step 4:
[0596] The server extracts text data from the acquired manuals and structures it hierarchically.
[0597] Input: Manual file
[0598] Output: Structured text data
[0599] Specific operation: The server extracts necessary text from the scanned manual using natural language processing and organizes it by headings and items.
[0600] Step 5:
[0601] The server supplies structured data to the generative model and prepares it for response generation.
[0602] Input: Structured text data
[0603] Output: Feed to the generative model
[0604] Specific operation: The server inputs data into the AI model and performs the necessary training to generate appropriate answers to user questions.
[0605] Step 6:
[0606] The user sends the question to the server via their device.
[0607] Input: User Question
[0608] Output: Query to the server
[0609] Specific operation: The user uses the app's interface to input questions about the product in natural language and sends them to the server.
[0610] Step 7:
[0611] The server uses a generative model to generate an answer to the question and sends it to the terminal.
[0612] Input: User Question Query
[0613] Output: Answer to the question
[0614] Specific operation: The server analyzes the received question using a generative AI model, generates an answer based on the trained data, and sends it to the user's terminal.
[0615] Step 8:
[0616] The terminal displays the response received from the server to the user.
[0617] Input: Response from the server
[0618] Output: Answer displayed on the user interface
[0619] Specific operation: The terminal displays the response provided by the server on the screen, providing the user with detailed information.
[0620] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0621] In embodiments of the present invention, the process begins with a user uploading a file from their terminal. This file contains digital data of a manual, and the server receives the file and verifies its format. If it is determined to be in a supported format, the server extracts text data from the file and structures it hierarchically. This process organizes the data by category, making it easy to search.
[0622] The hierarchically structured data is fed to a generative model by the server, where it is trained. This training enables the model to generate appropriate answers to user inquiries. Next, the server utilizes the trained model to build a chatbot. This chatbot has the functionality to respond to user questions and provides information through dialogue.
[0623] Furthermore, this invention incorporates an emotion engine. This engine has the function of recognizing emotions from the user's input and conversation tone. The chatbot determines the user's emotions using the emotion engine and adjusts the tone and content of its response based on that information.
[0624] As a concrete example, consider a scenario where a user asks for instructions on how to use a product for the first time and feels anxious. They might consult a chatbot via their device, saying, "This is my first time using this device, so I'm worried about whether I've set it up correctly." In this case, the server utilizes an emotion engine to recognize the user's anxiety and provides a reassuring response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence." This approach improves the user experience and enhances the quality of information provided.
[0625] The above describes an embodiment of the present invention that combines an emotion engine, enabling the construction of a system that provides users with information quickly and with consideration for their emotions.
[0626] The following describes the processing flow.
[0627] Step 1:
[0628] The user selects a manual file from their terminal and uploads it to the system. This file is a digitized version of the manual.
[0629] Step 2:
[0630] The server receives the uploaded file and checks its format. If the file is in a supported format, it proceeds to the next step.
[0631] Step 3:
[0632] The server extracts text data from the received file. During this process, OCR and file parsing technologies are used to accurately extract the necessary text information.
[0633] Step 4:
[0634] The server organizes the extracted text data into a hierarchical structure. It classifies the data by item and heading, making it easy for users to quickly access the information they need.
[0635] Step 5:
[0636] The server supplies hierarchically structured data to a generative model, which is then trained. This training enables the model to generate appropriate answers to questions.
[0637] Step 6:
[0638] The server builds an FAQ chatbot based on a pre-trained generative model. This chatbot provides answers to questions in real time.
[0639] Step 7:
[0640] Users access the chatbot through their device and input questions in natural language. They can ask questions in a conversational format.
[0641] Step 8:
[0642] The server analyzes user input and uses an emotion engine to determine the user's emotions. By extracting emotion data from the input, it understands the user's feelings.
[0643] Step 9:
[0644] The server uses a generative model to generate responses with a tone adjusted according to the user's emotions. It then sends the generated responses to the user's device.
[0645] Step 10:
[0646] Users can review the answers on their device and ask additional questions to the chatbot as needed. Through this process, users can obtain information smoothly. This system enables the delivery of more personalized and emotionally sensitive information to users.
[0647] (Example 2)
[0648] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0649] In modern information technology, there is a need for systems that can respond quickly and accurately to user inquiries. However, conventional systems struggle to generate responses that take user emotions into account, and there is a demand for technologies that can achieve more human-like interaction. Furthermore, there is a lack of mechanisms to improve response accuracy by utilizing feedback, and data integration to provide comprehensive information is insufficient.
[0650] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0651] In this invention, the server includes means for receiving a file and verifying its format, means for extracting information from the received file and structuring it hierarchically, means for supplying the hierarchically structured information to a generative model for training, and means for providing an appropriate response to a user inquiry. This enables the generation of responses that take into account the user's emotional perception, improvement of response accuracy through feedback, and comprehensive information provision through the integration of multiple information sources.
[0652] "Receiving a file" refers to the act of a server retrieving digital data provided by a user from their device.
[0653] "Verifying the format" is the process of verifying whether the received digital data is in a specific, compatible data format.
[0654] "Information extraction" is the process of extracting necessary text data and related information from received digital data.
[0655] "Hierarchical structuring" refers to classifying extracted information into chapters or items and organizing them hierarchically.
[0656] A "generative model" is an algorithm that is trained using large amounts of digital data and generates responses to queries through natural language processing.
[0657] "Training" refers to the process of improving the accuracy of a generative model's responses by using input data to learn from it.
[0658] An "automated response system" is software or a system that provides a mechanism for the system to automatically generate a response to a user's inquiry.
[0659] "Identifying emotions" is the process of analyzing user input and context to detect the emotions contained within it.
[0660] "Adjusting responses" refers to changing the content and tone of responses provided to the user based on identified emotional information.
[0661] "Integrating information sources" is the process of aggregating data from multiple related databases and information streams and structuring it into a single, comprehensive piece of information.
[0662] In implementing this invention, the process begins with a user uploading a file containing digital data using their own device. This file contains information such as product manuals and instructions. Once the user uploads the file, the server receives the data. The server first checks the format of the received file, verifying, for example, whether it is in PDF or DOCX format. This verification process can be automated using open-source libraries or proprietary programs.
[0663] If the server determines that the file is in a supported format, it begins data extraction. For example, it uses tools such as Apache Tika or Python's PDFMiner to extract the necessary text data from the file. Because the extracted data is difficult to handle directly, the server then performs a step to structure the data hierarchically based on chapters and headings. This organizes the information by category, making subsequent searches easier.
[0664] The hierarchically structured information is then fed to the generative AI model. The server trains the model using open-source machine learning libraries such as TensorFlow and PyTorch. This training enables the model to generate fast and appropriate responses to real-time user inquiries.
[0665] Using a pre-trained generative AI model, the server builds a chatbot that automatically generates responses to user questions. Furthermore, the invention integrates an emotion engine, allowing the server to identify emotions from user input. This enables the adjustment of response tone based on emotional information obtained through data analysis.
[0666] As a concrete example, consider a scenario where a user is unsure about the initial setup of a product. When the user enters a prompt message through their device such as, "This is my first time using this device, so I'm worried about whether I've set it up correctly," the server uses an emotion engine to detect the user's anxiety. As a result, the chatbot generates a response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence," and provides it to the user. This system allows users to receive thoughtful and emotionally sensitive information.
[0667] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0668] Step 1:
[0669] The user uploads a digital data file to the server using a terminal. The input is a digital data file, and this process saves the file to the server. The server receives the uploaded file and temporarily stores it for processing in the next step.
[0670] Step 2:
[0671] The server checks the format of the uploaded file. Specifically, it verifies whether the file is in a supported format such as PDF or DOCX. The input is the uploaded file, and the output is the result of the file format verification. If the format is not appropriate, the server sends an error message to the user and prompts them to re-upload the file.
[0672] Step 3:
[0673] The server extracts information from files in the correct format. This extraction process uses tools such as Apache Tika or PDFMiner. The input is a formatted file, and the output is extracted text data. The server stores the extracted text data in memory and proceeds to the next structuring step.
[0674] Step 4:
[0675] The server hierarchically structures the extracted text data. The data is primarily organized based on document chapters and headings. The input is extracted text data, and the output is hierarchically structured information. This hierarchical structure organizes the information, making it easier to search and use later.
[0676] Step 5:
[0677] The server supplies hierarchically structured information to a generative AI model for training. This process involves training the generative model with new data patterns to improve the accuracy of response generation. The input is hierarchically structured information, and the output is the trained model. Using this model enables accurate and appropriate responses.
[0678] Step 6:
[0679] The server uses a trained generative AI model to build a chatbot. This automated response system has the ability to instantly respond to user prompts. The input is the user's question, and the output is an automatically generated response based on the model.
[0680] Step 7:
[0681] The server integrates an emotion engine, adding the ability to identify the user's input emotions. This system analyzes the emotions the user expresses and adjusts the response accordingly. The input is the user's response, including the prompt, and the output is the response adjusted based on the emotion identification.
[0682] Through these steps, the server enables the provision of advanced information to improve the user experience.
[0683] (Application Example 2)
[0684] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0685] In systems that provide advanced information, users often find it difficult to easily understand operation manuals and setup procedures, which can be particularly psychologically stressful and anxiety-inducing. This invention aims to reduce user anxiety during the information provision process and enable smooth information access by adjusting responses to take into account the user's emotional state.
[0686] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0687] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for supplying the hierarchically structured data to a generative model and training it, means for constructing a dialogue system that generates answers to questions using the generative model, and means for analyzing the emotions of the user in response to a question, adjusting the answer, and presenting an appropriate answer. This enables effective information provision that takes into account the user's psychological state.
[0688] A "file" is a collection of data used to store digital information and process it within a computer system.
[0689] "Means of checking format" refers to a function that determines whether the received data conforms to the specifications that the system can process.
[0690] "Text data" refers to digital information that includes string information, and is data recorded in the form of sentences or words.
[0691] "Means of hierarchical structuring" refers to the function of organizing information structurally and clarifying the relationships between each element.
[0692] A "generative model" is an algorithm that uses machine learning techniques to generate new information or results based on input data.
[0693] "Training methods" refer to the process of using data to train a generative model and improve its accuracy and capabilities.
[0694] A "dialogue system" is a computer-based conversation management system that provides information and generates responses based on input from users.
[0695] A "user" is an individual or organization that interacts with the system and receives or provides information.
[0696] "Means of analyzing emotions" refers to a function that identifies and evaluates a user's psychological state based on input data.
[0697] "Methods for adjusting responses" refer to the process of optimizing the content and expression of generated responses based on analyzed emotional information.
[0698] Modes for carrying out the invention
[0699] To implement this invention, a server plays a central role. The server receives files uploaded by users using their terminals. After receiving the files, the server uses means to verify that their format conforms to predetermined specifications. Subsequently, text data is extracted from the received files, and the information is organized through hierarchical structuring. This hierarchical structuring clarifies the relationships between data, facilitating subsequent processing.
[0700] Next, the server supplies the hierarchically structured data to the generative model for training. The machine learning framework "TensorFlow" is used to train the generative model. As a result, the trained generative model has the ability to function as a dialogue system. This dialogue system provides information and generates responses based on user input.
[0701] Furthermore, the server uses an emotion analysis engine to analyze the input data received from the user and evaluate the user's emotional state. The "Hugging Face Transformers" library is utilized for emotion analysis. Based on this emotional information, the server optimizes the generated response by adjusting its content and tone.
[0702] As a concrete application example, consider a scenario where a first-time user of an autonomous vehicle asks, "I'm worried about operating the autonomous driving mode." In this case, the server uses sentiment analysis to identify the user's anxiety and provides a reassuring response such as, "The autonomous driving mode has been confirmed to be safe. Please follow the steps below to use it with confidence." Through this process, the user can access accurate information to safely operate the vehicle system.
[0703] An example of a prompt message is, "Generate a method for providing vehicle information to the server that takes user sentiment into consideration."
[0704] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0705] Step 1:
[0706] The server receives files sent from the terminal. It parses the received file to determine if it is in a supported format. The input is a digital file uploaded by the user via the terminal, and the output is a boolean value indicating whether the format is valid.
[0707] Step 2:
[0708] The server extracts text data from verified digital files. The extracted text data is then organized into higher and lower levels to create a hierarchical structure. The input is raw data extracted according to a verified file format, and the output is hierarchically structured data.
[0709] Step 3:
[0710] The server supplies hierarchically structured data to a generative model using the machine learning framework "TensorFlow" and trains the model. The input is hierarchically structured data, and the output is a trained generative model.
[0711] Step 4:
[0712] When a user enters a question using a terminal, the server receives the question. Using a trained generative model, it generates an appropriate answer to the question. The input is a question from the user in natural language, and the output is the generated answer.
[0713] Step 5:
[0714] The server uses the emotion analysis engine "Hugging Face Transformers" to evaluate the user's emotional state from the context of the question. The input is the user's question text, and the output is an identified emotional state (e.g., anxiety, joy).
[0715] Step 6:
[0716] The server adjusts the content and expression of the generated response based on the emotion assessment results and provides it to the user in the most appropriate form. The input is the emotional state analyzed and the generated response, and the output is the adjusted response.
[0717] Step 7:
[0718] The server responds to the user's question by sending a pre-configured answer to the terminal and presenting it to the user. The input is the optimized answer, and the output is the final information provided to the user.
[0719] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0720] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0721] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0722] [Fourth Embodiment]
[0723] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0724] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0725] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0726] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0727] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0728] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0729] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0730] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0731] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0732] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0733] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0734] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0735] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0736] The following systems are conceivable as embodiments for carrying out the present invention.
[0737] First, the user uploads a file from their device to the system. This file contains digital data of the manual, and the server checks the format of the received file before processing it. If it is in a supported format, it proceeds to the next step.
[0738] Next, the server extracts text data from the received file. Once the text data is extracted, it is structured hierarchically. At this stage, the data is organized by category and headings so that users can easily search for it.
[0739] The server then feeds this structured data to a generative model. This generative model uses AI to learn from the supplied data. This prepares the model to generate appropriate answers to queries.
[0740] Next, the server uses the trained generative model to build an FAQ chatbot. This chatbot is programmed to receive questions from users and generate answers in real time.
[0741] Through a conversational interface accessible from the terminal, users input questions in natural language. The server receives these questions, uses a generative model to generate appropriate answers, and sends them to the user's terminal, enabling rapid information delivery.
[0742] As a concrete example, consider a question about how to operate a new device. When a user asks the chatbot on their device, "How do I set up this device?", the server uses a generative model to find the relevant data and provides a specific answer such as, "Please follow these steps for the initial setup."
[0743] This system allows users to quickly access important information, eliminating the need to read the entire manual. This is an embodiment of the present invention.
[0744] The following describes the processing flow.
[0745] Step 1:
[0746] The user selects a manual file from their terminal and uploads it. The uploaded file is sent to the server.
[0747] Step 2:
[0748] The server checks the received file and verifies that it is in a supported format. If the format is not appropriate, it returns an error message to the user.
[0749] Step 3:
[0750] The server uses OCR (Optical Character Recognition) or other parsing techniques to extract text data from the appropriate files.
[0751] Step 4:
[0752] The server organizes and hierarchically structures the extracted text data. This process classifies and structures the data by headings and paragraphs.
[0753] Step 5:
[0754] The server supplies the generated hierarchical structured data to the generative model and performs training. This training builds the model's ability to answer future questions.
[0755] Step 6:
[0756] The server uses a trained generative model to build an FAQ chatbot, enabling automated responses to user questions.
[0757] Step 7:
[0758] Users access the chatbot through their device and input the information they want to obtain in natural language.
[0759] Step 8:
[0760] The server sends the question received from the user to a generative model, which generates the optimal answer. The generated answer is then returned to the user's device and displayed.
[0761] Through this series of steps, users can easily and quickly access the information in the manual.
[0762] (Example 1)
[0763] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0764] This invention aims to improve a system that quickly and accurately retrieves useful information from files containing large amounts of data and provides appropriate answers to user inquiries. Furthermore, it seeks to solve the problem of improving the accuracy of answers by utilizing user feedback and enhancing the comprehensiveness of information provision by integrating data from diverse information sources.
[0765] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0766] In this invention, the server includes means for receiving information and confirming its structure, means for extracting and hierarchizing data from the received information, and means for providing the hierarchical data to a learning model for training. This makes it possible to provide quick and accurate answers to user inquiries. Furthermore, by utilizing feedback, it is possible to improve the accuracy of the system and provide comprehensive information through information integration.
[0767] "Means of receiving information and verifying its structure" refers to the process of taking in information provided from external sources and checking whether that information conforms to a specific format or standard.
[0768] "Methods for extracting and hierarchizing data" refers to the process of extracting necessary items from received information and organizing that data based on a logical order or classification.
[0769] "Means of providing and training a learning model" refers to the process of inputting prepared data into an artificial intelligence model, allowing that model to learn patterns and knowledge from the data.
[0770] "Methods for constructing dialogue programs" refers to the process of developing a system that uses a trained artificial intelligence model to automatically communicate with users and generate appropriate answers to questions.
[0771] "A means of providing appropriate answers to user inquiries" refers to a function that analyzes questions received from users, searches for relevant information, and provides responses.
[0772] "Methods for collecting evaluations, using that information to improve the learning model, and enhancing the accuracy of responses" refers to methods of collecting feedback from users, retraining the model based on that feedback, and improving the system's performance.
[0773] "Means of connecting diverse information sources to provide comprehensive information" refers to methods of aggregating data from multiple different databases and information sets to provide users with consistent and detailed information.
[0774] To implement this invention, the system is primarily operated through the collaboration of three parties: a server, a terminal, and a user. The specific method is described below.
[0775] First, users upload digital information to the system using their own devices. These devices are expected to be computers or smartphones. Examples of uploaded information include product manuals and technical documents.
[0776] The server automatically analyzes the format and structure of the received digital information. This analysis can utilize software that recognizes specific file formats, such as PDF analysis tools or OCR (optical character recognition) technology. Based on the analysis results, the server extracts the information as text and organizes the data structurally. This organization process involves completing a hierarchical format using a database management system or similar method.
[0777] The organized data is fed into a generative AI model, and the learning process begins. An example of a generative AI model is a natural language processing model using open-source libraries. This model learns from the supplied data and prepares answers to subsequent user inquiries.
[0778] The completed model is implemented as an interactive program by the server. Users can use the terminal's interactive interface to input questions in natural language. For example, they can ask, "Please tell me the initial setup procedure for this product." This question is processed by the server, and the optimal answer is generated based on the generative model. The answer is then immediately sent to the user's terminal.
[0779] This system allows users to obtain information instantly and accurately. Furthermore, user questions and feedback are recorded on the server, and the model is regularly updated based on this information to improve its accuracy.
[0780] An example of a prompt would be, "Please tell me how to troubleshoot this." This would prepare the model to provide relevant solutions.
[0781] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0782] Step 1:
[0783] Users upload digital files from their devices. These files include, for example, product manuals and technical documents. They click the "Select File" button on their device, choose the desired file, and submit it. The input is a digital file, and the output is the transfer of the file to the server.
[0784] Step 2:
[0785] The server verifies the format of the received file. Specifically, the server checks the file extension and internal metadata to determine if it is a supported format (e.g., PDF, DOCX). The input for this step is the uploaded file, and the output is the result of the verification that the file is in the correct format.
[0786] Step 3:
[0787] The server extracts text data from the file. If the format is supported, it extracts the text using appropriate software (e.g., PDF analysis tools or OCR technology). The input for this step is a verified digital file, and the output is the extracted text data.
[0788] Step 4:
[0789] The server hierarchically structures the extracted text data. Using a database management system, it structures the data by section and topic. This process makes it easy for users to search for information. The input for this step is text data, and the output is hierarchically structured data.
[0790] Step 5:
[0791] The server supplies hierarchical data to a generative AI model for training. The model analyzes the supplied data and is trained to prepare appropriate answers to queries. The input for this step is hierarchical data, and the output is the trained generative model.
[0792] Step 6:
[0793] The server uses a trained generative model to build a conversational program (chatbot). This chatbot is designed to receive questions from users and generate answers in real time using the model. The input to this step is the trained generative model, and the output is a working chatbot.
[0794] Step 7:
[0795] The user enters a question through the terminal's interactive interface. For example, they might ask, "Please tell me how to set up this product." In this step, the input is the user's question, and the output is the sending of the question to the server.
[0796] Step 8:
[0797] The server analyzes the received question using a generative model and generates an appropriate answer. Based on the analysis results, it searches for relevant information, generates an answer, and sends it to the user's terminal. The input for this step is the user's question, and the output is the generated answer.
[0798] (Application Example 1)
[0799] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0800] Currently, many physical stores are required to provide customers with prompt information regarding product details and service information in response to their inquiries. However, there is a lack of efficient methods to achieve this, which can lead to decreased customer satisfaction. The objective of this invention is to solve this problem by providing a system that provides customers with the information they need quickly and accurately in physical stores.
[0801] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0802] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for constructing an interactive program that generates answers to questions using a generative model, means for users to access product information using identification information and for presenting appropriate answers to questions, and means for acquiring identification information through the user terminal and redirecting to the relevant information based on that information. This enables rapid information provision to customers in physical stores.
[0803] "Means for receiving files and verifying their format" refers to a function that receives digital data sent from a user to the system and determines whether the data format is supported by the system.
[0804] "A means of extracting text data from received files and structuring it hierarchically" refers to a function that extracts textual information from transmitted digital data and organizes and structures that information by item or heading.
[0805] "A means of building an interactive program that generates answers to questions using generative models" refers to a function that utilizes machine learning techniques to create a program that generates appropriate responses to user questions.
[0806] "A means for users to access product information using identification information and to be presented with appropriate answers to questions" refers to a function that allows users to access a digital system using a specific ID or code and automatically answers questions based on detailed information of the selected product.
[0807] "A means of acquiring identification information through the user's terminal and redirecting them to relevant information based on that information" refers to a function that collects identification data supplied via the user's device and uses that data to guide the user to related content or information.
[0808] To implement this invention, the following system is necessary. Specifically, a system is constructed that utilizes the user's terminal and a server installed in the store to provide product information and generate quick responses to user inquiries.
[0809] The server receives manual files uploaded by users. Next, it verifies the file format and extracts and hierarchically structures the text data. Natural language processing libraries (e.g., TensorFlow or PyTorch) are used for this text data processing. The server then feeds the organized data to a machine learning model, training a generative AI model to respond to user questions. This model forms the basis for real-time generated interactive programs.
[0810] Users can use their smartphones or other devices to scan QR codes attached to products in physical stores. This sends identifying information about the relevant product or service to a server. The server then uses a generative model based on this information to instantly provide appropriate answers to user questions. This entire process enhances the user's in-store shopping experience and allows them to receive quick and detailed information.
[0811] As a concrete example, consider a case where a user scans a QR code for a specific product—for example, a camera—and asks about its initial setup. In this case, the server inputs the prompt "How do I set up the camera?" into the AI model and provides detailed setup instructions to the user's device.
[0812] This system allows customers to instantly obtain the information they need without any hassle, effectively complementing in-store customer service.
[0813] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0814] Step 1:
[0815] The user scans the QR code included with the product using their device.
[0816] Input: QR code data
[0817] Output: Identification information
[0818] Specific operation: The device uses its camera function to read the QR code, analyzes its contents, and obtains specific identification information.
[0819] Step 2:
[0820] The terminal sends identification information to the server.
[0821] Input: Identification Information
[0822] Output: Formal request to the server
[0823] Specific operation: The terminal sends identification information to the server via the network and requests information related to the product.
[0824] Step 3:
[0825] Based on the identification information received by the server, it searches for and retrieves the relevant manual files.
[0826] Input: Identification Information
[0827] Output: Manual file
[0828] Specific operation: The server refers to an internal database and searches for and retrieves the digital manual corresponding to the specified identification information.
[0829] Step 4:
[0830] The server extracts text data from the acquired manuals and structures it hierarchically.
[0831] Input: Manual file
[0832] Output: Structured text data
[0833] Specific operation: The server extracts necessary text from the scanned manual using natural language processing and organizes it by headings and items.
[0834] Step 5:
[0835] The server supplies structured data to the generative model and prepares it for response generation.
[0836] Input: Structured text data
[0837] Output: Feed to the generative model
[0838] Specific operation: The server inputs data into the AI model and performs the necessary training to generate appropriate answers to user questions.
[0839] Step 6:
[0840] The user sends the question to the server via their device.
[0841] Input: User Question
[0842] Output: Query to the server
[0843] Specific operation: The user uses the app's interface to input questions about the product in natural language and sends them to the server.
[0844] Step 7:
[0845] The server uses a generative model to generate an answer to the question and sends it to the terminal.
[0846] Input: User Question Query
[0847] Output: Answer to the question
[0848] Specific operation: The server analyzes the received question using a generative AI model, generates an answer based on the trained data, and sends it to the user's terminal.
[0849] Step 8:
[0850] The terminal displays the response received from the server to the user.
[0851] Input: Response from the server
[0852] Output: Answer displayed on the user interface
[0853] Specific operation: The terminal displays the response provided by the server on the screen, providing the user with detailed information.
[0854] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0855] In embodiments of the present invention, the process begins with a user uploading a file from their terminal. This file contains digital data of a manual, and the server receives the file and verifies its format. If it is determined to be in a supported format, the server extracts text data from the file and structures it hierarchically. This process organizes the data by category, making it easy to search.
[0856] The hierarchically structured data is fed to a generative model by the server, where it is trained. This training enables the model to generate appropriate answers to user inquiries. Next, the server utilizes the trained model to build a chatbot. This chatbot has the functionality to respond to user questions and provides information through dialogue.
[0857] Furthermore, this invention incorporates an emotion engine. This engine has the function of recognizing emotions from the user's input and conversation tone. The chatbot determines the user's emotions using the emotion engine and adjusts the tone and content of its response based on that information.
[0858] As a concrete example, consider a scenario where a user asks for instructions on how to use a product for the first time and feels anxious. They might consult a chatbot via their device, saying, "This is my first time using this device, so I'm worried about whether I've set it up correctly." In this case, the server utilizes an emotion engine to recognize the user's anxiety and provides a reassuring response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence." This approach improves the user experience and enhances the quality of information provided.
[0859] The above describes an embodiment of the present invention that combines an emotion engine, enabling the construction of a system that provides users with information quickly and with consideration for their emotions.
[0860] The following describes the processing flow.
[0861] Step 1:
[0862] The user selects a manual file from their terminal and uploads it to the system. This file is a digitized version of the manual.
[0863] Step 2:
[0864] The server receives the uploaded file and checks its format. If the file is in a supported format, it proceeds to the next step.
[0865] Step 3:
[0866] The server extracts text data from the received file. During this process, OCR and file parsing technologies are used to accurately extract the necessary text information.
[0867] Step 4:
[0868] The server organizes the extracted text data into a hierarchical structure. It classifies the data by item and heading, making it easy for users to quickly access the information they need.
[0869] Step 5:
[0870] The server supplies hierarchically structured data to a generative model, which is then trained. This training enables the model to generate appropriate answers to questions.
[0871] Step 6:
[0872] The server builds an FAQ chatbot based on a pre-trained generative model. This chatbot provides answers to questions in real time.
[0873] Step 7:
[0874] Users access the chatbot through their device and input questions in natural language. They can ask questions in a conversational format.
[0875] Step 8:
[0876] The server analyzes user input and uses an emotion engine to determine the user's emotions. By extracting emotion data from the input, it understands the user's feelings.
[0877] Step 9:
[0878] The server uses a generative model to generate responses with a tone adjusted according to the user's emotions. It then sends the generated responses to the user's device.
[0879] Step 10:
[0880] Users can review the answers on their device and ask additional questions to the chatbot as needed. Through this process, users can obtain information smoothly. This system enables the delivery of more personalized and emotionally sensitive information to users.
[0881] (Example 2)
[0882] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0883] In modern information technology, there is a need for systems that can respond quickly and accurately to user inquiries. However, conventional systems struggle to generate responses that take user emotions into account, and there is a demand for technologies that can achieve more human-like interaction. Furthermore, there is a lack of mechanisms to improve response accuracy by utilizing feedback, and data integration to provide comprehensive information is insufficient.
[0884] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0885] In this invention, the server includes means for receiving a file and verifying its format, means for extracting information from the received file and structuring it hierarchically, means for supplying the hierarchically structured information to a generative model for training, and means for providing an appropriate response to a user inquiry. This enables the generation of responses that take into account the user's emotional perception, improvement of response accuracy through feedback, and comprehensive information provision through the integration of multiple information sources.
[0886] "Receiving a file" refers to the act of a server retrieving digital data provided by a user from their device.
[0887] "Verifying the format" is the process of verifying whether the received digital data is in a specific, compatible data format.
[0888] "Information extraction" is the process of extracting necessary text data and related information from received digital data.
[0889] "Hierarchical structuring" refers to classifying extracted information into chapters or items and organizing them hierarchically.
[0890] A "generative model" is an algorithm that is trained using large amounts of digital data and generates responses to queries through natural language processing.
[0891] "Training" refers to the process of improving the accuracy of a generative model's responses by using input data to learn from it.
[0892] An "automated response system" is software or a system that provides a mechanism for the system to automatically generate a response to a user's inquiry.
[0893] "Identifying emotions" is the process of analyzing user input and context to detect the emotions contained within it.
[0894] "Adjusting responses" refers to changing the content and tone of responses provided to the user based on identified emotional information.
[0895] "Integrating information sources" is the process of aggregating data from multiple related databases and information streams and structuring it into a single, comprehensive piece of information.
[0896] In implementing this invention, the process begins with a user uploading a file containing digital data using their own device. This file contains information such as product manuals and instructions. Once the user uploads the file, the server receives the data. The server first checks the format of the received file, verifying, for example, whether it is in PDF or DOCX format. This verification process can be automated using open-source libraries or proprietary programs.
[0897] If the server determines that the file is in a supported format, it begins data extraction. For example, it uses tools such as Apache Tika or Python's PDFMiner to extract the necessary text data from the file. Because the extracted data is difficult to handle directly, the server then performs a step to structure the data hierarchically based on chapters and headings. This organizes the information by category, making subsequent searches easier.
[0898] The hierarchically structured information is then fed to the generative AI model. The server trains the model using open-source machine learning libraries such as TensorFlow and PyTorch. This training enables the model to generate fast and appropriate responses to real-time user inquiries.
[0899] Using a pre-trained generative AI model, the server builds a chatbot that automatically generates responses to user questions. Furthermore, the invention integrates an emotion engine, allowing the server to identify emotions from user input. This enables the adjustment of response tone based on emotional information obtained through data analysis.
[0900] As a concrete example, consider a scenario where a user is unsure about the initial setup of a product. When the user enters a prompt message through their device such as, "This is my first time using this device, so I'm worried about whether I've set it up correctly," the server uses an emotion engine to detect the user's anxiety. As a result, the chatbot generates a response such as, "Don't worry. Detailed setup instructions are provided below, so please proceed with confidence," and provides it to the user. This system allows users to receive thoughtful and emotionally sensitive information.
[0901] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0902] Step 1:
[0903] The user uploads a digital data file to the server using a terminal. The input is a digital data file, and this process saves the file to the server. The server receives the uploaded file and temporarily stores it for processing in the next step.
[0904] Step 2:
[0905] The server checks the format of the uploaded file. Specifically, it verifies whether the file is in a supported format such as PDF or DOCX. The input is the uploaded file, and the output is the result of the file format verification. If the format is not appropriate, the server sends an error message to the user and prompts them to re-upload the file.
[0906] Step 3:
[0907] The server extracts information from files in the correct format. This extraction process uses tools such as Apache Tika or PDFMiner. The input is a formatted file, and the output is extracted text data. The server stores the extracted text data in memory and proceeds to the next structuring step.
[0908] Step 4:
[0909] The server hierarchically structures the extracted text data. The data is primarily organized based on document chapters and headings. The input is extracted text data, and the output is hierarchically structured information. This hierarchical structure organizes the information, making it easier to search and use later.
[0910] Step 5:
[0911] The server supplies hierarchically structured information to a generative AI model for training. This process involves training the generative model with new data patterns to improve the accuracy of response generation. The input is hierarchically structured information, and the output is the trained model. Using this model enables accurate and appropriate responses.
[0912] Step 6:
[0913] The server uses a trained generative AI model to build a chatbot. This automated response system has the ability to instantly respond to user prompts. The input is the user's question, and the output is an automatically generated response based on the model.
[0914] Step 7:
[0915] The server integrates an emotion engine, adding the ability to identify the user's input emotions. This system analyzes the emotions the user expresses and adjusts the response accordingly. The input is the user's response, including the prompt, and the output is the response adjusted based on the emotion identification.
[0916] Through these steps, the server enables the provision of advanced information to improve the user experience.
[0917] (Application Example 2)
[0918] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0919] In systems that provide advanced information, users often find it difficult to easily understand operation manuals and setup procedures, which can be particularly psychologically stressful and anxiety-inducing. This invention aims to reduce user anxiety during the information provision process and enable smooth information access by adjusting responses to take into account the user's emotional state.
[0920] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0921] In this invention, the server includes means for receiving a file and verifying its format, means for extracting text data from the received file and structuring it hierarchically, means for supplying the hierarchically structured data to a generative model and training it, means for constructing a dialogue system that generates answers to questions using the generative model, and means for analyzing the emotions of the user in response to a question, adjusting the answer, and presenting an appropriate answer. This enables effective information provision that takes into account the user's psychological state.
[0922] A "file" is a collection of data used to store digital information and process it within a computer system.
[0923] "Means of checking format" refers to a function that determines whether the received data conforms to the specifications that the system can process.
[0924] "Text data" refers to digital information that includes string information, and is data recorded in the form of sentences or words.
[0925] "Means of hierarchical structuring" refers to the function of organizing information structurally and clarifying the relationships between each element.
[0926] A "generative model" is an algorithm that uses machine learning techniques to generate new information or results based on input data.
[0927] "Training methods" refer to the process of using data to train a generative model and improve its accuracy and capabilities.
[0928] A "dialogue system" is a computer-based conversation management system that provides information and generates responses based on input from users.
[0929] A "user" is an individual or organization that interacts with the system and receives or provides information.
[0930] "Means of analyzing emotions" refers to a function that identifies and evaluates a user's psychological state based on input data.
[0931] "Methods for adjusting responses" refer to the process of optimizing the content and expression of generated responses based on analyzed emotional information.
[0932] Modes for carrying out the invention
[0933] To implement this invention, a server plays a central role. The server receives files uploaded by users using their terminals. After receiving the files, the server uses means to verify that their format conforms to predetermined specifications. Subsequently, text data is extracted from the received files, and the information is organized through hierarchical structuring. This hierarchical structuring clarifies the relationships between data, facilitating subsequent processing.
[0934] Next, the server supplies the hierarchically structured data to the generative model for training. The machine learning framework "TensorFlow" is used to train the generative model. As a result, the trained generative model has the ability to function as a dialogue system. This dialogue system provides information and generates responses based on user input.
[0935] Furthermore, the server uses an emotion analysis engine to analyze the input data received from the user and evaluate the user's emotional state. The "Hugging Face Transformers" library is utilized for emotion analysis. Based on this emotional information, the server optimizes the generated response by adjusting its content and tone.
[0936] As a concrete application example, consider a scenario where a first-time user of an autonomous vehicle asks, "I'm worried about operating the autonomous driving mode." In this case, the server uses sentiment analysis to identify the user's anxiety and provides a reassuring response such as, "The autonomous driving mode has been confirmed to be safe. Please follow the steps below to use it with confidence." Through this process, the user can access accurate information to safely operate the vehicle system.
[0937] An example of a prompt message is, "Generate a method for providing vehicle information to the server that takes user sentiment into consideration."
[0938] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0939] Step 1:
[0940] The server receives files sent from the terminal. It parses the received file to determine if it is in a supported format. The input is a digital file uploaded by the user via the terminal, and the output is a boolean value indicating whether the format is valid.
[0941] Step 2:
[0942] The server extracts text data from verified digital files. The extracted text data is then organized into higher and lower levels to create a hierarchical structure. The input is raw data extracted according to a verified file format, and the output is hierarchically structured data.
[0943] Step 3:
[0944] The server supplies hierarchically structured data to a generative model using the machine learning framework "TensorFlow" and trains the model. The input is hierarchically structured data, and the output is a trained generative model.
[0945] Step 4:
[0946] When a user enters a question using a terminal, the server receives the question. Using a trained generative model, it generates an appropriate answer to the question. The input is a question from the user in natural language, and the output is the generated answer.
[0947] Step 5:
[0948] The server uses the emotion analysis engine "Hugging Face Transformers" to evaluate the user's emotional state from the context of the question. The input is the user's question text, and the output is an identified emotional state (e.g., anxiety, joy).
[0949] Step 6:
[0950] The server adjusts the content and expression of the generated response based on the emotion assessment results and provides it to the user in the most appropriate form. The input is the emotional state analyzed and the generated response, and the output is the adjusted response.
[0951] Step 7:
[0952] The server responds to the user's question by sending a pre-configured answer to the terminal and presenting it to the user. The input is the optimized answer, and the output is the final information provided to the user.
[0953] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0954] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0955] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0956] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0957] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0958] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0959] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0960] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0961] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0962] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0963] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0964] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0965] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0966] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0967] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0968] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0969] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0970] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0971] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0972] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0973] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0974] The following is further disclosed regarding the embodiments described above.
[0975] (Claim 1)
[0976] A means of receiving a file and verifying its format,
[0977] A method for extracting text data from a received file and structuring it hierarchically,
[0978] A means of supplying hierarchically structured data to a generative model and performing training,
[0979] A means of building a chatbot that generates answers to questions using a generative model,
[0980] A system that includes means for providing appropriate answers to user questions.
[0981] (Claim 2)
[0982] The system according to claim 1, comprising means for recording user feedback via a chatbot, updating a generative model using that information, and improving the accuracy of responses.
[0983] (Claim 3)
[0984] The system according to claim 1, comprising means for integrating the generated hierarchical structured data with multiple data sources to enable comprehensive information provision.
[0985] "Example 1"
[0986] (Claim 1)
[0987] A means of receiving information and confirming its structure,
[0988] A means of extracting and hierarchically organizing data from received information,
[0989] A means of providing hierarchical data to a learning model and training it,
[0990] A means of constructing a dialogue program that generates answers to inquiries using a learning model,
[0991] A means of providing appropriate answers to inquiries from users,
[0992] A means of providing an interface for users to transmit information and recording that transmitted information,
[0993] A means of sending the generated response back to the user,
[0994] A system that includes this.
[0995] (Claim 2)
[0996] The system according to claim 1, comprising means for collecting user evaluations through the aforementioned dialogue program, improving the learning model using that information, and improving the accuracy of responses.
[0997] (Claim 3)
[0998] The system according to claim 1, comprising means for linking generated hierarchical data with diverse information sources to realize comprehensive information provision.
[0999] "Application Example 1"
[1000] (Claim 1)
[1001] A means of receiving a file and verifying its format,
[1002] A method for extracting text data from a received file and structuring it hierarchically,
[1003] A means of supplying hierarchically structured data to a generative model and performing training,
[1004] A means of constructing an interactive program that generates answers to questions using a generative model,
[1005] A means by which users can access product information using identification information and be presented with appropriate answers to their questions,
[1006] A means of obtaining identification information through the user's terminal and redirecting to the relevant information based on that information,
[1007] A system that includes this.
[1008] (Claim 2)
[1009] The system according to claim 1, comprising means for recording user feedback through an interactive program, updating a generative model using that information, and improving the accuracy of responses.
[1010] (Claim 3)
[1011] The system according to claim 1, comprising means for integrating the generated hierarchically structured data with multiple data sources to enable comprehensive information provision.
[1012] "Example 2 of combining an emotion engine"
[1013] (Claim 1)
[1014] A means of receiving a file and verifying its format,
[1015] A means of extracting information from a received file and structuring it hierarchically,
[1016] A means of supplying hierarchically structured information to a generative model and training it,
[1017] Means for constructing an automated response system that generates query responses using a trained generative model,
[1018] A means for identifying emotions from input information and adjusting responses based on that information,
[1019] A system that includes means for providing appropriate responses to inquiries from users.
[1020] (Claim 2)
[1021] The system according to claim 1, comprising means for recording user opinions via an automated response device, updating a generation model using that information, and improving response accuracy.
[1022] (Claim 3)
[1023] The system according to claim 1, comprising means for integrating generated hierarchical structured information with multiple information sources to enable comprehensive information provision.
[1024] "Application example 2 when combining with an emotional engine"
[1025] (Claim 1)
[1026] A means of receiving a file and verifying its format,
[1027] A method for extracting text data from a received file and structuring it hierarchically,
[1028] A means of supplying hierarchically structured data to a generative model and performing training,
[1029] A means of constructing a dialogue system that generates answers to questions using a generative model,
[1030] A system that includes means to analyze the emotions of users in response to their questions, adjust the answers accordingly, and provide appropriate responses.
[1031] (Claim 2)
[1032] The system according to claim 1, comprising means for recording user feedback via a dialogue system, updating a generative model using that information, and improving the accuracy of responses.
[1033] (Claim 3)
[1034] The system according to claim 1, comprising means for integrating the generated hierarchically structured data with multiple information sources to enable comprehensive information provision. [Explanation of Symbols]
[1035] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving a file and verifying its format, A method for extracting text data from a received file and structuring it hierarchically, A means of supplying hierarchically structured data to a generative model and performing training, A means of building a chatbot that generates answers to questions using a generative model, A system that includes means for providing appropriate answers to user questions.
2. The system according to claim 1, comprising means for recording user feedback via a chatbot, updating a generative model using that information, and improving the accuracy of responses.
3. The system according to claim 1, comprising means for integrating the generated hierarchical structured data with multiple data sources to enable comprehensive information provision.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A