System

The system addresses inefficiencies in manual-based troubleshooting by analyzing digital documents with OCR and NLP, generating solutions with generative AI, and providing visual aids, enhancing problem-solving efficiency.

JP2026017934APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118995
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Users face inefficiencies in resolving electronic device issues due to the time-consuming process of referring to manuals, which complicates quick problem resolution and often leads to cumbersome and inefficient troubleshooting.

Method used

A system that analyzes digital documents using optical character recognition (OCR) and natural language processing (NLP), generates solutions using a generative AI model, and presents them to users, including visual aids like images or diagrams.

Benefits of technology

Enables users to quickly and effectively solve problems without detailed manual reading, improving efficiency and understanding through visual and textual solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017934000001_ABST
    Figure 2026017934000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for analyzing a digital document of an instruction manual or a failure handling manual; means for receiving information on a problem or a phenomenon input from a user; means for extracting related information from the digital document; means including a generation model for generating a handling method based on the extracted information and input information from the user; and means for presenting the generated handling method to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, users use a wide variety of electronic devices, but referring to manuals for setting them up or troubleshooting them takes time and effort, making it difficult to respond quickly. Furthermore, many users often find the process of referring to manuals cumbersome, which can result in inefficient problem resolution. It is necessary to solve these issues and improve user convenience. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides a system including: means for analyzing a digital document such as an instruction manual or a troubleshooting manual; means for receiving information about a problem or phenomenon input by a user; means for extracting relevant information from the digital document; means for including a generative model for generating solutions based on the extracted information and the user's input information; and means for presenting the generated solutions to the user. Furthermore, the system includes means for converting the analyzed digital document into text using optical character recognition (OCR) technology, and also includes means for generating an image corresponding to the generated solution. This allows users to quickly and easily obtain solutions to problems, dramatically improving the efficiency of problem solving.

[0006] An "instruction manual" is a document that provides detailed instructions on how to set up, use, and maintain an electronic device.

[0007] A "failure response manual" is a document that describes how to identify the cause of an electronic device failure, repair procedures, and troubleshooting methods.

[0008] A "digital document" is a document that is stored in electronic form and can be accessed and viewed using a computer or electronic device.

[0009] "Analyzing" is the process of examining data in detail and extracting or retrieving specific information.

[0010] "User" means any person or entity using a system or electronic device.

[0011] "Information about the problem or phenomenon" refers to details of the trouble or malfunction the user is facing and information about specific symptoms.

[0012] "Receiving" refers to the process of taking in data or signals transmitted from an external source.

[0013] "Extraction" is the process of extracting specific elements from a large amount of data or information.

[0014] A "generative model" is a system that uses algorithms based on machine learning and artificial intelligence to generate new outputs based on input data.

[0015] "Generate" refers to the process by which a system creates a new result based on specified conditions and input data.

[0016] "Presenting" refers to the act of showing, displaying, or making information accessible to a user.

[0017] Optical character recognition (OCR) is a technology that scans printed or handwritten characters and converts them into text data.

[0018] "Converting to text" refers to the process of digitizing non-textual information, such as images or PDFs, into character data.

[0019] An "image diagram" is a diagram or image that visually represents the content or concept being explained. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] System Overview

[0042] The present invention provides a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual for electronic devices. This system can analyze the digital document of the instruction manual or troubleshooting manual and generate a solution to the problem or phenomenon input by the user.

[0043] composition

[0044] Users use their own devices (smartphones or PCs) to enter information about problems or issues.

[0045] The terminal transmits the user's input information to the server.

[0046] The server analyzes the digital document and extracts the required information.

[0047] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[0048] The server presents the generated solution and related image diagram to the user.

[0049] Explanation of program processing

[0050] The detailed operation of this system will be explained below.

[0051] Initial Setup

[0052] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[0053] The server converts the PDF file into text using optical character recognition (OCR) technology.

[0054] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[0055] Query reception

[0056] The user uses the device to input a description of the issue or problem in natural language, for example, "I can't connect to Wi-Fi on my smartphone."

[0057] The terminal sends the input query to the server.

[0058] Data Analysis and Information Extraction

[0059] The server receives the user's query and searches for relevant information in a database.

[0060] The server provides the search results to the generative model.

[0061] Generate a workaround

[0062] The generative model generates a solution based on the user's query and related information provided by the server.

[0063] For example, it generates specific solutions such as "Select the network name on the Wi-Fi settings screen and enter the password."

[0064] Generate image diagrams

[0065] The server generates an image related to the solution (for example, a screenshot of the setting screen).

[0066] Presenting solutions and illustrations

[0067] The server generated solution and image diagram are summarized on one page.

[0068] The server sends the compiled information to the terminal.

[0069] The terminal displays this to the user, allowing the user to easily resolve the problem.

[0070] Specific examples

[0071] Below is a specific example where the user enters "My TV won't turn on."

[0072] 1. The user types "TV won't turn on" into the device.

[0073] 2. The device sends this query to the server.

[0074] 3. The server analyzes the TV's instruction manual using OCR technology and extracts information related to the "won't turn on" problem.

[0075] 4. Based on the extracted information and user input, the generative model generates a solution such as "Make sure the power cord is connected correctly. Then, replace the batteries in the remote control."

[0076] 5. The server generates a diagram that visually explains how to resolve the issue (for example, a diagram showing where to connect the power cord).

[0077] 6. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[0078] 7. The device should display this to the user so that they can quickly and easily understand what to do.

[0079] This allows users to solve problems efficiently without having to read the instruction manual in detail.

[0080] The processing flow will be explained below.

[0081] Step 1:

[0082] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My TV won't turn on."

[0083] Step 2:

[0084] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[0085] Step 3:

[0086] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[0087] Step 4:

[0088] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance in storage, and uses OCR technology to convert the PDFs into text data.

[0089] Step 5:

[0090] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[0091] Step 6:

[0092] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[0093] Step 7:

[0094] The server inputs the extracted information and the user's input data into a generative AI model (e.g., ChatGPT) to generate an appropriate response. The generative AI model generates the most appropriate response in natural language based on the input data.

[0095] Step 8:

[0096] The server generates specific images (e.g., screenshots of power cord connection locations or settings screens) based on the generated solutions, using automatic image generation technology and images retrieved from a database.

[0097] Step 9:

[0098] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[0099] Step 10:

[0100] The server then sends the final information to the device, including specific solutions and related images.

[0101] Step 11:

[0102] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem, thus enabling the user to effectively respond without having to refer to an instruction manual or instruction manual.

[0103] Example 1

[0104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0105] In the past, when users wanted to refer to instruction manuals or troubleshooting manuals for electronic devices, they had to refer to paper documents or search websites, which was time-consuming and laborious. Furthermore, it was difficult to quickly obtain specific solutions, which reduced the efficiency of problem-solving. In particular, when dealing with electronic device malfunctions, accurate and prompt information provision is required.

[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0107] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for converting the digital document into text using optical character recognition technology, means for extracting relevant information from the digital document, means for searching for relevant information in a database, means for including a generative AI model that generates solutions based on the extracted information and information input by the user, means for generating an image corresponding to the generated solution, and means for presenting the generated solution and the image to the user. This allows the user to solve problems quickly and effectively without having to read the instruction manual in detail.

[0108] An "instruction manual or troubleshooting manual" is a document that describes how to use an electronic device and what to do if a problem occurs.

[0109] A "digital document" is a document in a form that is stored, displayed, and transmitted electronically.

[0110] "Means for analysis" refers to technology or devices that have the ability to extract necessary information from digital documents.

[0111] "Means for receiving information about a problem or phenomenon input by a user" refers to a technology or device that has the function of transmitting information input by a user through a terminal to a server and receiving that information.

[0112] "Optical Character Recognition (OCR)" is a technology that converts characters in an image into electronic text.

[0113] A "database" is a system for efficiently storing, retrieving, and managing structured data.

[0114] A "natural language processing engine" is an artificial intelligence technology that analyzes text data, extracts semantic information, and understands language.

[0115] A "generative AI model" is an artificial intelligence model that generates useful information, such as countermeasures, based on input information from users and related data.

[0116] An "image" is a visual illustration that supplements text information.

[0117] An "HTTP request" is a communication protocol used by a web browser or application to request data from a server.

[0118] A "server" is a computer system that processes queries, stores data, and communicates over a network.

[0119] A "terminal" is a device that a user uses to access a server, such as a smartphone or PC.

[0120] A "prompt" is a text question or instruction used as input to a generative AI model.

[0121] A "means for converting to text" is a technique or device that converts characters in an image into electronic text using optical character recognition technology.

[0122] System Overview

[0123] This invention is a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual of the electronic device. Specifically, it analyzes digital documents using optical character recognition (OCR) technology and a natural language processing engine (NLP), generates solutions using a generative AI model, and presents them to the user.

[0124] Hardware and software used

[0125] Server: A computer system for storing, analyzing, and storing digital documents in a database.

[0126] Storage engine: Amazon S3, etc.

[0127] OCR technology: Tesseract OCR

[0128] Database: MySQL

[0129] Natural language processing engine: spaCy

[0130] Generative AI models: ChatGPT, etc.

[0131] Terminal: A device through which a user accesses a server and enters queries.

[0132] Devices: Smartphones, PCs, etc.

[0133] Program processing explanation

[0134] Information storage and analysis

[0135] The server stores PDF files of instruction manuals and troubleshooting manuals in a storage engine, for example, in an Amazon S3 bucket.

[0136] The server converts the saved PDF file to text using Tesseract OCR, specifically by running the command tesseract manual.pdf output.txt.

[0137] The server stores the converted text data in a MySQL database, for example by executing the SQL statement INSERT INTO manuals (text_data) VALUES ('...').

[0138] The server analyzes the text data using the spaCy NLP engine to extract semantic information, for example by running the following code: nlp = spacy.load('en_core_web_sm').

[0139] Query reception and processing

[0140] The user inputs a symptom or problem in natural language through the terminal. For example, the user inputs "The TV won't turn on."

[0141] The device sends this query to the server via an HTTP POST request.

[0142] Data analysis and generation of action plans

[0143] The server receives the user's query and searches the database for relevant information using an SQL query, for example SELECT FROM manuals WHERE text_data LIKE '%cannot power on%'.

[0144] The server sends the relevant information to a generative AI model (e.g., ChatGPT) via an API request. An example prompt sentence is "The user explains that the TV won't turn on. What should I do?"

[0145] The generative AI model generates a solution and sends it back to the server, such as "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0146] Presenting solutions and illustrations

[0147] The server uses the Pillow library to generate an image related to the solution, for example, a diagram showing where to connect the power cord.

[0148] The server integrates the generated solutions and images into an HTML template and compiles them into a single page.

[0149] The server sends the compiled information to the terminal as an HTTP response.

[0150] The terminal displays the information to the user and provides a quick and effective way to deal with the problem.

[0151] Specific examples

[0152] For example, if the user inputs "TV won't turn on," the system will operate as follows:

[0153] 1. The user types "TV won't turn on" into the device.

[0154] 2. The device sends the query to the server via an HTTP POST request.

[0155] 3. The server receives the query and searches the database for relevant information using an SQL query.

[0156] 4. The server sends the relevant information to the generative AI model via an API request.

[0157] 5. The generative AI model generates a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0158] 6. The server generates a diagram showing where the power cords are connected.

[0159] 7. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[0160] 8. The device displays information to the user and provides a quick and effective way to deal with the problem.

[0161] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0162] Step 1: Initial Setup

[0163] The server stores PDF files of instruction manuals and troubleshooting manuals in the storage engine. As a concrete example, using Amazon S3, the command aws s3 cp manual.pdf s3: / / bucket-name / is executed as input, and the output is the PDF file stored in Amazon S3.

[0164] The server converts PDF files to text using Tesseract OCR. Specifically, it receives input from the command tesseract manual.pdf output.txt and generates an output file called output.txt.

[0165] The server stores the converted text data in a MySQL database. For example, it takes input as input and executes the SQL statement INSERT INTO manuals (text_data) VALUES ('...'), and gets output as text data stored in the database.

[0166] Step 2: Enter and accept user queries

[0167] The user inputs the phenomenon or problem in natural language using the terminal. For example, the input "The TV won't turn on" is received, and the output is the manually input data in text format.

[0168] The terminal sends the entered query to the server. Specifically, it receives input to send the entered query to the server via an HTTP POST request, and obtains output showing that the query has reached the server.

[0169] Step 3: Data analysis and information extraction

[0170] The server receives a user query and searches for relevant information in the database. Specifically, it receives input by executing the SQL query SELECT FROM manuals WHERE text_data LIKE '%TV won't turn on%' and obtains output that provides relevant information (for example, the contents of the corresponding instruction manual).

[0171] The server provides the relevant information to the generative AI model, taking the specific action of sending the results of a database search to the generative AI model via an API request as input, and obtaining an output ready for the generative AI model to generate a response.

[0172] Step 4: Generate a solution

[0173] The generative AI model generates a solution based on the user's query and information provided by the server. Specifically, the model receives a prompt, "The user explains that the TV won't turn on. What should I do?", and outputs a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0174] The server receives the generated solutions and performs filtering. Specifically, it receives the generated text, deletes unnecessary parts, and outputs solutions that are ready to be provided to the user.

[0175] Step 5: Generate an image

[0176] The server uses the Pillow library to generate an image diagram related to the solution. For example, for the command "Please check the power cord connection," input is made using an image editing tool, and an output is generated showing the power cord connection location.

[0177] Step 6: Present solutions and images

[0178] The server then combines the generated solutions and the resulting image onto a single page. Specifically, it uses an HTML template as input to integrate the solutions and the image, and outputs a single page that can be presented to the user.

[0179] The server sends the compiled information to the terminal. As a specific input, it performs the operation of sending the generated page as an HTTP response, and obtains the output that the page arrives at the user terminal.

[0180] The device displays the information to the user. Specifically, the device inputs the received page to display in a browser or dedicated app, and the user receives an output that allows them to visually confirm the solution to the problem.

[0181] (Application example 1)

[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0183] The present invention relates to a system for quickly and effectively resolving problems without referring to instruction manuals or troubleshooting manuals. However, conventional systems require users to accurately input details of the problem, which can lead to errors, especially when using voice input. Furthermore, solutions are presented only in text, which can be difficult for users to understand. Furthermore, it is difficult for factory workers to respond to equipment failures without visual information. It is necessary to solve these problems.

[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0185] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, means for converting voice input into text, and means for generating a 3D model or diagram of the solution. This allows the user to input a problem using voice input, and the solution is presented not only as text but also as a 3D model or diagram, enabling the user to solve the problem quickly and visually.

[0186] An "instruction manual" is a document that provides detailed instructions on how to use and operate an electronic device or machine.

[0187] A "troubleshooting manual" is a document that explains how to repair and resolve problems when electronic devices or machinery break down.

[0188] "Digital documents" are document data stored electronically, including PDFs and text files.

[0189] "Means of analysis" refers to the technology and devices used to read digital documents and understand and classify the necessary information.

[0190] "Means for receiving" refers to technology or devices for receiving information or data input by a user.

[0191] An "extraction means" is a technique or device used to identify and extract the required information from a digital document.

[0192] A "generative model" is an artificial intelligence model that automatically generates appropriate solutions or answers based on input information.

[0193] The "presentation means" refers to a technique or device for displaying the generated solutions and information to the user.

[0194] "Means for converting voice input into text" refers to technology or devices for converting voice input by a user into text data.

[0195] "Means for generating 3D models or diagrams of solutions" refers to technology or devices for expressing solutions in 3D models or figures to make them visually easier to understand.

[0196] "Server" means a computer system for processing, storing, and managing data.

[0197] This invention is a system that can quickly and effectively solve problems without the user having to refer to an instruction manual or a troubleshooting manual. This system is particularly useful for effectively dealing with factory robot failures. The specific system configuration and its operation will be described below.

[0198] System Configuration

[0199] 1. User Device

[0200] Wearable devices such as smart glasses.

[0201] It has voice input and display functions.

[0202] 2. Server

[0203] A cloud server for analyzing and storing digital documents.

[0204] It has high-performance computing resources and data storage.

[0205] 3. OCR technology

[0206] Convert PDF to text data using optical character recognition technology. Specifically, we use pytesseract.

[0207] 4. Voice Recognition Technology

[0208] It uses technology to convert voice input into text data, specifically, the speech_recognition library.

[0209] 5. Generative AI Models

[0210] It is used for natural language processing and countermeasure generation. Specifically, it uses the GPT-3 model from the transformers library.

[0211] 6. Digital Documents

[0212] PDF documents such as instruction manuals and troubleshooting manuals.

[0213] Data processing and calculation

[0214] 1. Digital Document Analysis

[0215] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using optical character recognition (OCR). This is where pytesseract comes in.

[0216] The converted text data is stored in a database on the server.

[0217] 2. Processing voice input

[0218] The user inputs the symptoms or problems by voice through the smart glasses.

[0219] The smart glasses convert the voice data into text data and send this data to the server, where the speech_recognition library is used.

[0220] 3. Extracting information and generating solutions

[0221] The server analyzes the received user query and searches for relevant information in a database.

[0222] Based on the search results, a generative model (e.g., GPT-3) is used to generate a solution. The transformers library is used.

[0223] 4. Providing solutions and visual information

[0224] Use software to convert the generated solutions into 3D models and diagrams.

[0225] The solution and 3D models or diagrams are sent to the smart glasses and visually presented to the user.

[0226] Specific examples

[0227] Example 1: Error handling in an automobile factory

[0228] User input: "The robot arm joints won't move."

[0229] Generated solution: "Oil the joints of the robot arm and check their operation. Also, check for any sensor error indications."

[0230] Prompt Sentence Examples

[0231] Problem: The robot arm joints won't move

[0232] How to respond:

[0233] This allows factory workers to quickly identify problems and respond effectively. The system is expected to improve work efficiency and reduce downtime through visual and audio assistance.

[0234] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0235] Step 1:

[0236] Digital document analysis

[0237] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using OCR technology. The server reads the PDF using pytesseract and converts each page into text data.

[0238] Input: PDF format instruction manual or troubleshooting manual

[0239] Output: Text data

[0240] Specific behavior: The server extracts text from the PDF and stores it in a database.

[0241] Step 2:

[0242] Receiving audio input

[0243] The user inputs the symptoms or problems by voice through the smart glasses.

[0244] The smart glasses capture voice data and convert it to text data using the speech_recognition library.

[0245] Input: Audio data of the phenomenon or problem

[0246] Output: Text data

[0247] How it works: The user speaks, "The robot arm's joints won't move," and the smart glasses convert the speech into text.

[0248] Step 3:

[0249] Submitting a query

[0250] The smart glasses send the converted text data to the server.

[0251] Input: Text data converted from speech

[0252] Output: Text data sent to the server

[0253] Specific operation: The smart glasses generate text data and send it to a server via the network.

[0254] Step 4:

[0255] Information extraction and analysis

[0256] The server analyzes the received text data (user queries) and searches for relevant information in the database using a natural language processing (NLP) engine.

[0257] Input: User query text data

[0258] Output: Search results containing relevant information

[0259] Specific operation: The server extracts information related to the query from instruction manuals and manuals in the database.

[0260] Step 5:

[0261] Generate a workaround

[0262] The server generates a solution using a generative model (GPT-3) based on relevant information. It uses the transformers library.

[0263] Input: Search results and user text query

[0264] Output: Text data of the solution

[0265] Specific operation: The server performs NLP processing and generates the optimal solution.

[0266] Step 6:

[0267] Visual information generation

[0268] The server generates 3D models and diagrams based on the generated solutions. 3D design software and image generation software are used to make the solutions easier to understand visually.

[0269] Input: Text data of the solution

[0270] Output: 3D models and diagrams

[0271] Specific operation: The server converts the solution into a 3D model or diagram, generating visual information.

[0272] Step 7:

[0273] Coping methods and visual information

[0274] The server sends the generated solutions and 3D models or diagrams to the smart glasses.

[0275] The smart glasses present the received action and visual information to the user.

[0276] Input: Text data and visual information of the solution

[0277] Output: Actions and visual information displayed on the smart glasses

[0278] Specific operation: The server sends information, and the smart glasses display it and present it to the user.

[0279] This allows the user to input a problem using voice input, and quickly check the generated solutions and visual information to help solve the problem.

[0280] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0281] System Overview

[0282] The present invention provides a system that allows users to quickly and effectively solve problems without referring to an instruction manual or troubleshooting manual. This system not only analyzes the digital document of the instruction manual or troubleshooting manual and generates solutions to problems or phenomena input by the user, but also recognizes the user's emotions and reflects them in the solutions it presents.

[0283] composition

[0284] The user uses their own device (smartphone or PC) to input information about the problem or phenomenon in natural language.

[0285] The terminal transmits the user's input information to the server.

[0286] The server analyzes the digital document and extracts the required information.

[0287] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[0288] The emotion engine recognizes emotions from user information and reflects them in the generated response methods.

[0289] The server presents the generated solution and related image diagram to the user.

[0290] Explanation of program processing

[0291] The detailed operation of this system will be described below.

[0292] Initial Setup

[0293] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[0294] The server converts the PDF file into text using optical character recognition (OCR) technology.

[0295] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[0296] Query reception

[0297] The user uses a terminal to input information about a phenomenon or problem in natural language.

[0298] The terminal sends the input query to the server.

[0299] Data Analysis and Information Extraction

[0300] The server receives the user's query and searches for relevant information in a database.

[0301] The server provides the search results to the generative model.

[0302] Emotion recognition

[0303] The server provides the user's input data to the emotion engine to analyze the user's emotions.

[0304] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and reflects the results in the generative model.

[0305] Generate a workaround

[0306] The generative model generates a response based on the user's query, related information provided by the server, and the recognized emotions.

[0307] For example, if the problem is "the TV won't turn on" and the emotion is "confused," you might preface the situation by saying, "Let's start with a simple method that anyone can do," and then explain, "First, check that the power cord is connected correctly."

[0308] Generate image diagrams

[0309] The server generates an image related to the solution (for example, a screenshot of the settings screen or a diagram of the connection points).

[0310] Presenting solutions and illustrations

[0311] The server generated solution and image diagram are summarized on one page.

[0312] The server sends the compiled information to the terminal.

[0313] The terminal displays this to the user, allowing the user to quickly and easily resolve the problem.

[0314] Specific examples

[0315] Below is a specific example of what happens when a user enters "My smartphone's Wi-Fi won't connect."

[0316] 1. The user types "My smartphone's Wi-Fi won't connect" into the device.

[0317] 2. The device sends this query to the server.

[0318] 3. The server analyzes the smartphone's instruction manual using OCR technology and extracts relevant information.

[0319] 4. The server provides the user's input data to the emotion engine and recognizes the emotion "confused."

[0320] 5. Based on the relevant information and the recognized emotion, the generative model generates specific countermeasures such as, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[0321] 6. The server generates a screenshot associated with this method.

[0322] 7. The server compiles a solution and an illustration and sends it to the terminal.

[0323] 8. The device displays this to the user, allowing them to quickly and easily resolve the issue.

[0324] This system allows users to find appropriate ways to deal with their emotions without having to refer to an instruction manual.

[0325] The processing flow will be explained below.

[0326] Step 1:

[0327] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My smartphone's Wi-Fi won't connect."

[0328] Step 2:

[0329] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[0330] Step 3:

[0331] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[0332] Step 4:

[0333] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance, and uses OCR technology to convert the PDFs into text data.

[0334] Step 5:

[0335] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[0336] Step 6:

[0337] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[0338] Step 7:

[0339] The server provides the user's input data to the emotion engine, which analyzes the user's emotions. The emotion engine recognizes the user's emotions based on the input content and past data.

[0340] Step 8:

[0341] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and returns the results to the server, which then inputs this information into the generative model.

[0342] Step 9:

[0343] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. For example, if the user's emotion is "confused" when the query is "Wi-Fi is not connecting," the model will explain, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[0344] Step 10:

[0345] Based on the solution generated by the server, a concrete image diagram (for example, a screenshot of the Wi-Fi setting screen or a diagram of the connection location) is generated.

[0346] Step 11:

[0347] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[0348] Step 12:

[0349] The server then sends the final information to the device, including specific solutions and related images.

[0350] Step 13:

[0351] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem. This allows the user to respond efficiently without having to refer to an instruction manual or instruction manual.

[0352] Example 2

[0353] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0354] There is a demand for systems that allow users to solve problems quickly and effectively without referring to instruction manuals or troubleshooting manuals. There is also a demand for systems that improve user satisfaction by presenting solutions that take into account the user's emotions.

[0355] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0356] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and means including an emotion recognition engine for recognizing the user's emotions and reflecting them in the generated solution. This allows the user to quickly obtain an appropriate solution based on their emotions without having to refer to the instruction manual.

[0357] An "instruction manual or troubleshooting manual" is a detailed guideline on how to use a product and what to do in the event of a malfunction.

[0358] "Digital documents" are documents that are stored and managed electronically, including PDF and TXT formats.

[0359] "Means of analysis" refers to the techniques and methods used to read digital documents and understand and process their contents.

[0360] "Information about the problem or phenomenon input by the user" refers to the content of the phenomenon input by the user and the problems related to it.

[0361] "Means for receiving" refers to an interface or method for obtaining input information from a user.

[0362] "Means for extracting relevant information" refers to techniques or methods for extracting specified information from a digital document.

[0363] A "generative model" is an algorithm or system that generates new data or information based on given data.

[0364] "Presentation means" refers to the method or interface for displaying the generated information to the user.

[0365] An "emotion recognition engine" is a technology or system that analyzes and recognizes emotions from user input data.

[0366] The present invention provides a system that allows users to quickly and effectively solve problems without referring to instruction manuals or troubleshooting manuals. This system analyzes digital documents, generates solutions to problems or phenomena input by the user, and recognizes the user's emotions and reflects them in the solutions it presents.

[0367] System Overview

[0368] The system includes the following components:

[0369] Server: PDF files of instruction manuals and troubleshooting manuals are stored in storage and converted to text using OCR technology (e.g., Tesseract OCR). The converted text data is stored in a database (e.g., MySQL) and analyzed using a natural language processing (NLP) engine (e.g., spaCy).

[0370] Terminal: Provides an interface for users to input information about problems or phenomena in natural language and sends the input information to the server.

[0371] Generative model: Generates a solution based on relevant information provided by the server and user input (e.g., OpenAI's ChatGPT).

[0372] Emotion recognition engine: Analyzes emotions based on user input data (e.g., IBM Watson Tone Analyzer) and reflects the results in a generative model.

[0373] Presentation method: The generated solution and related image diagram are presented to the user.

[0374] How it works

[0375] 1. The server converts PDF files of instruction manuals and troubleshooting manuals into text using optical character recognition (OCR) technology. For example, by using Tesseract OCR, the PDF files are output as text files and the contents are stored in a database.

[0376] 2. The user uses the device to input a natural language description of the symptom or problem, for example, "My TV won't turn on."

[0377] 3. The device sends the entered query to the server via an HTTP POST request, specifying the server's API endpoint to send the data.

[0378] 4. The server receives the user's query and searches the database for relevant information using an SQL query, such as SELECT content FROM manuals_table WHERE content LIKE '%won't turn on%'.

[0379] 5. The server provides the search results to the generative model, which processes the content provided as prompts for the generative model.

[0380] 6. The server provides the user's input data to the emotion engine to analyze the emotion, for example, using IBM Watson Tone Analyzer to recognize the user's emotion.

[0381] 7. The generative model generates specific countermeasures based on the user's query, related information provided by the server, and the recognized emotions.

[0382] 8. The server generates an image diagram related to the generated solution, for example, using the Pillow library to generate the related diagram.

[0383] 9. The server compiles the generated solutions and images in HTML format and sends them to the terminal, where they are displayed to the user, allowing them to solve the problem quickly and easily.

[0384] Specific examples

[0385] Below is an example of a specific prompt sentence when a user enters "My smartphone's Wi-Fi won't connect."

[0386] "The user inputs, 'My smartphone's Wi-Fi won't connect.' Provide relevant information from the instruction manual. Furthermore, the user's emotion, 'confused,' is recognized. Based on this information, generate a specific solution."

[0387] This invention allows users to quickly find appropriate solutions according to their emotions without having to refer to an instruction manual or troubleshooting manual. Furthermore, by using an emotion recognition engine, it becomes possible to respond to each user in a way that is tailored to their individual needs, thereby improving user satisfaction.

[0388] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0389] Step 1:

[0390] The server saves PDF files of instruction manuals and troubleshooting manuals to storage. First, it saves the PDF files provided by the user in a dedicated directory called "manuals." For example, the file path is / var / www / manuals / manual1.pdf. The file path after saving is output.

[0391] Step 2:

[0392] The server converts the PDF file to text using optical character recognition (OCR). It then uses Tesseract OCR to output the PDF file as a text file. The input is the path to the saved PDF file, and the output is the path to the converted text file. Example: / var / www / manuals / manual1.txt

[0393] Step 3:

[0394] The server saves the converted text data in the database. It uses a MySQL database and saves the text data in a table (e.g. manuals_table). The input is the path to the text file and its contents, and the output is the updated result in the database. Example SQL query: INSERT INTO manuals_table (manual_id, content) VALUES (1, LOAD_FILE(' / var / www / manuals / manual1.txt'))

[0395] Step 4:

[0396] The user uses the device to input information about the problem or issue in natural language, for example, "The TV won't turn on." This input information is used in the next step.

[0397] Step 5:

[0398] The device sends the entered query to the server via an HTTP POST request. Specify the server's API endpoint and send the query data. The input is the user-entered query, and the output is the result of the HTTP request sent to the server. Example: POST / api / query {"query": "The TV won't turn on"}

[0399] Step 6:

[0400] The server receives a user query and searches for relevant information in the database. It uses SQL queries to retrieve relevant information from the database based on the input query. The input is the user query, and the output is the search results for relevant information. SQL query example: SELECT content FROM manuals_table WHERE content LIKE '%Power won't turn on%'

[0401] Step 7:

[0402] The server provides the search results to the generative model. The provided information is processed as a prompt for the generative model. The input is the search result information, and the output is the formation of a prompt for the generative model. Example prompt: "The user entered 'The TV won't turn on.' Please generate a solution by referring to the contents of the instruction manual below: [Related information]."

[0403] Step 8:

[0404] The server provides the user's input data to the emotion engine to analyze the user's emotion. For example, IBM Watson Tone Analyzer is used. The input is the user query, and the output is the emotion recognition result. API example: POST / api / tone_analyzer {"text": "The TV won't turn on"}. Result format: {"emotion": "Confused"}

[0405] Step 9:

[0406] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. The input is the prompt and emotion recognition result for the generative model, and the output is the generated solution. Example: "Don't worry. First, make sure the power cord is properly connected."

[0407] Step 10:

[0408] The server generates an image related to the solution. It uses the Pillow library to generate an image containing specific content. The input is the solution content, and the output is the generated image. Example: image = Image.open("power_connection.png")

[0409] Step 11:

[0410] The server compiles the generated solutions and images in HTML format and sends them to the terminal. It uses an HTML template to compile information into one page. The input is the solutions and images, and the output is the content in HTML format. Example: {"html": " <h1> Solution< / h1> ...}"}

[0411] Step 12:

[0412] The terminal displays this to the user, allowing them to solve the problem quickly and easily. It uses a rendering engine to display HTML content. The input is the HTML formatted content, and the output is the displayed result.

[0413] (Application example 2)

[0414] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0415] There is a problem that users cannot find a quick and accurate solution on the spot when performing maintenance or troubleshooting on factory robots. In addition, a system that provides a uniform solution without considering the user's feelings makes it difficult to reduce the user's stress and confusion. Also, referring to the instruction manual or troubleshooting manual every time is time-consuming and does not allow for efficient problem solving.

[0416] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0417] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and emotion recognition means for recognizing the user's emotions and reflecting those emotions in the generation of a solution. This enables the user to quickly and effectively obtain a solution that takes into account the user's emotions, without having to refer to the instruction manual or troubleshooting manual.

[0418] An "instruction manual" is a document that details how to operate and configure equipment or software.

[0419] A "failure response manual" is a document that explains the response procedures and solutions required when equipment breaks down.

[0420] A "digital document" is a document that is stored and displayed electronically, including PDFs and text files.

[0421] "Means of analysis" refers to techniques and methods for understanding the content of a digital document and extracting the necessary information.

[0422] "Information about a problem or phenomenon input by a user" is detailed information about a malfunction or phenomenon of a device input by a user in natural language.

[0423] "Means for receiving" refers to the technology or method for obtaining and processing information sent by a user.

[0424] An "extraction means" is a technique or method for extracting relevant information from a digital document.

[0425] A "generative model" is an AI model that generates appropriate solutions based on user input and information extracted from digital documents.

[0426] The "presentation means" refers to a technique or method for displaying the generated solution to the user.

[0427] "Emotion recognition means" refers to a technique or method for analyzing emotions from information entered by a user and using the results to adapt a response method.

[0428] A "server" is a computer system that stores, analyzes, and processes data, and is a device that communicates with user terminals via a network.

[0429] The present invention provides a system that allows a user to quickly and accurately find a solution when performing maintenance or troubleshooting on a factory robot. Specific embodiments of the present invention will be described below.

[0430] Hardware and software used

[0431] 1. Server

[0432] Storage: Storage for saving instruction manuals and troubleshooting manuals. Specifically, it uses the server's HDD or SSD.

[0433] OCR Technology: The server converts the PDF file into text using optical character recognition (OCR) technology (e.g., Tesseract).

[0434] Database: A database is required to store the converted text data and analyze it with a natural language processing (NLP) engine. For example, MySQL or PostgreSQL will be used.

[0435] Generative AI models: Use generative AI (e.g., OpenAI models) to generate solutions based on the user query and extracted information.

[0436] Emotion recognition engine: Analyzes emotions from user information and reflects them in response. Uses natural language processing models such as BERT.

[0437] 2. Terminal

[0438] Smartphone or PC: Users use these devices to input information about problems or phenomena in natural language, and the devices then send this information to a server.

[0439] Operation overview

[0440] 1. Initial data preparation

[0441] The server stores PDF files of instruction manuals and troubleshooting manuals in storage and converts them into text data using OCR technology. The converted text data is then stored in a database and analyzed by a natural language processing engine.

[0442] 2. Accepting user queries

[0443] The user uses a terminal to input questions about the robot's problems or symptoms in natural language, and the queries are sent from the terminal to the server.

[0444] 3. Data analysis and information extraction

[0445] The server receives user queries and searches and extracts relevant information from a database.

[0446] 4. Emotional Recognition

[0447] The server provides the user's input data to the emotion recognition engine, which recognizes the user's emotions. The recognized emotions are reflected in the generative AI model.

[0448] 5. Generate solutions

[0449] The generative AI model generates a solution based on the user's query, extracted information, and recognized emotions, and presents guidance in the most appropriate language based on the user's emotions.

[0450] 6. Creating an image diagram

[0451] The server generates an image related to the solution (for example, a screenshot of the setting screen or a diagram of the connection points).

[0452] 7. Present solutions and illustrations

[0453] The server then compiles the generated solutions and images into a single page and sends it to the device, which displays it to the user, allowing them to quickly and easily solve the problem.

[0454] Specific examples

[0455] Here is a specific example where the user inputs "The robot's arm won't move."

[0456] 1. User query: "The robot arm won't move."

[0457] 2. Emotion recognition: The emotion recognition engine recognizes "confusion."

[0458] 3. Generating a solution: The generative AI model generated the following: "Don't be confused. First, turn off the robot and check if there is an abnormality in the arm connection. Next, turn it on again, and if an error code is displayed, refer to the procedures in the maintenance manual."

[0459] 4. Image generation: Image of the "arm connection part".

[0460] Prompt Sentence Examples

[0461] Please generate solutions to user queries based on the following manual.

[0462] manual:

[0463] [Instruction manual and troubleshooting manual text]

[0464] Query: Robot arm not moving

[0465] User Emotion: Confused

[0466] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0467] Step 1:

[0468] The server saves PDF files of instruction manuals and troubleshooting manuals in storage. The input of this step is the PDF files, and the output is the PDF files saved in storage. This prepares the documents to be analyzed.

[0469] Step 2:

[0470] The server converts the PDF file into text data using OCR technology (e.g., Tesseract). The input for this step is the PDF file, and the output is text data. Using OCR technology, character information is extracted from the image data and converted into text format.

[0471] Step 3:

[0472] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine. The input to this step is the text data generated by OCR, and the output is structured data of the document analyzed by NLP. As a result of the analysis, important information within the document is extracted and registered in the database.

[0473] Step 4:

[0474] The user uses a terminal to input information about the problem or phenomenon of the robot in natural language. The input of this step is a query entered by the user, and the output is a query sent from the terminal to the server. The user describes the problem in natural language, and it is sent to the server.

[0475] Step 5:

[0476] The server receives a user query and searches and extracts relevant information in the database. The input to this step is the user query, and the output is the relevant information extracted from the database. The server pulls the appropriate information from the database based on the query.

[0477] Step 6:

[0478] The server provides the user's query and extracted information to an emotion recognition engine to analyze the user's emotion. The input of this step is the user query and related information, and the output is analyzed emotion data. The emotion recognition engine is used to identify the user's emotion (e.g., confusion, anger), and the result is provided to the next step.

[0479] Step 7:

[0480] The server uses a generative AI model (e.g., OpenAI's model) to generate a solution based on the extracted information and the recognized emotion. The inputs for this step are the user query, the extracted information, and the emotion data, and the output is a solution generated by the generative AI model. The generated solution takes the user's emotion into consideration.

[0481] Step 8:

[0482] The server generates an image associated with the generated solution. The input to this step is the generated solution, and the output is an image associated with the solution. The server generates appropriate visual aids (screenshots and diagrams) to accompany the solution.

[0483] Step 9:

[0484] The server compiles the generated solutions and the image onto a single page and sends it to the terminal. The input to this step is the solutions and the image, and the output is the data sent to the terminal. The user receives this on their terminal and it is displayed.

[0485] Step 10:

[0486] The user can quickly and easily solve the problem by referring to the solution and image displayed on the terminal. The input of this step is the solution and image displayed on the terminal, and the output is the user's problem solution. The user proceeds with the solution according to the proposed method.

[0487] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0488] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0489] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0490] [Second embodiment]

[0491] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0492] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0493] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0494] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0495] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0496] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0497] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0498] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0499] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0500] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0501] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0502] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0503] System Overview

[0504] The present invention provides a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual for electronic devices. This system can analyze the digital document of the instruction manual or troubleshooting manual and generate a solution to the problem or phenomenon input by the user.

[0505] composition

[0506] Users use their own devices (smartphones or PCs) to enter information about problems or issues.

[0507] The terminal transmits the user's input information to the server.

[0508] The server analyzes the digital document and extracts the required information.

[0509] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[0510] The server presents the generated solution and related image diagram to the user.

[0511] Explanation of program processing

[0512] The detailed operation of this system will be explained below.

[0513] Initial Setup

[0514] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[0515] The server converts the PDF file into text using optical character recognition (OCR) technology.

[0516] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[0517] Query reception

[0518] The user uses the device to input a description of the issue or problem in natural language, for example, "I can't connect to Wi-Fi on my smartphone."

[0519] The terminal sends the input query to the server.

[0520] Data Analysis and Information Extraction

[0521] The server receives the user's query and searches for relevant information in a database.

[0522] The server provides the search results to the generative model.

[0523] Generate a workaround

[0524] The generative model generates a solution based on the user's query and related information provided by the server.

[0525] For example, it generates specific solutions such as "Select the network name on the Wi-Fi settings screen and enter the password."

[0526] Generate image diagrams

[0527] The server generates an image related to the solution (for example, a screenshot of the setting screen).

[0528] Presenting solutions and illustrations

[0529] The server generated solution and image diagram are summarized on one page.

[0530] The server sends the compiled information to the terminal.

[0531] The terminal displays this to the user, allowing the user to easily resolve the problem.

[0532] Specific examples

[0533] Below is a specific example where the user enters "My TV won't turn on."

[0534] 1. The user types "TV won't turn on" into the device.

[0535] 2. The device sends this query to the server.

[0536] 3. The server analyzes the TV's instruction manual using OCR technology and extracts information related to the "won't turn on" problem.

[0537] 4. Based on the extracted information and user input, the generative model generates a solution such as "Make sure the power cord is connected correctly. Then, replace the batteries in the remote control."

[0538] 5. The server generates a diagram that visually explains how to resolve the issue (for example, a diagram showing where to connect the power cord).

[0539] 6. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[0540] 7. The device should display this to the user so that they can quickly and easily understand what to do.

[0541] This allows users to solve problems efficiently without having to read the instruction manual in detail.

[0542] The processing flow will be explained below.

[0543] Step 1:

[0544] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My TV won't turn on."

[0545] Step 2:

[0546] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[0547] Step 3:

[0548] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[0549] Step 4:

[0550] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance in storage, and uses OCR technology to convert the PDFs into text data.

[0551] Step 5:

[0552] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[0553] Step 6:

[0554] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[0555] Step 7:

[0556] The server inputs the extracted information and the user's input data into a generative AI model (e.g., ChatGPT) to generate an appropriate response. The generative AI model generates the most appropriate response in natural language based on the input data.

[0557] Step 8:

[0558] The server generates specific images (e.g., screenshots of power cord connection locations or settings screens) based on the generated solutions, using automatic image generation technology and images retrieved from a database.

[0559] Step 9:

[0560] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[0561] Step 10:

[0562] The server then sends the final information to the device, including specific solutions and related images.

[0563] Step 11:

[0564] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem, thus enabling the user to effectively respond without having to refer to an instruction manual or instruction manual.

[0565] Example 1

[0566] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0567] In the past, when users wanted to refer to instruction manuals or troubleshooting manuals for electronic devices, they had to refer to paper documents or search websites, which was time-consuming and laborious. Furthermore, it was difficult to quickly obtain specific solutions, which reduced the efficiency of problem-solving. In particular, when dealing with electronic device malfunctions, accurate and prompt information provision is required.

[0568] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0569] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for converting the digital document into text using optical character recognition technology, means for extracting relevant information from the digital document, means for searching for relevant information in a database, means for including a generative AI model that generates solutions based on the extracted information and information input by the user, means for generating an image corresponding to the generated solution, and means for presenting the generated solution and the image to the user. This allows the user to solve problems quickly and effectively without having to read the instruction manual in detail.

[0570] An "instruction manual or troubleshooting manual" is a document that describes how to use an electronic device and what to do if a problem occurs.

[0571] A "digital document" is a document in a form that is stored, displayed, and transmitted electronically.

[0572] "Means for analysis" refers to technology or devices that have the ability to extract necessary information from digital documents.

[0573] "Means for receiving information about a problem or phenomenon input by a user" refers to a technology or device that has the function of transmitting information input by a user through a terminal to a server and receiving that information.

[0574] "Optical Character Recognition (OCR)" is a technology that converts characters in an image into electronic text.

[0575] A "database" is a system for efficiently storing, retrieving, and managing structured data.

[0576] A "natural language processing engine" is an artificial intelligence technology that analyzes text data, extracts semantic information, and understands language.

[0577] A "generative AI model" is an artificial intelligence model that generates useful information, such as countermeasures, based on input information from users and related data.

[0578] An "image" is a visual illustration that supplements text information.

[0579] An "HTTP request" is a communication protocol used by a web browser or application to request data from a server.

[0580] A "server" is a computer system that processes queries, stores data, and communicates over a network.

[0581] A "terminal" is a device that a user uses to access a server, such as a smartphone or PC.

[0582] A "prompt" is a text question or instruction used as input to a generative AI model.

[0583] A "means for converting to text" is a technique or device that converts characters in an image into electronic text using optical character recognition technology.

[0584] System Overview

[0585] This invention is a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual of the electronic device. Specifically, it analyzes digital documents using optical character recognition (OCR) technology and a natural language processing engine (NLP), generates solutions using a generative AI model, and presents them to the user.

[0586] Hardware and software used

[0587] Server: A computer system for storing, analyzing, and storing digital documents in a database.

[0588] Storage engine: Amazon S3, etc.

[0589] OCR technology: Tesseract OCR

[0590] Database: MySQL

[0591] Natural language processing engine: spaCy

[0592] Generative AI models: ChatGPT, etc.

[0593] Terminal: A device through which a user accesses a server and enters queries.

[0594] Devices: Smartphones, PCs, etc.

[0595] Program processing explanation

[0596] Information storage and analysis

[0597] The server stores PDF files of instruction manuals and troubleshooting manuals in a storage engine, for example, in an Amazon S3 bucket.

[0598] The server converts the saved PDF file to text using Tesseract OCR, specifically by running the command tesseract manual.pdf output.txt.

[0599] The server stores the converted text data in a MySQL database, for example by executing the SQL statement INSERT INTO manuals (text_data) VALUES ('...').

[0600] The server analyzes the text data using the spaCy NLP engine to extract semantic information, for example by running the following code: nlp = spacy.load('en_core_web_sm').

[0601] Query reception and processing

[0602] The user inputs a symptom or problem in natural language through the terminal. For example, the user inputs "The TV won't turn on."

[0603] The device sends this query to the server via an HTTP POST request.

[0604] Data analysis and generation of action plans

[0605] The server receives the user's query and searches the database for relevant information using an SQL query, for example SELECT FROM manuals WHERE text_data LIKE '%cannot power on%'.

[0606] The server sends the relevant information to a generative AI model (e.g., ChatGPT) via an API request. An example prompt sentence is "The user explains that the TV won't turn on. What should I do?"

[0607] The generative AI model generates a solution and sends it back to the server, such as "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0608] Presenting solutions and illustrations

[0609] The server uses the Pillow library to generate an image related to the solution, for example, a diagram showing where to connect the power cord.

[0610] The server integrates the generated solutions and images into an HTML template and compiles them into a single page.

[0611] The server sends the compiled information to the terminal as an HTTP response.

[0612] The terminal displays the information to the user and provides a quick and effective way to deal with the problem.

[0613] Specific examples

[0614] For example, if the user inputs "TV won't turn on," the system will operate as follows:

[0615] 1. The user types "TV won't turn on" into the device.

[0616] 2. The device sends the query to the server via an HTTP POST request.

[0617] 3. The server receives the query and searches the database for relevant information using an SQL query.

[0618] 4. The server sends the relevant information to the generative AI model via an API request.

[0619] 5. The generative AI model generates a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0620] 6. The server generates a diagram showing where the power cords are connected.

[0621] 7. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[0622] 8. The device displays information to the user and provides a quick and effective way to deal with the problem.

[0623] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0624] Step 1: Initial Setup

[0625] The server stores PDF files of instruction manuals and troubleshooting manuals in the storage engine. As a concrete example, using Amazon S3, the command aws s3 cp manual.pdf s3: / / bucket-name / is executed as input, and the output is the PDF file stored in Amazon S3.

[0626] The server converts PDF files to text using Tesseract OCR. Specifically, it receives input from the command tesseract manual.pdf output.txt and generates an output file called output.txt.

[0627] The server stores the converted text data in a MySQL database. For example, it takes input as input and executes the SQL statement INSERT INTO manuals (text_data) VALUES ('...'), and gets output as text data stored in the database.

[0628] Step 2: Enter and accept user queries

[0629] The user inputs the phenomenon or problem in natural language using the terminal. For example, the input "The TV won't turn on" is received, and the output is the manually input data in text format.

[0630] The terminal sends the entered query to the server. Specifically, it receives input to send the entered query to the server via an HTTP POST request, and obtains output showing that the query has reached the server.

[0631] Step 3: Data analysis and information extraction

[0632] The server receives a user query and searches for relevant information in the database. Specifically, it receives input by executing the SQL query SELECT FROM manuals WHERE text_data LIKE '%TV won't turn on%' and obtains output that provides relevant information (for example, the contents of the corresponding instruction manual).

[0633] The server provides the relevant information to the generative AI model, taking the specific action of sending the results of a database search to the generative AI model via an API request as input, and obtaining an output ready for the generative AI model to generate a response.

[0634] Step 4: Generate a solution

[0635] The generative AI model generates a solution based on the user's query and information provided by the server. Specifically, the model receives a prompt, "The user explains that the TV won't turn on. What should I do?", and outputs a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[0636] The server receives the generated solutions and performs filtering. Specifically, it receives the generated text, deletes unnecessary parts, and outputs solutions that are ready to be provided to the user.

[0637] Step 5: Generate an image

[0638] The server uses the Pillow library to generate an image diagram related to the solution. For example, for the command "Please check the power cord connection," input is made using an image editing tool, and an output is generated showing the power cord connection location.

[0639] Step 6: Present solutions and images

[0640] The server then combines the generated solutions and the resulting image onto a single page. Specifically, it uses an HTML template as input to integrate the solutions and the image, and outputs a single page that can be presented to the user.

[0641] The server sends the compiled information to the terminal. As a specific input, it performs the operation of sending the generated page as an HTTP response, and obtains the output that the page arrives at the user terminal.

[0642] The device displays the information to the user. Specifically, the device inputs the received page to display in a browser or dedicated app, and the user receives an output that allows them to visually confirm the solution to the problem.

[0643] (Application example 1)

[0644] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0645] The present invention relates to a system for quickly and effectively resolving problems without referring to instruction manuals or troubleshooting manuals. However, conventional systems require users to accurately input details of the problem, which can lead to errors, especially when using voice input. Furthermore, solutions are presented only in text, which can be difficult for users to understand. Furthermore, it is difficult for factory workers to respond to equipment failures without visual information. It is necessary to solve these problems.

[0646] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0647] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, means for converting voice input into text, and means for generating a 3D model or diagram of the solution. This allows the user to input a problem using voice input, and the solution is presented not only as text but also as a 3D model or diagram, enabling the user to solve the problem quickly and visually.

[0648] An "instruction manual" is a document that provides detailed instructions on how to use and operate an electronic device or machine.

[0649] A "troubleshooting manual" is a document that explains how to repair and resolve problems when electronic devices or machinery break down.

[0650] "Digital documents" are document data stored electronically, including PDFs and text files.

[0651] "Means of analysis" refers to the technology and devices used to read digital documents and understand and classify the necessary information.

[0652] "Means for receiving" refers to technology or devices for receiving information or data input by a user.

[0653] An "extraction means" is a technique or device used to identify and extract the required information from a digital document.

[0654] A "generative model" is an artificial intelligence model that automatically generates appropriate solutions or answers based on input information.

[0655] The "presentation means" refers to a technique or device for displaying the generated solutions and information to the user.

[0656] "Means for converting voice input into text" refers to technology or devices for converting voice input by a user into text data.

[0657] "Means for generating 3D models or diagrams of solutions" refers to technology or devices for expressing solutions in 3D models or figures to make them visually easier to understand.

[0658] "Server" means a computer system for processing, storing, and managing data.

[0659] This invention is a system that can quickly and effectively solve problems without the user having to refer to an instruction manual or a troubleshooting manual. This system is particularly useful for effectively dealing with factory robot failures. The specific system configuration and its operation will be described below.

[0660] System Configuration

[0661] 1. User Device

[0662] Wearable devices such as smart glasses.

[0663] It has voice input and display functions.

[0664] 2. Server

[0665] A cloud server for analyzing and storing digital documents.

[0666] It has high-performance computing resources and data storage.

[0667] 3. OCR technology

[0668] Convert PDF to text data using optical character recognition technology. Specifically, we use pytesseract.

[0669] 4. Voice Recognition Technology

[0670] It uses technology to convert voice input into text data, specifically, the speech_recognition library.

[0671] 5. Generative AI Models

[0672] It is used for natural language processing and countermeasure generation. Specifically, it uses the GPT-3 model from the transformers library.

[0673] 6. Digital Documents

[0674] PDF documents such as instruction manuals and troubleshooting manuals.

[0675] Data processing and calculation

[0676] 1. Digital Document Analysis

[0677] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using optical character recognition (OCR). This is where pytesseract comes in.

[0678] The converted text data is stored in a database on the server.

[0679] 2. Processing voice input

[0680] The user inputs the symptoms or problems by voice through the smart glasses.

[0681] The smart glasses convert the voice data into text data and send this data to the server, where the speech_recognition library is used.

[0682] 3. Extracting information and generating solutions

[0683] The server analyzes the received user query and searches for relevant information in a database.

[0684] Based on the search results, a generative model (e.g., GPT-3) is used to generate a solution. The transformers library is used.

[0685] 4. Providing solutions and visual information

[0686] Use software to convert the generated solutions into 3D models and diagrams.

[0687] The solution and 3D models or diagrams are sent to the smart glasses and visually presented to the user.

[0688] Specific examples

[0689] Example 1: Error handling in an automobile factory

[0690] User input: "The robot arm joints won't move."

[0691] Generated solution: "Oil the joints of the robot arm and check their operation. Also, check for any sensor error indications."

[0692] Prompt Sentence Examples

[0693] Problem: The robot arm joints won't move

[0694] How to respond:

[0695] This allows factory workers to quickly identify problems and respond effectively. The system is expected to improve work efficiency and reduce downtime through visual and audio assistance.

[0696] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0697] Step 1:

[0698] Digital document analysis

[0699] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using OCR technology. The server reads the PDF using pytesseract and converts each page into text data.

[0700] Input: PDF format instruction manual or troubleshooting manual

[0701] Output: Text data

[0702] Specific behavior: The server extracts text from the PDF and stores it in a database.

[0703] Step 2:

[0704] Receiving audio input

[0705] The user inputs the symptoms or problems by voice through the smart glasses.

[0706] The smart glasses capture voice data and convert it to text data using the speech_recognition library.

[0707] Input: Audio data of the phenomenon or problem

[0708] Output: Text data

[0709] How it works: The user speaks, "The robot arm's joints won't move," and the smart glasses convert the speech into text.

[0710] Step 3:

[0711] Submitting a query

[0712] The smart glasses send the converted text data to the server.

[0713] Input: Text data converted from speech

[0714] Output: Text data sent to the server

[0715] Specific operation: The smart glasses generate text data and send it to a server via the network.

[0716] Step 4:

[0717] Information extraction and analysis

[0718] The server analyzes the received text data (user queries) and searches for relevant information in the database using a natural language processing (NLP) engine.

[0719] Input: User query text data

[0720] Output: Search results containing relevant information

[0721] Specific operation: The server extracts information related to the query from instruction manuals and manuals in the database.

[0722] Step 5:

[0723] Generate a workaround

[0724] The server generates a solution using a generative model (GPT-3) based on relevant information. It uses the transformers library.

[0725] Input: Search results and user text query

[0726] Output: Text data of the solution

[0727] Specific operation: The server performs NLP processing and generates the optimal solution.

[0728] Step 6:

[0729] Visual information generation

[0730] The server generates 3D models and diagrams based on the generated solutions. 3D design software and image generation software are used to make the solutions easier to understand visually.

[0731] Input: Text data of the solution

[0732] Output: 3D models and diagrams

[0733] Specific operation: The server converts the solution into a 3D model or diagram, generating visual information.

[0734] Step 7:

[0735] Coping methods and visual information

[0736] The server sends the generated solutions and 3D models or diagrams to the smart glasses.

[0737] The smart glasses present the received action and visual information to the user.

[0738] Input: Text data and visual information of the solution

[0739] Output: Actions and visual information displayed on the smart glasses

[0740] Specific operation: The server sends information, and the smart glasses display it and present it to the user.

[0741] This allows the user to input a problem using voice input, and quickly check the generated solutions and visual information to help solve the problem.

[0742] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0743] System Overview

[0744] The present invention provides a system that allows users to quickly and effectively solve problems without referring to an instruction manual or troubleshooting manual. This system not only analyzes the digital document of the instruction manual or troubleshooting manual and generates solutions to problems or phenomena input by the user, but also recognizes the user's emotions and reflects them in the solutions it presents.

[0745] composition

[0746] The user uses their own device (smartphone or PC) to input information about the problem or phenomenon in natural language.

[0747] The terminal transmits the user's input information to the server.

[0748] The server analyzes the digital document and extracts the required information.

[0749] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[0750] The emotion engine recognizes emotions from user information and reflects them in the generated response methods.

[0751] The server presents the generated solution and related image diagram to the user.

[0752] Explanation of program processing

[0753] The detailed operation of this system will be described below.

[0754] Initial Setup

[0755] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[0756] The server converts the PDF file into text using optical character recognition (OCR) technology.

[0757] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[0758] Query reception

[0759] The user uses a terminal to input information about a phenomenon or problem in natural language.

[0760] The terminal sends the input query to the server.

[0761] Data Analysis and Information Extraction

[0762] The server receives the user's query and searches for relevant information in a database.

[0763] The server provides the search results to the generative model.

[0764] Emotion recognition

[0765] The server provides the user's input data to the emotion engine to analyze the user's emotions.

[0766] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and reflects the results in the generative model.

[0767] Generate a workaround

[0768] The generative model generates a response based on the user's query, related information provided by the server, and the recognized emotions.

[0769] For example, if the problem is "the TV won't turn on" and the emotion is "confused," you might preface the situation by saying, "Let's start with a simple method that anyone can do," and then explain, "First, check that the power cord is connected correctly."

[0770] Generate image diagrams

[0771] The server generates an image related to the solution (for example, a screenshot of the settings screen or a diagram of the connection points).

[0772] Presenting solutions and illustrations

[0773] The server generated solution and image diagram are summarized on one page.

[0774] The server sends the compiled information to the terminal.

[0775] The terminal displays this to the user, allowing the user to quickly and easily resolve the problem.

[0776] Specific examples

[0777] Below is a specific example of what happens when a user enters "My smartphone's Wi-Fi won't connect."

[0778] 1. The user types "My smartphone's Wi-Fi won't connect" into the device.

[0779] 2. The device sends this query to the server.

[0780] 3. The server analyzes the smartphone's instruction manual using OCR technology and extracts relevant information.

[0781] 4. The server provides the user's input data to the emotion engine and recognizes the emotion "confused."

[0782] 5. Based on the relevant information and the recognized emotion, the generative model generates specific countermeasures such as, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[0783] 6. The server generates a screenshot associated with this method.

[0784] 7. The server compiles a solution and an illustration and sends it to the terminal.

[0785] 8. The device displays this to the user, allowing them to quickly and easily resolve the issue.

[0786] This system allows users to find appropriate ways to deal with their emotions without having to refer to an instruction manual.

[0787] The processing flow will be explained below.

[0788] Step 1:

[0789] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My smartphone's Wi-Fi won't connect."

[0790] Step 2:

[0791] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[0792] Step 3:

[0793] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[0794] Step 4:

[0795] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance, and uses OCR technology to convert the PDFs into text data.

[0796] Step 5:

[0797] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[0798] Step 6:

[0799] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[0800] Step 7:

[0801] The server provides the user's input data to the emotion engine, which analyzes the user's emotions. The emotion engine recognizes the user's emotions based on the input content and past data.

[0802] Step 8:

[0803] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and returns the results to the server, which then inputs this information into the generative model.

[0804] Step 9:

[0805] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. For example, if the user's emotion is "confused" when the query is "Wi-Fi is not connecting," the model will explain, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[0806] Step 10:

[0807] Based on the solution generated by the server, a concrete image diagram (for example, a screenshot of the Wi-Fi setting screen or a diagram of the connection location) is generated.

[0808] Step 11:

[0809] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[0810] Step 12:

[0811] The server then sends the final information to the device, including specific solutions and related images.

[0812] Step 13:

[0813] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem. This allows the user to respond efficiently without having to refer to an instruction manual or instruction manual.

[0814] Example 2

[0815] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0816] There is a demand for systems that allow users to solve problems quickly and effectively without referring to instruction manuals or troubleshooting manuals. There is also a demand for systems that improve user satisfaction by presenting solutions that take into account the user's emotions.

[0817] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0818] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and means including an emotion recognition engine for recognizing the user's emotions and reflecting them in the generated solution. This allows the user to quickly obtain an appropriate solution based on their emotions without having to refer to the instruction manual.

[0819] An "instruction manual or troubleshooting manual" is a detailed guideline on how to use a product and what to do in the event of a malfunction.

[0820] "Digital documents" are documents that are stored and managed electronically, including PDF and TXT formats.

[0821] "Means of analysis" refers to the techniques and methods used to read digital documents and understand and process their contents.

[0822] "Information about the problem or phenomenon input by the user" refers to the content of the phenomenon input by the user and the problems related to it.

[0823] "Means for receiving" refers to an interface or method for obtaining input information from a user.

[0824] "Means for extracting relevant information" refers to techniques or methods for extracting specified information from a digital document.

[0825] A "generative model" is an algorithm or system that generates new data or information based on given data.

[0826] "Presentation means" refers to the method or interface for displaying the generated information to the user.

[0827] An "emotion recognition engine" is a technology or system that analyzes and recognizes emotions from user input data.

[0828] The present invention provides a system that allows users to quickly and effectively solve problems without referring to instruction manuals or troubleshooting manuals. This system analyzes digital documents, generates solutions to problems or phenomena input by the user, and recognizes the user's emotions and reflects them in the solutions it presents.

[0829] System Overview

[0830] The system includes the following components:

[0831] Server: PDF files of instruction manuals and troubleshooting manuals are stored in storage and converted to text using OCR technology (e.g., Tesseract OCR). The converted text data is stored in a database (e.g., MySQL) and analyzed using a natural language processing (NLP) engine (e.g., spaCy).

[0832] Terminal: Provides an interface for users to input information about problems or phenomena in natural language and sends the input information to the server.

[0833] Generative model: Generates a solution based on relevant information provided by the server and user input (e.g., OpenAI's ChatGPT).

[0834] Emotion recognition engine: Analyzes emotions based on user input data (e.g., IBM Watson Tone Analyzer) and reflects the results in a generative model.

[0835] Presentation method: The generated solution and related image diagram are presented to the user.

[0836] How it works

[0837] 1. The server converts PDF files of instruction manuals and troubleshooting manuals into text using optical character recognition (OCR) technology. For example, by using Tesseract OCR, the PDF files are output as text files and the contents are stored in a database.

[0838] 2. The user uses the device to input a natural language description of the symptom or problem, for example, "My TV won't turn on."

[0839] 3. The device sends the entered query to the server via an HTTP POST request, specifying the server's API endpoint to send the data.

[0840] 4. The server receives the user's query and searches the database for relevant information using an SQL query, such as SELECT content FROM manuals_table WHERE content LIKE '%won't turn on%'.

[0841] 5. The server provides the search results to the generative model, which processes the content provided as prompts for the generative model.

[0842] 6. The server provides the user's input data to the emotion engine to analyze the emotion, for example, using IBM Watson Tone Analyzer to recognize the user's emotion.

[0843] 7. The generative model generates specific countermeasures based on the user's query, related information provided by the server, and the recognized emotions.

[0844] 8. The server generates an image diagram related to the generated solution, for example, using the Pillow library to generate the related diagram.

[0845] 9. The server compiles the generated solutions and images in HTML format and sends them to the terminal, where they are displayed to the user, allowing them to solve the problem quickly and easily.

[0846] Specific examples

[0847] Below is an example of a specific prompt sentence when a user enters "My smartphone's Wi-Fi won't connect."

[0848] "The user inputs, 'My smartphone's Wi-Fi won't connect.' Provide relevant information from the instruction manual. Furthermore, the user's emotion, 'confused,' is recognized. Based on this information, generate a specific solution."

[0849] This invention allows users to quickly find appropriate solutions according to their emotions without having to refer to an instruction manual or troubleshooting manual. Furthermore, by using an emotion recognition engine, it becomes possible to respond to each user in a way that is tailored to their individual needs, thereby improving user satisfaction.

[0850] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0851] Step 1:

[0852] The server saves PDF files of instruction manuals and troubleshooting manuals to storage. First, it saves the PDF files provided by the user in a dedicated directory called "manuals." For example, the file path is / var / www / manuals / manual1.pdf. The file path after saving is output.

[0853] Step 2:

[0854] The server converts the PDF file to text using optical character recognition (OCR). It then uses Tesseract OCR to output the PDF file as a text file. The input is the path to the saved PDF file, and the output is the path to the converted text file. Example: / var / www / manuals / manual1.txt

[0855] Step 3:

[0856] The server saves the converted text data in the database. It uses a MySQL database and saves the text data in a table (e.g. manuals_table). The input is the path to the text file and its contents, and the output is the updated result in the database. Example SQL query: INSERT INTO manuals_table (manual_id, content) VALUES (1, LOAD_FILE(' / var / www / manuals / manual1.txt'))

[0857] Step 4:

[0858] The user uses the device to input information about the problem or issue in natural language, for example, "The TV won't turn on." This input information is used in the next step.

[0859] Step 5:

[0860] The device sends the entered query to the server via an HTTP POST request. Specify the server's API endpoint and send the query data. The input is the user-entered query, and the output is the result of the HTTP request sent to the server. Example: POST / api / query {"query": "The TV won't turn on"}

[0861] Step 6:

[0862] The server receives a user query and searches for relevant information in the database. It uses SQL queries to retrieve relevant information from the database based on the input query. The input is the user query, and the output is the search results for relevant information. SQL query example: SELECT content FROM manuals_table WHERE content LIKE '%Power won't turn on%'

[0863] Step 7:

[0864] The server provides the search results to the generative model. The provided information is processed as a prompt for the generative model. The input is the search result information, and the output is the formation of a prompt for the generative model. Example prompt: "The user entered 'The TV won't turn on.' Please generate a solution by referring to the contents of the instruction manual below: [Related information]."

[0865] Step 8:

[0866] The server provides the user's input data to the emotion engine to analyze the user's emotion. For example, IBM Watson Tone Analyzer is used. The input is the user query, and the output is the emotion recognition result. API example: POST / api / tone_analyzer {"text": "The TV won't turn on"}. Result format: {"emotion": "Confused"}

[0867] Step 9:

[0868] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. The input is the prompt and emotion recognition result for the generative model, and the output is the generated solution. Example: "Don't worry. First, make sure the power cord is properly connected."

[0869] Step 10:

[0870] The server generates an image related to the solution. It uses the Pillow library to generate an image containing specific content. The input is the solution content, and the output is the generated image. Example: image = Image.open("power_connection.png")

[0871] Step 11:

[0872] The server compiles the generated solutions and images in HTML format and sends them to the terminal. It uses an HTML template to compile information into one page. The input is the solutions and images, and the output is the content in HTML format. Example: {"html": " <h1> Solution< / h1> ...}"}

[0873] Step 12:

[0874] The terminal displays this to the user, allowing them to solve the problem quickly and easily. It uses a rendering engine to display HTML content. The input is the HTML formatted content, and the output is the displayed result.

[0875] (Application example 2)

[0876] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0877] There is a problem that users cannot find a quick and accurate solution on the spot when performing maintenance or troubleshooting on factory robots. In addition, a system that provides a uniform solution without considering the user's feelings makes it difficult to reduce the user's stress and confusion. Also, referring to the instruction manual or troubleshooting manual every time is time-consuming and does not allow for efficient problem solving.

[0878] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0879] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and emotion recognition means for recognizing the user's emotions and reflecting those emotions in the generation of a solution. This enables the user to quickly and effectively obtain a solution that takes into account the user's emotions, without having to refer to the instruction manual or troubleshooting manual.

[0880] An "instruction manual" is a document that details how to operate and configure equipment or software.

[0881] A "failure response manual" is a document that explains the response procedures and solutions required when equipment breaks down.

[0882] A "digital document" is a document that is stored and displayed electronically, including PDFs and text files.

[0883] "Means of analysis" refers to techniques and methods for understanding the content of a digital document and extracting the necessary information.

[0884] "Information about a problem or phenomenon input by a user" is detailed information about a malfunction or phenomenon of a device input by a user in natural language.

[0885] "Means for receiving" refers to the technology or method for obtaining and processing information sent by a user.

[0886] An "extraction means" is a technique or method for extracting relevant information from a digital document.

[0887] A "generative model" is an AI model that generates appropriate solutions based on user input and information extracted from digital documents.

[0888] The "presentation means" refers to a technique or method for displaying the generated solution to the user.

[0889] "Emotion recognition means" refers to a technique or method for analyzing emotions from information entered by a user and using the results to adapt a response method.

[0890] A "server" is a computer system that stores, analyzes, and processes data, and is a device that communicates with user terminals via a network.

[0891] The present invention provides a system that allows a user to quickly and accurately find a solution when performing maintenance or troubleshooting on a factory robot. Specific embodiments of the present invention will be described below.

[0892] Hardware and software used

[0893] 1. Server

[0894] Storage: Storage for saving instruction manuals and troubleshooting manuals. Specifically, it uses the server's HDD or SSD.

[0895] OCR Technology: The server converts the PDF file into text using optical character recognition (OCR) technology (e.g., Tesseract).

[0896] Database: A database is required to store the converted text data and analyze it with a natural language processing (NLP) engine. For example, MySQL or PostgreSQL will be used.

[0897] Generative AI models: Use generative AI (e.g., OpenAI models) to generate solutions based on the user query and extracted information.

[0898] Emotion recognition engine: Analyzes emotions from user information and reflects them in response. Uses natural language processing models such as BERT.

[0899] 2. Terminal

[0900] Smartphone or PC: Users use these devices to input information about problems or phenomena in natural language, and the devices then send this information to a server.

[0901] Operation overview

[0902] 1. Initial data preparation

[0903] The server stores PDF files of instruction manuals and troubleshooting manuals in storage and converts them into text data using OCR technology. The converted text data is then stored in a database and analyzed by a natural language processing engine.

[0904] 2. Accepting user queries

[0905] The user uses a terminal to input questions about the robot's problems or symptoms in natural language, and the queries are sent from the terminal to the server.

[0906] 3. Data analysis and information extraction

[0907] The server receives user queries and searches and extracts relevant information from a database.

[0908] 4. Emotional Recognition

[0909] The server provides the user's input data to the emotion recognition engine, which recognizes the user's emotions. The recognized emotions are reflected in the generative AI model.

[0910] 5. Generate solutions

[0911] The generative AI model generates a solution based on the user's query, extracted information, and recognized emotions, and presents guidance in the most appropriate language based on the user's emotions.

[0912] 6. Creating an image diagram

[0913] The server generates an image related to the solution (for example, a screenshot of the setting screen or a diagram of the connection points).

[0914] 7. Present solutions and illustrations

[0915] The server then compiles the generated solutions and images into a single page and sends it to the device, which displays it to the user, allowing them to quickly and easily solve the problem.

[0916] Specific examples

[0917] Here is a specific example where the user inputs "The robot's arm won't move."

[0918] 1. User query: "The robot arm won't move."

[0919] 2. Emotion recognition: The emotion recognition engine recognizes "confusion."

[0920] 3. Generating a solution: The generative AI model generated the following: "Don't be confused. First, turn off the robot and check if there is an abnormality in the arm connection. Next, turn it on again, and if an error code is displayed, refer to the procedures in the maintenance manual."

[0921] 4. Image generation: Image of the "arm connection part".

[0922] Prompt Sentence Examples

[0923] Please generate solutions to user queries based on the following manual.

[0924] manual:

[0925] [Instruction manual and troubleshooting manual text]

[0926] Query: Robot arm not moving

[0927] User Emotion: Confused

[0928] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0929] Step 1:

[0930] The server saves PDF files of instruction manuals and troubleshooting manuals in storage. The input of this step is the PDF files, and the output is the PDF files saved in storage. This prepares the documents to be analyzed.

[0931] Step 2:

[0932] The server converts the PDF file into text data using OCR technology (e.g., Tesseract). The input for this step is the PDF file, and the output is text data. Using OCR technology, character information is extracted from the image data and converted into text format.

[0933] Step 3:

[0934] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine. The input to this step is the text data generated by OCR, and the output is structured data of the document analyzed by NLP. As a result of the analysis, important information within the document is extracted and registered in the database.

[0935] Step 4:

[0936] The user uses a terminal to input information about the problem or phenomenon of the robot in natural language. The input of this step is a query entered by the user, and the output is a query sent from the terminal to the server. The user describes the problem in natural language, and it is sent to the server.

[0937] Step 5:

[0938] The server receives a user query and searches and extracts relevant information in the database. The input to this step is the user query, and the output is the relevant information extracted from the database. The server pulls the appropriate information from the database based on the query.

[0939] Step 6:

[0940] The server provides the user's query and extracted information to an emotion recognition engine to analyze the user's emotion. The input of this step is the user query and related information, and the output is analyzed emotion data. The emotion recognition engine is used to identify the user's emotion (e.g., confusion, anger), and the result is provided to the next step.

[0941] Step 7:

[0942] The server uses a generative AI model (e.g., OpenAI's model) to generate a solution based on the extracted information and the recognized emotion. The inputs for this step are the user query, the extracted information, and the emotion data, and the output is a solution generated by the generative AI model. The generated solution takes the user's emotion into consideration.

[0943] Step 8:

[0944] The server generates an image associated with the generated solution. The input to this step is the generated solution, and the output is an image associated with the solution. The server generates appropriate visual aids (screenshots and diagrams) to accompany the solution.

[0945] Step 9:

[0946] The server compiles the generated solutions and the image onto a single page and sends it to the terminal. The input to this step is the solutions and the image, and the output is the data sent to the terminal. The user receives this on their terminal and it is displayed.

[0947] Step 10:

[0948] The user can quickly and easily solve the problem by referring to the solution and image displayed on the terminal. The input of this step is the solution and image displayed on the terminal, and the output is the user's problem solution. The user proceeds with the solution according to the proposed method.

[0949] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0950] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0951] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0952] [Third embodiment]

[0953] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0954] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0955] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0956] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0957] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0958] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0959] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0960] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0961] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0962] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0963] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0964] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0965] System Overview

[0966] The present invention provides a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual for electronic devices. This system can analyze the digital document of the instruction manual or troubleshooting manual and generate a solution to the problem or phenomenon input by the user.

[0967] composition

[0968] Users use their own devices (smartphones or PCs) to enter information about problems or issues.

[0969] The terminal transmits the user's input information to the server.

[0970] The server analyzes the digital document and extracts the required information.

[0971] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[0972] The server presents the generated solution and related image diagram to the user.

[0973] Explanation of program processing

[0974] The detailed operation of this system will be explained below.

[0975] Initial Setup

[0976] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[0977] The server converts the PDF file into text using optical character recognition (OCR) technology.

[0978] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[0979] Query reception

[0980] The user uses the device to input a description of the issue or problem in natural language, for example, "I can't connect to Wi-Fi on my smartphone."

[0981] The terminal sends the input query to the server.

[0982] Data Analysis and Information Extraction

[0983] The server receives the user's query and searches for relevant information in a database.

[0984] The server provides the search results to the generative model.

[0985] Generate a workaround

[0986] The generative model generates a solution based on the user's query and related information provided by the server.

[0987] For example, it generates specific solutions such as "Select the network name on the Wi-Fi settings screen and enter the password."

[0988] Generate image diagrams

[0989] The server generates an image related to the solution (for example, a screenshot of the setting screen).

[0990] Presenting solutions and illustrations

[0991] The server generated solution and image diagram are summarized on one page.

[0992] The server sends the compiled information to the terminal.

[0993] The terminal displays this to the user, allowing the user to easily resolve the problem.

[0994] Specific examples

[0995] Below is a specific example where the user enters "My TV won't turn on."

[0996] 1. The user types "TV won't turn on" into the device.

[0997] 2. The device sends this query to the server.

[0998] 3. The server analyzes the TV's instruction manual using OCR technology and extracts information related to the "won't turn on" problem.

[0999] 4. Based on the extracted information and user input, the generative model generates a solution such as "Make sure the power cord is connected correctly. Then, replace the batteries in the remote control."

[1000] 5. The server generates a diagram that visually explains how to resolve the issue (for example, a diagram showing where to connect the power cord).

[1001] 6. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[1002] 7. The device should display this to the user so that they can quickly and easily understand what to do.

[1003] This allows users to solve problems efficiently without having to read the instruction manual in detail.

[1004] The processing flow will be explained below.

[1005] Step 1:

[1006] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My TV won't turn on."

[1007] Step 2:

[1008] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[1009] Step 3:

[1010] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[1011] Step 4:

[1012] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance in storage, and uses OCR technology to convert the PDFs into text data.

[1013] Step 5:

[1014] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[1015] Step 6:

[1016] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[1017] Step 7:

[1018] The server inputs the extracted information and the user's input data into a generative AI model (e.g., ChatGPT) to generate an appropriate response. The generative AI model generates the most appropriate response in natural language based on the input data.

[1019] Step 8:

[1020] The server generates specific images (e.g., screenshots of power cord connection locations or settings screens) based on the generated solutions, using automatic image generation technology and images retrieved from a database.

[1021] Step 9:

[1022] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[1023] Step 10:

[1024] The server then sends the final information to the device, including specific solutions and related images.

[1025] Step 11:

[1026] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem, thus enabling the user to effectively respond without having to refer to an instruction manual or instruction manual.

[1027] Example 1

[1028] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1029] In the past, when users wanted to refer to instruction manuals or troubleshooting manuals for electronic devices, they had to refer to paper documents or search websites, which was time-consuming and laborious. Furthermore, it was difficult to quickly obtain specific solutions, which reduced the efficiency of problem-solving. In particular, when dealing with electronic device malfunctions, accurate and prompt information provision is required.

[1030] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1031] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for converting the digital document into text using optical character recognition technology, means for extracting relevant information from the digital document, means for searching for relevant information in a database, means for including a generative AI model that generates solutions based on the extracted information and information input by the user, means for generating an image corresponding to the generated solution, and means for presenting the generated solution and the image to the user. This allows the user to solve problems quickly and effectively without having to read the instruction manual in detail.

[1032] An "instruction manual or troubleshooting manual" is a document that describes how to use an electronic device and what to do if a problem occurs.

[1033] A "digital document" is a document in a form that is stored, displayed, and transmitted electronically.

[1034] "Means for analysis" refers to technology or devices that have the ability to extract necessary information from digital documents.

[1035] "Means for receiving information about a problem or phenomenon input by a user" refers to a technology or device that has the function of transmitting information input by a user through a terminal to a server and receiving that information.

[1036] "Optical Character Recognition (OCR)" is a technology that converts characters in an image into electronic text.

[1037] A "database" is a system for efficiently storing, retrieving, and managing structured data.

[1038] A "natural language processing engine" is an artificial intelligence technology that analyzes text data, extracts semantic information, and understands language.

[1039] A "generative AI model" is an artificial intelligence model that generates useful information, such as countermeasures, based on input information from users and related data.

[1040] An "image" is a visual illustration that supplements text information.

[1041] An "HTTP request" is a communication protocol used by a web browser or application to request data from a server.

[1042] A "server" is a computer system that processes queries, stores data, and communicates over a network.

[1043] A "terminal" is a device that a user uses to access a server, such as a smartphone or PC.

[1044] A "prompt" is a text question or instruction used as input to a generative AI model.

[1045] A "means for converting to text" is a technique or device that converts characters in an image into electronic text using optical character recognition technology.

[1046] System Overview

[1047] This invention is a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual of the electronic device. Specifically, it analyzes digital documents using optical character recognition (OCR) technology and a natural language processing engine (NLP), generates solutions using a generative AI model, and presents them to the user.

[1048] Hardware and software used

[1049] Server: A computer system for storing, analyzing, and storing digital documents in a database.

[1050] Storage engine: Amazon S3, etc.

[1051] OCR technology: Tesseract OCR

[1052] Database: MySQL

[1053] Natural language processing engine: spaCy

[1054] Generative AI models: ChatGPT, etc.

[1055] Terminal: A device through which a user accesses a server and enters queries.

[1056] Devices: Smartphones, PCs, etc.

[1057] Program processing explanation

[1058] Information storage and analysis

[1059] The server stores PDF files of instruction manuals and troubleshooting manuals in a storage engine, for example, in an Amazon S3 bucket.

[1060] The server converts the saved PDF file to text using Tesseract OCR, specifically by running the command tesseract manual.pdf output.txt.

[1061] The server stores the converted text data in a MySQL database, for example by executing the SQL statement INSERT INTO manuals (text_data) VALUES ('...').

[1062] The server analyzes the text data using the spaCy NLP engine to extract semantic information, for example by running the following code: nlp = spacy.load('en_core_web_sm').

[1063] Query reception and processing

[1064] The user inputs a symptom or problem in natural language through the terminal. For example, the user inputs "The TV won't turn on."

[1065] The device sends this query to the server via an HTTP POST request.

[1066] Data analysis and generation of action plans

[1067] The server receives the user's query and searches the database for relevant information using an SQL query, for example SELECT FROM manuals WHERE text_data LIKE '%cannot power on%'.

[1068] The server sends the relevant information to a generative AI model (e.g., ChatGPT) via an API request. An example prompt sentence is "The user explains that the TV won't turn on. What should I do?"

[1069] The generative AI model generates a solution and sends it back to the server, such as "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1070] Presenting solutions and illustrations

[1071] The server uses the Pillow library to generate an image related to the solution, for example, a diagram showing where to connect the power cord.

[1072] The server integrates the generated solutions and images into an HTML template and compiles them into a single page.

[1073] The server sends the compiled information to the terminal as an HTTP response.

[1074] The terminal displays the information to the user and provides a quick and effective way to deal with the problem.

[1075] Specific examples

[1076] For example, if the user inputs "TV won't turn on," the system will operate as follows:

[1077] 1. The user types "TV won't turn on" into the device.

[1078] 2. The device sends the query to the server via an HTTP POST request.

[1079] 3. The server receives the query and searches the database for relevant information using an SQL query.

[1080] 4. The server sends the relevant information to the generative AI model via an API request.

[1081] 5. The generative AI model generates a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1082] 6. The server generates a diagram showing where the power cords are connected.

[1083] 7. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[1084] 8. The device displays information to the user and provides a quick and effective way to deal with the problem.

[1085] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1086] Step 1: Initial Setup

[1087] The server stores PDF files of instruction manuals and troubleshooting manuals in the storage engine. As a concrete example, using Amazon S3, the command aws s3 cp manual.pdf s3: / / bucket-name / is executed as input, and the output is the PDF file stored in Amazon S3.

[1088] The server converts PDF files to text using Tesseract OCR. Specifically, it receives input from the command tesseract manual.pdf output.txt and generates an output file called output.txt.

[1089] The server stores the converted text data in a MySQL database. For example, it takes input as input and executes the SQL statement INSERT INTO manuals (text_data) VALUES ('...'), and gets output as text data stored in the database.

[1090] Step 2: Enter and accept user queries

[1091] The user inputs the phenomenon or problem in natural language using the terminal. For example, the input "The TV won't turn on" is received, and the output is the manually input data in text format.

[1092] The terminal sends the entered query to the server. Specifically, it receives input to send the entered query to the server via an HTTP POST request, and obtains output showing that the query has reached the server.

[1093] Step 3: Data analysis and information extraction

[1094] The server receives a user query and searches for relevant information in the database. Specifically, it receives input by executing the SQL query SELECT FROM manuals WHERE text_data LIKE '%TV won't turn on%' and obtains output that provides relevant information (for example, the contents of the corresponding instruction manual).

[1095] The server provides the relevant information to the generative AI model, taking the specific action of sending the results of a database search to the generative AI model via an API request as input, and obtaining an output ready for the generative AI model to generate a response.

[1096] Step 4: Generate a solution

[1097] The generative AI model generates a solution based on the user's query and information provided by the server. Specifically, the model receives a prompt, "The user explains that the TV won't turn on. What should I do?", and outputs a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1098] The server receives the generated solutions and performs filtering. Specifically, it receives the generated text, deletes unnecessary parts, and outputs solutions that are ready to be provided to the user.

[1099] Step 5: Generate an image

[1100] The server uses the Pillow library to generate an image diagram related to the solution. For example, for the command "Please check the power cord connection," input is made using an image editing tool, and an output is generated showing the power cord connection location.

[1101] Step 6: Present solutions and images

[1102] The server then combines the generated solutions and the resulting image onto a single page. Specifically, it uses an HTML template as input to integrate the solutions and the image, and outputs a single page that can be presented to the user.

[1103] The server sends the compiled information to the terminal. As a specific input, it performs the operation of sending the generated page as an HTTP response, and obtains the output that the page arrives at the user terminal.

[1104] The device displays the information to the user. Specifically, the device inputs the received page to display in a browser or dedicated app, and the user receives an output that allows them to visually confirm the solution to the problem.

[1105] (Application example 1)

[1106] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1107] The present invention relates to a system for quickly and effectively resolving problems without referring to instruction manuals or troubleshooting manuals. However, conventional systems require users to accurately input details of the problem, which can lead to errors, especially when using voice input. Furthermore, solutions are presented only in text, which can be difficult for users to understand. Furthermore, it is difficult for factory workers to respond to equipment failures without visual information. It is necessary to solve these problems.

[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1109] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, means for converting voice input into text, and means for generating a 3D model or diagram of the solution. This allows the user to input a problem using voice input, and the solution is presented not only as text but also as a 3D model or diagram, enabling the user to solve the problem quickly and visually.

[1110] An "instruction manual" is a document that provides detailed instructions on how to use and operate an electronic device or machine.

[1111] A "troubleshooting manual" is a document that explains how to repair and resolve problems when electronic devices or machinery break down.

[1112] "Digital documents" are document data stored electronically, including PDFs and text files.

[1113] "Means of analysis" refers to the technology and devices used to read digital documents and understand and classify the necessary information.

[1114] "Means for receiving" refers to technology or devices for receiving information or data input by a user.

[1115] An "extraction means" is a technique or device used to identify and extract the required information from a digital document.

[1116] A "generative model" is an artificial intelligence model that automatically generates appropriate solutions or answers based on input information.

[1117] The "presentation means" refers to a technique or device for displaying the generated solutions and information to the user.

[1118] "Means for converting voice input into text" refers to technology or devices for converting voice input by a user into text data.

[1119] "Means for generating 3D models or diagrams of solutions" refers to technology or devices for expressing solutions in 3D models or figures to make them visually easier to understand.

[1120] "Server" means a computer system for processing, storing, and managing data.

[1121] This invention is a system that can quickly and effectively solve problems without the user having to refer to an instruction manual or a troubleshooting manual. This system is particularly useful for effectively dealing with factory robot failures. The specific system configuration and its operation will be described below.

[1122] System Configuration

[1123] 1. User Device

[1124] Wearable devices such as smart glasses.

[1125] It has voice input and display functions.

[1126] 2. Server

[1127] A cloud server for analyzing and storing digital documents.

[1128] It has high-performance computing resources and data storage.

[1129] 3. OCR technology

[1130] Convert PDF to text data using optical character recognition technology. Specifically, we use pytesseract.

[1131] 4. Voice Recognition Technology

[1132] It uses technology to convert voice input into text data, specifically, the speech_recognition library.

[1133] 5. Generative AI Models

[1134] It is used for natural language processing and countermeasure generation. Specifically, it uses the GPT-3 model from the transformers library.

[1135] 6. Digital Documents

[1136] PDF documents such as instruction manuals and troubleshooting manuals.

[1137] Data processing and calculation

[1138] 1. Digital Document Analysis

[1139] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using optical character recognition (OCR). This is where pytesseract comes in.

[1140] The converted text data is stored in a database on the server.

[1141] 2. Processing voice input

[1142] The user inputs the symptoms or problems by voice through the smart glasses.

[1143] The smart glasses convert the voice data into text data and send this data to the server, where the speech_recognition library is used.

[1144] 3. Extracting information and generating solutions

[1145] The server analyzes the received user query and searches for relevant information in a database.

[1146] Based on the search results, a generative model (e.g., GPT-3) is used to generate a solution. The transformers library is used.

[1147] 4. Providing solutions and visual information

[1148] Use software to convert the generated solutions into 3D models and diagrams.

[1149] The solution and 3D models or diagrams are sent to the smart glasses and visually presented to the user.

[1150] Specific examples

[1151] Example 1: Error handling in an automobile factory

[1152] User input: "The robot arm joints won't move."

[1153] Generated solution: "Oil the joints of the robot arm and check their operation. Also, check for any sensor error indications."

[1154] Prompt Sentence Examples

[1155] Problem: The robot arm joints won't move

[1156] How to respond:

[1157] This allows factory workers to quickly identify problems and respond effectively. The system is expected to improve work efficiency and reduce downtime through visual and audio assistance.

[1158] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1159] Step 1:

[1160] Digital document analysis

[1161] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using OCR technology. The server reads the PDF using pytesseract and converts each page into text data.

[1162] Input: PDF format instruction manual or troubleshooting manual

[1163] Output: Text data

[1164] Specific behavior: The server extracts text from the PDF and stores it in a database.

[1165] Step 2:

[1166] Receiving audio input

[1167] The user inputs the symptoms or problems by voice through the smart glasses.

[1168] The smart glasses capture voice data and convert it to text data using the speech_recognition library.

[1169] Input: Audio data of the phenomenon or problem

[1170] Output: Text data

[1171] How it works: The user speaks, "The robot arm's joints won't move," and the smart glasses convert the speech into text.

[1172] Step 3:

[1173] Submitting a query

[1174] The smart glasses send the converted text data to the server.

[1175] Input: Text data converted from speech

[1176] Output: Text data sent to the server

[1177] Specific operation: The smart glasses generate text data and send it to a server via the network.

[1178] Step 4:

[1179] Information extraction and analysis

[1180] The server analyzes the received text data (user queries) and searches for relevant information in the database using a natural language processing (NLP) engine.

[1181] Input: User query text data

[1182] Output: Search results containing relevant information

[1183] Specific operation: The server extracts information related to the query from instruction manuals and manuals in the database.

[1184] Step 5:

[1185] Generate a workaround

[1186] The server generates a solution using a generative model (GPT-3) based on relevant information. It uses the transformers library.

[1187] Input: Search results and user text query

[1188] Output: Text data of the solution

[1189] Specific operation: The server performs NLP processing and generates the optimal solution.

[1190] Step 6:

[1191] Visual information generation

[1192] The server generates 3D models and diagrams based on the generated solutions. 3D design software and image generation software are used to make the solutions easier to understand visually.

[1193] Input: Text data of the solution

[1194] Output: 3D models and diagrams

[1195] Specific operation: The server converts the solution into a 3D model or diagram, generating visual information.

[1196] Step 7:

[1197] Coping methods and visual information

[1198] The server sends the generated solutions and 3D models or diagrams to the smart glasses.

[1199] The smart glasses present the received action and visual information to the user.

[1200] Input: Text data and visual information of the solution

[1201] Output: Actions and visual information displayed on the smart glasses

[1202] Specific operation: The server sends information, and the smart glasses display it and present it to the user.

[1203] This allows the user to input a problem using voice input, and quickly check the generated solutions and visual information to help solve the problem.

[1204] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1205] System Overview

[1206] The present invention provides a system that allows users to quickly and effectively solve problems without referring to an instruction manual or troubleshooting manual. This system not only analyzes the digital document of the instruction manual or troubleshooting manual and generates solutions to problems or phenomena input by the user, but also recognizes the user's emotions and reflects them in the solutions it presents.

[1207] composition

[1208] The user uses their own device (smartphone or PC) to input information about the problem or phenomenon in natural language.

[1209] The terminal transmits the user's input information to the server.

[1210] The server analyzes the digital document and extracts the required information.

[1211] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[1212] The emotion engine recognizes emotions from user information and reflects them in the generated response methods.

[1213] The server presents the generated solution and related image diagram to the user.

[1214] Explanation of program processing

[1215] The detailed operation of this system will be described below.

[1216] Initial Setup

[1217] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[1218] The server converts the PDF file into text using optical character recognition (OCR) technology.

[1219] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[1220] Query reception

[1221] The user uses a terminal to input information about a phenomenon or problem in natural language.

[1222] The terminal sends the input query to the server.

[1223] Data Analysis and Information Extraction

[1224] The server receives the user's query and searches for relevant information in a database.

[1225] The server provides the search results to the generative model.

[1226] Emotion recognition

[1227] The server provides the user's input data to the emotion engine to analyze the user's emotions.

[1228] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and reflects the results in the generative model.

[1229] Generate a workaround

[1230] The generative model generates a response based on the user's query, related information provided by the server, and the recognized emotions.

[1231] For example, if the problem is "the TV won't turn on" and the emotion is "confused," you might preface the situation by saying, "Let's start with a simple method that anyone can do," and then explain, "First, check that the power cord is connected correctly."

[1232] Generate image diagrams

[1233] The server generates an image related to the solution (for example, a screenshot of the settings screen or a diagram of the connection points).

[1234] Presenting solutions and illustrations

[1235] The server generated solution and image diagram are summarized on one page.

[1236] The server sends the compiled information to the terminal.

[1237] The terminal displays this to the user, allowing the user to quickly and easily resolve the problem.

[1238] Specific examples

[1239] Below is a specific example of what happens when a user enters "My smartphone's Wi-Fi won't connect."

[1240] 1. The user types "My smartphone's Wi-Fi won't connect" into the device.

[1241] 2. The device sends this query to the server.

[1242] 3. The server analyzes the smartphone's instruction manual using OCR technology and extracts relevant information.

[1243] 4. The server provides the user's input data to the emotion engine and recognizes the emotion "confused."

[1244] 5. Based on the relevant information and the recognized emotion, the generative model generates specific countermeasures such as, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[1245] 6. The server generates a screenshot associated with this method.

[1246] 7. The server compiles a solution and an illustration and sends it to the terminal.

[1247] 8. The device displays this to the user, allowing them to quickly and easily resolve the issue.

[1248] This system allows users to find appropriate ways to deal with their emotions without having to refer to an instruction manual.

[1249] The processing flow will be explained below.

[1250] Step 1:

[1251] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My smartphone's Wi-Fi won't connect."

[1252] Step 2:

[1253] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[1254] Step 3:

[1255] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[1256] Step 4:

[1257] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance, and uses OCR technology to convert the PDFs into text data.

[1258] Step 5:

[1259] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[1260] Step 6:

[1261] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[1262] Step 7:

[1263] The server provides the user's input data to the emotion engine, which analyzes the user's emotions. The emotion engine recognizes the user's emotions based on the input content and past data.

[1264] Step 8:

[1265] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and returns the results to the server, which then inputs this information into the generative model.

[1266] Step 9:

[1267] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. For example, if the user's emotion is "confused" when the query is "Wi-Fi is not connecting," the model will explain, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[1268] Step 10:

[1269] Based on the solution generated by the server, a concrete image diagram (for example, a screenshot of the Wi-Fi setting screen or a diagram of the connection location) is generated.

[1270] Step 11:

[1271] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[1272] Step 12:

[1273] The server then sends the final information to the device, including specific solutions and related images.

[1274] Step 13:

[1275] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem. This allows the user to respond efficiently without having to refer to an instruction manual or instruction manual.

[1276] Example 2

[1277] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1278] There is a demand for systems that allow users to solve problems quickly and effectively without referring to instruction manuals or troubleshooting manuals. There is also a demand for systems that improve user satisfaction by presenting solutions that take into account the user's emotions.

[1279] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1280] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and means including an emotion recognition engine for recognizing the user's emotions and reflecting them in the generated solution. This allows the user to quickly obtain an appropriate solution based on their emotions without having to refer to the instruction manual.

[1281] An "instruction manual or troubleshooting manual" is a detailed guideline on how to use a product and what to do in the event of a malfunction.

[1282] "Digital documents" are documents that are stored and managed electronically, including PDF and TXT formats.

[1283] "Means of analysis" refers to the techniques and methods used to read digital documents and understand and process their contents.

[1284] "Information about the problem or phenomenon input by the user" refers to the content of the phenomenon input by the user and the problems related to it.

[1285] "Means for receiving" refers to an interface or method for obtaining input information from a user.

[1286] "Means for extracting relevant information" refers to techniques or methods for extracting specified information from a digital document.

[1287] A "generative model" is an algorithm or system that generates new data or information based on given data.

[1288] "Presentation means" refers to the method or interface for displaying the generated information to the user.

[1289] An "emotion recognition engine" is a technology or system that analyzes and recognizes emotions from user input data.

[1290] The present invention provides a system that allows users to quickly and effectively solve problems without referring to instruction manuals or troubleshooting manuals. This system analyzes digital documents, generates solutions to problems or phenomena input by the user, and recognizes the user's emotions and reflects them in the solutions it presents.

[1291] System Overview

[1292] The system includes the following components:

[1293] Server: PDF files of instruction manuals and troubleshooting manuals are stored in storage and converted to text using OCR technology (e.g., Tesseract OCR). The converted text data is stored in a database (e.g., MySQL) and analyzed using a natural language processing (NLP) engine (e.g., spaCy).

[1294] Terminal: Provides an interface for users to input information about problems or phenomena in natural language and sends the input information to the server.

[1295] Generative model: Generates a solution based on relevant information provided by the server and user input (e.g., OpenAI's ChatGPT).

[1296] Emotion recognition engine: Analyzes emotions based on user input data (e.g., IBM Watson Tone Analyzer) and reflects the results in a generative model.

[1297] Presentation method: The generated solution and related image diagram are presented to the user.

[1298] How it works

[1299] 1. The server converts PDF files of instruction manuals and troubleshooting manuals into text using optical character recognition (OCR) technology. For example, by using Tesseract OCR, the PDF files are output as text files and the contents are stored in a database.

[1300] 2. The user uses the device to input a natural language description of the symptom or problem, for example, "My TV won't turn on."

[1301] 3. The device sends the entered query to the server via an HTTP POST request, specifying the server's API endpoint to send the data.

[1302] 4. The server receives the user's query and searches the database for relevant information using an SQL query, such as SELECT content FROM manuals_table WHERE content LIKE '%won't turn on%'.

[1303] 5. The server provides the search results to the generative model, which processes the content provided as prompts for the generative model.

[1304] 6. The server provides the user's input data to the emotion engine to analyze the emotion, for example, using IBM Watson Tone Analyzer to recognize the user's emotion.

[1305] 7. The generative model generates specific countermeasures based on the user's query, related information provided by the server, and the recognized emotions.

[1306] 8. The server generates an image diagram related to the generated solution, for example, using the Pillow library to generate the related diagram.

[1307] 9. The server compiles the generated solutions and images in HTML format and sends them to the terminal, where they are displayed to the user, allowing them to solve the problem quickly and easily.

[1308] Specific examples

[1309] Below is an example of a specific prompt sentence when a user enters "My smartphone's Wi-Fi won't connect."

[1310] "The user inputs, 'My smartphone's Wi-Fi won't connect.' Provide relevant information from the instruction manual. Furthermore, the user's emotion, 'confused,' is recognized. Based on this information, generate a specific solution."

[1311] This invention allows users to quickly find appropriate solutions according to their emotions without having to refer to an instruction manual or troubleshooting manual. Furthermore, by using an emotion recognition engine, it becomes possible to respond to each user in a way that is tailored to their individual needs, thereby improving user satisfaction.

[1312] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1313] Step 1:

[1314] The server saves PDF files of instruction manuals and troubleshooting manuals to storage. First, it saves the PDF files provided by the user in a dedicated directory called "manuals." For example, the file path is / var / www / manuals / manual1.pdf. The file path after saving is output.

[1315] Step 2:

[1316] The server converts the PDF file to text using optical character recognition (OCR). It then uses Tesseract OCR to output the PDF file as a text file. The input is the path to the saved PDF file, and the output is the path to the converted text file. Example: / var / www / manuals / manual1.txt

[1317] Step 3:

[1318] The server saves the converted text data in the database. It uses a MySQL database and saves the text data in a table (e.g. manuals_table). The input is the path to the text file and its contents, and the output is the updated result in the database. Example SQL query: INSERT INTO manuals_table (manual_id, content) VALUES (1, LOAD_FILE(' / var / www / manuals / manual1.txt'))

[1319] Step 4:

[1320] The user uses the device to input information about the problem or issue in natural language, for example, "The TV won't turn on." This input information is used in the next step.

[1321] Step 5:

[1322] The device sends the entered query to the server via an HTTP POST request. Specify the server's API endpoint and send the query data. The input is the user-entered query, and the output is the result of the HTTP request sent to the server. Example: POST / api / query {"query": "The TV won't turn on"}

[1323] Step 6:

[1324] The server receives a user query and searches for relevant information in the database. It uses SQL queries to retrieve relevant information from the database based on the input query. The input is the user query, and the output is the search results for relevant information. SQL query example: SELECT content FROM manuals_table WHERE content LIKE '%Power won't turn on%'

[1325] Step 7:

[1326] The server provides the search results to the generative model. The provided information is processed as a prompt for the generative model. The input is the search result information, and the output is the formation of a prompt for the generative model. Example prompt: "The user entered 'The TV won't turn on.' Please generate a solution by referring to the contents of the instruction manual below: [Related information]."

[1327] Step 8:

[1328] The server provides the user's input data to the emotion engine to analyze the user's emotion. For example, IBM Watson Tone Analyzer is used. The input is the user query, and the output is the emotion recognition result. API example: POST / api / tone_analyzer {"text": "The TV won't turn on"}. Result format: {"emotion": "Confused"}

[1329] Step 9:

[1330] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. The input is the prompt and emotion recognition result for the generative model, and the output is the generated solution. Example: "Don't worry. First, make sure the power cord is properly connected."

[1331] Step 10:

[1332] The server generates an image related to the solution. It uses the Pillow library to generate an image containing specific content. The input is the solution content, and the output is the generated image. Example: image = Image.open("power_connection.png")

[1333] Step 11:

[1334] The server compiles the generated solutions and images in HTML format and sends them to the terminal. It uses an HTML template to compile information into one page. The input is the solutions and images, and the output is the content in HTML format. Example: {"html": " <h1> Solution< / h1> ...}"}

[1335] Step 12:

[1336] The terminal displays this to the user, allowing them to solve the problem quickly and easily. It uses a rendering engine to display HTML content. The input is the HTML formatted content, and the output is the displayed result.

[1337] (Application example 2)

[1338] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1339] There is a problem that users cannot find a quick and accurate solution on the spot when performing maintenance or troubleshooting on factory robots. In addition, a system that provides a uniform solution without considering the user's feelings makes it difficult to reduce the user's stress and confusion. Also, referring to the instruction manual or troubleshooting manual every time is time-consuming and does not allow for efficient problem solving.

[1340] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1341] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and emotion recognition means for recognizing the user's emotions and reflecting those emotions in the generation of a solution. This enables the user to quickly and effectively obtain a solution that takes into account the user's emotions, without having to refer to the instruction manual or troubleshooting manual.

[1342] An "instruction manual" is a document that details how to operate and configure equipment or software.

[1343] A "failure response manual" is a document that explains the response procedures and solutions required when equipment breaks down.

[1344] A "digital document" is a document that is stored and displayed electronically, including PDFs and text files.

[1345] "Means of analysis" refers to techniques and methods for understanding the content of a digital document and extracting the necessary information.

[1346] "Information about a problem or phenomenon input by a user" is detailed information about a malfunction or phenomenon of a device input by a user in natural language.

[1347] "Means for receiving" refers to the technology or method for obtaining and processing information sent by a user.

[1348] An "extraction means" is a technique or method for extracting relevant information from a digital document.

[1349] A "generative model" is an AI model that generates appropriate solutions based on user input and information extracted from digital documents.

[1350] The "presentation means" refers to a technique or method for displaying the generated solution to the user.

[1351] "Emotion recognition means" refers to a technique or method for analyzing emotions from information entered by a user and using the results to adapt a response method.

[1352] A "server" is a computer system that stores, analyzes, and processes data, and is a device that communicates with user terminals via a network.

[1353] The present invention provides a system that allows a user to quickly and accurately find a solution when performing maintenance or troubleshooting on a factory robot. Specific embodiments of the present invention will be described below.

[1354] Hardware and software used

[1355] 1. Server

[1356] Storage: Storage for saving instruction manuals and troubleshooting manuals. Specifically, it uses the server's HDD or SSD.

[1357] OCR Technology: The server converts the PDF file into text using optical character recognition (OCR) technology (e.g., Tesseract).

[1358] Database: A database is required to store the converted text data and analyze it with a natural language processing (NLP) engine. For example, MySQL or PostgreSQL will be used.

[1359] Generative AI models: Use generative AI (e.g., OpenAI models) to generate solutions based on the user query and extracted information.

[1360] Emotion recognition engine: Analyzes emotions from user information and reflects them in response. Uses natural language processing models such as BERT.

[1361] 2. Terminal

[1362] Smartphone or PC: Users use these devices to input information about problems or phenomena in natural language, and the devices then send this information to a server.

[1363] Operation overview

[1364] 1. Initial data preparation

[1365] The server stores PDF files of instruction manuals and troubleshooting manuals in storage and converts them into text data using OCR technology. The converted text data is then stored in a database and analyzed by a natural language processing engine.

[1366] 2. Accepting user queries

[1367] The user uses a terminal to input questions about the robot's problems or symptoms in natural language, and the queries are sent from the terminal to the server.

[1368] 3. Data analysis and information extraction

[1369] The server receives user queries and searches and extracts relevant information from a database.

[1370] 4. Emotional Recognition

[1371] The server provides the user's input data to the emotion recognition engine, which recognizes the user's emotions. The recognized emotions are reflected in the generative AI model.

[1372] 5. Generate solutions

[1373] The generative AI model generates a solution based on the user's query, extracted information, and recognized emotions, and presents guidance in the most appropriate language based on the user's emotions.

[1374] 6. Creating an image diagram

[1375] The server generates an image related to the solution (for example, a screenshot of the setting screen or a diagram of the connection points).

[1376] 7. Present solutions and illustrations

[1377] The server then compiles the generated solutions and images into a single page and sends it to the device, which displays it to the user, allowing them to quickly and easily solve the problem.

[1378] Specific examples

[1379] Here is a specific example where the user inputs "The robot's arm won't move."

[1380] 1. User query: "The robot arm won't move."

[1381] 2. Emotion recognition: The emotion recognition engine recognizes "confusion."

[1382] 3. Generating a solution: The generative AI model generated the following: "Don't be confused. First, turn off the robot and check if there is an abnormality in the arm connection. Next, turn it on again, and if an error code is displayed, refer to the procedures in the maintenance manual."

[1383] 4. Image generation: Image of the "arm connection part".

[1384] Prompt Sentence Examples

[1385] Please generate solutions to user queries based on the following manual.

[1386] manual:

[1387] [Instruction manual and troubleshooting manual text]

[1388] Query: Robot arm not moving

[1389] User Emotion: Confused

[1390] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1391] Step 1:

[1392] The server saves PDF files of instruction manuals and troubleshooting manuals in storage. The input of this step is the PDF files, and the output is the PDF files saved in storage. This prepares the documents to be analyzed.

[1393] Step 2:

[1394] The server converts the PDF file into text data using OCR technology (e.g., Tesseract). The input for this step is the PDF file, and the output is text data. Using OCR technology, character information is extracted from the image data and converted into text format.

[1395] Step 3:

[1396] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine. The input to this step is the text data generated by OCR, and the output is structured data of the document analyzed by NLP. As a result of the analysis, important information within the document is extracted and registered in the database.

[1397] Step 4:

[1398] The user uses a terminal to input information about the problem or phenomenon of the robot in natural language. The input of this step is a query entered by the user, and the output is a query sent from the terminal to the server. The user describes the problem in natural language, and it is sent to the server.

[1399] Step 5:

[1400] The server receives a user query and searches and extracts relevant information in the database. The input to this step is the user query, and the output is the relevant information extracted from the database. The server pulls the appropriate information from the database based on the query.

[1401] Step 6:

[1402] The server provides the user's query and extracted information to an emotion recognition engine to analyze the user's emotion. The input of this step is the user query and related information, and the output is analyzed emotion data. The emotion recognition engine is used to identify the user's emotion (e.g., confusion, anger), and the result is provided to the next step.

[1403] Step 7:

[1404] The server uses a generative AI model (e.g., OpenAI's model) to generate a solution based on the extracted information and the recognized emotion. The inputs for this step are the user query, the extracted information, and the emotion data, and the output is a solution generated by the generative AI model. The generated solution takes the user's emotion into consideration.

[1405] Step 8:

[1406] The server generates an image associated with the generated solution. The input to this step is the generated solution, and the output is an image associated with the solution. The server generates appropriate visual aids (screenshots and diagrams) to accompany the solution.

[1407] Step 9:

[1408] The server compiles the generated solutions and the image onto a single page and sends it to the terminal. The input to this step is the solutions and the image, and the output is the data sent to the terminal. The user receives this on their terminal and it is displayed.

[1409] Step 10:

[1410] The user can quickly and easily solve the problem by referring to the solution and image displayed on the terminal. The input of this step is the solution and image displayed on the terminal, and the output is the user's problem solution. The user proceeds with the solution according to the proposed method.

[1411] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1412] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1413] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1414] [Fourth embodiment]

[1415] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1416] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1417] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1418] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1419] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1420] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1421] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1422] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1423] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1424] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1425] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1426] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1427] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1428] System Overview

[1429] The present invention provides a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual for electronic devices. This system can analyze the digital document of the instruction manual or troubleshooting manual and generate a solution to the problem or phenomenon input by the user.

[1430] composition

[1431] Users use their own devices (smartphones or PCs) to enter information about problems or issues.

[1432] The terminal transmits the user's input information to the server.

[1433] The server analyzes the digital document and extracts the required information.

[1434] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[1435] The server presents the generated solution and related image diagram to the user.

[1436] Explanation of program processing

[1437] The detailed operation of this system will be explained below.

[1438] Initial Setup

[1439] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[1440] The server converts the PDF file into text using optical character recognition (OCR) technology.

[1441] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[1442] Query reception

[1443] The user uses the device to input a description of the issue or problem in natural language, for example, "I can't connect to Wi-Fi on my smartphone."

[1444] The terminal sends the input query to the server.

[1445] Data Analysis and Information Extraction

[1446] The server receives the user's query and searches for relevant information in a database.

[1447] The server provides the search results to the generative model.

[1448] Generate a workaround

[1449] The generative model generates a solution based on the user's query and related information provided by the server.

[1450] For example, it generates specific solutions such as "Select the network name on the Wi-Fi settings screen and enter the password."

[1451] Generate image diagrams

[1452] The server generates an image related to the solution (for example, a screenshot of the setting screen).

[1453] Presenting solutions and illustrations

[1454] The server generated solution and image diagram are summarized on one page.

[1455] The server sends the compiled information to the terminal.

[1456] The terminal displays this to the user, allowing the user to easily resolve the problem.

[1457] Specific examples

[1458] Below is a specific example where the user enters "My TV won't turn on."

[1459] 1. The user types "TV won't turn on" into the device.

[1460] 2. The device sends this query to the server.

[1461] 3. The server analyzes the TV's instruction manual using OCR technology and extracts information related to the "won't turn on" problem.

[1462] 4. Based on the extracted information and user input, the generative model generates a solution such as "Make sure the power cord is connected correctly. Then, replace the batteries in the remote control."

[1463] 5. The server generates a diagram that visually explains how to resolve the issue (for example, a diagram showing where to connect the power cord).

[1464] 6. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[1465] 7. The device should display this to the user so that they can quickly and easily understand what to do.

[1466] This allows users to solve problems efficiently without having to read the instruction manual in detail.

[1467] The processing flow will be explained below.

[1468] Step 1:

[1469] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My TV won't turn on."

[1470] Step 2:

[1471] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[1472] Step 3:

[1473] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[1474] Step 4:

[1475] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance in storage, and uses OCR technology to convert the PDFs into text data.

[1476] Step 5:

[1477] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[1478] Step 6:

[1479] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[1480] Step 7:

[1481] The server inputs the extracted information and the user's input data into a generative AI model (e.g., ChatGPT) to generate an appropriate response. The generative AI model generates the most appropriate response in natural language based on the input data.

[1482] Step 8:

[1483] The server generates specific images (e.g., screenshots of power cord connection locations or settings screens) based on the generated solutions, using automatic image generation technology and images retrieved from a database.

[1484] Step 9:

[1485] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[1486] Step 10:

[1487] The server then sends the final information to the device, including specific solutions and related images.

[1488] Step 11:

[1489] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem, thus enabling the user to effectively respond without having to refer to an instruction manual or instruction manual.

[1490] Example 1

[1491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1492] In the past, when users wanted to refer to instruction manuals or troubleshooting manuals for electronic devices, they had to refer to paper documents or search websites, which was time-consuming and laborious. Furthermore, it was difficult to quickly obtain specific solutions, which reduced the efficiency of problem-solving. In particular, when dealing with electronic device malfunctions, accurate and prompt information provision is required.

[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1494] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for converting the digital document into text using optical character recognition technology, means for extracting relevant information from the digital document, means for searching for relevant information in a database, means for including a generative AI model that generates solutions based on the extracted information and information input by the user, means for generating an image corresponding to the generated solution, and means for presenting the generated solution and the image to the user. This allows the user to solve problems quickly and effectively without having to read the instruction manual in detail.

[1495] An "instruction manual or troubleshooting manual" is a document that describes how to use an electronic device and what to do if a problem occurs.

[1496] A "digital document" is a document in a form that is stored, displayed, and transmitted electronically.

[1497] "Means for analysis" refers to technology or devices that have the ability to extract necessary information from digital documents.

[1498] "Means for receiving information about a problem or phenomenon input by a user" refers to a technology or device that has the function of transmitting information input by a user through a terminal to a server and receiving that information.

[1499] "Optical Character Recognition (OCR)" is a technology that converts characters in an image into electronic text.

[1500] A "database" is a system for efficiently storing, retrieving, and managing structured data.

[1501] A "natural language processing engine" is an artificial intelligence technology that analyzes text data, extracts semantic information, and understands language.

[1502] A "generative AI model" is an artificial intelligence model that generates useful information, such as countermeasures, based on input information from users and related data.

[1503] An "image" is a visual illustration that supplements text information.

[1504] An "HTTP request" is a communication protocol used by a web browser or application to request data from a server.

[1505] A "server" is a computer system that processes queries, stores data, and communicates over a network.

[1506] A "terminal" is a device that a user uses to access a server, such as a smartphone or PC.

[1507] A "prompt" is a text question or instruction used as input to a generative AI model.

[1508] A "means for converting to text" is a technique or device that converts characters in an image into electronic text using optical character recognition technology.

[1509] System Overview

[1510] This invention is a system that allows users to quickly and effectively solve problems without referring to the instruction manual or troubleshooting manual of the electronic device. Specifically, it analyzes digital documents using optical character recognition (OCR) technology and a natural language processing engine (NLP), generates solutions using a generative AI model, and presents them to the user.

[1511] Hardware and software used

[1512] Server: A computer system for storing, analyzing, and storing digital documents in a database.

[1513] Storage engine: Amazon S3, etc.

[1514] OCR technology: Tesseract OCR

[1515] Database: MySQL

[1516] Natural language processing engine: spaCy

[1517] Generative AI models: ChatGPT, etc.

[1518] Terminal: A device through which a user accesses a server and enters queries.

[1519] Devices: Smartphones, PCs, etc.

[1520] Program processing explanation

[1521] Information storage and analysis

[1522] The server stores PDF files of instruction manuals and troubleshooting manuals in a storage engine, for example, in an Amazon S3 bucket.

[1523] The server converts the saved PDF file to text using Tesseract OCR, specifically by running the command tesseract manual.pdf output.txt.

[1524] The server stores the converted text data in a MySQL database, for example by executing the SQL statement INSERT INTO manuals (text_data) VALUES ('...').

[1525] The server analyzes the text data using the spaCy NLP engine to extract semantic information, for example by running the following code: nlp = spacy.load('en_core_web_sm').

[1526] Query reception and processing

[1527] The user inputs a symptom or problem in natural language through the terminal. For example, the user inputs "The TV won't turn on."

[1528] The device sends this query to the server via an HTTP POST request.

[1529] Data analysis and generation of action plans

[1530] The server receives the user's query and searches the database for relevant information using an SQL query, for example SELECT FROM manuals WHERE text_data LIKE '%cannot power on%'.

[1531] The server sends the relevant information to a generative AI model (e.g., ChatGPT) via an API request. An example prompt sentence is "The user explains that the TV won't turn on. What should I do?"

[1532] The generative AI model generates a solution and sends it back to the server, such as "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1533] Presenting solutions and illustrations

[1534] The server uses the Pillow library to generate an image related to the solution, for example, a diagram showing where to connect the power cord.

[1535] The server integrates the generated solutions and images into an HTML template and compiles them into a single page.

[1536] The server sends the compiled information to the terminal as an HTTP response.

[1537] The terminal displays the information to the user and provides a quick and effective way to deal with the problem.

[1538] Specific examples

[1539] For example, if the user inputs "TV won't turn on," the system will operate as follows:

[1540] 1. The user types "TV won't turn on" into the device.

[1541] 2. The device sends the query to the server via an HTTP POST request.

[1542] 3. The server receives the query and searches the database for relevant information using an SQL query.

[1543] 4. The server sends the relevant information to the generative AI model via an API request.

[1544] 5. The generative AI model generates a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1545] 6. The server generates a diagram showing where the power cords are connected.

[1546] 7. The server compiles the solution and an illustration onto one page and sends it to the terminal.

[1547] 8. The device displays information to the user and provides a quick and effective way to deal with the problem.

[1548] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1549] Step 1: Initial Setup

[1550] The server stores PDF files of instruction manuals and troubleshooting manuals in the storage engine. As a concrete example, using Amazon S3, the command aws s3 cp manual.pdf s3: / / bucket-name / is executed as input, and the output is the PDF file stored in Amazon S3.

[1551] The server converts PDF files to text using Tesseract OCR. Specifically, it receives input from the command tesseract manual.pdf output.txt and generates an output file called output.txt.

[1552] The server stores the converted text data in a MySQL database. For example, it takes input as input and executes the SQL statement INSERT INTO manuals (text_data) VALUES ('...'), and gets output as text data stored in the database.

[1553] Step 2: Enter and accept user queries

[1554] The user inputs the phenomenon or problem in natural language using the terminal. For example, the input "The TV won't turn on" is received, and the output is the manually input data in text format.

[1555] The terminal sends the entered query to the server. Specifically, it receives input to send the entered query to the server via an HTTP POST request, and obtains output showing that the query has reached the server.

[1556] Step 3: Data analysis and information extraction

[1557] The server receives a user query and searches for relevant information in the database. Specifically, it receives input by executing the SQL query SELECT FROM manuals WHERE text_data LIKE '%TV won't turn on%' and obtains output that provides relevant information (for example, the contents of the corresponding instruction manual).

[1558] The server provides the relevant information to the generative AI model, taking the specific action of sending the results of a database search to the generative AI model via an API request as input, and obtaining an output ready for the generative AI model to generate a response.

[1559] Step 4: Generate a solution

[1560] The generative AI model generates a solution based on the user's query and information provided by the server. Specifically, the model receives a prompt, "The user explains that the TV won't turn on. What should I do?", and outputs a solution such as, "Make sure the power cord is connected properly. Then, replace the batteries in the remote control."

[1561] The server receives the generated solutions and performs filtering. Specifically, it receives the generated text, deletes unnecessary parts, and outputs solutions that are ready to be provided to the user.

[1562] Step 5: Generate an image

[1563] The server uses the Pillow library to generate an image diagram related to the solution. For example, for the command "Please check the power cord connection," input is made using an image editing tool, and an output is generated showing the power cord connection location.

[1564] Step 6: Present solutions and images

[1565] The server then combines the generated solutions and the resulting image onto a single page. Specifically, it uses an HTML template as input to integrate the solutions and the image, and outputs a single page that can be presented to the user.

[1566] The server sends the compiled information to the terminal. As a specific input, it performs the operation of sending the generated page as an HTTP response, and obtains the output that the page arrives at the user terminal.

[1567] The device displays the information to the user. Specifically, the device inputs the received page to display in a browser or dedicated app, and the user receives an output that allows them to visually confirm the solution to the problem.

[1568] (Application example 1)

[1569] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1570] The present invention relates to a system for quickly and effectively resolving problems without referring to instruction manuals or troubleshooting manuals. However, conventional systems require users to accurately input details of the problem, which can lead to errors, especially when using voice input. Furthermore, solutions are presented only in text, which can be difficult for users to understand. Furthermore, it is difficult for factory workers to respond to equipment failures without visual information. It is necessary to solve these problems.

[1571] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1572] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, means for converting voice input into text, and means for generating a 3D model or diagram of the solution. This allows the user to input a problem using voice input, and the solution is presented not only as text but also as a 3D model or diagram, enabling the user to solve the problem quickly and visually.

[1573] An "instruction manual" is a document that provides detailed instructions on how to use and operate an electronic device or machine.

[1574] A "troubleshooting manual" is a document that explains how to repair and resolve problems when electronic devices or machinery break down.

[1575] "Digital documents" are document data stored electronically, including PDFs and text files.

[1576] "Means of analysis" refers to the technology and devices used to read digital documents and understand and classify the necessary information.

[1577] "Means for receiving" refers to technology or devices for receiving information or data input by a user.

[1578] An "extraction means" is a technique or device used to identify and extract the required information from a digital document.

[1579] A "generative model" is an artificial intelligence model that automatically generates appropriate solutions or answers based on input information.

[1580] The "presentation means" refers to a technique or device for displaying the generated solutions and information to the user.

[1581] "Means for converting voice input into text" refers to technology or devices for converting voice input by a user into text data.

[1582] "Means for generating 3D models or diagrams of solutions" refers to technology or devices for expressing solutions in 3D models or figures to make them visually easier to understand.

[1583] "Server" means a computer system for processing, storing, and managing data.

[1584] This invention is a system that can quickly and effectively solve problems without the user having to refer to an instruction manual or a troubleshooting manual. This system is particularly useful for effectively dealing with factory robot failures. The specific system configuration and its operation will be described below.

[1585] System Configuration

[1586] 1. User Device

[1587] Wearable devices such as smart glasses.

[1588] It has voice input and display functions.

[1589] 2. Server

[1590] A cloud server for analyzing and storing digital documents.

[1591] It has high-performance computing resources and data storage.

[1592] 3. OCR technology

[1593] Convert PDF to text data using optical character recognition technology. Specifically, we use pytesseract.

[1594] 4. Voice Recognition Technology

[1595] It uses technology to convert voice input into text data, specifically, the speech_recognition library.

[1596] 5. Generative AI Models

[1597] It is used for natural language processing and countermeasure generation. Specifically, it uses the GPT-3 model from the transformers library.

[1598] 6. Digital Documents

[1599] PDF documents such as instruction manuals and troubleshooting manuals.

[1600] Data processing and calculation

[1601] 1. Digital Document Analysis

[1602] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using optical character recognition (OCR). This is where pytesseract comes in.

[1603] The converted text data is stored in a database on the server.

[1604] 2. Processing voice input

[1605] The user inputs the symptoms or problems by voice through the smart glasses.

[1606] The smart glasses convert the voice data into text data and send this data to the server, where the speech_recognition library is used.

[1607] 3. Extracting information and generating solutions

[1608] The server analyzes the received user query and searches for relevant information in a database.

[1609] Based on the search results, a generative model (e.g., GPT-3) is used to generate a solution. The transformers library is used.

[1610] 4. Providing solutions and visual information

[1611] Use software to convert the generated solutions into 3D models and diagrams.

[1612] The solution and 3D models or diagrams are sent to the smart glasses and visually presented to the user.

[1613] Specific examples

[1614] Example 1: Error handling in an automobile factory

[1615] User input: "The robot arm joints won't move."

[1616] Generated solution: "Oil the joints of the robot arm and check their operation. Also, check for any sensor error indications."

[1617] Prompt Sentence Examples

[1618] Problem: The robot arm joints won't move

[1619] How to respond:

[1620] This allows factory workers to quickly identify problems and respond effectively. The system is expected to improve work efficiency and reduce downtime through visual and audio assistance.

[1621] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1622] Step 1:

[1623] Digital document analysis

[1624] The server converts PDF files of instruction manuals and troubleshooting manuals into text data using OCR technology. The server reads the PDF using pytesseract and converts each page into text data.

[1625] Input: PDF format instruction manual or troubleshooting manual

[1626] Output: Text data

[1627] Specific behavior: The server extracts text from the PDF and stores it in a database.

[1628] Step 2:

[1629] Receiving audio input

[1630] The user inputs the symptoms or problems by voice through the smart glasses.

[1631] The smart glasses capture voice data and convert it to text data using the speech_recognition library.

[1632] Input: Audio data of the phenomenon or problem

[1633] Output: Text data

[1634] How it works: The user speaks, "The robot arm's joints won't move," and the smart glasses convert the speech into text.

[1635] Step 3:

[1636] Submitting a query

[1637] The smart glasses send the converted text data to the server.

[1638] Input: Text data converted from speech

[1639] Output: Text data sent to the server

[1640] Specific operation: The smart glasses generate text data and send it to a server via the network.

[1641] Step 4:

[1642] Information extraction and analysis

[1643] The server analyzes the received text data (user queries) and searches for relevant information in the database using a natural language processing (NLP) engine.

[1644] Input: User query text data

[1645] Output: Search results containing relevant information

[1646] Specific operation: The server extracts information related to the query from instruction manuals and manuals in the database.

[1647] Step 5:

[1648] Generate a workaround

[1649] The server generates a solution using a generative model (GPT-3) based on relevant information. It uses the transformers library.

[1650] Input: Search results and user text query

[1651] Output: Text data of the solution

[1652] Specific operation: The server performs NLP processing and generates the optimal solution.

[1653] Step 6:

[1654] Visual information generation

[1655] The server generates 3D models and diagrams based on the generated solutions. 3D design software and image generation software are used to make the solutions easier to understand visually.

[1656] Input: Text data of the solution

[1657] Output: 3D models and diagrams

[1658] Specific operation: The server converts the solution into a 3D model or diagram, generating visual information.

[1659] Step 7:

[1660] Coping methods and visual information

[1661] The server sends the generated solutions and 3D models or diagrams to the smart glasses.

[1662] The smart glasses present the received action and visual information to the user.

[1663] Input: Text data and visual information of the solution

[1664] Output: Actions and visual information displayed on the smart glasses

[1665] Specific operation: The server sends information, and the smart glasses display it and present it to the user.

[1666] This allows the user to input a problem using voice input, and quickly check the generated solutions and visual information to help solve the problem.

[1667] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1668] System Overview

[1669] The present invention provides a system that allows users to quickly and effectively solve problems without referring to an instruction manual or troubleshooting manual. This system not only analyzes the digital document of the instruction manual or troubleshooting manual and generates solutions to problems or phenomena input by the user, but also recognizes the user's emotions and reflects them in the solutions it presents.

[1670] composition

[1671] The user uses their own device (smartphone or PC) to input information about the problem or phenomenon in natural language.

[1672] The terminal transmits the user's input information to the server.

[1673] The server analyzes the digital document and extracts the required information.

[1674] A solution is generated based on the information extracted by the generative model (e.g., ChatGPT) and the user's input information.

[1675] The emotion engine recognizes emotions from user information and reflects them in the generated response methods.

[1676] The server presents the generated solution and related image diagram to the user.

[1677] Explanation of program processing

[1678] The detailed operation of this system will be described below.

[1679] Initial Setup

[1680] The server stores PDF files of instruction manuals and troubleshooting manuals in storage.

[1681] The server converts the PDF file into text using optical character recognition (OCR) technology.

[1682] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine.

[1683] Query reception

[1684] The user uses a terminal to input information about a phenomenon or problem in natural language.

[1685] The terminal sends the input query to the server.

[1686] Data Analysis and Information Extraction

[1687] The server receives the user's query and searches for relevant information in a database.

[1688] The server provides the search results to the generative model.

[1689] Emotion recognition

[1690] The server provides the user's input data to the emotion engine to analyze the user's emotions.

[1691] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and reflects the results in the generative model.

[1692] Generate a workaround

[1693] The generative model generates a response based on the user's query, related information provided by the server, and the recognized emotions.

[1694] For example, if the problem is "the TV won't turn on" and the emotion is "confused," you might preface the situation by saying, "Let's start with a simple method that anyone can do," and then explain, "First, check that the power cord is connected correctly."

[1695] Generate image diagrams

[1696] The server generates an image related to the solution (for example, a screenshot of the settings screen or a diagram of the connection points).

[1697] Presenting solutions and illustrations

[1698] The server generated solution and image diagram are summarized on one page.

[1699] The server sends the compiled information to the terminal.

[1700] The terminal displays this to the user, allowing the user to quickly and easily resolve the problem.

[1701] Specific examples

[1702] Below is a specific example of what happens when a user enters "My smartphone's Wi-Fi won't connect."

[1703] 1. The user types "My smartphone's Wi-Fi won't connect" into the device.

[1704] 2. The device sends this query to the server.

[1705] 3. The server analyzes the smartphone's instruction manual using OCR technology and extracts relevant information.

[1706] 4. The server provides the user's input data to the emotion engine and recognizes the emotion "confused."

[1707] 5. Based on the relevant information and the recognized emotion, the generative model generates specific countermeasures such as, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[1708] 6. The server generates a screenshot associated with this method.

[1709] 7. The server compiles a solution and an illustration and sends it to the terminal.

[1710] 8. The device displays this to the user, allowing them to quickly and easily resolve the issue.

[1711] This system allows users to find appropriate ways to deal with their emotions without having to refer to an instruction manual.

[1712] The processing flow will be explained below.

[1713] Step 1:

[1714] The user uses their own device (smartphone or PC) to input a description of the problem or phenomenon in natural language. For example, they might input "My smartphone's Wi-Fi won't connect."

[1715] Step 2:

[1716] The device receives the details of the problem or symptom entered by the user, converts it into an appropriate format, and sends it to the server via an HTTP POST request.

[1717] Step 3:

[1718] The server receives the HTTP POST request and prepares to parse the user's input data, which is then temporarily stored in a database on the server.

[1719] Step 4:

[1720] The server analyzes PDF files of instruction manuals and troubleshooting manuals that have been stored in advance, and uses OCR technology to convert the PDFs into text data.

[1721] Step 5:

[1722] The server then uses a natural language processing (NLP) engine to analyze the OCR-converted text data and categorize it into sections, a process that allows for efficient extraction of information related to specific problems or phenomena.

[1723] Step 6:

[1724] The server searches the analyzed text data in the database for information that matches the user's input, and the extracted information is used as input to the generative model.

[1725] Step 7:

[1726] The server provides the user's input data to the emotion engine, which analyzes the user's emotions. The emotion engine recognizes the user's emotions based on the input content and past data.

[1727] Step 8:

[1728] The emotion engine recognizes the user's emotions (e.g., anxiety, anger, confusion, etc.) and returns the results to the server, which then inputs this information into the generative model.

[1729] Step 9:

[1730] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. For example, if the user's emotion is "confused" when the query is "Wi-Fi is not connecting," the model will explain, "Don't worry. First, select the network name on the Wi-Fi settings screen and enter your password."

[1731] Step 10:

[1732] Based on the solution generated by the server, a concrete image diagram (for example, a screenshot of the Wi-Fi setting screen or a diagram of the connection location) is generated.

[1733] Step 11:

[1734] The server will compile the generated solutions and illustrations on one page, which will be formatted in a way that is easy for users to understand.

[1735] Step 12:

[1736] The server then sends the final information to the device, including specific solutions and related images.

[1737] Step 13:

[1738] The device displays the received information to the user, allowing the user to quickly and easily resolve the problem. This allows the user to respond efficiently without having to refer to an instruction manual or instruction manual.

[1739] Example 2

[1740] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1741] There is a demand for systems that allow users to solve problems quickly and effectively without referring to instruction manuals or troubleshooting manuals. There is also a demand for systems that improve user satisfaction by presenting solutions that take into account the user's emotions.

[1742] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1743] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and means including an emotion recognition engine for recognizing the user's emotions and reflecting them in the generated solution. This allows the user to quickly obtain an appropriate solution based on their emotions without having to refer to the instruction manual.

[1744] An "instruction manual or troubleshooting manual" is a detailed guideline on how to use a product and what to do in the event of a malfunction.

[1745] "Digital documents" are documents that are stored and managed electronically, including PDF and TXT formats.

[1746] "Means of analysis" refers to the techniques and methods used to read digital documents and understand and process their contents.

[1747] "Information about the problem or phenomenon input by the user" refers to the content of the phenomenon input by the user and the problems related to it.

[1748] "Means for receiving" refers to an interface or method for obtaining input information from a user.

[1749] "Means for extracting relevant information" refers to techniques or methods for extracting specified information from a digital document.

[1750] A "generative model" is an algorithm or system that generates new data or information based on given data.

[1751] "Presentation means" refers to the method or interface for displaying the generated information to the user.

[1752] An "emotion recognition engine" is a technology or system that analyzes and recognizes emotions from user input data.

[1753] The present invention provides a system that allows users to quickly and effectively solve problems without referring to instruction manuals or troubleshooting manuals. This system analyzes digital documents, generates solutions to problems or phenomena input by the user, and recognizes the user's emotions and reflects them in the solutions it presents.

[1754] System Overview

[1755] The system includes the following components:

[1756] Server: PDF files of instruction manuals and troubleshooting manuals are stored in storage and converted to text using OCR technology (e.g., Tesseract OCR). The converted text data is stored in a database (e.g., MySQL) and analyzed using a natural language processing (NLP) engine (e.g., spaCy).

[1757] Terminal: Provides an interface for users to input information about problems or phenomena in natural language and sends the input information to the server.

[1758] Generative model: Generates a solution based on relevant information provided by the server and user input (e.g., OpenAI's ChatGPT).

[1759] Emotion recognition engine: Analyzes emotions based on user input data (e.g., IBM Watson Tone Analyzer) and reflects the results in a generative model.

[1760] Presentation method: The generated solution and related image diagram are presented to the user.

[1761] How it works

[1762] 1. The server converts PDF files of instruction manuals and troubleshooting manuals into text using optical character recognition (OCR) technology. For example, by using Tesseract OCR, the PDF files are output as text files and the contents are stored in a database.

[1763] 2. The user uses the device to input a natural language description of the symptom or problem, for example, "My TV won't turn on."

[1764] 3. The device sends the entered query to the server via an HTTP POST request, specifying the server's API endpoint to send the data.

[1765] 4. The server receives the user's query and searches the database for relevant information using an SQL query, such as SELECT content FROM manuals_table WHERE content LIKE '%won't turn on%'.

[1766] 5. The server provides the search results to the generative model, which processes the content provided as prompts for the generative model.

[1767] 6. The server provides the user's input data to the emotion engine to analyze the emotion, for example, using IBM Watson Tone Analyzer to recognize the user's emotion.

[1768] 7. The generative model generates specific countermeasures based on the user's query, related information provided by the server, and the recognized emotions.

[1769] 8. The server generates an image diagram related to the generated solution, for example, using the Pillow library to generate the related diagram.

[1770] 9. The server compiles the generated solutions and images in HTML format and sends them to the terminal, where they are displayed to the user, allowing them to solve the problem quickly and easily.

[1771] Specific examples

[1772] Below is an example of a specific prompt sentence when a user enters "My smartphone's Wi-Fi won't connect."

[1773] "The user inputs, 'My smartphone's Wi-Fi won't connect.' Provide relevant information from the instruction manual. Furthermore, the user's emotion, 'confused,' is recognized. Based on this information, generate a specific solution."

[1774] This invention allows users to quickly find appropriate solutions according to their emotions without having to refer to an instruction manual or troubleshooting manual. Furthermore, by using an emotion recognition engine, it becomes possible to respond to each user in a way that is tailored to their individual needs, thereby improving user satisfaction.

[1775] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1776] Step 1:

[1777] The server saves PDF files of instruction manuals and troubleshooting manuals to storage. First, it saves the PDF files provided by the user in a dedicated directory called "manuals." For example, the file path is / var / www / manuals / manual1.pdf. The file path after saving is output.

[1778] Step 2:

[1779] The server converts the PDF file to text using optical character recognition (OCR). It then uses Tesseract OCR to output the PDF file as a text file. The input is the path to the saved PDF file, and the output is the path to the converted text file. Example: / var / www / manuals / manual1.txt

[1780] Step 3:

[1781] The server saves the converted text data in the database. It uses a MySQL database and saves the text data in a table (e.g. manuals_table). The input is the path to the text file and its contents, and the output is the updated result in the database. Example SQL query: INSERT INTO manuals_table (manual_id, content) VALUES (1, LOAD_FILE(' / var / www / manuals / manual1.txt'))

[1782] Step 4:

[1783] The user uses the device to input information about the problem or issue in natural language, for example, "The TV won't turn on." This input information is used in the next step.

[1784] Step 5:

[1785] The device sends the entered query to the server via an HTTP POST request. Specify the server's API endpoint and send the query data. The input is the user-entered query, and the output is the result of the HTTP request sent to the server. Example: POST / api / query {"query": "The TV won't turn on"}

[1786] Step 6:

[1787] The server receives a user query and searches for relevant information in the database. It uses SQL queries to retrieve relevant information from the database based on the input query. The input is the user query, and the output is the search results for relevant information. SQL query example: SELECT content FROM manuals_table WHERE content LIKE '%Power won't turn on%'

[1788] Step 7:

[1789] The server provides the search results to the generative model. The provided information is processed as a prompt for the generative model. The input is the search result information, and the output is the formation of a prompt for the generative model. Example prompt: "The user entered 'The TV won't turn on.' Please generate a solution by referring to the contents of the instruction manual below: [Related information]."

[1790] Step 8:

[1791] The server provides the user's input data to the emotion engine to analyze the user's emotion. For example, IBM Watson Tone Analyzer is used. The input is the user query, and the output is the emotion recognition result. API example: POST / api / tone_analyzer {"text": "The TV won't turn on"}. Result format: {"emotion": "Confused"}

[1792] Step 9:

[1793] The generative model generates a solution based on the user's query, related information provided by the server, and the recognized emotion. The input is the prompt and emotion recognition result for the generative model, and the output is the generated solution. Example: "Don't worry. First, make sure the power cord is properly connected."

[1794] Step 10:

[1795] The server generates an image related to the solution. It uses the Pillow library to generate an image containing specific content. The input is the solution content, and the output is the generated image. Example: image = Image.open("power_connection.png")

[1796] Step 11:

[1797] The server compiles the generated solutions and images in HTML format and sends them to the terminal. It uses an HTML template to compile information into one page. The input is the solutions and images, and the output is the content in HTML format. Example: {"html": " <h1> Solution< / h1> ...}"}

[1798] Step 12:

[1799] The terminal displays this to the user, allowing them to solve the problem quickly and easily. It uses a rendering engine to display HTML content. The input is the HTML formatted content, and the output is the displayed result.

[1800] (Application example 2)

[1801] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1802] There is a problem that users cannot find a quick and accurate solution on the spot when performing maintenance or troubleshooting on factory robots. In addition, a system that provides a uniform solution without considering the user's feelings makes it difficult to reduce the user's stress and confusion. Also, referring to the instruction manual or troubleshooting manual every time is time-consuming and does not allow for efficient problem solving.

[1803] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1804] In this invention, the server includes means for analyzing a digital document such as an instruction manual or a troubleshooting manual, means for receiving information about a problem or phenomenon input by a user, means for extracting relevant information from the digital document, means including a generative model for generating a solution based on the extracted information and information input by the user, means for presenting the generated solution to the user, and emotion recognition means for recognizing the user's emotions and reflecting those emotions in the generation of a solution. This enables the user to quickly and effectively obtain a solution that takes into account the user's emotions, without having to refer to the instruction manual or troubleshooting manual.

[1805] An "instruction manual" is a document that details how to operate and configure equipment or software.

[1806] A "failure response manual" is a document that explains the response procedures and solutions required when equipment breaks down.

[1807] A "digital document" is a document that is stored and displayed electronically, including PDFs and text files.

[1808] "Means of analysis" refers to techniques and methods for understanding the content of a digital document and extracting the necessary information.

[1809] "Information about a problem or phenomenon input by a user" is detailed information about a malfunction or phenomenon of a device input by a user in natural language.

[1810] "Means for receiving" refers to the technology or method for obtaining and processing information sent by a user.

[1811] An "extraction means" is a technique or method for extracting relevant information from a digital document.

[1812] A "generative model" is an AI model that generates appropriate solutions based on user input and information extracted from digital documents.

[1813] The "presentation means" refers to a technique or method for displaying the generated solution to the user.

[1814] "Emotion recognition means" refers to a technique or method for analyzing emotions from information entered by a user and using the results to adapt a response method.

[1815] A "server" is a computer system that stores, analyzes, and processes data, and is a device that communicates with user terminals via a network.

[1816] The present invention provides a system that allows a user to quickly and accurately find a solution when performing maintenance or troubleshooting on a factory robot. Specific embodiments of the present invention will be described below.

[1817] Hardware and software used

[1818] 1. Server

[1819] Storage: Storage for saving instruction manuals and troubleshooting manuals. Specifically, it uses the server's HDD or SSD.

[1820] OCR Technology: The server converts the PDF file into text using optical character recognition (OCR) technology (e.g., Tesseract).

[1821] Database: A database is required to store the converted text data and analyze it with a natural language processing (NLP) engine. For example, MySQL or PostgreSQL will be used.

[1822] Generative AI models: Use generative AI (e.g., OpenAI models) to generate solutions based on the user query and extracted information.

[1823] Emotion recognition engine: Analyzes emotions from user information and reflects them in response. Uses natural language processing models such as BERT.

[1824] 2. Terminal

[1825] Smartphone or PC: Users use these devices to input information about problems or phenomena in natural language, and the devices then send this information to a server.

[1826] Operation overview

[1827] 1. Initial data preparation

[1828] The server stores PDF files of instruction manuals and troubleshooting manuals in storage and converts them into text data using OCR technology. The converted text data is then stored in a database and analyzed by a natural language processing engine.

[1829] 2. Accepting user queries

[1830] The user uses a terminal to input questions about the robot's problems or symptoms in natural language, and the queries are sent from the terminal to the server.

[1831] 3. Data analysis and information extraction

[1832] The server receives user queries and searches and extracts relevant information from a database.

[1833] 4. Emotional Recognition

[1834] The server provides the user's input data to the emotion recognition engine, which recognizes the user's emotions. The recognized emotions are reflected in the generative AI model.

[1835] 5. Generate solutions

[1836] The generative AI model generates a solution based on the user's query, extracted information, and recognized emotions, and presents guidance in the most appropriate language based on the user's emotions.

[1837] 6. Creating an image diagram

[1838] The server generates an image related to the solution (for example, a screenshot of the setting screen or a diagram of the connection points).

[1839] 7. Present solutions and illustrations

[1840] The server then compiles the generated solutions and images into a single page and sends it to the device, which displays it to the user, allowing them to quickly and easily solve the problem.

[1841] Specific examples

[1842] Here is a specific example where the user inputs "The robot's arm won't move."

[1843] 1. User query: "The robot arm won't move."

[1844] 2. Emotion recognition: The emotion recognition engine recognizes "confusion."

[1845] 3. Generating a solution: The generative AI model generated the following: "Don't be confused. First, turn off the robot and check if there is an abnormality in the arm connection. Next, turn it on again, and if an error code is displayed, refer to the procedures in the maintenance manual."

[1846] 4. Image generation: Image of the "arm connection part".

[1847] Prompt Sentence Examples

[1848] Please generate solutions to user queries based on the following manual.

[1849] manual:

[1850] [Instruction manual and troubleshooting manual text]

[1851] Query: Robot arm not moving

[1852] User Emotion: Confused

[1853] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1854] Step 1:

[1855] The server saves PDF files of instruction manuals and troubleshooting manuals in storage. The input of this step is the PDF files, and the output is the PDF files saved in storage. This prepares the documents to be analyzed.

[1856] Step 2:

[1857] The server converts the PDF file into text data using OCR technology (e.g., Tesseract). The input for this step is the PDF file, and the output is text data. Using OCR technology, character information is extracted from the image data and converted into text format.

[1858] Step 3:

[1859] The server stores the converted text data in a database and analyzes it using a natural language processing (NLP) engine. The input to this step is the text data generated by OCR, and the output is structured data of the document analyzed by NLP. As a result of the analysis, important information within the document is extracted and registered in the database.

[1860] Step 4:

[1861] The user uses a terminal to input information about the problem or phenomenon of the robot in natural language. The input of this step is a query entered by the user, and the output is a query sent from the terminal to the server. The user describes the problem in natural language, and it is sent to the server.

[1862] Step 5:

[1863] The server receives a user query and searches and extracts relevant information in the database. The input to this step is the user query, and the output is the relevant information extracted from the database. The server pulls the appropriate information from the database based on the query.

[1864] Step 6:

[1865] The server provides the user's query and extracted information to an emotion recognition engine to analyze the user's emotion. The input of this step is the user query and related information, and the output is analyzed emotion data. The emotion recognition engine is used to identify the user's emotion (e.g., confusion, anger), and the result is provided to the next step.

[1866] Step 7:

[1867] The server uses a generative AI model (e.g., OpenAI's model) to generate a solution based on the extracted information and the recognized emotion. The inputs for this step are the user query, the extracted information, and the emotion data, and the output is a solution generated by the generative AI model. The generated solution takes the user's emotion into consideration.

[1868] Step 8:

[1869] The server generates an image associated with the generated solution. The input to this step is the generated solution, and the output is an image associated with the solution. The server generates appropriate visual aids (screenshots and diagrams) to accompany the solution.

[1870] Step 9:

[1871] The server compiles the generated solutions and the image onto a single page and sends it to the terminal. The input to this step is the solutions and the image, and the output is the data sent to the terminal. The user receives this on their terminal and it is displayed.

[1872] Step 10:

[1873] The user can quickly and easily solve the problem by referring to the solution and image displayed on the terminal. The input of this step is the solution and image displayed on the terminal, and the output is the user's problem solution. The user proceeds with the solution according to the proposed method.

[1874] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1875] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1876] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1877] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1878] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1879] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1880] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1881] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1882] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1883] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1884] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1885] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1886] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1887] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1888] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1889] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1890] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1891] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1892] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1893] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1894] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1895] The following is further disclosed regarding the above embodiment.

[1896] (Claim 1)

[1897] A means for analyzing a digital document of an instruction manual or a troubleshooting manual;

[1898] means for receiving information about a problem or symptom input from a user;

[1899] means for extracting relevant information from said digital document;

[1900] means including a generative model for generating a countermeasure based on the extracted information and input information from a user;

[1901] means for presenting the generated solution to a user;

[1902] A system including:

[1903] (Claim 2)

[1904] 10. The system of claim 1, further comprising means for converting the parsed digital document into text using optical character recognition techniques.

[1905] (Claim 3)

[1906] 10. The system of claim 1, further comprising means for generating an image corresponding to the generated solution.

[1907] "Example 1"

[1908] (Claim 1)

[1909] A means for analyzing a digital document of an instruction manual or a troubleshooting manual;

[1910] means for receiving information about a problem or symptom input from a user;

[1911] means for converting the digital document into text using optical character recognition technology;

[1912] means for extracting relevant information from said digital document;

[1913] a means for searching for relevant information in a database;

[1914] A means including a generative AI model that generates a countermeasure based on the extracted information and input information from a user;

[1915] means for generating an image diagram corresponding to the generated solution;

[1916] means for presenting the generated solution and image diagram to a user;

[1917] A system including:

[1918] (Claim 2)

[1919] 10. The system of claim 1, further comprising: means for extracting semantic information from the analyzed digital document using a natural language processing engine.

[1920] (Claim 3)

[1921] 2. The system according to claim 1, further comprising means for integrating the generated solutions and the image diagram, compiling them into one page, and transmitting the page to the user terminal.

[1922] "Application Example 1"

[1923] (Claim 1)

[1924] A means for analyzing a digital document of an instruction manual or a troubleshooting manual;

[1925] means for receiving information about a problem or symptom input from a user;

[1926] means for extracting relevant information from said digital document;

[1927] means including a generative model for generating a countermeasure based on the extracted information and input information from a user;

[1928] means for presenting the generated solution to a user;

[1929] a means for converting voice input into text;

[1930] A means to generate 3D models or diagrams of the solution;

[1931] A system including:

[1932] (Claim 2)

[1933] 10. The system of claim 1, further comprising means for converting the parsed digital document into text using optical character recognition techniques.

[1934] (Claim 3)

[1935] 10. The system of claim 1, further comprising means for generating an image and a 3D model corresponding to the solution.

[1936] "Example 2: Combining Emotion Engines"

[1937] (Claim 1)

[1938] A means for analyzing a digital document of an instruction manual or a troubleshooting manual;

[1939] means for receiving information about a problem or symptom input from a user;

[1940] means for extracting relevant information from said digital document;

[1941] means including a generative model for generating a countermeasure based on the extracted information and input information from a user;

[1942] means for presenting the generated solution to a user;

[1943] means including an emotion recognition engine for recognizing the user's emotion and reflecting it in the generated response;

[1944] A system including:

[1945] (Claim 2)

[1946] 10. The system of claim 1, further comprising means for converting the parsed digital document into text using optical character recognition techniques.

[1947] (Claim 3)

[1948] 10. The system of claim 1, further comprising means for generating an image corresponding to the generated solution.

[1949] "Application example 2 when combining emotion engines"

[1950] (Claim 1)

[1951] A means for analyzing a digital document of an instruction manual or a troubleshooting manual;

[1952] means for receiving information about a problem or symptom input from a user;

[1953] means for extracting relevant information from said digital document;

[1954] means including a generative model for generating a countermeasure based on the extracted information and input information from a user;

[1955] means for presenting the generated solution to a user;

[1956] an emotion recognition means for recognizing an emotion of the user and reflecting the emotion in generating a countermeasure;

[1957] A system including:

[1958] (Claim 2)

[1959] 10. The system of claim 1, further comprising means for converting the parsed digital document into text using optical character recognition techniques.

[1960] (Claim 3)

[1961] 10. The system of claim 1, further comprising means for generating an image corresponding to the generated solution. [Explanation of symbols]

[1962] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for analyzing a digital document of an instruction manual or a troubleshooting manual; means for receiving information about a problem or symptom input from a user; means for extracting relevant information from said digital document; means including a generative model for generating a countermeasure based on the extracted information and input information from a user; means for presenting the generated solution to a user; A system including:

2. 10. The system of claim 1, further comprising means for converting the parsed digital document into text using optical character recognition techniques.

3. The system of claim 1 further comprising means for generating an image corresponding to the generated solution.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A