system
The system addresses the immediacy and cost challenges of conventional learning support by using OCR and generative AI to extract and solve learning problems from images, improving user efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Conventional learning support systems, such as private tutors and cram schools, face challenges in providing immediate and cost-effective solutions to users' learning problems, particularly in qualification examinations and specialized learning environments.
A system that uses optical character recognition (OCR) to extract text from images, combines it with user-entered questions, and employs a generative artificial intelligence model for natural language processing to generate accurate answers, enabling quick and cost-effective problem-solving.
The system allows users to quickly and accurately understand and resolve learning problems, enhancing learning efficiency by providing immediate and affordable solutions compared to conventional methods.
Smart Images

Figure 2026063815000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a real learning environment, it is difficult for users to quickly and accurately understand the problems during learning and obtain solutions thereto. Particularly in qualification examinations and specialized learning, it is required to immediately understand the content of problems and receive appropriate advice. However, current conventional learning support systems such as private tutors and cram schools have limitations in terms of immediacy and cost. Therefore, there is a need for a technology that allows users to quickly and at low cost solve problems during learning.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides the following means. First, it receives an image posted by a user and extracts text information from the image using optical character recognition (OCR) technology. Next, it combines the extracted text information with the question entered by the user, interprets this combined text using natural language processing (NLP), and generates an appropriate answer. Finally, it sends the generated answer to the user so that the user can immediately obtain a solution to the problem. The system of the present invention solves the problems of immediacy and cost by using an optical character recognition library and a generative artificial intelligence model to extract information from images and present solutions using natural language processing.
[0006] "Submitted images" refer to images that users upload or send to the system, depicting problems or issues.
[0007] "Means of receiving" refers to the interface or function used to capture image data sent by the user.
[0008] Optical character recognition (OCR) is a technology that analyzes character information contained within an image and converts it into text data.
[0009] "Method of extraction" refers to the process of extracting text information from an image using optical character recognition technology.
[0010] "Method of combining" refers to the process of linking extracted text information with questions entered by the user to create a single text data.
[0011] "Natural language processing" is a technology that enables computers to understand, generate, and respond to human language.
[0012] "Means for generating answers" refers to a function that uses natural language processing technology to create appropriate answers and advice based on combined questions and textual information.
[0013] A "generative artificial intelligence model" refers to an algorithm or model that learns from large amounts of text data and performs natural language understanding and generation.
[0014] "Means of transmission" refers to the communication interface or protocol used to deliver the generated answer to the user's device. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention relates to a system for quickly and accurately resolving learning problems using AI. The following describes specific embodiments of this invention.
[0037] First, the user takes a picture of the problem they are studying with a smartphone or other device and sends the image to the system. This is the "submitted image." At the same time, the user also enters a question about the problem in text format. For example, a question like, "Please tell me how to solve this problem."
[0038] Next, the terminal sends the received image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the Python pytesseract library to extract text information such as mathematical formulas and sentences from the image as text data.
[0039] The extracted text is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "x^2 + 2x + 1 = 0" and the user's question is "How do I solve this equation?", the server combines them to generate the following text.
[0040] Problem statement: x^2 + 2x + 1 = 0
[0041] Question: How do I solve this equation?
[0042] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3(registered trademark).5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model generates a response like the following:
[0043] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0044] The generated answers are temporarily stored by the server and then sent to the user's device. The device receives the response from the server and displays the answer on the user interface. The user reviews this answer and uses it to aid in learning.
[0045] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of this equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image and generates a solution using a generative artificial intelligence model. As a result, the user receives the answer "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0046] As described above, the system of the present invention combines optical character recognition technology and a generative artificial intelligence model to enable users to quickly and accurately solve problems they are learning. This system provides learning support that is more immediate and cost-effective than conventional learning support methods.
[0047] The following describes the processing flow.
[0048] Step 1:
[0049] Users take pictures of problems or assignments they are studying using a smartphone or other device. At the same time, users also input questions about the problem in text format. For example, they might input a question like, "Please explain how to solve this problem."
[0050] Step 2:
[0051] The terminal receives the captured image and the entered question, and sends them to the server as a single data package. An HTTP request is used for this process.
[0052] Step 3:
[0053] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0054] Step 4:
[0055] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data.
[0056] Step 5:
[0057] The server temporarily stores the extracted text information. Let's assume the extracted text is "x^2 + 2x + 1 = 0".
[0058] Step 6:
[0059] The server combines the extracted text information with the user's question text. For example, it may be formatted as follows:
[0060] Problem statement: x^2 + 2x + 1 = 0
[0061] Question: How do I solve this equation?
[0062] Step 7:
[0063] The server sends the combined text to a generative artificial intelligence model. The generative AI model analyzes the combined text and generates an appropriate response.
[0064] Step 8:
[0065] A generative artificial intelligence model (e.g., GPT-3.5) generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1."
[0066] Step 9:
[0067] The server receives the generated response and sends it to the terminal. The HTTP protocol is used again for communication.
[0068] Step 10:
[0069] The terminal receives an HTTP response from the server and extracts the answer data. Then, it displays the answer on the user interface.
[0070] Step 11:
[0071] The user checks the answer displayed on their device. This allows the user to quickly and accurately obtain a solution to the problem.
[0072] The above steps provide answers to user questions. This system becomes a powerful tool for users to solve problems they are learning in real time.
[0073] (Example 1)
[0074] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0075] Conventional learning support systems have the challenge of not being able to quickly and accurately solve problems that users encounter during their studies. In particular, extracting problems from images containing text information, interpreting them appropriately, and providing immediate solutions is difficult, which reduces the user's learning efficiency. Furthermore, if an appropriate solution cannot be obtained, users are forced to interrupt their learning.
[0076] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0077] In this invention, the server includes means for receiving an image taken by the user, means for extracting character information from the received image using optical character recognition, means for combining the extracted character information with an input question, means for interpreting the combined text using natural language processing based on a generative artificial intelligence model and generating an answer, means for sending the generated answer to the user's terminal, and means for displaying the answer on the user's terminal. This makes it possible to quickly and accurately solve problems that the user faces during learning.
[0078] A "user" refers to an individual who uses this system to try to solve a problem they are learning.
[0079] "Device" refers to electronic devices such as smartphones and tablets used by users.
[0080] "Captured images" refers to image data containing the learning problems acquired by the user using the camera function of their device.
[0081] "Means of receiving" refers to the communication interface used to transfer images and questions taken by the user to the server.
[0082] "Optical character recognition" refers to a technology that analyzes character information within an image and extracts it as text data.
[0083] "Means of extraction" refers to software and hardware used to analyze text information within a received image and obtain it as text data.
[0084] "Method of combining" refers to the process of combining extracted character information and user-entered questions into a single text data file.
[0085] A "generative artificial intelligence model" refers to an algorithm or model that performs natural language processing based on given data to generate appropriate responses.
[0086] "Natural language processing" refers to the technology of processing human language using computers, and in this context, it refers to the technology of generating appropriate answers based on questions.
[0087] "Means of transmission" refers to the communication interface used to transfer the generated response to the user's device.
[0088] "Means of display" refers to a user interface used to visually show the answers obtained on the device.
[0089] This invention is a system for quickly and accurately resolving problems that users encounter during learning. The specific configuration and operation of the system are described as embodiments for carrying out this invention.
[0090] The user takes a picture of the problem they are studying with their device's camera and obtains the image data. For example, if the problem is written in their notebook as the equation "x^2 + 2x + 1 = 0", they would take a picture of that notebook. Then, the user uses the application on their device to type the question in text format, such as "Please tell me how to solve this equation."
[0091] The device sends the captured image and text question to the server as an HTTP request. The server receives this HTTP request and then processes the image data.
[0092] The server uses the Python library pytesseract to extract text information from the image using optical character recognition (OCR) technology. The OCR process extracts the text data "x^2 + 2x + 1 = 0".
[0093] The extracted text data and the questions entered by the user are combined to generate text like the following.
[0094] Problem statement: x^2 + 2x + 1 = 0
[0095] Question: How do I solve this equation?
[0096] Next, the server sends the generated text to a generative artificial intelligence model (for example, GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates an appropriate response. For example, the following response may be generated.
[0097] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0098] The generated response is temporarily stored on the server. The server then sends the generated response to the user's device as an HTTP response.
[0099] The device displays the received responses on the user interface. The user reviews these responses to aid their learning.
[0100] Thus, the system of the present invention enables users to quickly and accurately solve learning problems they face. By combining optical character recognition technology with a generative artificial intelligence model, the immediacy and accuracy are improved compared to conventional learning support methods, thereby enhancing the user's learning efficiency.
[0101] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0102] Step 1:
[0103] The user takes a picture of the problem they are studying using their device's camera. For example, they might take a picture of a notebook page with the equation "x^2 + 2x + 1 = 0" written on it. Simultaneously, they use the application to input a question. For example, they might input "Please tell me how to solve this equation." The input here consists of image data and text data. The output is an image file and a text question generated by the user's actions.
[0104] Step 2:
[0105] The terminal packages the image data and text data acquired by the user as an HTTP request and sends it to the server. The input consists of the image acquired by the user and the text question entered; this data is processed for transmission as an HTTP request to the server, and the output is the data sent to the server.
[0106] Step 3:
[0107] The server receives an HTTP request sent from the terminal. The input consists of image and text data sent from the terminal. The server uses the Python pytesseract library to analyze the image data and extracts text information from the image using optical character recognition (OCR) technology. Specifically, the mathematical expression "x^2 + 2x + 1 = 0" is extracted as text. This process outputs the input image as text data.
[0108] Step 4:
[0109] The server combines the text data extracted by OCR processing with the question text data entered by the user. For example, it generates a prompt message like the following:
[0110] Problem statement: x^2 + 2x + 1 = 0
[0111] Question: How do I solve this equation?
[0112] The input consists of extracted character information and question text, and the output is generated through a data merging process called prompt text generation.
[0113] Step 5:
[0114] The server sends the generated prompt to a generative artificial intelligence model, which then uses natural language processing (NLP) techniques to generate an appropriate response. For example, the following response might be generated:
[0115] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0116] The input here is a generated prompt sentence, and the output is the response text obtained from a generative artificial intelligence model.
[0117] Step 6:
[0118] The server temporarily stores the generated response and then sends it to the user's terminal as an HTTP response. The input is the generated response text, which is then sent to the terminal as an HTTP response. The output is the data sent to the terminal.
[0119] Step 7:
[0120] The terminal analyzes the response received from the server and displays it on the user interface. The input is the response text from the server, and the output is the visual information displayed on the user interface. The user looks at this to aid in learning.
[0121] Through the above series of processing steps, users can quickly and accurately solve problems they encounter during their learning process.
[0122] (Application Example 1)
[0123] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0124] Machine maintenance in modern factories is complex and requires quick and accurate solutions. However, many workplaces lack the necessary specialized technicians, often resulting in inefficient maintenance. To address this problem, a system is needed that accurately identifies machine problems and provides appropriate repair procedures.
[0125] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0126] In this invention, the server includes means for receiving posted images, means for extracting text information from the received images using optical character recognition, means for combining the extracted text information with an entered question, means for interpreting the combined text using natural language processing and generating an answer, means for transmitting the generated answer, and means for providing identification and repair procedures for damaged machine parts to support machine maintenance work in the factory. This enables factory workers to quickly and accurately identify machine problems and obtain appropriate repair procedures using smartphones or tablets.
[0127] "Means for receiving images" refers to a device or software that has the function of receiving and storing image data transmitted by a user.
[0128] "Means of extraction using optical character recognition" refers to devices or software that use technology to analyze character information within an image and convert it into text data.
[0129] "Means of combining" refers to a function that integrates extracted text information and user-entered questions into a single text file.
[0130] "Means of interpretation using natural language processing" refers to techniques that analyze combined texts and generate answers in an easily understandable format.
[0131] "Means for generating answers" refers to a device or software that has the function of providing appropriate answers from information analyzed by natural language processing.
[0132] "Means of sending responses" refers to a device or software for sending generated responses to a user's device.
[0133] "Means of supporting machine maintenance work" refers to devices or software that have the function of identifying problems with machinery and providing repair methods and procedures.
[0134] "Identifying damaged machine parts" refers to the process of recognizing damaged parts from captured images and extracting their details.
[0135] "Means of providing repair procedures" refers to a device or software that has the function of notifying the user of the appropriate repair method or process for an identified problem.
[0136] This invention relates to a system for supporting machine maintenance work in a factory. The following describes specific embodiments for carrying out this invention.
[0137] First, the user takes a picture of the problematic machine in the factory using a smartphone or tablet and sends the image to the system. At the same time, the user also enters a question about the machine in text format, such as, "How do I replace this part?"
[0138] Next, the device sends the captured image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the pytesseract library to extract text data such as character information and part names from the image.
[0139] The extracted text data is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "bolt" and "gear," and the user's question is "How do I replace this part?", the server combines these to generate the following text.
[0140] Problem statement: How to replace bolts and gears
[0141] Question: How do I replace this part?
[0142] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0143] To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt.
[0144] The generated response is temporarily stored by the server and then sent to the user's terminal. The terminal receives the response from the server and displays the response on the user interface. The user reviews this response and uses it as a reference for maintenance work.
[0145] As a concrete example, the process involves a user sending an image of a problematic machine part to the system along with the question, "How do I replace this bolt and gear?" The server extracts information about the bolt and gear from the image and generates a replacement method using a generative artificial intelligence model. As a result, the user receives the answer, "To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt."
[0146] This system combines optical character recognition technology with generative artificial intelligence models to enable factory workers to perform machine maintenance quickly and accurately. This improves on-site work efficiency and allows for proper maintenance even in situations where there is a shortage of specialized technicians.
[0147] Example of a prompt:
[0148] Problem statement: How to replace bolts and gears
[0149] Question: How do I replace this part?
[0150] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0151] Step 1:
[0152] Users take pictures of problematic machinery in the factory using their smartphones or tablets and send the images to the system. At the same time, they also enter questions about the machinery in text format. This process sends both the image and text data from the device to the server.
[0153] Step 2:
[0154] The device sends the captured image and the entered question to the server as an HTTP request. Specifically, the device sends the image file and text as a POST request to the specified endpoint on the server. The input consists of image data and text data, and the output is the request data sent to the server.
[0155] Step 3:
[0156] The server receives the transmitted image. The received image data is temporarily stored on the server. In this step, the input is the image data, and the output is the stored image data.
[0157] Step 4:
[0158] The server uses Optical Character Recognition (OCR) technology to extract text information from the received image. The server uses the pytesseract library to obtain the text information within the image. The input is image data, and the output is the extracted text data.
[0159] Step 5:
[0160] The server combines the extracted text data with the question entered by the user. This is done to generate the problem statement and question as a single text file. The input is the extracted text data and the question data, and the output is the combined text data.
[0161] Step 6:
[0162] The server sends the combined text data to a generative artificial intelligence model (e.g., GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates a response. In this step, prompts are sent to the generative AI model via an API to generate the response text. The input is the combined text data, and the output is the generated response text.
[0163] Step 7:
[0164] The generated response is temporarily stored on the server and then sent to the user's device. Specifically, the server returns a response containing the generated response to the user's device. The input is the generated response text, and the output is the response data sent to the user's device.
[0165] Step 8:
[0166] The terminal receives a response from the server and displays the answer on the user interface. The user reviews this answer and uses it as a reference for machine maintenance work. The input is the response data from the server, and the output is the displayed answer data.
[0167] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0168] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0169] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0170] Next, the device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and the user's question contained within the request.
[0171] First, the server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0172] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0173] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0174] Furthermore, this invention includes an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when a user enters text for a question, the emotional tone is analyzed from the content and style of the text.
[0175] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, it can add more detailed explanations if the user is confused, or add words of encouragement if the server detects that the user is stressed.
[0176] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. The user reviews the displayed answers and uses them to aid in learning.
[0177] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user receives an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0178] Through the steps described above, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0179] The following describes the processing flow.
[0180] Step 1:
[0181] Users use their smartphones or tablets to take pictures of problems or assignments they are studying. At the same time, users also input questions about the problems in text format. For example, they might input a question like, "Please explain how to solve this equation."
[0182] Step 2:
[0183] The device combines the captured image and the user's input into a single data package. This data package is then sent to the server as an HTTP request.
[0184] Step 3:
[0185] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0186] Step 4:
[0187] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, it extracts the equation "x^2 + 2x + 1 = 0" from the image.
[0188] Step 5:
[0189] The server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. The generated text will take the form of, for example, "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0190] Step 6:
[0191] The server sends the combined text data to the sentiment engine. The sentiment engine analyzes the content and format of the questions entered by the user and evaluates the user's emotions. For example, it recognizes emotions such as the user being nervous, confused, or irritated.
[0192] Step 7:
[0193] Based on the emotion data recognized by the emotion engine, the server sends the combined text to a generative artificial intelligence model. The generative AI model (e.g., GPT-3.5) analyzes the combined text and generates an appropriate response.
[0194] Step 8:
[0195] A generative artificial intelligence model generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0196] Step 9:
[0197] The server receives the generated response and adjusts it. Based on the user's emotions recognized by the emotion engine, the server modifies how the response is expressed. For example, if the user is confused, it may add a more detailed explanation or use gentler language.
[0198] Step 10:
[0199] The server sends the final response data to the terminal.
[0200] Step 11:
[0201] The terminal receives an HTTP response from the server and displays the response data on the user interface.
[0202] Step 12:
[0203] Users review the answers displayed on their device to aid their learning. Because these answers are adjusted by an emotion engine, users receive effective learning support that takes their emotional state into account.
[0204] Through the steps described above, users receive immediate and appropriate answers to their questions. This system allows users to quickly and accurately resolve learning problems, providing learning support that is more immediate and cost-effective than traditional methods.
[0205] (Example 2)
[0206] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0207] Conventional learning support systems can interpret questions from image data and generate answers, but they cannot take into account the user's emotions, which means some users may not receive appropriate support. For example, if a user is emotionally stressed, the answer provided may reduce their learning efficiency. In such situations, appropriate feedback that takes the user's emotions into account is necessary.
[0208] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0209] In this invention, the server includes means for receiving posted images, means for extracting textual information from the received images using optical character recognition, means for combining the extracted textual information with the entered question, means for interpreting the combined text using natural language processing and generating an answer, means for analyzing the user's entered question and dialogue content and recognizing their emotional state, means for adjusting the generated answer based on the recognized emotion, and means for transmitting the generated answer. This makes it possible not only to provide appropriate answers according to the user's learning progress but also to provide feedback according to the user's emotional state.
[0210] "Submitted images" refer to image data containing learning problems that users capture using their devices and send to the server.
[0211] "Textual information" refers to text data contained within the posted image, specifically including strings of characters such as equations and mathematical problems.
[0212] Optical character recognition (OCR) is a technology that extracts character information from images as text data, and is sometimes abbreviated as OCR.
[0213] "Means of receiving" refers to the interface or function that a server uses to receive data sent from a client, such as an HTTP request.
[0214] "Means of extraction" refers to software or libraries used to extract textual information from images. For example, optical character recognition libraries fall into this category.
[0215] A "question" is text data that a user enters regarding a learning problem they want to solve.
[0216] "Means of combining" refers to a process or function for integrating extracted character information and user-entered questions into a single text data file.
[0217] Natural language processing (NLP) is the technology that enables computers to understand and generate human language. This processing includes text analysis and generation.
[0218] A "generative artificial intelligence model" is an artificial intelligence technology that generates natural language responses based on provided input data, and for example, it utilizes natural language processing technology.
[0219] "Means of recognizing emotional states" refers to algorithms and engines that analyze user question text and dialogue content to determine the user's emotions.
[0220] "Means of adjustment" refers to a process or function for appropriately modifying or supplementing responses generated based on perceived emotional states.
[0221] "Means of transmission" refers to an interface or function for sending generated response data from the server to the client.
[0222] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0223] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0224] The device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and user question contained within the request.
[0225] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0226] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0227] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0228] Furthermore, the system of the present invention is equipped with an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when the user enters the text of a question, it analyzes the emotional tone from the content and style of the text.
[0229] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, if it detects that the user is feeling stressed due to a question, it can add words of encouragement.
[0230] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. As another example, if the system recognizes that the user has a question, it can add a detailed explanation. The user reviews the displayed answers and uses them to aid in their learning.
[0231] Specific example
[0232] This illustrates the process for a user who wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends it to the system with the question, "Please tell me how to solve this equation." The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user can receive an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0233] Example of a prompt
[0234] Problem statement: x^2 + 2x + 1 = 0
[0235] Question: How do I solve this equation?
[0236] In the above configuration, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0237] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0238] Step 1:
[0239] The user takes a picture of the problem and enters the question.
[0240] Users take photos of the problems they are learning using a device such as a smartphone or tablet. For example, they might take a picture of a piece of paper with "x^2 + 2x + 1 = 0" written on it. At the same time, they input a question into the device, such as "Please tell me how to solve this equation." The input is done through a text box.
[0241] Input: Captured image and text question
[0242] Output: Combine images and text questions into a single data package.
[0243] Step 2:
[0244] The terminal sends the data package to the server.
[0245] The device generates a single data package containing the captured image and the entered question. This data package is sent to the server via an HTTP request. The device packages the image data and text question as a payload and sends it over the network.
[0246] Input: Data package (including images and text questions)
[0247] Output: Data package sent to the server
[0248] Step 3:
[0249] The server analyzes the image data using OCR technology.
[0250] The server extracts image data from the received data package. Optical Character Recognition (OCR) technology is used to extract text information from the image as text data. The specific software used in this process is the Python library pytesseract. For example, the string "x^2 + 2x + 1 = 0" is extracted from the image.
[0251] Input: Image data
[0252] Output: Text information (text data) within the image
[0253] Step 4:
[0254] The server combines the text information and the question to generate text data.
[0255] The server combines the character information extracted by OCR with the user's question text to generate a single text file. This text file contains the problem statement and the user's question. For example, it might take the form of: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0256] Input: Text information (text data) and question text
[0257] Output: Combined text data
[0258] Step 5:
[0259] The server sends text data to an AI model that generates an answer.
[0260] The server sends the combined text data to a generative artificial intelligence model. This model uses natural language processing techniques to generate an appropriate response. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0261] Input: Combined text data
[0262] Output: Text response
[0263] Step 6:
[0264] The server uses an emotion engine to recognize the user's emotions and adjust the generated response accordingly.
[0265] The server is equipped with an emotion engine that analyzes the user's input questions and conversations to recognize their emotional state. Based on the recognized emotion, the generated response is appropriately adjusted. For example, if the server detects that the user is feeling stressed, it will add encouraging words such as, "It's okay, you can do it!"
[0266] Input: Text question and generated text answer
[0267] Output: Adjusted text response
[0268] Step 7:
[0269] The server sends the response to the terminal, and the terminal displays it to the user.
[0270] The server sends the final answer to the terminal. The terminal displays the received answer on the user interface. The user checks the displayed answer and uses it to aid in learning. For example, the displayed message might be: "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Don't worry, you can do it!"
[0271] Input: Adjusted text response
[0272] Output: Answer displayed on the terminal
[0273] (Application Example 2)
[0274] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0275] In today's educational environment, students often face numerous problems and questions. However, systems that allow them to immediately resolve their doubts during self-study or in class are limited. Furthermore, there are very few systems that understand students' emotional states and provide learning support accordingly. As a result, students' learning efficiency may decrease, potentially leading to a decline in their motivation to learn.
[0276] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving posted images, means for extracting character information from the received images using optical character recognition, means for combining the extracted character information with the input question, means for interpreting the combined text using natural language processing and generating an answer, means for recognizing the user's emotions and adjusting the generated answer according to the recognized emotions, and means for transmitting the generated answer and playing it back as audio. As a result, students can solve problems they are learning in real time and receive detailed learning support tailored to their emotional state.
[0277] "Means for receiving posted images" refers to a function that allows the server to receive image data taken by users using their smart devices.
[0278] "Means for extracting character information from a received image using optical character recognition" refers to a function that extracts character and numerical information contained within an image as text data using optical character recognition technology.
[0279] "Means for combining extracted text information and entered questions" refers to a function that integrates text information extracted from an image with questions submitted by the user into a single text data file.
[0280] "Means for interpreting combined text using natural language processing and generating answers" refers to a function that analyzes combined text using a generative artificial intelligence model and creates an appropriate answer based on that analysis.
[0281] The "means for recognizing the user's emotions and adjusting the response generated according to the recognized emotions" is a function that analyzes the emotional state from the user's text or voice tone and makes changes or adjustments to the response generated based on the result.
[0282] The "means for transmitting the generated response and playing it as audio" is a function that transmits the response from the server to the user's terminal and outputs the response as audio using text-to-speech technology.
[0283] The "optical character recognition library" is a software library used to extract character information in an image using optical character recognition technology.
[0284] The "generative artificial intelligence model" is an AI model that analyzes text data using natural language processing and generates appropriate responses or information.
[0285] The system that realizes this invention is composed of multiple components such as smart glasses, a server, and text-to-speech technology. Specific embodiments will be described below.
[0286] First, the user uses smart glasses to capture an image of the problem during learning. The user also submits a question to the system using voice input. Specifically, the built-in camera of the smart glasses takes a picture of the image, and the microphone records the question.
[0287] Next, the smart glasses send the captured image and voice data to the server as one data package. This transmission is carried out through an HTTP request.
[0288] The server uses optical character recognition (OCR) technology to process the received image data. Specifically, for OCR, the pytesseract library in Python is used to extract the character information in the image. At this stage, an equation such as "x^2 + 2x + 1 = 0" is obtained in text form.
[0289] The extracted character information and the text data converted from the user's voice input are combined and organized into a single text file. For example, it might take the form of: "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0290] Next, the server sends this combined text data to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text data. The generated response might be something like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0291] Furthermore, this invention also incorporates an emotion engine. The server analyzes the user's voice data and uses the emotion engine to recognize the user's emotional state. For example, if the tone and content of the user's questions indicate that they are stressed, the server can add words of encouragement to the generated response. "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Keep going!"
[0292] Finally, the generated response is sent from the server to the user's smart glasses and played back as speech using speech synthesis technology. For speech synthesis, the gTTS (Google® Text-to-Speech) library is used, for example.
[0293] The following shows an example of input to a generative AI model.
[0294] "Problem statement: x^2 + 2x + 1 = 0
[0295] Question: How do I solve this equation?
[0296] Emotion: Confusion.”
[0297] In this way, users can quickly and appropriately resolve learning problems and receive personalized learning support that is tailored to their emotions.
[0298] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0299] Step 1:
[0300] The user takes a picture of the problem they are studying with the camera on their smart glasses and uses the smart glasses' microphone to voice-input the question. Image data and audio data are generated from the input. The image data is treated as the user's visual information, and the audio data is treated as voice input information.
[0301] Step 2:
[0302] The smart glasses send captured image and audio data to the server using HTTP requests. The HTTP request module is used to transfer this data to the server as a single data package. The output is the data package received by the server.
[0303] Step 3:
[0304] The server analyzes the received image data using the Python pytesseract library and converts the character information within the image into text data. The input is image data, which is processed using optical character recognition technology, and the output is text data. For example, the text "x^2 + 2x + 1 = 0" is obtained.
[0305] Step 4:
[0306] The server uses speech recognition technology to convert speech data into text. This process uses speech recognition software, with the input being speech data and the output being the user's question text. For example, the text "Please teach me how to solve this equation." can be obtained.
[0307] Step 5:
[0308] The server combines the extracted text data and the question text converted by speech recognition to generate a prompt sentence. The input is the text "x^2 + 2x + 1 = 0" and the question text "Please teach me how to solve this equation." The output is the combined text "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please teach me how to solve this equation."
[0309] Step 6:
[0310] The server inputs the generated prompt sentence into the emotion engine to analyze the user's emotional state. The emotion engine estimates the emotion from the tone and content of the input text. For example, it detects that the user is confused. The input is "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please teach me how to solve this equation.", and the output is "Emotion: Confused".
[0311] Step 7:
[0312] The server inputs the combined text and the emotional state into the generative artificial intelligence model to generate an appropriate answer. The input is "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please teach me how to solve this equation. Emotion: Confused", and the output is the answer text "This equation can be solved as follows. First, factor both sides. x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Keep it up."
[0313] Step 8:
[0314] The server sends the generated response to the user's smart glasses. The response data is transferred to the smart glasses using an HTTP response. The input is the generated response text, and the output is the response text received by the smart glasses.
[0315] Step 9:
[0316] The smart glasses play back the received response text as speech using speech synthesis technology. The gTTS (Google Text-to-Speech) library is used for speech synthesis. The input is the response text, and the output is the speech played back to the user.
[0317] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0318] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0319] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0320] [Second Embodiment]
[0321] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0322] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0323] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0324] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0325] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0326] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0327] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0328] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0329] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0330] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0331] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0332] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0333] This invention relates to a system for quickly and accurately resolving learning problems using AI. The following describes specific embodiments of this invention.
[0334] First, the user takes a picture of the problem they are studying with a smartphone or other device and sends the image to the system. This is the "submitted image." At the same time, the user also enters a question about the problem in text format. For example, a question like, "Please tell me how to solve this problem."
[0335] Next, the terminal sends the received image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the Python pytesseract library to extract text information such as mathematical formulas and sentences from the image as text data.
[0336] The extracted text is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "x^2 + 2x + 1 = 0" and the user's question is "How do I solve this equation?", the server combines them to generate the following text.
[0337] Problem statement: x^2 + 2x + 1 = 0
[0338] Question: How do I solve this equation?
[0339] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0340] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0341] The generated answers are temporarily stored by the server and then sent to the user's device. The device receives the response from the server and displays the answer on the user interface. The user reviews this answer and uses it to aid in learning.
[0342] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of this equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image and generates a solution using a generative artificial intelligence model. As a result, the user receives the answer "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0343] As described above, the system of the present invention combines optical character recognition technology and a generative artificial intelligence model to enable users to quickly and accurately solve problems they are learning. This system provides learning support that is more immediate and cost-effective than conventional learning support methods.
[0344] The following describes the processing flow.
[0345] Step 1:
[0346] Users take pictures of problems or assignments they are studying using a smartphone or other device. At the same time, users also input questions about the problem in text format. For example, they might input a question like, "Please explain how to solve this problem."
[0347] Step 2:
[0348] The terminal receives the captured image and the entered question, and sends them to the server as a single data package. An HTTP request is used for this process.
[0349] Step 3:
[0350] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0351] Step 4:
[0352] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data.
[0353] Step 5:
[0354] The server temporarily stores the extracted text information. Let's assume the extracted text is "x^2 + 2x + 1 = 0".
[0355] Step 6:
[0356] The server combines the extracted text information with the user's question text. For example, it may be formatted as follows:
[0357] Problem statement: x^2 + 2x + 1 = 0
[0358] Question: How do I solve this equation?
[0359] Step 7:
[0360] The server sends the combined text to a generative artificial intelligence model. The generative AI model analyzes the combined text and generates an appropriate response.
[0361] Step 8:
[0362] A generative artificial intelligence model (e.g., GPT-3.5) generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1."
[0363] Step 9:
[0364] The server receives the generated response and sends it to the terminal. The HTTP protocol is used again for communication.
[0365] Step 10:
[0366] The terminal receives an HTTP response from the server and extracts the answer data. Then, it displays the answer on the user interface.
[0367] Step 11:
[0368] The user checks the answer displayed on their device. This allows the user to quickly and accurately obtain a solution to the problem.
[0369] The above steps provide answers to user questions. This system becomes a powerful tool for users to solve problems they are learning in real time.
[0370] (Example 1)
[0371] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0372] Conventional learning support systems have the challenge of not being able to quickly and accurately solve problems that users encounter during their studies. In particular, extracting problems from images containing text information, interpreting them appropriately, and providing immediate solutions is difficult, which reduces the user's learning efficiency. Furthermore, if an appropriate solution cannot be obtained, users are forced to interrupt their learning.
[0373] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0374] In this invention, the server includes means for receiving an image taken by the user, means for extracting character information from the received image using optical character recognition, means for combining the extracted character information with an input question, means for interpreting the combined text using natural language processing based on a generative artificial intelligence model and generating an answer, means for sending the generated answer to the user's terminal, and means for displaying the answer on the user's terminal. This makes it possible to quickly and accurately solve problems that the user faces during learning.
[0375] A "user" refers to an individual who uses this system to try to solve a problem they are learning.
[0376] "Device" refers to electronic devices such as smartphones and tablets used by users.
[0377] "Captured images" refers to image data containing the learning problems acquired by the user using the camera function of their device.
[0378] "Means of receiving" refers to the communication interface used to transfer images and questions taken by the user to the server.
[0379] "Optical character recognition" refers to a technology that analyzes character information within an image and extracts it as text data.
[0380] "Means of extraction" refers to software and hardware used to analyze text information within a received image and obtain it as text data.
[0381] "Method of combining" refers to the process of combining extracted character information and user-entered questions into a single text data file.
[0382] A "generative artificial intelligence model" refers to an algorithm or model that performs natural language processing based on given data to generate appropriate responses.
[0383] "Natural language processing" refers to the technology of processing human language using computers, and in this context, it refers to the technology of generating appropriate answers based on questions.
[0384] "Means of transmission" refers to the communication interface used to transfer the generated response to the user's device.
[0385] "Means of display" refers to a user interface used to visually show the answers obtained on the device.
[0386] This invention is a system for quickly and accurately resolving problems that users encounter during learning. The specific configuration and operation of the system are described as embodiments for carrying out this invention.
[0387] The user takes a picture of the problem they are studying with their device's camera and obtains the image data. For example, if the problem is written in their notebook as the equation "x^2 + 2x + 1 = 0", they would take a picture of that notebook. Then, the user uses the application on their device to type the question in text format, such as "Please tell me how to solve this equation."
[0388] The device sends the captured image and text question to the server as an HTTP request. The server receives this HTTP request and then processes the image data.
[0389] The server uses the Python library pytesseract to extract text information from the image using optical character recognition (OCR) technology. The OCR process extracts the text data "x^2 + 2x + 1 = 0".
[0390] The extracted text data and the questions entered by the user are combined to generate text like the following.
[0391] Problem statement: x^2 + 2x + 1 = 0
[0392] Question: How do I solve this equation?
[0393] Next, the server sends the generated text to a generative artificial intelligence model (for example, GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates an appropriate response. For example, the following response may be generated.
[0394] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0395] The generated response is temporarily stored on the server. The server then sends the generated response to the user's device as an HTTP response.
[0396] The device displays the received responses on the user interface. The user reviews these responses to aid their learning.
[0397] Thus, the system of the present invention enables users to quickly and accurately solve learning problems they face. By combining optical character recognition technology with a generative artificial intelligence model, the immediacy and accuracy are improved compared to conventional learning support methods, thereby enhancing the user's learning efficiency.
[0398] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0399] Step 1:
[0400] The user takes a picture of the problem they are studying using their device's camera. For example, they might take a picture of a notebook page with the equation "x^2 + 2x + 1 = 0" written on it. Simultaneously, they use the application to input a question. For example, they might input "Please tell me how to solve this equation." The input here consists of image data and text data. The output is an image file and a text question generated by the user's actions.
[0401] Step 2:
[0402] The terminal packages the image data and text data acquired by the user as an HTTP request and sends it to the server. The input consists of the image acquired by the user and the text question entered; this data is processed for transmission as an HTTP request to the server, and the output is the data sent to the server.
[0403] Step 3:
[0404] The server receives an HTTP request sent from the terminal. The input consists of image and text data sent from the terminal. The server uses the Python pytesseract library to analyze the image data and extracts text information from the image using optical character recognition (OCR) technology. Specifically, the mathematical expression "x^2 + 2x + 1 = 0" is extracted as text. This process outputs the input image as text data.
[0405] Step 4:
[0406] The server combines the text data extracted by OCR processing with the question text data entered by the user. For example, it generates a prompt message like the following:
[0407] Problem statement: x^2 + 2x + 1 = 0
[0408] Question: How do I solve this equation?
[0409] The input consists of extracted character information and question text, and the output is generated through a data merging process called prompt text generation.
[0410] Step 5:
[0411] The server sends the generated prompt to a generative artificial intelligence model, which then uses natural language processing (NLP) techniques to generate an appropriate response. For example, the following response might be generated:
[0412] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0413] The input here is a generated prompt sentence, and the output is the response text obtained from a generative artificial intelligence model.
[0414] Step 6:
[0415] The server temporarily stores the generated response and then sends it to the user's terminal as an HTTP response. The input is the generated response text, which is then sent to the terminal as an HTTP response. The output is the data sent to the terminal.
[0416] Step 7:
[0417] The terminal analyzes the response received from the server and displays it on the user interface. The input is the response text from the server, and the output is the visual information displayed on the user interface. The user looks at this to aid in learning.
[0418] Through the above series of processing steps, users can quickly and accurately solve problems they encounter during their learning process.
[0419] (Application Example 1)
[0420] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0421] Machine maintenance in modern factories is complex and requires quick and accurate solutions. However, many workplaces lack the necessary specialized technicians, often resulting in inefficient maintenance. To address this problem, a system is needed that accurately identifies machine problems and provides appropriate repair procedures.
[0422] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0423] In this invention, the server includes means for receiving posted images, means for extracting text information from the received images using optical character recognition, means for combining the extracted text information with an entered question, means for interpreting the combined text using natural language processing and generating an answer, means for transmitting the generated answer, and means for providing identification and repair procedures for damaged machine parts to support machine maintenance work in the factory. This enables factory workers to quickly and accurately identify machine problems and obtain appropriate repair procedures using smartphones or tablets.
[0424] "Means for receiving images" refers to a device or software that has the function of receiving and storing image data transmitted by a user.
[0425] "Means of extraction using optical character recognition" refers to devices or software that use technology to analyze character information within an image and convert it into text data.
[0426] "Means of combining" refers to a function that integrates extracted text information and user-entered questions into a single text file.
[0427] "Means of interpretation using natural language processing" refers to techniques that analyze combined texts and generate answers in an easily understandable format.
[0428] "Means for generating answers" refers to a device or software that has the function of providing appropriate answers from information analyzed by natural language processing.
[0429] "Means of sending responses" refers to a device or software for sending generated responses to a user's device.
[0430] "Means of supporting machine maintenance work" refers to devices or software that have the function of identifying problems with machinery and providing repair methods and procedures.
[0431] "Identifying damaged machine parts" refers to the process of recognizing damaged parts from captured images and extracting their details.
[0432] "Means of providing repair procedures" refers to a device or software that has the function of notifying the user of the appropriate repair method or process for an identified problem.
[0433] This invention relates to a system for supporting machine maintenance work in a factory. The following describes specific embodiments for carrying out this invention.
[0434] First, the user takes a picture of the problematic machine in the factory using a smartphone or tablet and sends the image to the system. At the same time, the user also enters a question about the machine in text format, such as, "How do I replace this part?"
[0435] Next, the device sends the captured image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the pytesseract library to extract text data such as character information and part names from the image.
[0436] The extracted text data is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "bolt" and "gear," and the user's question is "How do I replace this part?", the server combines these to generate the following text.
[0437] Problem statement: How to replace bolts and gears
[0438] Question: How do I replace this part?
[0439] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0440] To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt.
[0441] The generated response is temporarily stored by the server and then sent to the user's terminal. The terminal receives the response from the server and displays the response on the user interface. The user reviews this response and uses it as a reference for maintenance work.
[0442] As a concrete example, the process involves a user sending an image of a problematic machine part to the system along with the question, "How do I replace this bolt and gear?" The server extracts information about the bolt and gear from the image and generates a replacement method using a generative artificial intelligence model. As a result, the user receives the answer, "To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt."
[0443] This system combines optical character recognition technology with generative artificial intelligence models to enable factory workers to perform machine maintenance quickly and accurately. This improves on-site work efficiency and allows for proper maintenance even in situations where there is a shortage of specialized technicians.
[0444] Example of a prompt:
[0445] Problem statement: How to replace bolts and gears
[0446] Question: How do I replace this part?
[0447] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0448] Step 1:
[0449] Users take pictures of problematic machinery in the factory using their smartphones or tablets and send the images to the system. At the same time, they also enter questions about the machinery in text format. This process sends both the image and text data from the device to the server.
[0450] Step 2:
[0451] The device sends the captured image and the entered question to the server as an HTTP request. Specifically, the device sends the image file and text as a POST request to the specified endpoint on the server. The input consists of image data and text data, and the output is the request data sent to the server.
[0452] Step 3:
[0453] The server receives the transmitted image. The received image data is temporarily stored on the server. In this step, the input is the image data, and the output is the stored image data.
[0454] Step 4:
[0455] The server uses Optical Character Recognition (OCR) technology to extract text information from the received image. The server uses the pytesseract library to obtain the text information within the image. The input is image data, and the output is the extracted text data.
[0456] Step 5:
[0457] The server combines the extracted text data with the question entered by the user. This is done to generate the problem statement and question as a single text file. The input is the extracted text data and the question data, and the output is the combined text data.
[0458] Step 6:
[0459] The server sends the combined text data to a generative artificial intelligence model (e.g., GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates a response. In this step, prompts are sent to the generative AI model via an API to generate the response text. The input is the combined text data, and the output is the generated response text.
[0460] Step 7:
[0461] The generated response is temporarily stored on the server and then sent to the user's device. Specifically, the server returns a response containing the generated response to the user's device. The input is the generated response text, and the output is the response data sent to the user's device.
[0462] Step 8:
[0463] The terminal receives a response from the server and displays the answer on the user interface. The user reviews this answer and uses it as a reference for machine maintenance work. The input is the response data from the server, and the output is the displayed answer data.
[0464] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0465] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0466] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0467] Next, the device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and the user's question contained within the request.
[0468] First, the server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0469] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0470] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0471] Furthermore, this invention includes an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when a user enters text for a question, the emotional tone is analyzed from the content and style of the text.
[0472] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, it can add more detailed explanations if the user is confused, or add words of encouragement if the server detects that the user is stressed.
[0473] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. The user reviews the displayed answers and uses them to aid in learning.
[0474] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user receives an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0475] Through the steps described above, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0476] The following describes the processing flow.
[0477] Step 1:
[0478] Users use their smartphones or tablets to take pictures of problems or assignments they are studying. At the same time, users also input questions about the problems in text format. For example, they might input a question like, "Please explain how to solve this equation."
[0479] Step 2:
[0480] The device combines the captured image and the user's input into a single data package. This data package is then sent to the server as an HTTP request.
[0481] Step 3:
[0482] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0483] Step 4:
[0484] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, it extracts the equation "x^2 + 2x + 1 = 0" from the image.
[0485] Step 5:
[0486] The server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. The generated text will take the form of, for example, "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0487] Step 6:
[0488] The server sends the combined text data to the sentiment engine. The sentiment engine analyzes the content and format of the questions entered by the user and evaluates the user's emotions. For example, it recognizes emotions such as the user being nervous, confused, or irritated.
[0489] Step 7:
[0490] Based on the emotion data recognized by the emotion engine, the server sends the combined text to a generative artificial intelligence model. The generative AI model (e.g., GPT-3.5) analyzes the combined text and generates an appropriate response.
[0491] Step 8:
[0492] A generative artificial intelligence model generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0493] Step 9:
[0494] The server receives the generated response and adjusts it. Based on the user's emotions recognized by the emotion engine, the server modifies how the response is expressed. For example, if the user is confused, it may add a more detailed explanation or use gentler language.
[0495] Step 10:
[0496] The server sends the final response data to the terminal.
[0497] Step 11:
[0498] The terminal receives an HTTP response from the server and displays the response data on the user interface.
[0499] Step 12:
[0500] Users review the answers displayed on their device to aid their learning. Because these answers are adjusted by an emotion engine, users receive effective learning support that takes their emotional state into account.
[0501] Through the steps described above, users receive immediate and appropriate answers to their questions. This system allows users to quickly and accurately resolve learning problems, providing learning support that is more immediate and cost-effective than traditional methods.
[0502] (Example 2)
[0503] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0504] Conventional learning support systems can interpret questions from image data and generate answers, but they cannot take into account the user's emotions, which means some users may not receive appropriate support. For example, if a user is emotionally stressed, the answer provided may reduce their learning efficiency. In such situations, appropriate feedback that takes the user's emotions into account is necessary.
[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0506] In this invention, the server includes means for receiving posted images, means for extracting textual information from the received images using optical character recognition, means for combining the extracted textual information with the entered question, means for interpreting the combined text using natural language processing and generating an answer, means for analyzing the user's entered question and dialogue content and recognizing their emotional state, means for adjusting the generated answer based on the recognized emotion, and means for transmitting the generated answer. This makes it possible not only to provide appropriate answers according to the user's learning progress but also to provide feedback according to the user's emotional state.
[0507] "Submitted images" refer to image data containing learning problems that users capture using their devices and send to the server.
[0508] "Textual information" refers to text data contained within the posted image, specifically including strings of characters such as equations and mathematical problems.
[0509] Optical character recognition (OCR) is a technology that extracts character information from images as text data, and is sometimes abbreviated as OCR.
[0510] "Means of receiving" refers to the interface or function that a server uses to receive data sent from a client, such as an HTTP request.
[0511] "Means of extraction" refers to software or libraries used to extract textual information from images. For example, optical character recognition libraries fall into this category.
[0512] A "question" is text data that a user enters regarding a learning problem they want to solve.
[0513] "Means of combining" refers to a process or function for integrating extracted character information and user-entered questions into a single text data file.
[0514] Natural language processing (NLP) is the technology that enables computers to understand and generate human language. This processing includes text analysis and generation.
[0515] A "generative artificial intelligence model" is an artificial intelligence technology that generates natural language responses based on provided input data, and for example, it utilizes natural language processing technology.
[0516] "Means of recognizing emotional states" refers to algorithms and engines that analyze user question text and dialogue content to determine the user's emotions.
[0517] "Means of adjustment" refers to a process or function for appropriately modifying or supplementing responses generated based on perceived emotional states.
[0518] "Means of transmission" refers to an interface or function for sending generated response data from the server to the client.
[0519] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0520] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0521] The device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and user question contained within the request.
[0522] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0523] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0524] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0525] Furthermore, the system of the present invention is equipped with an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when the user enters the text of a question, it analyzes the emotional tone from the content and style of the text.
[0526] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, if it detects that the user is feeling stressed due to a question, it can add words of encouragement.
[0527] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. As another example, if the system recognizes that the user has a question, it can add a detailed explanation. The user reviews the displayed answers and uses them to aid in their learning.
[0528] Specific example
[0529] This illustrates the process for a user who wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends it to the system with the question, "Please tell me how to solve this equation." The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user can receive an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0530] Example of a prompt
[0531] Problem statement: x^2 + 2x + 1 = 0
[0532] Question: How do I solve this equation?
[0533] In the above configuration, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0534] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0535] Step 1:
[0536] The user takes a picture of the problem and enters the question.
[0537] Users take photos of the problems they are learning using a device such as a smartphone or tablet. For example, they might take a picture of a piece of paper with "x^2 + 2x + 1 = 0" written on it. At the same time, they input a question into the device, such as "Please tell me how to solve this equation." The input is done through a text box.
[0538] Input: Captured image and text question
[0539] Output: Combine images and text questions into a single data package.
[0540] Step 2:
[0541] The terminal sends the data package to the server.
[0542] The device generates a single data package containing the captured image and the entered question. This data package is sent to the server via an HTTP request. The device packages the image data and text question as a payload and sends it over the network.
[0543] Input: Data package (including images and text questions)
[0544] Output: Data package sent to the server
[0545] Step 3:
[0546] The server analyzes the image data using OCR technology.
[0547] The server extracts image data from the received data package. Optical Character Recognition (OCR) technology is used to extract text information from the image as text data. The specific software used in this process is the Python library pytesseract. For example, the string "x^2 + 2x + 1 = 0" is extracted from the image.
[0548] Input: Image data
[0549] Output: Text information (text data) within the image
[0550] Step 4:
[0551] The server combines the text information and the question to generate text data.
[0552] The server combines the character information extracted by OCR with the user's question text to generate a single text file. This text file contains the problem statement and the user's question. For example, it might take the form of: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0553] Input: Text information (text data) and question text
[0554] Output: Combined text data
[0555] Step 5:
[0556] The server sends text data to an AI model that generates an answer.
[0557] The server sends the combined text data to a generative artificial intelligence model. This model uses natural language processing techniques to generate an appropriate response. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0558] Input: Combined text data
[0559] Output: Text response
[0560] Step 6:
[0561] The server uses an emotion engine to recognize the user's emotions and adjust the generated response accordingly.
[0562] The server is equipped with an emotion engine that analyzes the user's input questions and conversations to recognize their emotional state. Based on the recognized emotion, the generated response is appropriately adjusted. For example, if the server detects that the user is feeling stressed, it will add encouraging words such as, "It's okay, you can do it!"
[0563] Input: Text question and generated text answer
[0564] Output: Adjusted text response
[0565] Step 7:
[0566] The server sends the response to the terminal, and the terminal displays it to the user.
[0567] The server sends the final answer to the terminal. The terminal displays the received answer on the user interface. The user checks the displayed answer and uses it to aid in learning. For example, the displayed message might be: "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Don't worry, you can do it!"
[0568] Input: Adjusted text response
[0569] Output: Answer displayed on the terminal
[0570] (Application Example 2)
[0571] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0572] In today's educational environment, students often face numerous problems and questions. However, systems that allow them to immediately resolve their doubts during self-study or in class are limited. Furthermore, there are very few systems that understand students' emotional states and provide learning support accordingly. As a result, students' learning efficiency may decrease, potentially leading to a decline in their motivation to learn.
[0573] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving posted images, means for extracting character information from the received images using optical character recognition, means for combining the extracted character information with the input question, means for interpreting the combined text using natural language processing and generating an answer, means for recognizing the user's emotions and adjusting the generated answer according to the recognized emotions, and means for transmitting the generated answer and playing it back as audio. As a result, students can solve problems they are learning in real time and receive detailed learning support tailored to their emotional state.
[0574] "Means for receiving posted images" refers to a function that allows the server to receive image data taken by users using their smart devices.
[0575] "Means for extracting character information from a received image using optical character recognition" refers to a function that extracts character and numerical information contained within an image as text data using optical character recognition technology.
[0576] "Means for combining extracted text information and entered questions" refers to a function that integrates text information extracted from an image with questions submitted by the user into a single text data file.
[0577] "Means for interpreting combined text using natural language processing and generating answers" refers to a function that analyzes combined text using a generative artificial intelligence model and creates an appropriate answer based on that analysis.
[0578] "Means for recognizing user emotions and adjusting generated responses accordingly" refers to a function that analyzes the user's emotional state from their text or voice tone and makes changes or adjustments to the generated responses based on the results.
[0579] "A means of sending generated answers and playing them back as audio" refers to a function that sends answers from a server to the user's terminal and outputs those answers as audio using speech synthesis technology.
[0580] An "optical character recognition library" is a software library used to extract character information from images using optical character recognition technology.
[0581] A "generative artificial intelligence model" is an AI model that uses natural language processing to analyze text data and generate appropriate responses or information.
[0582] The system that realizes this invention is composed of multiple components, including smart glasses, a server, and speech synthesis technology. Specific embodiments are described below.
[0583] First, the user uses smart glasses to capture images of the problems they are studying. The user also submits questions to the system using voice input. Specifically, they take pictures with the smart glasses' built-in camera and record their questions with the microphone.
[0584] Next, the smart glasses send the captured image and audio data to the server as a single data package. This transmission is done via an HTTP request.
[0585] The server uses optical character recognition (OCR) technology to process the received image data. Specifically, the Python library pytesseract is used for OCR to extract text information from the image. At this stage, equations such as "x^2 + 2x + 1 = 0" are obtained in text format.
[0586] The extracted character information and the text data converted from the user's voice input are combined and organized into a single text file. For example, it might take the form of: "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0587] Next, the server sends this combined text data to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text data. The generated response might be something like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0588] Furthermore, this invention also incorporates an emotion engine. The server analyzes the user's voice data and uses the emotion engine to recognize the user's emotional state. For example, if the tone and content of the user's questions indicate that they are stressed, the server can add words of encouragement to the generated response. "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Keep going!"
[0589] Finally, the generated response is sent from the server to the user's smart glasses and played back as speech using speech synthesis technology. For speech synthesis, the gTTS (Google Text-to-Speech) library is used, for example.
[0590] The following shows an example of input to a generative AI model.
[0591] "Problem statement: x^2 + 2x + 1 = 0
[0592] Question: How do I solve this equation?
[0593] Emotion: Confusion.”
[0594] In this way, users can quickly and appropriately resolve learning problems and receive personalized learning support that is tailored to their emotions.
[0595] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0596] Step 1:
[0597] The user takes a picture of the problem they are studying with the camera on their smart glasses and uses the smart glasses' microphone to voice-input the question. Image data and audio data are generated from the input. The image data is treated as the user's visual information, and the audio data is treated as voice input information.
[0598] Step 2:
[0599] The smart glasses send captured image and audio data to the server using HTTP requests. The HTTP request module is used to transfer this data to the server as a single data package. The output is the data package received by the server.
[0600] Step 3:
[0601] The server analyzes the received image data using the Python pytesseract library and converts the character information within the image into text data. The input is image data, which is processed using optical character recognition technology, and the output is text data. For example, the text "x^2 + 2x + 1 = 0" is obtained.
[0602] Step 4:
[0603] The server uses speech recognition technology to convert speech data into text. This process uses speech recognition software, with speech data as input and the user's question text as output. For example, it might produce the text, "Please tell me how to solve this equation."
[0604] Step 5:
[0605] The server combines the extracted text data with the question text converted by speech recognition to generate a prompt. The input is the text "x^2 + 2x + 1 = 0" and the question text "Please tell me how to solve this equation." The output is the combined text "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0606] Step 6:
[0607] The server inputs the generated prompt text into the emotion engine, which analyzes the user's emotional state. The emotion engine estimates the emotion from the tone and content of the input text. For example, it can detect that the user is confused. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: How do I solve this equation?" and the output is "Emotion: Confused".
[0608] Step 7:
[0609] The server combines the text and emotional state and inputs it into a generative artificial intelligence model to generate an appropriate response. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation. Emotion: Confused", and the output is the response text "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1. Good luck."
[0610] Step 8:
[0611] The server sends the generated response to the user's smart glasses. The response data is transferred to the smart glasses using an HTTP response. The input is the generated response text, and the output is the response text received by the smart glasses.
[0612] Step 9:
[0613] The smart glasses play back the received response text as speech using speech synthesis technology. The gTTS (Google Text-to-Speech) library is used for speech synthesis. The input is the response text, and the output is the speech played back to the user.
[0614] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0615] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0616] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0617] [Third Embodiment]
[0618] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0619] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0620] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0621] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0622] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0623] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0624] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0625] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0626] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0627] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0628] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0629] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0630] This invention relates to a system for quickly and accurately resolving learning problems using AI. The following describes specific embodiments of this invention.
[0631] First, the user takes a picture of the problem they are studying with a smartphone or other device and sends the image to the system. This is the "submitted image." At the same time, the user also enters a question about the problem in text format. For example, a question like, "Please tell me how to solve this problem."
[0632] Next, the terminal sends the received image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the Python pytesseract library to extract text information such as mathematical formulas and sentences from the image as text data.
[0633] The extracted text is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "x^2 + 2x + 1 = 0" and the user's question is "How do I solve this equation?", the server combines them to generate the following text.
[0634] Problem statement: x^2 + 2x + 1 = 0
[0635] Question: How do I solve this equation?
[0636] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0637] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0638] The generated answers are temporarily stored by the server and then sent to the user's device. The device receives the response from the server and displays the answer on the user interface. The user reviews this answer and uses it to aid in learning.
[0639] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of this equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image and generates a solution using a generative artificial intelligence model. As a result, the user receives the answer "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0640] As described above, the system of the present invention combines optical character recognition technology and a generative artificial intelligence model to enable users to quickly and accurately solve problems they are learning. This system provides learning support that is more immediate and cost-effective than conventional learning support methods.
[0641] The following describes the processing flow.
[0642] Step 1:
[0643] Users take pictures of problems or assignments they are studying using a smartphone or other device. At the same time, users also input questions about the problem in text format. For example, they might input a question like, "Please explain how to solve this problem."
[0644] Step 2:
[0645] The terminal receives the captured image and the entered question, and sends them to the server as a single data package. An HTTP request is used for this process.
[0646] Step 3:
[0647] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0648] Step 4:
[0649] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data.
[0650] Step 5:
[0651] The server temporarily stores the extracted text information. Let's assume the extracted text is "x^2 + 2x + 1 = 0".
[0652] Step 6:
[0653] The server combines the extracted text information with the user's question text. For example, it may be formatted as follows:
[0654] Problem statement: x^2 + 2x + 1 = 0
[0655] Question: How do I solve this equation?
[0656] Step 7:
[0657] The server sends the combined text to a generative artificial intelligence model. The generative AI model analyzes the combined text and generates an appropriate response.
[0658] Step 8:
[0659] A generative artificial intelligence model (e.g., GPT-3.5) generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1."
[0660] Step 9:
[0661] The server receives the generated response and sends it to the terminal. The HTTP protocol is used again for communication.
[0662] Step 10:
[0663] The terminal receives an HTTP response from the server and extracts the answer data. Then, it displays the answer on the user interface.
[0664] Step 11:
[0665] The user checks the answer displayed on their device. This allows the user to quickly and accurately obtain a solution to the problem.
[0666] The above steps provide answers to user questions. This system becomes a powerful tool for users to solve problems they are learning in real time.
[0667] (Example 1)
[0668] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0669] Conventional learning support systems have the challenge of not being able to quickly and accurately solve problems that users encounter during their studies. In particular, extracting problems from images containing text information, interpreting them appropriately, and providing immediate solutions is difficult, which reduces the user's learning efficiency. Furthermore, if an appropriate solution cannot be obtained, users are forced to interrupt their learning.
[0670] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0671] In this invention, the server includes means for receiving an image taken by the user, means for extracting character information from the received image using optical character recognition, means for combining the extracted character information with an input question, means for interpreting the combined text using natural language processing based on a generative artificial intelligence model and generating an answer, means for sending the generated answer to the user's terminal, and means for displaying the answer on the user's terminal. This makes it possible to quickly and accurately solve problems that the user faces during learning.
[0672] A "user" refers to an individual who uses this system to try to solve a problem they are learning.
[0673] "Device" refers to electronic devices such as smartphones and tablets used by users.
[0674] "Captured images" refers to image data containing the learning problems acquired by the user using the camera function of their device.
[0675] "Means of receiving" refers to the communication interface used to transfer images and questions taken by the user to the server.
[0676] "Optical character recognition" refers to a technology that analyzes character information within an image and extracts it as text data.
[0677] "Means of extraction" refers to software and hardware used to analyze text information within a received image and obtain it as text data.
[0678] "Method of combining" refers to the process of combining extracted character information and user-entered questions into a single text data file.
[0679] A "generative artificial intelligence model" refers to an algorithm or model that performs natural language processing based on given data to generate appropriate responses.
[0680] "Natural language processing" refers to the technology of processing human language using computers, and in this context, it refers to the technology of generating appropriate answers based on questions.
[0681] "Means of transmission" refers to the communication interface used to transfer the generated response to the user's device.
[0682] "Means of display" refers to a user interface used to visually show the answers obtained on the device.
[0683] This invention is a system for quickly and accurately resolving problems that users encounter during learning. The specific configuration and operation of the system are described as embodiments for carrying out this invention.
[0684] The user takes a picture of the problem they are studying with their device's camera and obtains the image data. For example, if the problem is written in their notebook as the equation "x^2 + 2x + 1 = 0", they would take a picture of that notebook. Then, the user uses the application on their device to type the question in text format, such as "Please tell me how to solve this equation."
[0685] The device sends the captured image and text question to the server as an HTTP request. The server receives this HTTP request and then processes the image data.
[0686] The server uses the Python library pytesseract to extract text information from the image using optical character recognition (OCR) technology. The OCR process extracts the text data "x^2 + 2x + 1 = 0".
[0687] The extracted text data and the questions entered by the user are combined to generate text like the following.
[0688] Problem statement: x^2 + 2x + 1 = 0
[0689] Question: How do I solve this equation?
[0690] Next, the server sends the generated text to a generative artificial intelligence model (for example, GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates an appropriate response. For example, the following response may be generated.
[0691] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0692] The generated response is temporarily stored on the server. The server then sends the generated response to the user's device as an HTTP response.
[0693] The device displays the received responses on the user interface. The user reviews these responses to aid their learning.
[0694] Thus, the system of the present invention enables users to quickly and accurately solve learning problems they face. By combining optical character recognition technology with a generative artificial intelligence model, the immediacy and accuracy are improved compared to conventional learning support methods, thereby enhancing the user's learning efficiency.
[0695] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0696] Step 1:
[0697] The user takes a picture of the problem they are studying using their device's camera. For example, they might take a picture of a notebook page with the equation "x^2 + 2x + 1 = 0" written on it. Simultaneously, they use the application to input a question. For example, they might input "Please tell me how to solve this equation." The input here consists of image data and text data. The output is an image file and a text question generated by the user's actions.
[0698] Step 2:
[0699] The terminal packages the image data and text data acquired by the user as an HTTP request and sends it to the server. The input consists of the image acquired by the user and the text question entered; this data is processed for transmission as an HTTP request to the server, and the output is the data sent to the server.
[0700] Step 3:
[0701] The server receives an HTTP request sent from the terminal. The input consists of image and text data sent from the terminal. The server uses the Python pytesseract library to analyze the image data and extracts text information from the image using optical character recognition (OCR) technology. Specifically, the mathematical expression "x^2 + 2x + 1 = 0" is extracted as text. This process outputs the input image as text data.
[0702] Step 4:
[0703] The server combines the text data extracted by OCR processing with the question text data entered by the user. For example, it generates a prompt message like the following:
[0704] Problem statement: x^2 + 2x + 1 = 0
[0705] Question: How do I solve this equation?
[0706] The input consists of extracted character information and question text, and the output is generated through a data merging process called prompt text generation.
[0707] Step 5:
[0708] The server sends the generated prompt to a generative artificial intelligence model, which then uses natural language processing (NLP) techniques to generate an appropriate response. For example, the following response might be generated:
[0709] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0710] The input here is a generated prompt sentence, and the output is the response text obtained from a generative artificial intelligence model.
[0711] Step 6:
[0712] The server temporarily stores the generated response and then sends it to the user's terminal as an HTTP response. The input is the generated response text, which is then sent to the terminal as an HTTP response. The output is the data sent to the terminal.
[0713] Step 7:
[0714] The terminal analyzes the response received from the server and displays it on the user interface. The input is the response text from the server, and the output is the visual information displayed on the user interface. The user looks at this to aid in learning.
[0715] Through the above series of processing steps, users can quickly and accurately solve problems they encounter during their learning process.
[0716] (Application Example 1)
[0717] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0718] Machine maintenance in modern factories is complex and requires quick and accurate solutions. However, many workplaces lack the necessary specialized technicians, often resulting in inefficient maintenance. To address this problem, a system is needed that accurately identifies machine problems and provides appropriate repair procedures.
[0719] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0720] In this invention, the server includes means for receiving posted images, means for extracting text information from the received images using optical character recognition, means for combining the extracted text information with an entered question, means for interpreting the combined text using natural language processing and generating an answer, means for transmitting the generated answer, and means for providing identification and repair procedures for damaged machine parts to support machine maintenance work in the factory. This enables factory workers to quickly and accurately identify machine problems and obtain appropriate repair procedures using smartphones or tablets.
[0721] "Means for receiving images" refers to a device or software that has the function of receiving and storing image data transmitted by a user.
[0722] "Means of extraction using optical character recognition" refers to devices or software that use technology to analyze character information within an image and convert it into text data.
[0723] "Means of combining" refers to a function that integrates extracted text information and user-entered questions into a single text file.
[0724] "Means of interpretation using natural language processing" refers to techniques that analyze combined texts and generate answers in an easily understandable format.
[0725] "Means for generating answers" refers to a device or software that has the function of providing appropriate answers from information analyzed by natural language processing.
[0726] "Means of sending responses" refers to a device or software for sending generated responses to a user's device.
[0727] "Means of supporting machine maintenance work" refers to devices or software that have the function of identifying problems with machinery and providing repair methods and procedures.
[0728] "Identifying damaged machine parts" refers to the process of recognizing damaged parts from captured images and extracting their details.
[0729] "Means of providing repair procedures" refers to a device or software that has the function of notifying the user of the appropriate repair method or process for an identified problem.
[0730] This invention relates to a system for supporting machine maintenance work in a factory. The following describes specific embodiments for carrying out this invention.
[0731] First, the user takes a picture of the problematic machine in the factory using a smartphone or tablet and sends the image to the system. At the same time, the user also enters a question about the machine in text format, such as, "How do I replace this part?"
[0732] Next, the device sends the captured image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the pytesseract library to extract text data such as character information and part names from the image.
[0733] The extracted text data is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "bolt" and "gear," and the user's question is "How do I replace this part?", the server combines these to generate the following text.
[0734] Problem statement: How to replace bolts and gears
[0735] Question: How do I replace this part?
[0736] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0737] To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt.
[0738] The generated response is temporarily stored by the server and then sent to the user's terminal. The terminal receives the response from the server and displays the response on the user interface. The user reviews this response and uses it as a reference for maintenance work.
[0739] As a concrete example, the process involves a user sending an image of a problematic machine part to the system along with the question, "How do I replace this bolt and gear?" The server extracts information about the bolt and gear from the image and generates a replacement method using a generative artificial intelligence model. As a result, the user receives the answer, "To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt."
[0740] This system combines optical character recognition technology with generative artificial intelligence models to enable factory workers to perform machine maintenance quickly and accurately. This improves on-site work efficiency and allows for proper maintenance even in situations where there is a shortage of specialized technicians.
[0741] Example of a prompt:
[0742] Problem statement: How to replace bolts and gears
[0743] Question: How do I replace this part?
[0744] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0745] Step 1:
[0746] Users take pictures of problematic machinery in the factory using their smartphones or tablets and send the images to the system. At the same time, they also enter questions about the machinery in text format. This process sends both the image and text data from the device to the server.
[0747] Step 2:
[0748] The device sends the captured image and the entered question to the server as an HTTP request. Specifically, the device sends the image file and text as a POST request to the specified endpoint on the server. The input consists of image data and text data, and the output is the request data sent to the server.
[0749] Step 3:
[0750] The server receives the transmitted image. The received image data is temporarily stored on the server. In this step, the input is the image data, and the output is the stored image data.
[0751] Step 4:
[0752] The server uses Optical Character Recognition (OCR) technology to extract text information from the received image. The server uses the pytesseract library to obtain the text information within the image. The input is image data, and the output is the extracted text data.
[0753] Step 5:
[0754] The server combines the extracted text data with the question entered by the user. This is done to generate the problem statement and question as a single text file. The input is the extracted text data and the question data, and the output is the combined text data.
[0755] Step 6:
[0756] The server sends the combined text data to a generative artificial intelligence model (e.g., GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates a response. In this step, prompts are sent to the generative AI model via an API to generate the response text. The input is the combined text data, and the output is the generated response text.
[0757] Step 7:
[0758] The generated response is temporarily stored on the server and then sent to the user's device. Specifically, the server returns a response containing the generated response to the user's device. The input is the generated response text, and the output is the response data sent to the user's device.
[0759] Step 8:
[0760] The terminal receives a response from the server and displays the answer on the user interface. The user reviews this answer and uses it as a reference for machine maintenance work. The input is the response data from the server, and the output is the displayed answer data.
[0761] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0762] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0763] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0764] Next, the device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and the user's question contained within the request.
[0765] First, the server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0766] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0767] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0768] Furthermore, this invention includes an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when a user enters text for a question, the emotional tone is analyzed from the content and style of the text.
[0769] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, it can add more detailed explanations if the user is confused, or add words of encouragement if the server detects that the user is stressed.
[0770] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. The user reviews the displayed answers and uses them to aid in learning.
[0771] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user receives an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0772] Through the steps described above, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0773] The following describes the processing flow.
[0774] Step 1:
[0775] Users use their smartphones or tablets to take pictures of problems or assignments they are studying. At the same time, users also input questions about the problems in text format. For example, they might input a question like, "Please explain how to solve this equation."
[0776] Step 2:
[0777] The device combines the captured image and the user's input into a single data package. This data package is then sent to the server as an HTTP request.
[0778] Step 3:
[0779] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0780] Step 4:
[0781] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, it extracts the equation "x^2 + 2x + 1 = 0" from the image.
[0782] Step 5:
[0783] The server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. The generated text will take the form of, for example, "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0784] Step 6:
[0785] The server sends the combined text data to the sentiment engine. The sentiment engine analyzes the content and format of the questions entered by the user and evaluates the user's emotions. For example, it recognizes emotions such as the user being nervous, confused, or irritated.
[0786] Step 7:
[0787] Based on the emotion data recognized by the emotion engine, the server sends the combined text to a generative artificial intelligence model. The generative AI model (e.g., GPT-3.5) analyzes the combined text and generates an appropriate response.
[0788] Step 8:
[0789] A generative artificial intelligence model generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0790] Step 9:
[0791] The server receives the generated response and adjusts it. Based on the user's emotions recognized by the emotion engine, the server modifies how the response is expressed. For example, if the user is confused, it may add a more detailed explanation or use gentler language.
[0792] Step 10:
[0793] The server sends the final response data to the terminal.
[0794] Step 11:
[0795] The terminal receives an HTTP response from the server and displays the response data on the user interface.
[0796] Step 12:
[0797] Users review the answers displayed on their device to aid their learning. Because these answers are adjusted by an emotion engine, users receive effective learning support that takes their emotional state into account.
[0798] Through the steps described above, users receive immediate and appropriate answers to their questions. This system allows users to quickly and accurately resolve learning problems, providing learning support that is more immediate and cost-effective than traditional methods.
[0799] (Example 2)
[0800] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0801] Conventional learning support systems can interpret questions from image data and generate answers, but they cannot take into account the user's emotions, which means some users may not receive appropriate support. For example, if a user is emotionally stressed, the answer provided may reduce their learning efficiency. In such situations, appropriate feedback that takes the user's emotions into account is necessary.
[0802] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0803] In this invention, the server includes means for receiving posted images, means for extracting textual information from the received images using optical character recognition, means for combining the extracted textual information with the entered question, means for interpreting the combined text using natural language processing and generating an answer, means for analyzing the user's entered question and dialogue content and recognizing their emotional state, means for adjusting the generated answer based on the recognized emotion, and means for transmitting the generated answer. This makes it possible not only to provide appropriate answers according to the user's learning progress but also to provide feedback according to the user's emotional state.
[0804] "Submitted images" refer to image data containing learning problems that users capture using their devices and send to the server.
[0805] "Textual information" refers to text data contained within the posted image, specifically including strings of characters such as equations and mathematical problems.
[0806] Optical character recognition (OCR) is a technology that extracts character information from images as text data, and is sometimes abbreviated as OCR.
[0807] "Means of receiving" refers to the interface or function that a server uses to receive data sent from a client, such as an HTTP request.
[0808] "Means of extraction" refers to software or libraries used to extract textual information from images. For example, optical character recognition libraries fall into this category.
[0809] A "question" is text data that a user enters regarding a learning problem they want to solve.
[0810] "Means of combining" refers to a process or function for integrating extracted character information and user-entered questions into a single text data file.
[0811] Natural language processing (NLP) is the technology that enables computers to understand and generate human language. This processing includes text analysis and generation.
[0812] A "generative artificial intelligence model" is an artificial intelligence technology that generates natural language responses based on provided input data, and for example, it utilizes natural language processing technology.
[0813] "Means of recognizing emotional states" refers to algorithms and engines that analyze user question text and dialogue content to determine the user's emotions.
[0814] "Means of adjustment" refers to a process or function for appropriately modifying or supplementing responses generated based on perceived emotional states.
[0815] "Means of transmission" refers to an interface or function for sending generated response data from the server to the client.
[0816] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[0817] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[0818] The device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and user question contained within the request.
[0819] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[0820] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0821] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0822] Furthermore, the system of the present invention is equipped with an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when the user enters the text of a question, it analyzes the emotional tone from the content and style of the text.
[0823] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, if it detects that the user is feeling stressed due to a question, it can add words of encouragement.
[0824] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. As another example, if the system recognizes that the user has a question, it can add a detailed explanation. The user reviews the displayed answers and uses them to aid in their learning.
[0825] Specific example
[0826] This illustrates the process for a user who wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends it to the system with the question, "Please tell me how to solve this equation." The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user can receive an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0827] Example of a prompt
[0828] Problem statement: x^2 + 2x + 1 = 0
[0829] Question: How do I solve this equation?
[0830] In the above configuration, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[0831] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0832] Step 1:
[0833] The user takes a picture of the problem and enters the question.
[0834] Users take photos of the problems they are learning using a device such as a smartphone or tablet. For example, they might take a picture of a piece of paper with "x^2 + 2x + 1 = 0" written on it. At the same time, they input a question into the device, such as "Please tell me how to solve this equation." The input is done through a text box.
[0835] Input: Captured image and text question
[0836] Output: Combine images and text questions into a single data package.
[0837] Step 2:
[0838] The terminal sends the data package to the server.
[0839] The device generates a single data package containing the captured image and the entered question. This data package is sent to the server via an HTTP request. The device packages the image data and text question as a payload and sends it over the network.
[0840] Input: Data package (including images and text questions)
[0841] Output: Data package sent to the server
[0842] Step 3:
[0843] The server analyzes the image data using OCR technology.
[0844] The server extracts image data from the received data package. Optical Character Recognition (OCR) technology is used to extract text information from the image as text data. The specific software used in this process is the Python library pytesseract. For example, the string "x^2 + 2x + 1 = 0" is extracted from the image.
[0845] Input: Image data
[0846] Output: Text information (text data) within the image
[0847] Step 4:
[0848] The server combines the text information and the question to generate text data.
[0849] The server combines the character information extracted by OCR with the user's question text to generate a single text file. This text file contains the problem statement and the user's question. For example, it might take the form of: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0850] Input: Text information (text data) and question text
[0851] Output: Combined text data
[0852] Step 5:
[0853] The server sends text data to an AI model that generates an answer.
[0854] The server sends the combined text data to a generative artificial intelligence model. This model uses natural language processing techniques to generate an appropriate response. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0855] Input: Combined text data
[0856] Output: Text response
[0857] Step 6:
[0858] The server uses an emotion engine to recognize the user's emotions and adjust the generated response accordingly.
[0859] The server is equipped with an emotion engine that analyzes the user's input questions and conversations to recognize their emotional state. Based on the recognized emotion, the generated response is appropriately adjusted. For example, if the server detects that the user is feeling stressed, it will add encouraging words such as, "It's okay, you can do it!"
[0860] Input: Text question and generated text answer
[0861] Output: Adjusted text response
[0862] Step 7:
[0863] The server sends the response to the terminal, and the terminal displays it to the user.
[0864] The server sends the final answer to the terminal. The terminal displays the received answer on the user interface. The user checks the displayed answer and uses it to aid in learning. For example, the displayed message might be: "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Don't worry, you can do it!"
[0865] Input: Adjusted text response
[0866] Output: Answer displayed on the terminal
[0867] (Application Example 2)
[0868] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0869] In today's educational environment, students often face numerous problems and questions. However, systems that allow them to immediately resolve their doubts during self-study or in class are limited. Furthermore, there are very few systems that understand students' emotional states and provide learning support accordingly. As a result, students' learning efficiency may decrease, potentially leading to a decline in their motivation to learn.
[0870] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving posted images, means for extracting character information from the received images using optical character recognition, means for combining the extracted character information with the input question, means for interpreting the combined text using natural language processing and generating an answer, means for recognizing the user's emotions and adjusting the generated answer according to the recognized emotions, and means for transmitting the generated answer and playing it back as audio. As a result, students can solve problems they are learning in real time and receive detailed learning support tailored to their emotional state.
[0871] "Means for receiving posted images" refers to a function that allows the server to receive image data taken by users using their smart devices.
[0872] "Means for extracting character information from a received image using optical character recognition" refers to a function that extracts character and numerical information contained within an image as text data using optical character recognition technology.
[0873] "Means for combining extracted text information and entered questions" refers to a function that integrates text information extracted from an image with questions submitted by the user into a single text data file.
[0874] "Means for interpreting combined text using natural language processing and generating answers" refers to a function that analyzes combined text using a generative artificial intelligence model and creates an appropriate answer based on that analysis.
[0875] "Means for recognizing user emotions and adjusting generated responses accordingly" refers to a function that analyzes the user's emotional state from their text or voice tone and makes changes or adjustments to the generated responses based on the results.
[0876] "A means of sending generated answers and playing them back as audio" refers to a function that sends answers from a server to the user's terminal and outputs those answers as audio using speech synthesis technology.
[0877] An "optical character recognition library" is a software library used to extract character information from images using optical character recognition technology.
[0878] A "generative artificial intelligence model" is an AI model that uses natural language processing to analyze text data and generate appropriate responses or information.
[0879] The system that realizes this invention is composed of multiple components, including smart glasses, a server, and speech synthesis technology. Specific embodiments are described below.
[0880] First, the user uses smart glasses to capture images of the problems they are studying. The user also submits questions to the system using voice input. Specifically, they take pictures with the smart glasses' built-in camera and record their questions with the microphone.
[0881] Next, the smart glasses send the captured image and audio data to the server as a single data package. This transmission is done via an HTTP request.
[0882] The server uses optical character recognition (OCR) technology to process the received image data. Specifically, the Python library pytesseract is used for OCR to extract text information from the image. At this stage, equations such as "x^2 + 2x + 1 = 0" are obtained in text format.
[0883] The extracted character information and the text data converted from the user's voice input are combined and organized into a single text file. For example, it might take the form of: "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0884] Next, the server sends this combined text data to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text data. The generated response might be something like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0885] Furthermore, this invention also incorporates an emotion engine. The server analyzes the user's voice data and uses the emotion engine to recognize the user's emotional state. For example, if the tone and content of the user's questions indicate that they are stressed, the server can add words of encouragement to the generated response. "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Keep going!"
[0886] Finally, the generated response is sent from the server to the user's smart glasses and played back as speech using speech synthesis technology. For speech synthesis, the gTTS (Google Text-to-Speech) library is used, for example.
[0887] The following shows an example of input to a generative AI model.
[0888] "Problem statement: x^2 + 2x + 1 = 0
[0889] Question: How do I solve this equation?
[0890] Emotion: Confusion.”
[0891] In this way, users can quickly and appropriately resolve learning problems and receive personalized learning support that is tailored to their emotions.
[0892] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0893] Step 1:
[0894] The user takes a picture of the problem they are studying with the camera on their smart glasses and uses the smart glasses' microphone to voice-input the question. Image data and audio data are generated from the input. The image data is treated as the user's visual information, and the audio data is treated as voice input information.
[0895] Step 2:
[0896] The smart glasses send captured image and audio data to the server using HTTP requests. The HTTP request module is used to transfer this data to the server as a single data package. The output is the data package received by the server.
[0897] Step 3:
[0898] The server analyzes the received image data using the Python pytesseract library and converts the character information within the image into text data. The input is image data, which is processed using optical character recognition technology, and the output is text data. For example, the text "x^2 + 2x + 1 = 0" is obtained.
[0899] Step 4:
[0900] The server uses speech recognition technology to convert speech data into text. This process uses speech recognition software, with speech data as input and the user's question text as output. For example, it might produce the text, "Please tell me how to solve this equation."
[0901] Step 5:
[0902] The server combines the extracted text data with the question text converted by speech recognition to generate a prompt. The input is the text "x^2 + 2x + 1 = 0" and the question text "Please tell me how to solve this equation." The output is the combined text "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[0903] Step 6:
[0904] The server inputs the generated prompt text into the emotion engine, which analyzes the user's emotional state. The emotion engine estimates the emotion from the tone and content of the input text. For example, it can detect that the user is confused. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: How do I solve this equation?" and the output is "Emotion: Confused".
[0905] Step 7:
[0906] The server combines the text and emotional state and inputs it into a generative artificial intelligence model to generate an appropriate response. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation. Emotion: Confused", and the output is the response text "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1. Good luck."
[0907] Step 8:
[0908] The server sends the generated response to the user's smart glasses. The response data is transferred to the smart glasses using an HTTP response. The input is the generated response text, and the output is the response text received by the smart glasses.
[0909] Step 9:
[0910] The smart glasses play back the received response text as speech using speech synthesis technology. The gTTS (Google Text-to-Speech) library is used for speech synthesis. The input is the response text, and the output is the speech played back to the user.
[0911] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0912] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0913] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0914] [Fourth Embodiment]
[0915] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0916] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0917] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0918] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0919] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0920] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0921] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0922] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0923] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0924] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0925] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0926] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0927] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0928] This invention relates to a system for quickly and accurately resolving learning problems using AI. The following describes specific embodiments of this invention.
[0929] First, the user takes a picture of the problem they are studying with a smartphone or other device and sends the image to the system. This is the "submitted image." At the same time, the user also enters a question about the problem in text format. For example, a question like, "Please tell me how to solve this problem."
[0930] Next, the terminal sends the received image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the Python pytesseract library to extract text information such as mathematical formulas and sentences from the image as text data.
[0931] The extracted text is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "x^2 + 2x + 1 = 0" and the user's question is "How do I solve this equation?", the server combines them to generate the following text.
[0932] Problem statement: x^2 + 2x + 1 = 0
[0933] Question: How do I solve this equation?
[0934] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[0935] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0936] The generated answers are temporarily stored by the server and then sent to the user's device. The device receives the response from the server and displays the answer on the user interface. The user reviews this answer and uses it to aid in learning.
[0937] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of this equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image and generates a solution using a generative artificial intelligence model. As a result, the user receives the answer "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[0938] As described above, the system of the present invention combines optical character recognition technology and a generative artificial intelligence model to enable users to quickly and accurately solve problems they are learning. This system provides learning support that is more immediate and cost-effective than conventional learning support methods.
[0939] The following describes the processing flow.
[0940] Step 1:
[0941] Users take pictures of problems or assignments they are studying using a smartphone or other device. At the same time, users also input questions about the problem in text format. For example, they might input a question like, "Please explain how to solve this problem."
[0942] Step 2:
[0943] The terminal receives the captured image and the entered question, and sends them to the server as a single data package. An HTTP request is used for this process.
[0944] Step 3:
[0945] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[0946] Step 4:
[0947] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data.
[0948] Step 5:
[0949] The server temporarily stores the extracted text information. Let's assume the extracted text is "x^2 + 2x + 1 = 0".
[0950] Step 6:
[0951] The server combines the extracted text information with the user's question text. For example, it may be formatted as follows:
[0952] Problem statement: x^2 + 2x + 1 = 0
[0953] Question: How do I solve this equation?
[0954] Step 7:
[0955] The server sends the combined text to a generative artificial intelligence model. The generative AI model analyzes the combined text and generates an appropriate response.
[0956] Step 8:
[0957] A generative artificial intelligence model (e.g., GPT-3.5) generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1."
[0958] Step 9:
[0959] The server receives the generated response and sends it to the terminal. The HTTP protocol is used again for communication.
[0960] Step 10:
[0961] The terminal receives an HTTP response from the server and extracts the answer data. Then, it displays the answer on the user interface.
[0962] Step 11:
[0963] The user checks the answer displayed on their device. This allows the user to quickly and accurately obtain a solution to the problem.
[0964] The above steps provide answers to user questions. This system becomes a powerful tool for users to solve problems they are learning in real time.
[0965] (Example 1)
[0966] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0967] Conventional learning support systems have the challenge of not being able to quickly and accurately solve problems that users encounter during their studies. In particular, extracting problems from images containing text information, interpreting them appropriately, and providing immediate solutions is difficult, which reduces the user's learning efficiency. Furthermore, if an appropriate solution cannot be obtained, users are forced to interrupt their learning.
[0968] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0969] In this invention, the server includes means for receiving an image taken by the user, means for extracting character information from the received image using optical character recognition, means for combining the extracted character information with an input question, means for interpreting the combined text using natural language processing based on a generative artificial intelligence model and generating an answer, means for sending the generated answer to the user's terminal, and means for displaying the answer on the user's terminal. This makes it possible to quickly and accurately solve problems that the user faces during learning.
[0970] A "user" refers to an individual who uses this system to try to solve a problem they are learning.
[0971] "Device" refers to electronic devices such as smartphones and tablets used by users.
[0972] "Captured images" refers to image data containing the learning problems acquired by the user using the camera function of their device.
[0973] "Means of receiving" refers to the communication interface used to transfer images and questions taken by the user to the server.
[0974] "Optical character recognition" refers to a technology that analyzes character information within an image and extracts it as text data.
[0975] "Means of extraction" refers to software and hardware used to analyze text information within a received image and obtain it as text data.
[0976] "Method of combining" refers to the process of combining extracted character information and user-entered questions into a single text data file.
[0977] A "generative artificial intelligence model" refers to an algorithm or model that performs natural language processing based on given data to generate appropriate responses.
[0978] "Natural language processing" refers to the technology of processing human language using computers, and in this context, it refers to the technology of generating appropriate answers based on questions.
[0979] "Means of transmission" refers to the communication interface used to transfer the generated response to the user's device.
[0980] "Means of display" refers to a user interface used to visually show the answers obtained on the device.
[0981] This invention is a system for quickly and accurately resolving problems that users encounter during learning. The specific configuration and operation of the system are described as embodiments for carrying out this invention.
[0982] The user takes a picture of the problem they are studying with their device's camera and obtains the image data. For example, if the problem is written in their notebook as the equation "x^2 + 2x + 1 = 0", they would take a picture of that notebook. Then, the user uses the application on their device to type the question in text format, such as "Please tell me how to solve this equation."
[0983] The device sends the captured image and text question to the server as an HTTP request. The server receives this HTTP request and then processes the image data.
[0984] The server uses the Python library pytesseract to extract text information from the image using optical character recognition (OCR) technology. The OCR process extracts the text data "x^2 + 2x + 1 = 0".
[0985] The extracted text data and the questions entered by the user are combined to generate text like the following.
[0986] Problem statement: x^2 + 2x + 1 = 0
[0987] Question: How do I solve this equation?
[0988] Next, the server sends the generated text to a generative artificial intelligence model (for example, GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates an appropriate response. For example, the following response may be generated.
[0989] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[0990] The generated response is temporarily stored on the server. The server then sends the generated response to the user's device as an HTTP response.
[0991] The device displays the received responses on the user interface. The user reviews these responses to aid their learning.
[0992] Thus, the system of the present invention enables users to quickly and accurately solve learning problems they face. By combining optical character recognition technology with a generative artificial intelligence model, the immediacy and accuracy are improved compared to conventional learning support methods, thereby enhancing the user's learning efficiency.
[0993] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0994] Step 1:
[0995] The user takes a picture of the problem they are studying using their device's camera. For example, they might take a picture of a notebook page with the equation "x^2 + 2x + 1 = 0" written on it. Simultaneously, they use the application to input a question. For example, they might input "Please tell me how to solve this equation." The input here consists of image data and text data. The output is an image file and a text question generated by the user's actions.
[0996] Step 2:
[0997] The terminal packages the image data and text data acquired by the user as an HTTP request and sends it to the server. The input consists of the image acquired by the user and the text question entered; this data is processed for transmission as an HTTP request to the server, and the output is the data sent to the server.
[0998] Step 3:
[0999] The server receives an HTTP request sent from the terminal. The input consists of image and text data sent from the terminal. The server uses the Python pytesseract library to analyze the image data and extracts text information from the image using optical character recognition (OCR) technology. Specifically, the mathematical expression "x^2 + 2x + 1 = 0" is extracted as text. This process outputs the input image as text data.
[1000] Step 4:
[1001] The server combines the text data extracted by OCR processing with the question text data entered by the user. For example, it generates a prompt message like the following:
[1002] Problem statement: x^2 + 2x + 1 = 0
[1003] Question: How do I solve this equation?
[1004] The input consists of extracted character information and question text, and the output is generated through a data merging process called prompt text generation.
[1005] Step 5:
[1006] The server sends the generated prompt to a generative artificial intelligence model, which then uses natural language processing (NLP) techniques to generate an appropriate response. For example, the following response might be generated:
[1007] This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1.
[1008] The input here is a generated prompt sentence, and the output is the response text obtained from a generative artificial intelligence model.
[1009] Step 6:
[1010] The server temporarily stores the generated response and then sends it to the user's terminal as an HTTP response. The input is the generated response text, which is then sent to the terminal as an HTTP response. The output is the data sent to the terminal.
[1011] Step 7:
[1012] The terminal analyzes the response received from the server and displays it on the user interface. The input is the response text from the server, and the output is the visual information displayed on the user interface. The user looks at this to aid in learning.
[1013] Through the above series of processing steps, users can quickly and accurately solve problems they encounter during their learning process.
[1014] (Application Example 1)
[1015] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1016] Machine maintenance in modern factories is complex and requires quick and accurate solutions. However, many workplaces lack the necessary specialized technicians, often resulting in inefficient maintenance. To address this problem, a system is needed that accurately identifies machine problems and provides appropriate repair procedures.
[1017] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1018] In this invention, the server includes means for receiving posted images, means for extracting text information from the received images using optical character recognition, means for combining the extracted text information with an entered question, means for interpreting the combined text using natural language processing and generating an answer, means for transmitting the generated answer, and means for providing identification and repair procedures for damaged machine parts to support machine maintenance work in the factory. This enables factory workers to quickly and accurately identify machine problems and obtain appropriate repair procedures using smartphones or tablets.
[1019] "Means for receiving images" refers to a device or software that has the function of receiving and storing image data transmitted by a user.
[1020] "Means of extraction using optical character recognition" refers to devices or software that use technology to analyze character information within an image and convert it into text data.
[1021] "Means of combining" refers to a function that integrates extracted text information and user-entered questions into a single text file.
[1022] "Means of interpretation using natural language processing" refers to techniques that analyze combined texts and generate answers in an easily understandable format.
[1023] "Means for generating answers" refers to a device or software that has the function of providing appropriate answers from information analyzed by natural language processing.
[1024] "Means of sending responses" refers to a device or software for sending generated responses to a user's device.
[1025] "Means of supporting machine maintenance work" refers to devices or software that have the function of identifying problems with machinery and providing repair methods and procedures.
[1026] "Identifying damaged machine parts" refers to the process of recognizing damaged parts from captured images and extracting their details.
[1027] "Means of providing repair procedures" refers to a device or software that has the function of notifying the user of the appropriate repair method or process for an identified problem.
[1028] This invention relates to a system for supporting machine maintenance work in a factory. The following describes specific embodiments for carrying out this invention.
[1029] First, the user takes a picture of the problematic machine in the factory using a smartphone or tablet and sends the image to the system. At the same time, the user also enters a question about the machine in text format, such as, "How do I replace this part?"
[1030] Next, the device sends the captured image and the user's question to the server as an HTTP request. The server receives the submitted image and extracts the text information within it using optical character recognition (OCR) technology. Specifically, the server uses the pytesseract library to extract text data such as character information and part names from the image.
[1031] The extracted text data is temporarily stored on the server. The server then combines the extracted information with the question entered by the user. For example, if the text extracted from the image is "bolt" and "gear," and the user's question is "How do I replace this part?", the server combines these to generate the following text.
[1032] Problem statement: How to replace bolts and gears
[1033] Question: How do I replace this part?
[1034] Next, the server sends the generated text to a generative artificial intelligence model (e.g., GPT-3.5), which interprets the text using natural language processing (NLP) techniques and generates a response. For example, the generative artificial intelligence model might generate a response like the following:
[1035] To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt.
[1036] The generated response is temporarily stored by the server and then sent to the user's terminal. The terminal receives the response from the server and displays the response on the user interface. The user reviews this response and uses it as a reference for maintenance work.
[1037] As a concrete example, the process involves a user sending an image of a problematic machine part to the system along with the question, "How do I replace this bolt and gear?" The server extracts information about the bolt and gear from the image and generates a replacement method using a generative artificial intelligence model. As a result, the user receives the answer, "To replace this part, first remove the old bolt and replace it with a new one. Then, properly align the gear and secure it with the bolt."
[1038] This system combines optical character recognition technology with generative artificial intelligence models to enable factory workers to perform machine maintenance quickly and accurately. This improves on-site work efficiency and allows for proper maintenance even in situations where there is a shortage of specialized technicians.
[1039] Example of a prompt:
[1040] Problem statement: How to replace bolts and gears
[1041] Question: How do I replace this part?
[1042] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1043] Step 1:
[1044] Users take pictures of problematic machinery in the factory using their smartphones or tablets and send the images to the system. At the same time, they also enter questions about the machinery in text format. This process sends both the image and text data from the device to the server.
[1045] Step 2:
[1046] The device sends the captured image and the entered question to the server as an HTTP request. Specifically, the device sends the image file and text as a POST request to the specified endpoint on the server. The input consists of image data and text data, and the output is the request data sent to the server.
[1047] Step 3:
[1048] The server receives the transmitted image. The received image data is temporarily stored on the server. In this step, the input is the image data, and the output is the stored image data.
[1049] Step 4:
[1050] The server uses Optical Character Recognition (OCR) technology to extract text information from the received image. The server uses the pytesseract library to obtain the text information within the image. The input is image data, and the output is the extracted text data.
[1051] Step 5:
[1052] The server combines the extracted text data with the question entered by the user. This is done to generate the problem statement and question as a single text file. The input is the extracted text data and the question data, and the output is the combined text data.
[1053] Step 6:
[1054] The server sends the combined text data to a generative artificial intelligence model (e.g., GPT-3.5), which interprets it using natural language processing (NLP) techniques and generates a response. In this step, prompts are sent to the generative AI model via an API to generate the response text. The input is the combined text data, and the output is the generated response text.
[1055] Step 7:
[1056] The generated response is temporarily stored on the server and then sent to the user's device. Specifically, the server returns a response containing the generated response to the user's device. The input is the generated response text, and the output is the response data sent to the user's device.
[1057] Step 8:
[1058] The terminal receives a response from the server and displays the answer on the user interface. The user reviews this answer and uses it as a reference for machine maintenance work. The input is the response data from the server, and the output is the displayed answer data.
[1059] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1060] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[1061] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[1062] Next, the device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and the user's question contained within the request.
[1063] First, the server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[1064] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1065] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1066] Furthermore, this invention includes an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when a user enters text for a question, the emotional tone is analyzed from the content and style of the text.
[1067] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, it can add more detailed explanations if the user is confused, or add words of encouragement if the server detects that the user is stressed.
[1068] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. The user reviews the displayed answers and uses them to aid in learning.
[1069] As a concrete example, let's consider the process when a user wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends the question "Please tell me how to solve this equation" to the system. The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user receives an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1070] Through the steps described above, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[1071] The following describes the processing flow.
[1072] Step 1:
[1073] Users use their smartphones or tablets to take pictures of problems or assignments they are studying. At the same time, users also input questions about the problems in text format. For example, they might input a question like, "Please explain how to solve this equation."
[1074] Step 2:
[1075] The device combines the captured image and the user's input into a single data package. This data package is then sent to the server as an HTTP request.
[1076] Step 3:
[1077] The server receives an HTTP request sent from the terminal. The request contains image data and the user's question text.
[1078] Step 4:
[1079] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, it extracts the equation "x^2 + 2x + 1 = 0" from the image.
[1080] Step 5:
[1081] The server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. The generated text will take the form of, for example, "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1082] Step 6:
[1083] The server sends the combined text data to the sentiment engine. The sentiment engine analyzes the content and format of the questions entered by the user and evaluates the user's emotions. For example, it recognizes emotions such as the user being nervous, confused, or irritated.
[1084] Step 7:
[1085] Based on the emotion data recognized by the emotion engine, the server sends the combined text to a generative artificial intelligence model. The generative AI model (e.g., GPT-3.5) analyzes the combined text and generates an appropriate response.
[1086] Step 8:
[1087] A generative artificial intelligence model generates an answer based on the combined text. For example, it might generate an answer like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1088] Step 9:
[1089] The server receives the generated response and adjusts it. Based on the user's emotions recognized by the emotion engine, the server modifies how the response is expressed. For example, if the user is confused, it may add a more detailed explanation or use gentler language.
[1090] Step 10:
[1091] The server sends the final response data to the terminal.
[1092] Step 11:
[1093] The terminal receives an HTTP response from the server and displays the response data on the user interface.
[1094] Step 12:
[1095] Users review the answers displayed on their device to aid their learning. Because these answers are adjusted by an emotion engine, users receive effective learning support that takes their emotional state into account.
[1096] Through the steps described above, users receive immediate and appropriate answers to their questions. This system allows users to quickly and accurately resolve learning problems, providing learning support that is more immediate and cost-effective than traditional methods.
[1097] (Example 2)
[1098] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1099] Conventional learning support systems can interpret questions from image data and generate answers, but they cannot take into account the user's emotions, which means some users may not receive appropriate support. For example, if a user is emotionally stressed, the answer provided may reduce their learning efficiency. In such situations, appropriate feedback that takes the user's emotions into account is necessary.
[1100] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1101] In this invention, the server includes means for receiving posted images, means for extracting textual information from the received images using optical character recognition, means for combining the extracted textual information with the entered question, means for interpreting the combined text using natural language processing and generating an answer, means for analyzing the user's entered question and dialogue content and recognizing their emotional state, means for adjusting the generated answer based on the recognized emotion, and means for transmitting the generated answer. This makes it possible not only to provide appropriate answers according to the user's learning progress but also to provide feedback according to the user's emotional state.
[1102] "Submitted images" refer to image data containing learning problems that users capture using their devices and send to the server.
[1103] "Textual information" refers to text data contained within the posted image, specifically including strings of characters such as equations and mathematical problems.
[1104] Optical character recognition (OCR) is a technology that extracts character information from images as text data, and is sometimes abbreviated as OCR.
[1105] "Means of receiving" refers to the interface or function that a server uses to receive data sent from a client, such as an HTTP request.
[1106] "Means of extraction" refers to software or libraries used to extract textual information from images. For example, optical character recognition libraries fall into this category.
[1107] A "question" is text data that a user enters regarding a learning problem they want to solve.
[1108] "Means of combining" refers to a process or function for integrating extracted character information and user-entered questions into a single text data file.
[1109] Natural language processing (NLP) is the technology that enables computers to understand and generate human language. This processing includes text analysis and generation.
[1110] A "generative artificial intelligence model" is an artificial intelligence technology that generates natural language responses based on provided input data, and for example, it utilizes natural language processing technology.
[1111] "Means of recognizing emotional states" refers to algorithms and engines that analyze user question text and dialogue content to determine the user's emotions.
[1112] "Means of adjustment" refers to a process or function for appropriately modifying or supplementing responses generated based on perceived emotional states.
[1113] "Means of transmission" refers to an interface or function for sending generated response data from the server to the client.
[1114] This invention combines a system that extracts learning problems based on submitted images and generates appropriate answers to user questions with an emotion engine that recognizes the user's emotions. The following describes specific embodiments for implementing this invention.
[1115] First, the user takes a picture of the problem they are studying using a device such as a smartphone or tablet. At this time, the user can also input a question about the problem in text. The question the user inputs might be something like, "Please tell me how to solve this equation."
[1116] The device sends the captured image and the entered question to the server as a single data package. This transmission is done via an HTTP request. The server receives the HTTP request sent from the device and analyzes the image data and user question contained within the request.
[1117] The server analyzes the received image data using optical character recognition (OCR) technology. Specifically, it uses the Python pytesseract library to extract text information from the image as text data. For example, an equation like "x^2 + 2x + 1 = 0" might be extracted from the image.
[1118] Next, the server temporarily stores the extracted text information. It then combines the extracted text information with the user's question text to generate a single text file. This combined text will look like this: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1119] The server then sends this combined text to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1120] Furthermore, the system of the present invention is equipped with an emotion engine that recognizes the user's emotions. The emotion engine is a technology that analyzes the questions and dialogue content entered by the user and recognizes the user's emotional state. For example, when the user enters the text of a question, it analyzes the emotional tone from the content and style of the text.
[1121] The server can adjust the generated responses based on the emotions recognized by the emotion engine. For example, if it detects that the user is feeling stressed due to a question, it can add words of encouragement.
[1122] The generated answers are sent from the server to the user's terminal. The terminal receives the response from the server and displays the answer data on the user interface. As another example, if the system recognizes that the user has a question, it can add a detailed explanation. The user reviews the displayed answers and uses them to aid in their learning.
[1123] Specific example
[1124] This illustrates the process for a user who wants to know how to solve the problem "x^2 + 2x + 1 = 0". The user takes a picture of the equation and sends it to the system with the question, "Please tell me how to solve this equation." The server extracts the equation from the image, analyzes the user's emotions using an emotion engine, and then generates a solution using a generative artificial intelligence model. As a result, the user can receive an answer such as, "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1125] Example of a prompt
[1126] Problem statement: x^2 + 2x + 1 = 0
[1127] Question: How do I solve this equation?
[1128] In the above configuration, the system of the present invention helps users quickly and appropriately solve problems they are learning. By combining optical character recognition technology, a generative artificial intelligence model, and an emotion engine, this system provides superior immediacy and cost-effectiveness compared to conventional learning support methods.
[1129] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1130] Step 1:
[1131] The user takes a picture of the problem and enters the question.
[1132] Users take photos of the problems they are learning using a device such as a smartphone or tablet. For example, they might take a picture of a piece of paper with "x^2 + 2x + 1 = 0" written on it. At the same time, they input a question into the device, such as "Please tell me how to solve this equation." The input is done through a text box.
[1133] Input: Captured image and text question
[1134] Output: Combine images and text questions into a single data package.
[1135] Step 2:
[1136] The terminal sends the data package to the server.
[1137] The device generates a single data package containing the captured image and the entered question. This data package is sent to the server via an HTTP request. The device packages the image data and text question as a payload and sends it over the network.
[1138] Input: Data package (including images and text questions)
[1139] Output: Data package sent to the server
[1140] Step 3:
[1141] The server analyzes the image data using OCR technology.
[1142] The server extracts image data from the received data package. Optical Character Recognition (OCR) technology is used to extract text information from the image as text data. The specific software used in this process is the Python library pytesseract. For example, the string "x^2 + 2x + 1 = 0" is extracted from the image.
[1143] Input: Image data
[1144] Output: Text information (text data) within the image
[1145] Step 4:
[1146] The server combines the text information and the question to generate text data.
[1147] The server combines the character information extracted by OCR with the user's question text to generate a single text file. This text file contains the problem statement and the user's question. For example, it might take the form of: "Problem statement: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1148] Input: Text information (text data) and question text
[1149] Output: Combined text data
[1150] Step 5:
[1151] The server sends text data to an AI model that generates an answer.
[1152] The server sends the combined text data to a generative artificial intelligence model. This model uses natural language processing techniques to generate an appropriate response. For example, it might generate a response like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1153] Input: Combined text data
[1154] Output: Text response
[1155] Step 6:
[1156] The server uses an emotion engine to recognize the user's emotions and adjust the generated response accordingly.
[1157] The server is equipped with an emotion engine that analyzes the user's input questions and conversations to recognize their emotional state. Based on the recognized emotion, the generated response is appropriately adjusted. For example, if the server detects that the user is feeling stressed, it will add encouraging words such as, "It's okay, you can do it!"
[1158] Input: Text question and generated text answer
[1159] Output: Adjusted text response
[1160] Step 7:
[1161] The server sends the response to the terminal, and the terminal displays it to the user.
[1162] The server sends the final answer to the terminal. The terminal displays the received answer on the user interface. The user checks the displayed answer and uses it to aid in learning. For example, the displayed message might be: "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Don't worry, you can do it!"
[1163] Input: Adjusted text response
[1164] Output: Answer displayed on the terminal
[1165] (Application Example 2)
[1166] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1167] In today's educational environment, students often face numerous problems and questions. However, systems that allow them to immediately resolve their doubts during self-study or in class are limited. Furthermore, there are very few systems that understand students' emotional states and provide learning support accordingly. As a result, students' learning efficiency may decrease, potentially leading to a decline in their motivation to learn.
[1168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving posted images, means for extracting character information from the received images using optical character recognition, means for combining the extracted character information with the input question, means for interpreting the combined text using natural language processing and generating an answer, means for recognizing the user's emotions and adjusting the generated answer according to the recognized emotions, and means for transmitting the generated answer and playing it back as audio. As a result, students can solve problems they are learning in real time and receive detailed learning support tailored to their emotional state.
[1169] "Means for receiving posted images" refers to a function that allows the server to receive image data taken by users using their smart devices.
[1170] "Means for extracting character information from a received image using optical character recognition" refers to a function that extracts character and numerical information contained within an image as text data using optical character recognition technology.
[1171] "Means for combining extracted text information and entered questions" refers to a function that integrates text information extracted from an image with questions submitted by the user into a single text data file.
[1172] "Means for interpreting combined text using natural language processing and generating answers" refers to a function that analyzes combined text using a generative artificial intelligence model and creates an appropriate answer based on that analysis.
[1173] "Means for recognizing user emotions and adjusting generated responses accordingly" refers to a function that analyzes the user's emotional state from their text or voice tone and makes changes or adjustments to the generated responses based on the results.
[1174] "A means of sending generated answers and playing them back as audio" refers to a function that sends answers from a server to the user's terminal and outputs those answers as audio using speech synthesis technology.
[1175] An "optical character recognition library" is a software library used to extract character information from images using optical character recognition technology.
[1176] A "generative artificial intelligence model" is an AI model that uses natural language processing to analyze text data and generate appropriate responses or information.
[1177] The system that realizes this invention is composed of multiple components, including smart glasses, a server, and speech synthesis technology. Specific embodiments are described below.
[1178] First, the user uses smart glasses to capture images of the problems they are studying. The user also submits questions to the system using voice input. Specifically, they take pictures with the smart glasses' built-in camera and record their questions with the microphone.
[1179] Next, the smart glasses send the captured image and audio data to the server as a single data package. This transmission is done via an HTTP request.
[1180] The server uses optical character recognition (OCR) technology to process the received image data. Specifically, the Python library pytesseract is used for OCR to extract text information from the image. At this stage, equations such as "x^2 + 2x + 1 = 0" are obtained in text format.
[1181] The extracted character information and the text data converted from the user's voice input are combined and organized into a single text file. For example, it might take the form of: "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1182] Next, the server sends this combined text data to a generative artificial intelligence model. The generative AI model uses natural language processing (NLP) techniques to generate an appropriate response based on the combined text data. The generated response might be something like, "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1."
[1183] Furthermore, this invention also incorporates an emotion engine. The server analyzes the user's voice data and uses the emotion engine to recognize the user's emotional state. For example, if the tone and content of the user's questions indicate that they are stressed, the server can add words of encouragement to the generated response. "This equation can be solved as follows: First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0. Therefore, x = -1. Keep going!"
[1184] Finally, the generated response is sent from the server to the user's smart glasses and played back as speech using speech synthesis technology. For speech synthesis, the gTTS (Google Text-to-Speech) library is used, for example.
[1185] The following shows an example of input to a generative AI model.
[1186] "Problem statement: x^2 + 2x + 1 = 0
[1187] Question: How do I solve this equation?
[1188] Emotion: Confusion.”
[1189] In this way, users can quickly and appropriately resolve learning problems and receive personalized learning support that is tailored to their emotions.
[1190] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1191] Step 1:
[1192] The user takes a picture of the problem they are studying with the camera on their smart glasses and uses the smart glasses' microphone to voice-input the question. Image data and audio data are generated from the input. The image data is treated as the user's visual information, and the audio data is treated as voice input information.
[1193] Step 2:
[1194] The smart glasses send captured image and audio data to the server using HTTP requests. The HTTP request module is used to transfer this data to the server as a single data package. The output is the data package received by the server.
[1195] Step 3:
[1196] The server analyzes the received image data using the Python pytesseract library and converts the character information within the image into text data. The input is image data, which is processed using optical character recognition technology, and the output is text data. For example, the text "x^2 + 2x + 1 = 0" is obtained.
[1197] Step 4:
[1198] The server uses speech recognition technology to convert speech data into text. This process uses speech recognition software, with speech data as input and the user's question text as output. For example, it might produce the text, "Please tell me how to solve this equation."
[1199] Step 5:
[1200] The server combines the extracted text data with the question text converted by speech recognition to generate a prompt. The input is the text "x^2 + 2x + 1 = 0" and the question text "Please tell me how to solve this equation." The output is the combined text "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation."
[1201] Step 6:
[1202] The server inputs the generated prompt text into the emotion engine, which analyzes the user's emotional state. The emotion engine estimates the emotion from the tone and content of the input text. For example, it can detect that the user is confused. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: How do I solve this equation?" and the output is "Emotion: Confused".
[1203] Step 7:
[1204] The server combines the text and emotional state and inputs it into a generative artificial intelligence model to generate an appropriate response. The input is "Problem: x^2 + 2x + 1 = 0\nQuestion: Please tell me how to solve this equation. Emotion: Confused", and the output is the response text "This equation can be solved as follows. First, factorize both sides: x^2 + 2x + 1 = (x + 1)^2 = 0 Therefore, x = -1. Good luck."
[1205] Step 8:
[1206] The server sends the generated response to the user's smart glasses. The response data is transferred to the smart glasses using an HTTP response. The input is the generated response text, and the output is the response text received by the smart glasses.
[1207] Step 9:
[1208] The smart glasses play back the received response text as speech using speech synthesis technology. The gTTS (Google Text-to-Speech) library is used for speech synthesis. The input is the response text, and the output is the speech played back to the user.
[1209] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1210] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1211] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1212] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1213] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1214] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1215] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1216] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1217] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1218] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1219] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1220] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1221] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1222] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1223] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1224] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1225] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1226] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1227] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1228] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1229] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1230] The following is further disclosed regarding the embodiments described above.
[1231] (Claim 1)
[1232] A means of receiving posted images,
[1233] A means for extracting text information from a received image using optical character recognition,
[1234] A means for combining extracted text information with the entered question,
[1235] A means for interpreting combined texts using natural language processing and generating an answer,
[1236] A means of sending the generated response,
[1237] A system that includes this.
[1238] (Claim 2)
[1239] The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
[1240] (Claim 3)
[1241] The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model.
[1242] "Example 1"
[1243] (Claim 1)
[1244] A means of receiving images taken by the user,
[1245] A means for extracting text information from a received image using optical character recognition,
[1246] A means for combining extracted text information with the entered question,
[1247] A means for interpreting combined texts using natural language processing based on a generative artificial intelligence model and generating an answer,
[1248] A means of sending the generated response to the user's device,
[1249] A means of displaying the answer on the user's device,
[1250] A system that includes this.
[1251] (Claim 2)
[1252] The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
[1253] (Claim 3)
[1254] The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model.
[1255] "Application Example 1"
[1256] (Claim 1)
[1257] A means of receiving posted images,
[1258] A means for extracting text information from a received image using optical character recognition,
[1259] A means for combining extracted text information with the entered question,
[1260] A means for interpreting combined texts using natural language processing and generating an answer,
[1261] A means of sending the generated response,
[1262] To support machine maintenance work within a factory, means of providing procedures for identifying and repairing damaged machine parts,
[1263] A system that includes this.
[1264] (Claim 2)
[1265] The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
[1266] (Claim 3)
[1267] The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model.
[1268] "Example 2 of combining an emotion engine"
[1269] (Claim 1)
[1270] A means of receiving posted images,
[1271] A means for extracting text information from a received image using optical character recognition,
[1272] A means for combining extracted text information with the entered question,
[1273] A means for interpreting combined texts using natural language processing and generating an answer,
[1274] A means of analyzing the questions and conversation content entered by the user to recognize their emotional state,
[1275] Means for adjusting responses generated based on recognized emotions,
[1276] A means of sending the generated response,
[1277] A system that includes this.
[1278] (Claim 2)
[1279] The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
[1280] (Claim 3)
[1281] The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model.
[1282] "Application example 2 of combining emotional engines"
[1283] (Claim 1)
[1284] A means of receiving posted images,
[1285] A means for extracting text information from a received image using optical character recognition,
[1286] A means for combining extracted text information with the entered question,
[1287] A means for interpreting combined texts using natural language processing and generating an answer,
[1288] A means of recognizing the user's emotions and adjusting the generated response according to the recognized emotions,
[1289] A means of sending the generated response and playing it back as audio,
[1290] A system that includes this.
[1291] (Claim 2)
[1292] The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
[1293] (Claim 3)
[1294] The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model. [Explanation of symbols]
[1295] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving posted images, A means for extracting text information from a received image using optical character recognition, A means for combining extracted text information with the entered question, A means for interpreting combined texts using natural language processing and generating an answer, A means of sending the generated response, A system that includes this.
2. The system according to claim 1, wherein the optical character recognition means extracts character information from an image using an optical character recognition library.
3. The system according to claim 1, wherein the means for interpreting and generating a response using natural language processing is a generative artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A