system

An automated system using image and natural language processing technologies addresses manual correction inefficiencies in educational institutions, offering rapid and personalized feedback to enhance learning efficiency.

JP2026069175APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional educational institutions face challenges with manual, time-consuming, and inconsistent answer correction processes that delay feedback, leading to decreased learning efficiency and inadequate individualized support.

Method used

An automated system integrating image processing, character recognition, and natural language processing technologies to quickly and accurately evaluate and provide feedback on student answers, utilizing AI for rapid and consistent correction.

Benefits of technology

The system reduces grading burdens, enhances learning efficiency by providing immediate, accurate, and personalized feedback, supporting individual student improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069175000001_ABST
    Figure 2026069175000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Image processing means for processing captured images and extracting answer information, Character recognition means for converting the aforementioned answer information into text data, An evaluation means that performs corrections based on the aforementioned text data and generates evaluation information, A system including a communication means for transmitting the aforementioned evaluation information to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional educational institutions, the work of correcting answers depends on manual labor, which is time - consuming and costly, and there is also a problem that the quality varies. In addition, since the answers are not returned quickly, the time for students to receive feedback is long, leading to a decrease in learning efficiency. Furthermore, it is difficult to provide detailed feedback to individual students by a limited number of teachers or correctors. [[ID=第37]]

Means for Solving the Problems

[0005] In this invention, image processing means and character recognition means are used to extract answer information from captured images and convert it into text data. Furthermore, evaluation means utilizing natural language processing technology are used to quickly perform accurate corrections and evaluations based on the text data. This effectively solves conventional problems by providing a system that generates stable and consistent feedback using AI in a short time and transmits it to the user via communication means.

[0006] "Image processing means" refers to a device or function that performs processing to remove unnecessary information from a captured image and convert the answer information into a state in which it can be identified.

[0007] "Character recognition means" refers to a device or technology for identifying characters from image data and converting them into text data.

[0008] "Evaluation means" refers to a device or process that analyzes text data and evaluates the accuracy and logical structure of the answer.

[0009] "Communication means" refers to the means for transmitting processed evaluation information to the user, and is a function that sends and receives information via network communication. [Brief explanation of the drawing]

[0010] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6]This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0011] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0012] First, let's explain the terminology used in the following explanation.

[0013] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0014] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0015] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0016] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0017] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0018] [First Embodiment]

[0019] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0022] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0023] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0024] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0027] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0028] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0029] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0030] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0031] This invention relates to an automated proofreading system that integrates image recognition technology and natural language processing technology. This system receives images of exam answers taken by the user using a smartphone or tablet as input, and has the function of quickly and accurately proofreading and providing feedback.

[0032] When the user takes a picture of the answer sheet, the device performs image processing, optimizing the resolution and removing noise. This makes the image suitable for character recognition. Next, the device sends the processed image data to the server.

[0033] The server extracts text data from the received image using OCR (Optical Character Recognition). This text data is then fed into an AI-powered natural language processing model to evaluate the accuracy, expression, and logic of the answer. Existing correct answer databases and grammatical rules are used for evaluation, automatically identifying errors and generating optimal feedback.

[0034] The generated feedback is formatted as text and sent back to the device. The device displays this information to the user, clearly indicating which parts of the answer are correct and which parts need improvement. This entire process allows students to efficiently check their understanding and move forward with their learning.

[0035] As a concrete example, let's consider a historical question. For instance, suppose a user creates a detailed answer to the question, "Explain the characteristics of the Nara period." When the user takes a picture of this answer and sends it to the system, the AI ​​analyzes its content and provides specific feedback, such as, "The description of the influence of Buddhism is appropriate, but the description of the relocation of the capital is inaccurate." This feedback gives the user an opportunity to improve the accuracy of their descriptions and their understanding of the subject matter.

[0036] The system of this invention automates this rapid and accurate feedback process, thereby reducing the burden of grading work in educational settings and improving the quality of individualized learning.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user uses a smartphone or tablet to take a picture of the sheet of paper containing the question and its answer. The user reviews the image through the application, confirms that there are no quality issues, and then presses the submit button.

[0040] Step 2:

[0041] The device receives the captured image data and performs image processing such as improving resolution, adjusting contrast, and cropping unnecessary parts. This makes the image optimal for character recognition.

[0042] Step 3:

[0043] The terminal uploads the processed image data to the server. The data is transmitted in an encrypted state for security purposes.

[0044] Step 4:

[0045] The server receives the transmitted image data and uses an OCR (Optical Character Recognition) engine to extract text information from the image. OCR processing converts the characters in the image into text data that can be analyzed by a computer.

[0046] Step 5:

[0047] The server analyzes the text data obtained by OCR and passes it to a natural language processing model. The model evaluates the accuracy, expressiveness, and logical structure of the answer based on the text data and generates correction comments.

[0048] Step 6:

[0049] The server structures the generated ratings and feedback, editing them into a user-friendly format. The feedback presented clearly indicates which parts of the answer are correct and which parts could be improved.

[0050] Step 7:

[0051] The terminal displays feedback received from the server on the user's device. This allows the user to check the evaluation of their answers and immediately understand areas for improvement in their learning.

[0052] (Example 1)

[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0054] The conventional process of grading exam answers is time-consuming and labor-intensive, and has been particularly problematic in providing efficient feedback for individual learners. Furthermore, manual grading is subjective, making it difficult to maintain consistency in evaluation. This invention aims to solve these problems and support the improvement of individual learners' understanding by providing rapid and accurate feedback.

[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0056] In this invention, the server includes means for acquiring captured images and transmitting them to an information processing device; image processing means for optimizing the resolution and removing noise from the images; character recognition means for converting the optimized images into text data using optical character recognition technology; evaluation means for analyzing the text data based on natural language processing technology, evaluating the accuracy and logic of the answers, and generating feedback; and communication means for transmitting the generated feedback to the user via a data transmission device. This enables rapid and objective correction and feedback provision.

[0057] "Means for acquiring captured images and transmitting them to an information processing device" refers to a component for receiving image data captured by a user and transmitting it to an information processing device via a network.

[0058] "Resolution optimization" is a process that appropriately adjusts the amount of data while preserving the visual information of an image.

[0059] "Noise reduction" is the process of removing unwanted visual data and distortions from an image, making the image clearer.

[0060] "Image processing means" refers to a processing device or program that has the function of applying a specific algorithm to input image data to optimize resolution and remove noise.

[0061] "Optical character recognition technology" is a technology that recognizes characters from images and extracts them as text data.

[0062] "Character recognition means" refers to a function that uses optical character recognition technology to extract character information from image data and analyze it as electronic text.

[0063] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language.

[0064] "Evaluation means" refers to a component that analyzes text data using natural language processing technology to evaluate the accuracy and logic of the answers.

[0065] A "data transmission device" is a device or mechanism for transmitting information and feedback generated on a server to a user terminal via a network.

[0066] "Communication means" refers to the technology or device that enables the sending and receiving of data between a server and a user terminal.

[0067] This invention begins with a user taking a picture of their exam answers using a device such as a smartphone or tablet. The device receives this image and performs image processing. Image processing includes resolution optimization and noise reduction, which makes character recognition easier. Specifically, the device adjusts the image sharpness using the OpenCV library and reduces noise through a filter.

[0068] Next, the processed image is sent from the terminal to the server. The terminal encrypts the data using the HTTPS protocol to ensure secure communication with the server. The server extracts text data from the image using optical character recognition (OCR) technology such as Tesseract. The extracted text data is processed in text format and evaluated by the server's AI engine.

[0069] The server evaluates the accuracy, expression, and logic of the answers based on natural language processing (NLP) techniques. Here, for example, it judges the quality of the answers by referring to pre-entered model answers and grammar checkers. The AI ​​includes models commonly used as generative AI models.

[0070] The generated feedback is sent back from the server to the device, where the user can view it on the display screen. The device's UI is designed using the React Native framework, allowing for intuitive touch-based viewing of detailed feedback.

[0071] As a concrete example, consider a case where a user answers the question, "Explain the characteristics of the Nara period," takes a photo of their answer, and uploads it to the system. The AI ​​analyzes the content and provides feedback that the description of "the influence of Buddhism" is appropriate, but the description of "the relocation of the capital" is inaccurate. In this way, the user can review their answer and obtain useful information to deepen their understanding.

[0072] An example of a prompt would be, "I have taken a picture of an answer explaining the characteristics of the Nara period. Please evaluate the accuracy of the historical facts and the way they are expressed in this answer, and provide feedback on areas for improvement." This specific prompt allows the AI ​​to provide the analysis results the user is looking for more accurately.

[0073] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0074] Step 1:

[0075] The user takes a picture of their exam answer sheet using a device. In this step, the exam answer sheet is used as visual input, and the device activates its camera function to acquire image data. The output obtained here is a digital image file.

[0076] Step 2:

[0077] The device performs image processing on the captured image. Specifically, it uses the OpenCV library to optimize the image resolution and apply a noise reduction filter. The input here is image data captured by the user, and the output is clear image data with optimized resolution and noise removed.

[0078] Step 3:

[0079] The terminal sends the processed image data to the server. The data is encrypted using the HTTPS protocol for secure transmission. The input is optimized image data, and the output is ready to be passed to the data receiver module on the server.

[0080] Step 4:

[0081] The server processes the received image using optical character recognition (OCR) technology and extracts character data. It uses the Tesseract engine to identify characters in the image and convert them into text format. The input is the image data received by the server, and the output is the extracted character data.

[0082] Step 5:

[0083] The server analyzes the recognized character data using natural language processing (NLP) techniques. A generative AI model is used to evaluate the accuracy and logic of the answers. This analysis compares the text structure and content with a database of known answers. The input is character data, and the output is feedback data including the evaluation results.

[0084] Step 6:

[0085] The server generates feedback data and sends it to the terminal. Communication takes place through a data transmission device, and the transmission is encrypted. The input here is feedback data including evaluation results, and the output is data sent back to the user's terminal.

[0086] Step 7:

[0087] The device displays feedback received from the server in the user interface. Using the React Native framework, users can view detailed feedback through touch gestures. The input is feedback data sent from the server, and the output is information displayed in a format that the user can see.

[0088] (Application Example 1)

[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0090] There is a need to improve operational efficiency and ensure picking accuracy in logistics centers. However, manual work verification can lead to human error and time loss, compromising the accuracy of inventory management. Therefore, a system that can perform evaluation and provide feedback in real time using automated methods is necessary.

[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0092] In this invention, the server includes a visual data processing means for processing captured visual data and extracting information; a data conversion means for converting the information into text data; a data evaluation means for performing evaluations based on the text data and generating evaluation results; and a verification means for matching the data with stored record information and evaluating its accuracy. This enables rapid and accurate evaluation of work performed within the logistics center and provides appropriate feedback to workers.

[0093] "Visual data processing means" refers to a device or program for optimizing captured images and extracting necessary information.

[0094] "Data conversion means" refers to a device or program for converting information extracted by visual data processing means into text data.

[0095] "Data evaluation means" refers to a device or program that evaluates the content of information based on converted text data and generates evaluation results.

[0096] "Information and communication means" refers to a communication device or program for transmitting the evaluated results to the user's terminal.

[0097] "Verification means" refers to a device or program that compares the generated evaluation results with the recorded and stored information to guarantee their accuracy.

[0098] "Record storage information" refers to a set of reference data that is used to verify evaluation results.

[0099] "Evaluation results" refer to evaluation information generated by data evaluation methods and are feedback provided to users.

[0100] The system that implements this application is intended for efficient evaluation and feedback provision for use in logistics centers and similar environments. A smartphone or tablet is responsible for processing the images of the picking list. First, the device uses OpenCV to optimize the image resolution and remove noise. Next, it extracts text information using Tesseract OCR and sends that information to the server.

[0101] The server analyzes the received text data using a natural language processing model based on TENSORFLOW®. This analysis evaluates whether the text data matches existing record-keeping information and determines whether the picking was carried out as planned. The evaluation results are confirmed by a verification means and then sent back to the terminal via an information and communication means.

[0102] Based on feedback received on the device, users can verify the accuracy of their picking work in real time and correct their work as needed. For example, a worker might take a photo of a list that says "5 boxes of oranges," and the AI ​​compares it with inventory data to confirm that 5 boxes were actually picked. If there is a discrepancy, the user is immediately notified.

[0103] An example of a prompt message is: "Design an application that takes a picture of a picking list used in a logistics center and instantly verifies the accuracy of the inventory using OCR and natural language processing."

[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0105] Step 1:

[0106] The terminal receives an image of the picking list taken by the user. OpenCV is used to optimize the resolution and remove noise. The input is the captured image data, and the output is the processed, clean image.

[0107] Step 2:

[0108] The device uses Tesseract OCR to extract text information from a clean image. This process takes an optimized image as input and generates text data as output.

[0109] Step 3:

[0110] The terminal sends the generated text data to the server. This data serves as basic information for analysis in the next processing step.

[0111] Step 4:

[0112] The server analyzes the received text data using a natural language processing model based on TensorFlow. The input is text data, and the output is an evaluation result assessing the accuracy of the picking. The data analysis verifies whether the text matches pre-recorded storage information.

[0113] Step 5:

[0114] The server cross-references the information to verify the accuracy of the picking and generates a final evaluation result. In this verification step, existing record-keeping information is used as input data.

[0115] Step 6:

[0116] The server sends the verified evaluation results to the terminal. The terminal displays these results to the user and provides feedback on the accuracy of the picking.

[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0118] This invention relates to a system that provides more personalized feedback to users by combining an emotion engine with a correction system. This system takes a picture of the paper on which the user has submitted their answers and uses that image to perform efficient analysis of the answers and evaluate their emotions.

[0119] Users use their smartphones or tablets to take photos of their completed answer sheets. The device receives these images and performs image processing such as adjusting the resolution and cropping. The processed images are then sent from the device to the server. The server uses OCR technology to extract text information from the images and formats it as text data.

[0120] Next, the server applies evaluation methods for automated correction. Utilizing natural language processing technology, it evaluates the logic and syntactic accuracy of the answers, generating specific corrections and feedback. Based on this data, an emotion engine is activated, analyzing the user's answers and writing style to infer their emotional state. Based on the inferred emotions, it generates appropriate advice and encouraging messages.

[0121] For example, consider a user who has created an answer to a history exam question asking them to "name important figures of the Sengoku period." When the user submits the answer to the system, the AI ​​analyzes its content and provides specific feedback, such as, "Oda Nobunaga is correct, but the explanation of Tokugawa Ieyasu's role is insufficient." Furthermore, the emotion engine considers the case where the user is frequently corrected and presents motivational messages such as, "Be confident and move on to the next challenge."

[0122] Thus, this system is provided not only as a tool for correcting answers, but also as a new form of educational support tool that empathizes with the user's emotions and supports their learning. As a result, students can not only correct their mistakes, but also positively advance their own learning process.

[0123] The following describes the processing flow.

[0124] Step 1:

[0125] The user uses a smartphone or tablet to take a picture of the answer sheet. After taking the picture, the user checks a preview of the image on the app and takes another picture if necessary. By pressing the submit button, the information is sent to the device.

[0126] Step 2:

[0127] The device receives the captured image data, optimizes the image resolution, crops it to remove excess white space, and removes image noise to prepare it for OCR processing.

[0128] Step 3:

[0129] The terminal sends the processed image data to the server. During this process, the data is encrypted before transmission, ensuring the security of the communication.

[0130] Step 4:

[0131] The server passes the received image data to the OCR process, which extracts text information from the image. The OCR process converts the answer into text data that can be handled by a computer.

[0132] Step 5:

[0133] The server inputs the extracted text data into a natural language processing model to evaluate the accuracy and logical validity of the answers. This model then generates appropriate correction feedback for the answers, comparing them against a database.

[0134] Step 6:

[0135] The server utilizes an emotion engine to analyze the user's responses and textual expressions to infer their emotional state. Based on this emotional state, it generates the most appropriate message and advice for the user.

[0136] Step 7:

[0137] The server sends generated editing feedback and emotion-based messages to the device. The feedback is formatted and presented in a way that is easy for the user to understand.

[0138] Step 8:

[0139] The device displays the received feedback to the user. Users can see not only the correction results but also messages that resonate with their emotional state. Through this information, they can maintain their motivation to learn and use it to improve their next learning session.

[0140] (Example 2)

[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0142] While conventional correction systems can evaluate the syntactic and logical accuracy of answers, they have limitations in providing emotional support to users during their learning process. In particular, there was a need for a system that enhances learning effectiveness by providing feedback that takes into account the emotional state users experience when answering. Furthermore, improvements in efficient image processing and information extraction methods were required to improve the accuracy of answer analysis.

[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0144] In this invention, the server includes optical recognition means, character conversion means, analysis means, sentiment analysis means, and communication means. This enables accurate extraction of answer information, generation of feedback based on syntactic and logical evaluation, and provision of personalized support according to the user's emotional state.

[0145] "Optical recognition means" refers to a function that uses technology to extract information such as characters and symbols from captured images.

[0146] A "character conversion method" is a function that uses technology to convert extracted character information into digital text format.

[0147] The "analysis tool" is a function that evaluates the syntactic and logical accuracy of the answer content based on digital text and generates appropriate feedback.

[0148] "Emotional analysis tools" refer to functions that use technology to infer the user's emotional state from their answers and written text, and to generate feedback that corresponds to those emotions.

[0149] "Communication methods" refer to functions that utilize communication technologies to deliver generated feedback and evaluation information to users.

[0150] The system of this invention begins with the user using a smartphone or tablet to take a picture of the answer sheet. The device receives the captured image and performs image processing such as adjusting the resolution and cropping out unnecessary parts. This processing includes a module that utilizes image recognition technology.

[0151] The processed images are sent from the terminal to the server. The server uses optical character recognition (OCR) technology to extract character information from the received images and converts them into digital text data. In this conversion process, noise reduction is performed as a pre-processing step to improve the accuracy of character recognition.

[0152] Next, the server utilizes natural language processing technology to evaluate the syntactic and logical accuracy of the answers based on the converted text data. This evaluation process detects the appropriateness and errors of the answers and generates feedback containing correction instructions and explanations.

[0153] The sentiment analysis engine operates to analyze the user's answers and infer the emotional state contained within them. This engine takes into account multiple parameters, such as the frequency of errors and the time taken to answer, to generate advice messages that are appropriate to the user's emotional state.

[0154] For example, consider a scenario where, while answering a history exam question, "Name an important figure from the Sengoku period," the user answers "Oda Nobunaga." The system provides feedback stating, "This is correct, but the explanation is insufficient." Furthermore, sentiment analysis adds an encouraging message such as, "Even if you make frequent mistakes, don't give up and keep trying."

[0155] As an example of a prompt, the generation AI model is instructed to "provide feedback on the answer and generate an encouraging message based on the user's emotions." This system allows users to receive emotionally responsive support that goes beyond simple correction instructions, thereby improving the quality of their learning.

[0156] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0157] Step 1:

[0158] The user takes a picture of the answer sheet using a smartphone or tablet. The input is a physical image of the answer sheet. The device acquires this image and temporarily stores it for subsequent image processing.

[0159] Step 2:

[0160] The device adjusts the resolution of the acquired image and crops out unnecessary parts. The input is the image captured in step 1, and the output is the processed, clear image. Image processing software is used to perform pre-processing to enable more accurate character recognition. As a result of this processing, the image quality is improved and the accuracy of character recognition is increased.

[0161] Step 3:

[0162] The terminal sends the processed image to the server. The input is the processed image from step 2, and the output is the receipt of image data on the server side. The image data is transmitted to the server securely and quickly via a communication protocol.

[0163] Step 4:

[0164] The server applies OCR technology to the received image data to extract text information from the image. The input is an image sent from the terminal, and the output is string data. The OCR software analyzes the shape of the characters and generates continuous text data.

[0165] Step 5:

[0166] The server uses natural language processing techniques based on text data to evaluate the syntax and logic of the answers. The input is the text data obtained in step 4, and the output is the evaluation result and feedback information. The natural language processing engine analyzes the text, detects errors, and generates correction instructions.

[0167] Step 6:

[0168] The server operates an emotion analysis engine based on the generated evaluation results to infer the user's emotional state. The input is the evaluation results from step 5, and the output is emotion evaluation data. The emotion analysis software infers emotions from the user's response patterns and frequency.

[0169] Step 7:

[0170] The server utilizes sentiment evaluation data to generate personalized feedback messages tailored to the user's emotions. The input is sentiment evaluation data, and the output is the final feedback information for the user. A generative AI model creates messages containing appropriate encouragement and advice.

[0171] Step 8:

[0172] The server sends the final feedback information to the terminal and presents it to the user. The input is the generated feedback information, and the output is the feedback displayed on the user's terminal. Information is sent to the terminal using communication means so that the user can easily check it.

[0173] (Application Example 2)

[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0175] Traditional grading systems focus solely on determining the correctness of answers, lacking sufficient emotional support for learners. Therefore, it is necessary to provide personalized feedback to help learners maintain their motivation and continue their studies.

[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0177] In this invention, the server includes data processing means for processing captured images and extracting answer information; character recognition means for converting the answer information into character information; evaluation means for correcting based on the character information and generating evaluation data; emotion estimation means for inferring the emotional state from the evaluation data and generating an appropriate feedback message; and transmission means for sending the evaluation data and feedback message to the user. This makes it possible for learners to receive feedback that takes their own emotional state into consideration.

[0178] "Data processing means" refers to technical elements used to analyze captured images and extract necessary answer information.

[0179] "Character recognition means" refers to a technical element for converting answer information obtained from an image into character information.

[0180] "Evaluation means" refers to technical elements for correcting answers based on textual information and generating evaluation data.

[0181] "Emotion inference means" refers to a technological element for inferring a learner's emotional state from evaluation data and generating feedback messages based on that inference.

[0182] "Means of transmission" refers to the technical elements used to send generated evaluation data and feedback messages to the user.

[0183] This invention realizes a system that digitizes photographed answer sheets and provides personalized feedback by linking a terminal and a server. Users take pictures of their answer sheets for learning assignments using a terminal such as a smartphone. The application on the terminal has the function of cropping and adjusting the resolution of this image using OpenCV.

[0184] After image processing, the data is converted into text information using Tesseract's OCR technology and sent to the server. The server analyzes the text information using a Python natural language processing library and generates evaluation data based on the logical consistency and syntactic accuracy of the answer.

[0185] Furthermore, the server uses sentiment analysis tools such as Google's Natural Language API to infer the learner's emotional state from this evaluation data. Based on this, the system generates and sends feedback messages to the learner, including encouragement and advice.

[0186] As a concrete example, consider a scenario where a user fills in an English composition question and submits it to the system. The system can not only provide specific feedback such as "there is an error in the causative usage of the verb," ​​but also offer emotionally considerate messages such as "improvement is a sign of growth. Let's do our best next time!"

[0187] Example of a prompt:

[0188] "Take photos of the students' answers, read the answers from the images, correct the content of the answers, infer the emotional state of the students, and then provide a feedback message."

[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0190] Step 1:

[0191] The user takes a picture of the answer sheet using their smartphone. The input is the raw image obtained from the camera, and this image is not suitable for analysis in its original state.

[0192] Step 2:

[0193] Using OpenCV on the terminal, the captured image is cropped and its resolution optimized. The input is the raw image from step 1, and the output is the cropped image with adjusted resolution. This process improves the accuracy of character recognition.

[0194] Step 3:

[0195] Using Tesseract on a terminal, text information is extracted from a processed image. The input is the processed image from step 2, and the output is the answer information in text format. This allows the text within the image to be treated as digital data.

[0196] Step 4:

[0197] The terminal sends text data to the server. The input is the text data obtained in step 3, which prepares the server for further analysis.

[0198] Step 5:

[0199] The server uses a Python natural language processing library to analyze the received text data and generate evaluation data based on its logical consistency and syntactic accuracy. The input is the text data from step 4, and the output is evaluation data including corrections. This analysis evaluates the quality of the answer content.

[0200] Step 6:

[0201] The server uses the Google Natural Language API to infer the learner's emotions from the evaluation data. The input is the evaluation data from step 5, and the output is data indicating the emotional state. At this step, the system is ready to provide feedback based on the learner's emotions.

[0202] Step 7:

[0203] The server generates appropriate feedback messages based on emotion inference data. The input is the emotion data inferred in step 6, and the output is an emotion-sensitive feedback message. This message aims to improve the learner's motivation.

[0204] Step 8:

[0205] The server sends evaluation data and feedback messages to the user's terminal. The input is the feedback message created in step 7, which is provided to the learner. The excuse information reaches the user in the form of the precise feedback and sentiment message provided.

[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0209] [Second Embodiment]

[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0222] This invention relates to an automated proofreading system that integrates image recognition technology and natural language processing technology. This system receives images of exam answers taken by the user using a smartphone or tablet as input, and has the function of quickly and accurately proofreading and providing feedback.

[0223] When the user takes a picture of the answer sheet, the device performs image processing, optimizing the resolution and removing noise. This makes the image suitable for character recognition. Next, the device sends the processed image data to the server.

[0224] The server extracts text data from the received image using OCR (Optical Character Recognition). This text data is then fed into an AI-powered natural language processing model to evaluate the accuracy, expression, and logic of the answer. Existing correct answer databases and grammatical rules are used for evaluation, automatically identifying errors and generating optimal feedback.

[0225] The generated feedback is formatted as text and sent back to the device. The device displays this information to the user, clearly indicating which parts of the answer are correct and which parts need improvement. This entire process allows students to efficiently check their understanding and move forward with their learning.

[0226] As a concrete example, let's consider a historical question. For instance, suppose a user creates a detailed answer to the question, "Explain the characteristics of the Nara period." When the user takes a picture of this answer and sends it to the system, the AI ​​analyzes its content and provides specific feedback, such as, "The description of the influence of Buddhism is appropriate, but the description of the relocation of the capital is inaccurate." This feedback gives the user an opportunity to improve the accuracy of their descriptions and their understanding of the subject matter.

[0227] The system of this invention automates this rapid and accurate feedback process, thereby reducing the burden of grading work in educational settings and improving the quality of individualized learning.

[0228] The following describes the processing flow.

[0229] Step 1:

[0230] The user uses a smartphone or tablet to take a picture of the sheet of paper containing the question and its answer. The user reviews the image through the application, confirms that there are no quality issues, and then presses the submit button.

[0231] Step 2:

[0232] The device receives the captured image data and performs image processing such as improving resolution, adjusting contrast, and cropping unnecessary parts. This makes the image optimal for character recognition.

[0233] Step 3:

[0234] The terminal uploads the processed image data to the server. The data is transmitted in an encrypted state for security purposes.

[0235] Step 4:

[0236] The server receives the transmitted image data and uses an OCR (Optical Character Recognition) engine to extract text information from the image. OCR processing converts the characters in the image into text data that can be analyzed by a computer.

[0237] Step 5:

[0238] The server analyzes the text data obtained by OCR and passes it to a natural language processing model. The model evaluates the accuracy, expressiveness, and logical structure of the answer based on the text data and generates correction comments.

[0239] Step 6:

[0240] The server structures the generated ratings and feedback, editing them into a user-friendly format. The feedback presented clearly indicates which parts of the answer are correct and which parts could be improved.

[0241] Step 7:

[0242] The terminal displays feedback received from the server on the user's device. This allows the user to check the evaluation of their answers and immediately understand areas for improvement in their learning.

[0243] (Example 1)

[0244] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0245] The conventional process of grading exam answers is time-consuming and labor-intensive, and has been particularly problematic in providing efficient feedback for individual learners. Furthermore, manual grading is subjective, making it difficult to maintain consistency in evaluation. This invention aims to solve these problems and support the improvement of individual learners' understanding by providing rapid and accurate feedback.

[0246] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0247] In this invention, the server includes means for acquiring captured images and transmitting them to an information processing device; image processing means for optimizing the resolution and removing noise from the images; character recognition means for converting the optimized images into text data using optical character recognition technology; evaluation means for analyzing the text data based on natural language processing technology, evaluating the accuracy and logic of the answers, and generating feedback; and communication means for transmitting the generated feedback to the user via a data transmission device. This enables rapid and objective correction and feedback provision.

[0248] "Means for acquiring captured images and transmitting them to an information processing device" refers to a component for receiving image data captured by a user and transmitting it to an information processing device via a network.

[0249] "Resolution optimization" is a process that appropriately adjusts the amount of data while preserving the visual information of an image.

[0250] "Noise reduction" is the process of removing unwanted visual data and distortions from an image, making the image clearer.

[0251] "Image processing means" refers to a processing device or program that has the function of applying a specific algorithm to input image data to optimize resolution and remove noise.

[0252] "Optical character recognition technology" is a technology that recognizes characters from images and extracts them as text data.

[0253] "Character recognition means" refers to a function that uses optical character recognition technology to extract character information from image data and analyze it as electronic text.

[0254] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language.

[0255] "Evaluation means" refers to a component that analyzes text data using natural language processing technology to evaluate the accuracy and logic of the answers.

[0256] A "data transmission device" is a device or mechanism for transmitting information and feedback generated on a server to a user terminal via a network.

[0257] "Communication means" refers to the technology or device that enables the sending and receiving of data between a server and a user terminal.

[0258] This invention begins with a user taking a picture of their exam answers using a device such as a smartphone or tablet. The device receives this image and performs image processing. Image processing includes resolution optimization and noise reduction, which makes character recognition easier. Specifically, the device adjusts the image sharpness using the OpenCV library and reduces noise through a filter.

[0259] Next, the processed image is sent from the terminal to the server. The terminal encrypts the data using the HTTPS protocol to ensure secure communication with the server. The server extracts text data from the image using optical character recognition (OCR) technology such as Tesseract. The extracted text data is processed in text format and evaluated by the server's AI engine.

[0260] The server evaluates the accuracy, expression, and logic of the answers based on natural language processing (NLP) techniques. Here, for example, it judges the quality of the answers by referring to pre-entered model answers and grammar checkers. The AI ​​includes models commonly used as generative AI models.

[0261] The generated feedback is sent back from the server to the device, where the user can view it on the display screen. The device's UI is designed using the React Native framework, allowing for intuitive touch-based viewing of detailed feedback.

[0262] As a concrete example, consider a case where a user answers the question, "Explain the characteristics of the Nara period," takes a photo of their answer, and uploads it to the system. The AI ​​analyzes the content and provides feedback that the description of "the influence of Buddhism" is appropriate, but the description of "the relocation of the capital" is inaccurate. In this way, the user can review their answer and obtain useful information to deepen their understanding.

[0263] An example of a prompt would be, "I have taken a picture of an answer explaining the characteristics of the Nara period. Please evaluate the accuracy of the historical facts and the way they are expressed in this answer, and provide feedback on areas for improvement." This specific prompt allows the AI ​​to provide the analysis results the user is looking for more accurately.

[0264] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0265] Step 1:

[0266] The user takes a picture of their exam answer sheet using a device. In this step, the exam answer sheet is used as visual input, and the device activates its camera function to acquire image data. The output obtained here is a digital image file.

[0267] Step 2:

[0268] The device performs image processing on the captured image. Specifically, it uses the OpenCV library to optimize the image resolution and apply a noise reduction filter. The input here is image data captured by the user, and the output is clear image data with optimized resolution and noise removed.

[0269] Step 3:

[0270] The terminal sends the processed image data to the server. The data is encrypted using the HTTPS protocol for secure transmission. The input is optimized image data, and the output is ready to be passed to the data receiver module on the server.

[0271] Step 4:

[0272] The server processes the received image using optical character recognition (OCR) technology and extracts character data. It uses the Tesseract engine to identify characters in the image and convert them into text format. The input is the image data received by the server, and the output is the extracted character data.

[0273] Step 5:

[0274] The server analyzes the recognized character data using natural language processing (NLP) techniques. A generative AI model is used to evaluate the accuracy and logic of the answers. This analysis compares the text structure and content with a database of known answers. The input is character data, and the output is feedback data including the evaluation results.

[0275] Step 6:

[0276] The server generates feedback data and sends it to the terminal. Communication takes place through a data transmission device, and the transmission is encrypted. The input here is feedback data including evaluation results, and the output is data sent back to the user's terminal.

[0277] Step 7:

[0278] The device displays feedback received from the server in the user interface. Using the React Native framework, users can view detailed feedback through touch gestures. The input is feedback data sent from the server, and the output is information displayed in a format that the user can see.

[0279] (Application Example 1)

[0280] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0281] There is a need to improve operational efficiency and ensure picking accuracy in logistics centers. However, manual work verification can lead to human error and time loss, compromising the accuracy of inventory management. Therefore, a system that can perform evaluation and provide feedback in real time using automated methods is necessary.

[0282] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0283] In this invention, the server includes visual data processing means for processing the captured visual data and extracting information, data conversion means for converting the information into text data, data evaluation means for performing an evaluation based on the text data and generating an evaluation result, and verification means for comparing with the recorded storage information and evaluating the accuracy. Thereby, the work content within the logistics center can be evaluated quickly and accurately, and appropriate feedback can be provided to the operator.

[0284] The "visual data processing means" is a device or program for optimizing the captured image and extracting the required information.

[0285] [[ID=,7]] The "data conversion means" is a device or program for converting the information extracted by the visual data processing means into text data.

[0286] The "data evaluation means" is a device or program for evaluating the content of the information based on the converted text data and generating an evaluation result.

[0287] The "information communication means" is a communication device or program for transmitting the evaluated result to the user's terminal.

[0288] The "verification means" is a device or program for comparing the generated evaluation result with the recorded storage information and ensuring the accuracy.

[0289] The "recorded storage information" is a set of reference data for verifying the evaluation result.

[0290] The "evaluation result" is the evaluation information generated by the data evaluation means and is the feedback provided to the user.

[0291] The system that implements this application is intended for efficient evaluation and feedback provision for use in logistics centers and similar environments. A smartphone or tablet is responsible for processing the images of the picking list. First, the device uses OpenCV to optimize the image resolution and remove noise. Next, it extracts text information using Tesseract OCR and sends that information to the server.

[0292] The server analyzes the received text data using a natural language processing model based on TensorFlow. This analysis evaluates whether the text data matches existing record-keeping information and determines whether the picking was carried out as planned. The evaluation results are confirmed by a verification mechanism and then sent back to the terminal via an information and communication mechanism.

[0293] Based on feedback received on the device, users can verify the accuracy of their picking work in real time and correct their work as needed. For example, a worker might take a photo of a list that says "5 boxes of oranges," and the AI ​​compares it with inventory data to confirm that 5 boxes were actually picked. If there is a discrepancy, the user is immediately notified.

[0294] An example of a prompt message is: "Design an application that takes a picture of a picking list used in a logistics center and instantly verifies the accuracy of the inventory using OCR and natural language processing."

[0295] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0296] Step 1:

[0297] The terminal receives an image of the picking list taken by the user. OpenCV is used to optimize the resolution and remove noise. The input is the captured image data, and the output is the processed, clean image.

[0298] Step 2:

[0299] The terminal uses Tesseract OCR to extract character information from the clean image. In this process, an image optimized as input is provided, and text data is generated as output.

[0300] Step 3:

[0301] The terminal transmits the text data it generated to the server. This data serves as the basic information for analysis in the subsequent process.

[0302] Step 4:

[0303] The server analyzes the received text data using a natural language processing model based on TensorFlow. The input is the text data, and the output is the evaluation result that assesses the accuracy of picking. Through data analysis, it is confirmed whether the text matches the pre-recorded storage information.

[0304] Step 5:

[0305] The server compares with the reference information, verifies the accuracy of picking, and generates the final evaluation result. In this verification step, the existing recorded storage information is used as the input data.

[0306] Step 6:

[0307] The server transmits the verified evaluation result to the terminal. The terminal displays this result to the user and provides feedback regarding the accuracy of picking.

[0308] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.

[0309] This invention relates to a system that provides more personalized feedback to users by combining an emotion engine with a correction system. This system takes a picture of the paper on which the user has submitted their answers and uses that image to perform efficient analysis of the answers and evaluate their emotions.

[0310] Users use their smartphones or tablets to take photos of their completed answer sheets. The device receives these images and performs image processing such as adjusting the resolution and cropping. The processed images are then sent from the device to the server. The server uses OCR technology to extract text information from the images and formats it as text data.

[0311] Next, the server applies evaluation methods for automated correction. Utilizing natural language processing technology, it evaluates the logic and syntactic accuracy of the answers, generating specific corrections and feedback. Based on this data, an emotion engine is activated, analyzing the user's answers and writing style to infer their emotional state. Based on the inferred emotions, it generates appropriate advice and encouraging messages.

[0312] For example, consider a user who has created an answer to a history exam question asking them to "name important figures of the Sengoku period." When the user submits the answer to the system, the AI ​​analyzes its content and provides specific feedback, such as, "Oda Nobunaga is correct, but the explanation of Tokugawa Ieyasu's role is insufficient." Furthermore, the emotion engine considers the case where the user is frequently corrected and presents motivational messages such as, "Be confident and move on to the next challenge."

[0313] Thus, this system is provided not only as a tool for correcting answers, but also as a new form of educational support tool that empathizes with the user's emotions and supports their learning. As a result, students can not only correct their mistakes, but also positively advance their own learning process.

[0314] The following describes the processing flow.

[0315] Step 1:

[0316] The user uses a smartphone or tablet to take a picture of the answer sheet. After taking the picture, the user checks a preview of the image on the app and takes another picture if necessary. By pressing the submit button, the information is sent to the device.

[0317] Step 2:

[0318] The device receives the captured image data, optimizes the image resolution, crops it to remove excess white space, and removes image noise to prepare it for OCR processing.

[0319] Step 3:

[0320] The terminal sends the processed image data to the server. During this process, the data is encrypted before transmission, ensuring the security of the communication.

[0321] Step 4:

[0322] The server passes the received image data to the OCR process, which extracts text information from the image. The OCR process converts the answer into text data that can be handled by a computer.

[0323] Step 5:

[0324] The server inputs the extracted text data into a natural language processing model to evaluate the accuracy and logical validity of the answers. This model then generates appropriate correction feedback for the answers, comparing them against a database.

[0325] Step 6:

[0326] The server utilizes an emotion engine to analyze the user's responses and textual expressions to infer their emotional state. Based on this emotional state, it generates the most appropriate message and advice for the user.

[0327] Step 7:

[0328] The server sends generated editing feedback and emotion-based messages to the device. The feedback is formatted and presented in a way that is easy for the user to understand.

[0329] Step 8:

[0330] The device displays the received feedback to the user. Users can see not only the correction results but also messages that resonate with their emotional state. Through this information, they can maintain their motivation to learn and use it to improve their next learning session.

[0331] (Example 2)

[0332] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0333] While conventional correction systems can evaluate the syntactic and logical accuracy of answers, they have limitations in providing emotional support to users during their learning process. In particular, there was a need for a system that enhances learning effectiveness by providing feedback that takes into account the emotional state users experience when answering. Furthermore, improvements in efficient image processing and information extraction methods were required to improve the accuracy of answer analysis.

[0334] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0335] In this invention, the server includes optical recognition means, character conversion means, analysis means, sentiment analysis means, and communication means. This enables accurate extraction of answer information, generation of feedback based on syntactic and logical evaluation, and provision of personalized support according to the user's emotional state.

[0336] "Optical recognition means" refers to a function that uses technology to extract information such as characters and symbols from captured images.

[0337] A "character conversion method" is a function that uses technology to convert extracted character information into digital text format.

[0338] The "analysis tool" is a function that evaluates the syntactic and logical accuracy of the answer content based on digital text and generates appropriate feedback.

[0339] "Emotional analysis tools" refer to functions that use technology to infer the user's emotional state from their answers and written text, and to generate feedback that corresponds to those emotions.

[0340] "Communication methods" refer to functions that utilize communication technologies to deliver generated feedback and evaluation information to users.

[0341] The system of this invention begins with the user using a smartphone or tablet to take a picture of the answer sheet. The device receives the captured image and performs image processing such as adjusting the resolution and cropping out unnecessary parts. This processing includes a module that utilizes image recognition technology.

[0342] The processed images are sent from the terminal to the server. The server uses optical character recognition (OCR) technology to extract character information from the received images and converts them into digital text data. In this conversion process, noise reduction is performed as a pre-processing step to improve the accuracy of character recognition.

[0343] Next, the server utilizes natural language processing technology to evaluate the syntactic and logical accuracy of the answers based on the converted text data. This evaluation process detects the appropriateness and errors of the answers and generates feedback containing correction instructions and explanations.

[0344] The sentiment analysis engine operates to analyze the user's answers and infer the emotional state contained within them. This engine takes into account multiple parameters, such as the frequency of errors and the time taken to answer, to generate advice messages that are appropriate to the user's emotional state.

[0345] For example, consider a scenario where, while answering a history exam question, "Name an important figure from the Sengoku period," the user answers "Oda Nobunaga." The system provides feedback stating, "This is correct, but the explanation is insufficient." Furthermore, sentiment analysis adds an encouraging message such as, "Even if you make frequent mistakes, don't give up and keep trying."

[0346] As an example of a prompt, the generation AI model is instructed to "provide feedback on the answer and generate an encouraging message based on the user's emotions." This system allows users to receive emotionally responsive support that goes beyond simple correction instructions, thereby improving the quality of their learning.

[0347] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0348] Step 1:

[0349] The user takes a picture of the answer sheet using a smartphone or tablet. The input is a physical image of the answer sheet. The device acquires this image and temporarily stores it for subsequent image processing.

[0350] Step 2:

[0351] The device adjusts the resolution of the acquired image and crops out unnecessary parts. The input is the image captured in step 1, and the output is the processed, clear image. Image processing software is used to perform pre-processing to enable more accurate character recognition. As a result of this processing, the image quality is improved and the accuracy of character recognition is increased.

[0352] Step 3:

[0353] The terminal sends the processed image to the server. The input is the processed image from step 2, and the output is the receipt of image data on the server side. The image data is transmitted to the server securely and quickly via a communication protocol.

[0354] Step 4:

[0355] The server applies OCR technology to the received image data to extract text information from the image. The input is an image sent from the terminal, and the output is string data. The OCR software analyzes the shape of the characters and generates continuous text data.

[0356] Step 5:

[0357] The server uses natural language processing techniques based on text data to evaluate the syntax and logic of the answers. The input is the text data obtained in step 4, and the output is the evaluation result and feedback information. The natural language processing engine analyzes the text, detects errors, and generates correction instructions.

[0358] Step 6:

[0359] The server operates an emotion analysis engine based on the generated evaluation results to infer the user's emotional state. The input is the evaluation results from step 5, and the output is emotion evaluation data. The emotion analysis software infers emotions from the user's response patterns and frequency.

[0360] Step 7:

[0361] The server utilizes sentiment evaluation data to generate personalized feedback messages tailored to the user's emotions. The input is sentiment evaluation data, and the output is the final feedback information for the user. A generative AI model creates messages containing appropriate encouragement and advice.

[0362] Step 8:

[0363] The server sends the final feedback information to the terminal and presents it to the user. The input is the generated feedback information, and the output is the feedback displayed on the user's terminal. Information is sent to the terminal using communication means so that the user can easily check it.

[0364] (Application Example 2)

[0365] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0366] Traditional grading systems focus solely on determining the correctness of answers, lacking sufficient emotional support for learners. Therefore, it is necessary to provide personalized feedback to help learners maintain their motivation and continue their studies.

[0367] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0368] In this invention, the server includes data processing means for processing captured images and extracting answer information; character recognition means for converting the answer information into character information; evaluation means for correcting based on the character information and generating evaluation data; emotion estimation means for inferring the emotional state from the evaluation data and generating an appropriate feedback message; and transmission means for sending the evaluation data and feedback message to the user. This makes it possible for learners to receive feedback that takes their own emotional state into consideration.

[0369] "Data processing means" refers to technical elements used to analyze captured images and extract necessary answer information.

[0370] "Character recognition means" refers to a technical element for converting answer information obtained from an image into character information.

[0371] "Evaluation means" refers to technical elements for correcting answers based on textual information and generating evaluation data.

[0372] "Emotion inference means" refers to a technological element for inferring a learner's emotional state from evaluation data and generating feedback messages based on that inference.

[0373] "Means of transmission" refers to the technical elements used to send generated evaluation data and feedback messages to the user.

[0374] This invention realizes a system that digitizes photographed answer sheets and provides personalized feedback by linking a terminal and a server. Users take pictures of their answer sheets for learning assignments using a terminal such as a smartphone. The application on the terminal has the function of cropping and adjusting the resolution of this image using OpenCV.

[0375] After image processing, the data is converted into text information using Tesseract's OCR technology and sent to the server. The server analyzes the text information using a Python natural language processing library and generates evaluation data based on the logical consistency and syntactic accuracy of the answer.

[0376] Furthermore, the server uses sentiment analysis tools such as the Google Natural Language API to infer the learner's emotional state from this evaluation data. Based on this, the system generates and sends feedback messages to the learner, including encouragement and advice.

[0377] As a concrete example, consider a scenario where a user fills in an English composition question and submits it to the system. The system can not only provide specific feedback such as "there is an error in the causative usage of the verb," ​​but also offer emotionally considerate messages such as "improvement is a sign of growth. Let's do our best next time!"

[0378] Example of a prompt:

[0379] "Take photos of the students' answers, read the answers from the images, correct the content of the answers, infer the emotional state of the students, and then provide a feedback message."

[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0381] Step 1:

[0382] The user takes a picture of the answer sheet using their smartphone. The input is the raw image obtained from the camera, and this image is not suitable for analysis in its original state.

[0383] Step 2:

[0384] Using OpenCV on the terminal, the captured image is cropped and its resolution optimized. The input is the raw image from step 1, and the output is the cropped image with adjusted resolution. This process improves the accuracy of character recognition.

[0385] Step 3:

[0386] Using Tesseract on a terminal, text information is extracted from a processed image. The input is the processed image from step 2, and the output is the answer information in text format. This allows the text within the image to be treated as digital data.

[0387] Step 4:

[0388] The terminal sends text data to the server. The input is the text data obtained in step 3, which prepares the server for further analysis.

[0389] Step 5:

[0390] The server uses a Python natural language processing library to analyze the received text data and generate evaluation data based on its logical consistency and syntactic accuracy. The input is the text data from step 4, and the output is evaluation data including corrections. This analysis evaluates the quality of the answer content.

[0391] Step 6:

[0392] The server uses the Google Natural Language API to infer the learner's emotions from the evaluation data. The input is the evaluation data from step 5, and the output is data indicating the emotional state. At this step, the system is ready to provide feedback based on the learner's emotions.

[0393] Step 7:

[0394] The server generates appropriate feedback messages based on emotion inference data. The input is the emotion data inferred in step 6, and the output is an emotion-sensitive feedback message. This message aims to improve the learner's motivation.

[0395] Step 8:

[0396] The server sends evaluation data and feedback messages to the user's terminal. The input is the feedback message created in step 7, which is provided to the learner. The excuse information reaches the user in the form of the precise feedback and sentiment message provided.

[0397] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0398] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0399] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0400] [Third Embodiment]

[0401] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0402] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0403] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0404] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0405] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0406] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0407] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0408] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0409] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0410] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0411] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0412] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0413] This invention relates to an automated proofreading system that integrates image recognition technology and natural language processing technology. This system receives images of exam answers taken by the user using a smartphone or tablet as input, and has the function of quickly and accurately proofreading and providing feedback.

[0414] When the user takes a picture of the answer sheet, the device performs image processing, optimizing the resolution and removing noise. This makes the image suitable for character recognition. Next, the device sends the processed image data to the server.

[0415] The server extracts text data from the received image using OCR (Optical Character Recognition). This text data is then fed into an AI-powered natural language processing model to evaluate the accuracy, expression, and logic of the answer. Existing correct answer databases and grammatical rules are used for evaluation, automatically identifying errors and generating optimal feedback.

[0416] The generated feedback is formatted as text and sent back to the device. The device displays this information to the user, clearly indicating which parts of the answer are correct and which parts need improvement. This entire process allows students to efficiently check their understanding and move forward with their learning.

[0417] As a concrete example, let's consider a historical question. For instance, suppose a user creates a detailed answer to the question, "Explain the characteristics of the Nara period." When the user takes a picture of this answer and sends it to the system, the AI ​​analyzes its content and provides specific feedback, such as, "The description of the influence of Buddhism is appropriate, but the description of the relocation of the capital is inaccurate." This feedback gives the user an opportunity to improve the accuracy of their descriptions and their understanding of the subject matter.

[0418] The system of this invention automates this rapid and accurate feedback process, thereby reducing the burden of grading work in educational settings and improving the quality of individualized learning.

[0419] The following describes the processing flow.

[0420] Step 1:

[0421] The user uses a smartphone or tablet to take a picture of the sheet of paper containing the question and its answer. The user reviews the image through the application, confirms that there are no quality issues, and then presses the submit button.

[0422] Step 2:

[0423] The device receives the captured image data and performs image processing such as improving resolution, adjusting contrast, and cropping unnecessary parts. This makes the image optimal for character recognition.

[0424] Step 3:

[0425] The terminal uploads the processed image data to the server. The data is transmitted in an encrypted state for security purposes.

[0426] Step 4:

[0427] The server receives the transmitted image data and uses an OCR (Optical Character Recognition) engine to extract text information from the image. OCR processing converts the characters in the image into text data that can be analyzed by a computer.

[0428] Step 5:

[0429] The server analyzes the text data obtained by OCR and passes it to a natural language processing model. The model evaluates the accuracy, expressiveness, and logical structure of the answer based on the text data and generates correction comments.

[0430] Step 6:

[0431] The server structures the generated ratings and feedback, editing them into a user-friendly format. The feedback presented clearly indicates which parts of the answer are correct and which parts could be improved.

[0432] Step 7:

[0433] The terminal displays feedback received from the server on the user's device. This allows the user to check the evaluation of their answers and immediately understand areas for improvement in their learning.

[0434] (Example 1)

[0435] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0436] The conventional process of grading exam answers is time-consuming and labor-intensive, and has been particularly problematic in providing efficient feedback for individual learners. Furthermore, manual grading is subjective, making it difficult to maintain consistency in evaluation. This invention aims to solve these problems and support the improvement of individual learners' understanding by providing rapid and accurate feedback.

[0437] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0438] In this invention, the server includes means for acquiring captured images and transmitting them to an information processing device; image processing means for optimizing the resolution and removing noise from the images; character recognition means for converting the optimized images into text data using optical character recognition technology; evaluation means for analyzing the text data based on natural language processing technology, evaluating the accuracy and logic of the answers, and generating feedback; and communication means for transmitting the generated feedback to the user via a data transmission device. This enables rapid and objective correction and feedback provision.

[0439] "Means for acquiring captured images and transmitting them to an information processing device" refers to a component for receiving image data captured by a user and transmitting it to an information processing device via a network.

[0440] "Resolution optimization" is a process that appropriately adjusts the amount of data while preserving the visual information of an image.

[0441] "Noise reduction" is the process of removing unwanted visual data and distortions from an image, making the image clearer.

[0442] "Image processing means" refers to a processing device or program that has the function of applying a specific algorithm to input image data to optimize resolution and remove noise.

[0443] "Optical character recognition technology" is a technology that recognizes characters from images and extracts them as text data.

[0444] "Character recognition means" refers to a function that uses optical character recognition technology to extract character information from image data and analyze it as electronic text.

[0445] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language.

[0446] "Evaluation means" refers to a component that analyzes text data using natural language processing technology to evaluate the accuracy and logic of the answers.

[0447] A "data transmission device" is a device or mechanism for transmitting information and feedback generated on a server to a user terminal via a network.

[0448] "Communication means" refers to the technology or device that enables the sending and receiving of data between a server and a user terminal.

[0449] This invention begins with a user taking a picture of their exam answers using a device such as a smartphone or tablet. The device receives this image and performs image processing. Image processing includes resolution optimization and noise reduction, which makes character recognition easier. Specifically, the device adjusts the image sharpness using the OpenCV library and reduces noise through a filter.

[0450] Next, the processed image is sent from the terminal to the server. The terminal encrypts the data using the HTTPS protocol to ensure secure communication with the server. The server extracts text data from the image using optical character recognition (OCR) technology such as Tesseract. The extracted text data is processed in text format and evaluated by the server's AI engine.

[0451] The server evaluates the accuracy, expression, and logic of the answers based on natural language processing (NLP) techniques. Here, for example, it judges the quality of the answers by referring to pre-entered model answers and grammar checkers. The AI ​​includes models commonly used as generative AI models.

[0452] The generated feedback is sent back from the server to the device, where the user can view it on the display screen. The device's UI is designed using the React Native framework, allowing for intuitive touch-based viewing of detailed feedback.

[0453] As a concrete example, consider a case where a user answers the question, "Explain the characteristics of the Nara period," takes a photo of their answer, and uploads it to the system. The AI ​​analyzes the content and provides feedback that the description of "the influence of Buddhism" is appropriate, but the description of "the relocation of the capital" is inaccurate. In this way, the user can review their answer and obtain useful information to deepen their understanding.

[0454] An example of a prompt would be, "I have taken a picture of an answer explaining the characteristics of the Nara period. Please evaluate the accuracy of the historical facts and the way they are expressed in this answer, and provide feedback on areas for improvement." This specific prompt allows the AI ​​to provide the analysis results the user is looking for more accurately.

[0455] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0456] Step 1:

[0457] The user takes a picture of their exam answer sheet using a device. In this step, the exam answer sheet is used as visual input, and the device activates its camera function to acquire image data. The output obtained here is a digital image file.

[0458] Step 2:

[0459] The device performs image processing on the captured image. Specifically, it uses the OpenCV library to optimize the image resolution and apply a noise reduction filter. The input here is image data captured by the user, and the output is clear image data with optimized resolution and noise removed.

[0460] Step 3:

[0461] The terminal sends the processed image data to the server. The data is encrypted using the HTTPS protocol for secure transmission. The input is optimized image data, and the output is ready to be passed to the data receiver module on the server.

[0462] Step 4:

[0463] The server processes the received image using optical character recognition (OCR) technology and extracts character data. It uses the Tesseract engine to identify characters in the image and convert them into text format. The input is the image data received by the server, and the output is the extracted character data.

[0464] Step 5:

[0465] The server analyzes the recognized character data using natural language processing (NLP) techniques. A generative AI model is used to evaluate the accuracy and logic of the answers. This analysis compares the text structure and content with a database of known answers. The input is character data, and the output is feedback data including the evaluation results.

[0466] Step 6:

[0467] The server generates feedback data and sends it to the terminal. Communication takes place through a data transmission device, and the transmission is encrypted. The input here is feedback data including evaluation results, and the output is data sent back to the user's terminal.

[0468] Step 7:

[0469] The device displays feedback received from the server in the user interface. Using the React Native framework, users can view detailed feedback through touch gestures. The input is feedback data sent from the server, and the output is information displayed in a format that the user can see.

[0470] (Application Example 1)

[0471] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0472] There is a need to improve operational efficiency and ensure picking accuracy in logistics centers. However, manual work verification can lead to human error and time loss, compromising the accuracy of inventory management. Therefore, a system that can perform evaluation and provide feedback in real time using automated methods is necessary.

[0473] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0474] In this invention, the server includes a visual data processing means for processing captured visual data and extracting information; a data conversion means for converting the information into text data; a data evaluation means for performing evaluations based on the text data and generating evaluation results; and a verification means for matching the data with stored record information and evaluating its accuracy. This enables rapid and accurate evaluation of work performed within the logistics center and provides appropriate feedback to workers.

[0475] "Visual data processing means" refers to a device or program for optimizing captured images and extracting necessary information.

[0476] "Data conversion means" refers to a device or program for converting information extracted by visual data processing means into text data.

[0477] "Data evaluation means" refers to a device or program that evaluates the content of information based on converted text data and generates evaluation results.

[0478] "Information and communication means" refers to a communication device or program for transmitting the evaluated results to the user's terminal.

[0479] "Verification means" refers to a device or program that compares the generated evaluation results with the recorded and stored information to guarantee their accuracy.

[0480] "Record storage information" refers to a set of reference data that is used to verify evaluation results.

[0481] "Evaluation results" refer to evaluation information generated by data evaluation methods and are feedback provided to users.

[0482] The system that implements this application is intended for efficient evaluation and feedback provision for use in logistics centers and similar environments. A smartphone or tablet is responsible for processing the images of the picking list. First, the device uses OpenCV to optimize the image resolution and remove noise. Next, it extracts text information using Tesseract OCR and sends that information to the server.

[0483] The server analyzes the received text data using a natural language processing model based on TensorFlow. This analysis evaluates whether the text data matches existing record-keeping information and determines whether the picking was carried out as planned. The evaluation results are confirmed by a verification mechanism and then sent back to the terminal via an information and communication mechanism.

[0484] Based on feedback received on the device, users can verify the accuracy of their picking work in real time and correct their work as needed. For example, a worker might take a photo of a list that says "5 boxes of oranges," and the AI ​​compares it with inventory data to confirm that 5 boxes were actually picked. If there is a discrepancy, the user is immediately notified.

[0485] An example of a prompt message is: "Design an application that takes a picture of a picking list used in a logistics center and instantly verifies the accuracy of the inventory using OCR and natural language processing."

[0486] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0487] Step 1:

[0488] The terminal receives an image of the picking list taken by the user. OpenCV is used to optimize the resolution and remove noise. The input is the captured image data, and the output is the processed, clean image.

[0489] Step 2:

[0490] The device uses Tesseract OCR to extract text information from a clean image. This process takes an optimized image as input and generates text data as output.

[0491] Step 3:

[0492] The terminal sends the generated text data to the server. This data serves as basic information for analysis in the next processing step.

[0493] Step 4:

[0494] The server analyzes the received text data using a natural language processing model based on TensorFlow. The input is text data, and the output is an evaluation result assessing the accuracy of the picking. The data analysis verifies whether the text matches pre-recorded storage information.

[0495] Step 5:

[0496] The server cross-references the information to verify the accuracy of the picking and generates a final evaluation result. In this verification step, existing record-keeping information is used as input data.

[0497] Step 6:

[0498] The server sends the verified evaluation results to the terminal. The terminal displays these results to the user and provides feedback on the accuracy of the picking.

[0499] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0500] This invention relates to a system that provides more personalized feedback to users by combining an emotion engine with a correction system. This system takes a picture of the paper on which the user has submitted their answers and uses that image to perform efficient analysis of the answers and evaluate their emotions.

[0501] Users use their smartphones or tablets to take photos of their completed answer sheets. The device receives these images and performs image processing such as adjusting the resolution and cropping. The processed images are then sent from the device to the server. The server uses OCR technology to extract text information from the images and formats it as text data.

[0502] Next, the server applies evaluation methods for automated correction. Utilizing natural language processing technology, it evaluates the logic and syntactic accuracy of the answers, generating specific corrections and feedback. Based on this data, an emotion engine is activated, analyzing the user's answers and writing style to infer their emotional state. Based on the inferred emotions, it generates appropriate advice and encouraging messages.

[0503] For example, consider a user who has created an answer to a history exam question asking them to "name important figures of the Sengoku period." When the user submits the answer to the system, the AI ​​analyzes its content and provides specific feedback, such as, "Oda Nobunaga is correct, but the explanation of Tokugawa Ieyasu's role is insufficient." Furthermore, the emotion engine considers the case where the user is frequently corrected and presents motivational messages such as, "Be confident and move on to the next challenge."

[0504] Thus, this system is provided not only as a tool for correcting answers, but also as a new form of educational support tool that empathizes with the user's emotions and supports their learning. As a result, students can not only correct their mistakes, but also positively advance their own learning process.

[0505] The following describes the processing flow.

[0506] Step 1:

[0507] The user uses a smartphone or tablet to take a picture of the answer sheet. After taking the picture, the user checks a preview of the image on the app and takes another picture if necessary. By pressing the submit button, the information is sent to the device.

[0508] Step 2:

[0509] The device receives the captured image data, optimizes the image resolution, crops it to remove excess white space, and removes image noise to prepare it for OCR processing.

[0510] Step 3:

[0511] The terminal sends the processed image data to the server. During this process, the data is encrypted before transmission, ensuring the security of the communication.

[0512] Step 4:

[0513] The server passes the received image data to the OCR process, which extracts text information from the image. The OCR process converts the answer into text data that can be handled by a computer.

[0514] Step 5:

[0515] The server inputs the extracted text data into a natural language processing model to evaluate the accuracy and logical validity of the answers. This model then generates appropriate correction feedback for the answers, comparing them against a database.

[0516] Step 6:

[0517] The server utilizes an emotion engine to analyze the user's responses and textual expressions to infer their emotional state. Based on this emotional state, it generates the most appropriate message and advice for the user.

[0518] Step 7:

[0519] The server sends generated editing feedback and emotion-based messages to the device. The feedback is formatted and presented in a way that is easy for the user to understand.

[0520] Step 8:

[0521] The device displays the received feedback to the user. Users can see not only the correction results but also messages that resonate with their emotional state. Through this information, they can maintain their motivation to learn and use it to improve their next learning session.

[0522] (Example 2)

[0523] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0524] While conventional correction systems can evaluate the syntactic and logical accuracy of answers, they have limitations in providing emotional support to users during their learning process. In particular, there was a need for a system that enhances learning effectiveness by providing feedback that takes into account the emotional state users experience when answering. Furthermore, improvements in efficient image processing and information extraction methods were required to improve the accuracy of answer analysis.

[0525] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0526] In this invention, the server includes optical recognition means, character conversion means, analysis means, sentiment analysis means, and communication means. This enables accurate extraction of answer information, generation of feedback based on syntactic and logical evaluation, and provision of personalized support according to the user's emotional state.

[0527] "Optical recognition means" refers to a function that uses technology to extract information such as characters and symbols from captured images.

[0528] A "character conversion method" is a function that uses technology to convert extracted character information into digital text format.

[0529] The "analysis tool" is a function that evaluates the syntactic and logical accuracy of the answer content based on digital text and generates appropriate feedback.

[0530] "Emotional analysis tools" refer to functions that use technology to infer the user's emotional state from their answers and written text, and to generate feedback that corresponds to those emotions.

[0531] "Communication methods" refer to functions that utilize communication technologies to deliver generated feedback and evaluation information to users.

[0532] The system of this invention begins with the user using a smartphone or tablet to take a picture of the answer sheet. The device receives the captured image and performs image processing such as adjusting the resolution and cropping out unnecessary parts. This processing includes a module that utilizes image recognition technology.

[0533] The processed images are sent from the terminal to the server. The server uses optical character recognition (OCR) technology to extract character information from the received images and converts them into digital text data. In this conversion process, noise reduction is performed as a pre-processing step to improve the accuracy of character recognition.

[0534] Next, the server utilizes natural language processing technology to evaluate the syntactic and logical accuracy of the answers based on the converted text data. This evaluation process detects the appropriateness and errors of the answers and generates feedback containing correction instructions and explanations.

[0535] The sentiment analysis engine operates to analyze the user's answers and infer the emotional state contained within them. This engine takes into account multiple parameters, such as the frequency of errors and the time taken to answer, to generate advice messages that are appropriate to the user's emotional state.

[0536] For example, consider a scenario where, while answering a history exam question, "Name an important figure from the Sengoku period," the user answers "Oda Nobunaga." The system provides feedback stating, "This is correct, but the explanation is insufficient." Furthermore, sentiment analysis adds an encouraging message such as, "Even if you make frequent mistakes, don't give up and keep trying."

[0537] As an example of a prompt, the generation AI model is instructed to "provide feedback on the answer and generate an encouraging message based on the user's emotions." This system allows users to receive emotionally responsive support that goes beyond simple correction instructions, thereby improving the quality of their learning.

[0538] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0539] Step 1:

[0540] The user takes a picture of the answer sheet using a smartphone or tablet. The input is a physical image of the answer sheet. The device acquires this image and temporarily stores it for subsequent image processing.

[0541] Step 2:

[0542] The device adjusts the resolution of the acquired image and crops out unnecessary parts. The input is the image captured in step 1, and the output is the processed, clear image. Image processing software is used to perform pre-processing to enable more accurate character recognition. As a result of this processing, the image quality is improved and the accuracy of character recognition is increased.

[0543] Step 3:

[0544] The terminal sends the processed image to the server. The input is the processed image from step 2, and the output is the receipt of image data on the server side. The image data is transmitted to the server securely and quickly via a communication protocol.

[0545] Step 4:

[0546] The server applies OCR technology to the received image data to extract text information from the image. The input is an image sent from the terminal, and the output is string data. The OCR software analyzes the shape of the characters and generates continuous text data.

[0547] Step 5:

[0548] The server uses natural language processing techniques based on text data to evaluate the syntax and logic of the answers. The input is the text data obtained in step 4, and the output is the evaluation result and feedback information. The natural language processing engine analyzes the text, detects errors, and generates correction instructions.

[0549] Step 6:

[0550] The server operates an emotion analysis engine based on the generated evaluation results to infer the user's emotional state. The input is the evaluation results from step 5, and the output is emotion evaluation data. The emotion analysis software infers emotions from the user's response patterns and frequency.

[0551] Step 7:

[0552] The server utilizes sentiment evaluation data to generate personalized feedback messages tailored to the user's emotions. The input is sentiment evaluation data, and the output is the final feedback information for the user. A generative AI model creates messages containing appropriate encouragement and advice.

[0553] Step 8:

[0554] The server sends the final feedback information to the terminal and presents it to the user. The input is the generated feedback information, and the output is the feedback displayed on the user's terminal. Information is sent to the terminal using communication means so that the user can easily check it.

[0555] (Application Example 2)

[0556] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0557] Traditional grading systems focus solely on determining the correctness of answers, lacking sufficient emotional support for learners. Therefore, it is necessary to provide personalized feedback to help learners maintain their motivation and continue their studies.

[0558] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0559] In this invention, the server includes data processing means for processing captured images and extracting answer information; character recognition means for converting the answer information into character information; evaluation means for correcting based on the character information and generating evaluation data; emotion estimation means for inferring the emotional state from the evaluation data and generating an appropriate feedback message; and transmission means for sending the evaluation data and feedback message to the user. This makes it possible for learners to receive feedback that takes their own emotional state into consideration.

[0560] "Data processing means" refers to technical elements used to analyze captured images and extract necessary answer information.

[0561] "Character recognition means" refers to a technical element for converting answer information obtained from an image into character information.

[0562] "Evaluation means" refers to technical elements for correcting answers based on textual information and generating evaluation data.

[0563] "Emotion inference means" refers to a technological element for inferring a learner's emotional state from evaluation data and generating feedback messages based on that inference.

[0564] "Means of transmission" refers to the technical elements used to send generated evaluation data and feedback messages to the user.

[0565] This invention realizes a system that digitizes photographed answer sheets and provides personalized feedback by linking a terminal and a server. Users take pictures of their answer sheets for learning assignments using a terminal such as a smartphone. The application on the terminal has the function of cropping and adjusting the resolution of this image using OpenCV.

[0566] After image processing, the data is converted into text information using Tesseract's OCR technology and sent to the server. The server analyzes the text information using a Python natural language processing library and generates evaluation data based on the logical consistency and syntactic accuracy of the answer.

[0567] Furthermore, the server uses sentiment analysis tools such as the Google Natural Language API to infer the learner's emotional state from this evaluation data. Based on this, the system generates and sends feedback messages to the learner, including encouragement and advice.

[0568] As a concrete example, consider a scenario where a user fills in an English composition question and submits it to the system. The system can not only provide specific feedback such as "there is an error in the causative usage of the verb," ​​but also offer emotionally considerate messages such as "improvement is a sign of growth. Let's do our best next time!"

[0569] Example of a prompt:

[0570] "Take photos of the students' answers, read the answers from the images, correct the content of the answers, infer the emotional state of the students, and then provide a feedback message."

[0571] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0572] Step 1:

[0573] The user takes a picture of the answer sheet using their smartphone. The input is the raw image obtained from the camera, and this image is not suitable for analysis in its original state.

[0574] Step 2:

[0575] Using OpenCV on the terminal, the captured image is cropped and its resolution optimized. The input is the raw image from step 1, and the output is the cropped image with adjusted resolution. This process improves the accuracy of character recognition.

[0576] Step 3:

[0577] Using Tesseract on a terminal, text information is extracted from a processed image. The input is the processed image from step 2, and the output is the answer information in text format. This allows the text within the image to be treated as digital data.

[0578] Step 4:

[0579] The terminal sends text data to the server. The input is the text data obtained in step 3, which prepares the server for further analysis.

[0580] Step 5:

[0581] The server uses a Python natural language processing library to analyze the received text data and generate evaluation data based on its logical consistency and syntactic accuracy. The input is the text data from step 4, and the output is evaluation data including corrections. This analysis evaluates the quality of the answer content.

[0582] Step 6:

[0583] The server uses the Google Natural Language API to infer the learner's emotions from the evaluation data. The input is the evaluation data from step 5, and the output is data indicating the emotional state. At this step, the system is ready to provide feedback based on the learner's emotions.

[0584] Step 7:

[0585] The server generates appropriate feedback messages based on emotion inference data. The input is the emotion data inferred in step 6, and the output is an emotion-sensitive feedback message. This message aims to improve the learner's motivation.

[0586] Step 8:

[0587] The server sends evaluation data and feedback messages to the user's terminal. The input is the feedback message created in step 7, which is provided to the learner. The excuse information reaches the user in the form of the precise feedback and sentiment message provided.

[0588] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0589] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0590] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0591] [Fourth Embodiment]

[0592] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0593] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0594] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0595] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0596] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0597] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0598] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0599] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0600] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0601] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0602] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0603] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0604] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0605] This invention relates to an automated proofreading system that integrates image recognition technology and natural language processing technology. This system receives images of exam answers taken by the user using a smartphone or tablet as input, and has the function of quickly and accurately proofreading and providing feedback.

[0606] When the user takes a picture of the answer sheet, the device performs image processing, optimizing the resolution and removing noise. This makes the image suitable for character recognition. Next, the device sends the processed image data to the server.

[0607] The server extracts text data from the received image using OCR (Optical Character Recognition). This text data is then fed into an AI-powered natural language processing model to evaluate the accuracy, expression, and logic of the answer. Existing correct answer databases and grammatical rules are used for evaluation, automatically identifying errors and generating optimal feedback.

[0608] The generated feedback is formatted as text and sent back to the device. The device displays this information to the user, clearly indicating which parts of the answer are correct and which parts need improvement. This entire process allows students to efficiently check their understanding and move forward with their learning.

[0609] As a concrete example, let's consider a historical question. For instance, suppose a user creates a detailed answer to the question, "Explain the characteristics of the Nara period." When the user takes a picture of this answer and sends it to the system, the AI ​​analyzes its content and provides specific feedback, such as, "The description of the influence of Buddhism is appropriate, but the description of the relocation of the capital is inaccurate." This feedback gives the user an opportunity to improve the accuracy of their descriptions and their understanding of the subject matter.

[0610] The system of this invention automates this rapid and accurate feedback process, thereby reducing the burden of grading work in educational settings and improving the quality of individualized learning.

[0611] The following describes the processing flow.

[0612] Step 1:

[0613] The user uses a smartphone or tablet to take a picture of the sheet of paper containing the question and its answer. The user reviews the image through the application, confirms that there are no quality issues, and then presses the submit button.

[0614] Step 2:

[0615] The device receives the captured image data and performs image processing such as improving resolution, adjusting contrast, and cropping unnecessary parts. This makes the image optimal for character recognition.

[0616] Step 3:

[0617] The terminal uploads the processed image data to the server. The data is transmitted in an encrypted state for security purposes.

[0618] Step 4:

[0619] The server receives the transmitted image data and uses an OCR (Optical Character Recognition) engine to extract text information from the image. OCR processing converts the characters in the image into text data that can be analyzed by a computer.

[0620] Step 5:

[0621] The server analyzes the text data obtained by OCR and passes it to a natural language processing model. The model evaluates the accuracy, expressiveness, and logical structure of the answer based on the text data and generates correction comments.

[0622] Step 6:

[0623] The server structures the generated ratings and feedback, editing them into a user-friendly format. The feedback presented clearly indicates which parts of the answer are correct and which parts could be improved.

[0624] Step 7:

[0625] The terminal displays feedback received from the server on the user's device. This allows the user to check the evaluation of their answers and immediately understand areas for improvement in their learning.

[0626] (Example 1)

[0627] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0628] The conventional process of grading exam answers is time-consuming and labor-intensive, and has been particularly problematic in providing efficient feedback for individual learners. Furthermore, manual grading is subjective, making it difficult to maintain consistency in evaluation. This invention aims to solve these problems and support the improvement of individual learners' understanding by providing rapid and accurate feedback.

[0629] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0630] In this invention, the server includes means for acquiring captured images and transmitting them to an information processing device; image processing means for optimizing the resolution and removing noise from the images; character recognition means for converting the optimized images into text data using optical character recognition technology; evaluation means for analyzing the text data based on natural language processing technology, evaluating the accuracy and logic of the answers, and generating feedback; and communication means for transmitting the generated feedback to the user via a data transmission device. This enables rapid and objective correction and feedback provision.

[0631] "Means for acquiring captured images and transmitting them to an information processing device" refers to a component for receiving image data captured by a user and transmitting it to an information processing device via a network.

[0632] "Resolution optimization" is a process that appropriately adjusts the amount of data while preserving the visual information of an image.

[0633] "Noise reduction" is the process of removing unwanted visual data and distortions from an image, making the image clearer.

[0634] "Image processing means" refers to a processing device or program that has the function of applying a specific algorithm to input image data to optimize resolution and remove noise.

[0635] "Optical character recognition technology" is a technology that recognizes characters from images and extracts them as text data.

[0636] "Character recognition means" refers to a function that uses optical character recognition technology to extract character information from image data and analyze it as electronic text.

[0637] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language.

[0638] "Evaluation means" refers to a component that analyzes text data using natural language processing technology to evaluate the accuracy and logic of the answers.

[0639] A "data transmission device" is a device or mechanism for transmitting information and feedback generated on a server to a user terminal via a network.

[0640] "Communication means" refers to the technology or device that enables the sending and receiving of data between a server and a user terminal.

[0641] This invention begins with a user taking a picture of their exam answers using a device such as a smartphone or tablet. The device receives this image and performs image processing. Image processing includes resolution optimization and noise reduction, which makes character recognition easier. Specifically, the device adjusts the image sharpness using the OpenCV library and reduces noise through a filter.

[0642] Next, the processed image is sent from the terminal to the server. The terminal encrypts the data using the HTTPS protocol to ensure secure communication with the server. The server extracts text data from the image using optical character recognition (OCR) technology such as Tesseract. The extracted text data is processed in text format and evaluated by the server's AI engine.

[0643] The server evaluates the accuracy, expression, and logic of the answers based on natural language processing (NLP) techniques. Here, for example, it judges the quality of the answers by referring to pre-entered model answers and grammar checkers. The AI ​​includes models commonly used as generative AI models.

[0644] The generated feedback is sent back from the server to the device, where the user can view it on the display screen. The device's UI is designed using the React Native framework, allowing for intuitive touch-based viewing of detailed feedback.

[0645] As a concrete example, consider a case where a user answers the question, "Explain the characteristics of the Nara period," takes a photo of their answer, and uploads it to the system. The AI ​​analyzes the content and provides feedback that the description of "the influence of Buddhism" is appropriate, but the description of "the relocation of the capital" is inaccurate. In this way, the user can review their answer and obtain useful information to deepen their understanding.

[0646] An example of a prompt would be, "I have taken a picture of an answer explaining the characteristics of the Nara period. Please evaluate the accuracy of the historical facts and the way they are expressed in this answer, and provide feedback on areas for improvement." This specific prompt allows the AI ​​to provide the analysis results the user is looking for more accurately.

[0647] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0648] Step 1:

[0649] The user takes a picture of their exam answer sheet using a device. In this step, the exam answer sheet is used as visual input, and the device activates its camera function to acquire image data. The output obtained here is a digital image file.

[0650] Step 2:

[0651] The device performs image processing on the captured image. Specifically, it uses the OpenCV library to optimize the image resolution and apply a noise reduction filter. The input here is image data captured by the user, and the output is clear image data with optimized resolution and noise removed.

[0652] Step 3:

[0653] The terminal sends the processed image data to the server. The data is encrypted using the HTTPS protocol for secure transmission. The input is optimized image data, and the output is ready to be passed to the data receiver module on the server.

[0654] Step 4:

[0655] The server processes the received image using optical character recognition (OCR) technology and extracts character data. It uses the Tesseract engine to identify characters in the image and convert them into text format. The input is the image data received by the server, and the output is the extracted character data.

[0656] Step 5:

[0657] The server analyzes the recognized character data using natural language processing (NLP) techniques. A generative AI model is used to evaluate the accuracy and logic of the answers. This analysis compares the text structure and content with a database of known answers. The input is character data, and the output is feedback data including the evaluation results.

[0658] Step 6:

[0659] The server generates feedback data and sends it to the terminal. Communication takes place through a data transmission device, and the transmission is encrypted. The input here is feedback data including evaluation results, and the output is data sent back to the user's terminal.

[0660] Step 7:

[0661] The device displays feedback received from the server in the user interface. Using the React Native framework, users can view detailed feedback through touch gestures. The input is feedback data sent from the server, and the output is information displayed in a format that the user can see.

[0662] (Application Example 1)

[0663] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0664] There is a need to improve operational efficiency and ensure picking accuracy in logistics centers. However, manual work verification can lead to human error and time loss, compromising the accuracy of inventory management. Therefore, a system that can perform evaluation and provide feedback in real time using automated methods is necessary.

[0665] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0666] In this invention, the server includes a visual data processing means for processing captured visual data and extracting information; a data conversion means for converting the information into text data; a data evaluation means for performing evaluations based on the text data and generating evaluation results; and a verification means for matching the data with stored record information and evaluating its accuracy. This enables rapid and accurate evaluation of work performed within the logistics center and provides appropriate feedback to workers.

[0667] "Visual data processing means" refers to a device or program for optimizing captured images and extracting necessary information.

[0668] "Data conversion means" refers to a device or program for converting information extracted by visual data processing means into text data.

[0669] "Data evaluation means" refers to a device or program that evaluates the content of information based on converted text data and generates evaluation results.

[0670] "Information and communication means" refers to a communication device or program for transmitting the evaluated results to the user's terminal.

[0671] "Verification means" refers to a device or program that compares the generated evaluation results with the recorded and stored information to guarantee their accuracy.

[0672] "Record storage information" refers to a set of reference data that is used to verify evaluation results.

[0673] "Evaluation results" refer to evaluation information generated by data evaluation methods and are feedback provided to users.

[0674] The system that implements this application is intended for efficient evaluation and feedback provision for use in logistics centers and similar environments. A smartphone or tablet is responsible for processing the images of the picking list. First, the device uses OpenCV to optimize the image resolution and remove noise. Next, it extracts text information using Tesseract OCR and sends that information to the server.

[0675] The server analyzes the received text data using a natural language processing model based on TensorFlow. This analysis evaluates whether the text data matches existing record-keeping information and determines whether the picking was carried out as planned. The evaluation results are confirmed by a verification mechanism and then sent back to the terminal via an information and communication mechanism.

[0676] Based on feedback received on the device, users can verify the accuracy of their picking work in real time and correct their work as needed. For example, a worker might take a photo of a list that says "5 boxes of oranges," and the AI ​​compares it with inventory data to confirm that 5 boxes were actually picked. If there is a discrepancy, the user is immediately notified.

[0677] An example of a prompt message is: "Design an application that takes a picture of a picking list used in a logistics center and instantly verifies the accuracy of the inventory using OCR and natural language processing."

[0678] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0679] Step 1:

[0680] The terminal receives an image of the picking list taken by the user. OpenCV is used to optimize the resolution and remove noise. The input is the captured image data, and the output is the processed, clean image.

[0681] Step 2:

[0682] The device uses Tesseract OCR to extract text information from a clean image. This process takes an optimized image as input and generates text data as output.

[0683] Step 3:

[0684] The terminal sends the generated text data to the server. This data serves as basic information for analysis in the next processing step.

[0685] Step 4:

[0686] The server analyzes the received text data using a natural language processing model based on TensorFlow. The input is text data, and the output is an evaluation result assessing the accuracy of the picking. The data analysis verifies whether the text matches pre-recorded storage information.

[0687] Step 5:

[0688] The server cross-references the information to verify the accuracy of the picking and generates a final evaluation result. In this verification step, existing record-keeping information is used as input data.

[0689] Step 6:

[0690] The server sends the verified evaluation results to the terminal. The terminal displays these results to the user and provides feedback on the accuracy of the picking.

[0691] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0692] This invention relates to a system that provides more personalized feedback to users by combining an emotion engine with a correction system. This system takes a picture of the paper on which the user has submitted their answers and uses that image to perform efficient analysis of the answers and evaluate their emotions.

[0693] Users use their smartphones or tablets to take photos of their completed answer sheets. The device receives these images and performs image processing such as adjusting the resolution and cropping. The processed images are then sent from the device to the server. The server uses OCR technology to extract text information from the images and formats it as text data.

[0694] Next, the server applies evaluation methods for automated correction. Utilizing natural language processing technology, it evaluates the logic and syntactic accuracy of the answers, generating specific corrections and feedback. Based on this data, an emotion engine is activated, analyzing the user's answers and writing style to infer their emotional state. Based on the inferred emotions, it generates appropriate advice and encouraging messages.

[0695] For example, consider a user who has created an answer to a history exam question asking them to "name important figures of the Sengoku period." When the user submits the answer to the system, the AI ​​analyzes its content and provides specific feedback, such as, "Oda Nobunaga is correct, but the explanation of Tokugawa Ieyasu's role is insufficient." Furthermore, the emotion engine considers the case where the user is frequently corrected and presents motivational messages such as, "Be confident and move on to the next challenge."

[0696] Thus, this system is provided not only as a tool for correcting answers, but also as a new form of educational support tool that empathizes with the user's emotions and supports their learning. As a result, students can not only correct their mistakes, but also positively advance their own learning process.

[0697] The following describes the processing flow.

[0698] Step 1:

[0699] The user uses a smartphone or tablet to take a picture of the answer sheet. After taking the picture, the user checks a preview of the image on the app and takes another picture if necessary. By pressing the submit button, the information is sent to the device.

[0700] Step 2:

[0701] The device receives the captured image data, optimizes the image resolution, crops it to remove excess white space, and removes image noise to prepare it for OCR processing.

[0702] Step 3:

[0703] The terminal sends the processed image data to the server. During this process, the data is encrypted before transmission, ensuring the security of the communication.

[0704] Step 4:

[0705] The server passes the received image data to the OCR process, which extracts text information from the image. The OCR process converts the answer into text data that can be handled by a computer.

[0706] Step 5:

[0707] The server inputs the extracted text data into a natural language processing model to evaluate the accuracy and logical validity of the answers. This model then generates appropriate correction feedback for the answers, comparing them against a database.

[0708] Step 6:

[0709] The server utilizes an emotion engine to analyze the user's responses and textual expressions to infer their emotional state. Based on this emotional state, it generates the most appropriate message and advice for the user.

[0710] Step 7:

[0711] The server sends generated editing feedback and emotion-based messages to the device. The feedback is formatted and presented in a way that is easy for the user to understand.

[0712] Step 8:

[0713] The device displays the received feedback to the user. Users can see not only the correction results but also messages that resonate with their emotional state. Through this information, they can maintain their motivation to learn and use it to improve their next learning session.

[0714] (Example 2)

[0715] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0716] While conventional correction systems can evaluate the syntactic and logical accuracy of answers, they have limitations in providing emotional support to users during their learning process. In particular, there was a need for a system that enhances learning effectiveness by providing feedback that takes into account the emotional state users experience when answering. Furthermore, improvements in efficient image processing and information extraction methods were required to improve the accuracy of answer analysis.

[0717] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0718] In this invention, the server includes optical recognition means, character conversion means, analysis means, sentiment analysis means, and communication means. This enables accurate extraction of answer information, generation of feedback based on syntactic and logical evaluation, and provision of personalized support according to the user's emotional state.

[0719] "Optical recognition means" refers to a function that uses technology to extract information such as characters and symbols from captured images.

[0720] A "character conversion method" is a function that uses technology to convert extracted character information into digital text format.

[0721] The "analysis tool" is a function that evaluates the syntactic and logical accuracy of the answer content based on digital text and generates appropriate feedback.

[0722] "Emotional analysis tools" refer to functions that use technology to infer the user's emotional state from their answers and written text, and to generate feedback that corresponds to those emotions.

[0723] "Communication methods" refer to functions that utilize communication technologies to deliver generated feedback and evaluation information to users.

[0724] The system of this invention begins with the user using a smartphone or tablet to take a picture of the answer sheet. The device receives the captured image and performs image processing such as adjusting the resolution and cropping out unnecessary parts. This processing includes a module that utilizes image recognition technology.

[0725] The processed images are sent from the terminal to the server. The server uses optical character recognition (OCR) technology to extract character information from the received images and converts them into digital text data. In this conversion process, noise reduction is performed as a pre-processing step to improve the accuracy of character recognition.

[0726] Next, the server utilizes natural language processing technology to evaluate the syntactic and logical accuracy of the answers based on the converted text data. This evaluation process detects the appropriateness and errors of the answers and generates feedback containing correction instructions and explanations.

[0727] The sentiment analysis engine operates to analyze the user's answers and infer the emotional state contained within them. This engine takes into account multiple parameters, such as the frequency of errors and the time taken to answer, to generate advice messages that are appropriate to the user's emotional state.

[0728] For example, consider a scenario where, while answering a history exam question, "Name an important figure from the Sengoku period," the user answers "Oda Nobunaga." The system provides feedback stating, "This is correct, but the explanation is insufficient." Furthermore, sentiment analysis adds an encouraging message such as, "Even if you make frequent mistakes, don't give up and keep trying."

[0729] As an example of a prompt, the generation AI model is instructed to "provide feedback on the answer and generate an encouraging message based on the user's emotions." This system allows users to receive emotionally responsive support that goes beyond simple correction instructions, thereby improving the quality of their learning.

[0730] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0731] Step 1:

[0732] The user takes a picture of the answer sheet using a smartphone or tablet. The input is a physical image of the answer sheet. The device acquires this image and temporarily stores it for subsequent image processing.

[0733] Step 2:

[0734] The device adjusts the resolution of the acquired image and crops out unnecessary parts. The input is the image captured in step 1, and the output is the processed, clear image. Image processing software is used to perform pre-processing to enable more accurate character recognition. As a result of this processing, the image quality is improved and the accuracy of character recognition is increased.

[0735] Step 3:

[0736] The terminal sends the processed image to the server. The input is the processed image from step 2, and the output is the receipt of image data on the server side. The image data is transmitted to the server securely and quickly via a communication protocol.

[0737] Step 4:

[0738] The server applies OCR technology to the received image data to extract text information from the image. The input is an image sent from the terminal, and the output is string data. The OCR software analyzes the shape of the characters and generates continuous text data.

[0739] Step 5:

[0740] The server uses natural language processing techniques based on text data to evaluate the syntax and logic of the answers. The input is the text data obtained in step 4, and the output is the evaluation result and feedback information. The natural language processing engine analyzes the text, detects errors, and generates correction instructions.

[0741] Step 6:

[0742] The server operates an emotion analysis engine based on the generated evaluation results to infer the user's emotional state. The input is the evaluation results from step 5, and the output is emotion evaluation data. The emotion analysis software infers emotions from the user's response patterns and frequency.

[0743] Step 7:

[0744] The server utilizes sentiment evaluation data to generate personalized feedback messages tailored to the user's emotions. The input is sentiment evaluation data, and the output is the final feedback information for the user. A generative AI model creates messages containing appropriate encouragement and advice.

[0745] Step 8:

[0746] The server sends the final feedback information to the terminal and presents it to the user. The input is the generated feedback information, and the output is the feedback displayed on the user's terminal. Information is sent to the terminal using communication means so that the user can easily check it.

[0747] (Application Example 2)

[0748] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0749] Traditional grading systems focus solely on determining the correctness of answers, lacking sufficient emotional support for learners. Therefore, it is necessary to provide personalized feedback to help learners maintain their motivation and continue their studies.

[0750] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0751] In this invention, the server includes data processing means for processing captured images and extracting answer information; character recognition means for converting the answer information into character information; evaluation means for correcting based on the character information and generating evaluation data; emotion estimation means for inferring the emotional state from the evaluation data and generating an appropriate feedback message; and transmission means for sending the evaluation data and feedback message to the user. This makes it possible for learners to receive feedback that takes their own emotional state into consideration.

[0752] "Data processing means" refers to technical elements used to analyze captured images and extract necessary answer information.

[0753] "Character recognition means" refers to a technical element for converting answer information obtained from an image into character information.

[0754] "Evaluation means" refers to technical elements for correcting answers based on textual information and generating evaluation data.

[0755] "Emotion inference means" refers to a technological element for inferring a learner's emotional state from evaluation data and generating feedback messages based on that inference.

[0756] "Means of transmission" refers to the technical elements used to send generated evaluation data and feedback messages to the user.

[0757] This invention realizes a system that digitizes photographed answer sheets and provides personalized feedback by linking a terminal and a server. Users take pictures of their answer sheets for learning assignments using a terminal such as a smartphone. The application on the terminal has the function of cropping and adjusting the resolution of this image using OpenCV.

[0758] After image processing, the data is converted into text information using Tesseract's OCR technology and sent to the server. The server analyzes the text information using a Python natural language processing library and generates evaluation data based on the logical consistency and syntactic accuracy of the answer.

[0759] Furthermore, the server uses sentiment analysis tools such as the Google Natural Language API to infer the learner's emotional state from this evaluation data. Based on this, the system generates and sends feedback messages to the learner, including encouragement and advice.

[0760] As a concrete example, consider a scenario where a user fills in an English composition question and submits it to the system. The system can not only provide specific feedback such as "there is an error in the causative usage of the verb," ​​but also offer emotionally considerate messages such as "improvement is a sign of growth. Let's do our best next time!"

[0761] Example of a prompt:

[0762] "Take photos of the students' answers, read the answers from the images, correct the content of the answers, infer the emotional state of the students, and then provide a feedback message."

[0763] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0764] Step 1:

[0765] The user takes a picture of the answer sheet using their smartphone. The input is the raw image obtained from the camera, and this image is not suitable for analysis in its original state.

[0766] Step 2:

[0767] Using OpenCV on the terminal, the captured image is cropped and its resolution optimized. The input is the raw image from step 1, and the output is the cropped image with adjusted resolution. This process improves the accuracy of character recognition.

[0768] Step 3:

[0769] Using Tesseract on a terminal, text information is extracted from a processed image. The input is the processed image from step 2, and the output is the answer information in text format. This allows the text within the image to be treated as digital data.

[0770] Step 4:

[0771] The terminal sends text data to the server. The input is the text data obtained in step 3, which prepares the server for further analysis.

[0772] Step 5:

[0773] The server uses a Python natural language processing library to analyze the received text data and generate evaluation data based on its logical consistency and syntactic accuracy. The input is the text data from step 4, and the output is evaluation data including corrections. This analysis evaluates the quality of the answer content.

[0774] Step 6:

[0775] The server uses the Google Natural Language API to infer the learner's emotions from the evaluation data. The input is the evaluation data from step 5, and the output is data indicating the emotional state. At this step, the system is ready to provide feedback based on the learner's emotions.

[0776] Step 7:

[0777] The server generates appropriate feedback messages based on emotion inference data. The input is the emotion data inferred in step 6, and the output is an emotion-sensitive feedback message. This message aims to improve the learner's motivation.

[0778] Step 8:

[0779] The server sends evaluation data and feedback messages to the user's terminal. The input is the feedback message created in step 7, which is provided to the learner. The excuse information reaches the user in the form of the precise feedback and sentiment message provided.

[0780] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0781] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0782] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0783] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0784] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0785] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0786] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0787] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0788] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0789] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0790] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0791] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0792] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0793] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0794] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0795] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0796] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0797] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0798] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0799] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0800] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0801] The following is further disclosed regarding the embodiments described above.

[0802] (Claim 1)

[0803] Image processing means for processing captured images and extracting answer information,

[0804] Character recognition means for converting the aforementioned answer information into text data,

[0805] An evaluation means that performs corrections based on the aforementioned text data and generates evaluation information,

[0806] A system including a communication means for transmitting the aforementioned evaluation information to a user.

[0807] (Claim 2)

[0808] The system according to claim 1, wherein the image processing means has a function to adjust the resolution of the answer information and crop out unnecessary parts.

[0809] (Claim 3)

[0810] The system according to claim 1, wherein the evaluation means has a function to evaluate the syntactic and logical accuracy of the answer using natural language processing technology.

[0811] "Example 1"

[0812] (Claim 1)

[0813] A means for acquiring a captured image and transmitting it to an information processing device,

[0814] Image processing means for optimizing the resolution and removing noise from the aforementioned image,

[0815] A character recognition means that converts optimized images into text data using optical character recognition technology,

[0816] An evaluation means that analyzes the aforementioned text data based on natural language processing technology, evaluates the accuracy and logic of the answer, and generates feedback,

[0817] A system including communication means for transmitting the generated feedback to a user via a data transmission device.

[0818] (Claim 2)

[0819] The system according to claim 1, wherein the image processing means has the function of optimizing the resolution of the image and removing noise.

[0820] (Claim 3)

[0821] The system according to claim 1, wherein the evaluation means has a function to evaluate the syntactic and logical accuracy of the answer using natural language processing technology.

[0822] "Application Example 1"

[0823] (Claim 1)

[0824] A visual data processing means for processing captured images and extracting information,

[0825] A data conversion means for converting the aforementioned information into text data,

[0826] A data evaluation means that performs an evaluation based on the aforementioned text data and generates an evaluation result,

[0827] Information communication means for transmitting the evaluation results to a terminal,

[0828] A verification method that compares the recorded information with the actual information and evaluates its accuracy,

[0829] A system that includes...

[0830] (Claim 2)

[0831] The system according to claim 1, wherein the visual data processing means has a function to adjust the resolution of the information and remove unnecessary parts.

[0832] (Claim 3)

[0833] The system according to claim 1, wherein the data evaluation means has a function to evaluate structural and logical accuracy using natural language processing technology.

[0834] "Example 2 of combining an emotion engine"

[0835] (Claim 1)

[0836] An optical recognition means that processes captured images and extracts information,

[0837] A character conversion means for converting the aforementioned information into text data,

[0838] An analysis means that performs corrections based on the aforementioned text data and generates evaluation information,

[0839] A sentiment analysis means that infers emotions based on the aforementioned evaluation information and generates feedback corresponding to those emotions,

[0840] A system including means of communication for providing the aforementioned feedback to the user.

[0841] (Claim 2)

[0842] The system according to claim 1, wherein the optical recognition means has a function to adjust the resolution of the information and remove unnecessary parts.

[0843] (Claim 3)

[0844] The system according to claim 1, wherein the analysis means has a function to evaluate syntactic and logical accuracy using language processing technology.

[0845] "Application example 2 when combining with an emotional engine"

[0846] (Claim 1)

[0847] A data processing means for processing captured images and extracting answer information,

[0848] A character recognition means that converts the aforementioned answer information into character information,

[0849] An evaluation means that performs corrections based on the aforementioned textual information and generates evaluation data,

[0850] An emotion inference means that infers the emotional state from the aforementioned evaluation data and generates an appropriate feedback message,

[0851] A system including means for transmitting the aforementioned evaluation data and feedback messages to the user.

[0852] (Claim 2)

[0853] The system according to claim 1, wherein the data processing means has a function to adjust the resolution of the answer information and trim unnecessary information.

[0854] (Claim 3)

[0855] The system according to claim 1, wherein the evaluation means has a function to evaluate the syntactic and logical accuracy of the answer using natural language processing technology, and the sentiment inference means has a function to infer the learner's emotional state based on the evaluated data. [Explanation of Symbols]

[0856] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Image processing means for processing captured images and extracting answer information, Character recognition means for converting the aforementioned answer information into text data, An evaluation means that performs corrections based on the aforementioned text data and generates evaluation information, A system including a communication means for transmitting the aforementioned evaluation information to a user.

2. The system according to claim 1, wherein the image processing means has a function to adjust the resolution of the answer information and crop out unnecessary parts.

3. The system according to claim 1, wherein the evaluation means has a function to evaluate the syntactic and logical accuracy of the answer using natural language processing technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A