system
The system addresses inefficiencies in paper-based information processing by converting paper media to digital data using OCR and NLP, selecting optimal processing methods with a generative AI model, and handling errors in real time, thereby enhancing operational efficiency and user satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Existing information processing systems using paper media in enterprises and local governments face inefficiencies due to manual input errors, increased time costs, and handling costs, necessitating a system that can efficiently digitize paper media information, provide optimal processing methods, and perform real-time error handling.
A system utilizing optical character recognition (OCR) to convert paper media into digital data, natural language processing (NLP) to extract necessary information, and a generative AI model to select optimal processing methods while handling errors in real time, generating and outputting processing results as documents.
This system enhances operational efficiency by reducing errors and expediting operations through automated digitization, real-time error handling, and optimal processing methods, improving user satisfaction and reducing manual work inefficiencies.
Smart Images

Figure 2026069064000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In information processing operations using paper media in enterprises and local governments, there is a problem of reduced work efficiency because it causes manual input errors, time costs, and increased handling costs. Therefore, there is a demand for a system that can efficiently digitize paper media information, provide an optimal processing method, and perform real-time error handling.
Means for Solving the Problems
[0005] This invention provides a means for converting information on paper media into digital data using optical character recognition technology. Furthermore, it includes means for extracting necessary processing information from the digital data using natural language processing technology and selecting the optimal processing method based on this information. In addition, it provides a system that enables efficient business processing by handling errors that occur during the processing process in real time and generating and outputting the processing results as a document.
[0006] "Paper media" refers to information formats that are printed or written on physical paper.
[0007] "Digital data" refers to information that has been converted into a format that can be processed by computers and electronic devices.
[0008] "Optical character recognition technology" is a method of scanning printed characters and converting them into digital text.
[0009] "Natural language processing technology" is a technology that uses computers to analyze and understand human language.
[0010] "Processing information" refers to the data and instructions necessary to perform a specific task.
[0011] Error handling refers to the procedures for detecting and appropriately processing errors that occur within a system.
[0012] A "document" is a text or recording medium created for the purpose of recording and transmitting information.
[0013] A "system" is a totality formed by combining multiple elements to perform a specific function. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to a system that efficiently digitizes information from paper documents and processes it appropriately. This system converts information from paper documents into digital data using optical character recognition technology and extracts necessary processing information from the digital data using natural language processing technology. Based on this, the user, terminal, or server selects the optimal processing method and handles processing errors in real time.
[0036] System Implementation Example
[0037] The user places a paper invoice in the scanner and presses the scan button. The scanned information is sent to the server via the terminal.
[0038] The server applies optical character recognition (OCR) technology to the received scan data and extracts text data from the image.
[0039] The extracted text data is parsed on the server using natural language processing technology, and the information necessary for processing (e.g., amount, payment due date, recipient information, etc.) is extracted.
[0040] Based on this information, the server uses a multimodal generative AI model to determine the optimal processing method. For example, it can select a method for processing a large number of small payments in a batch.
[0041] Furthermore, the server has a real-time error handling function to prepare for errors that may occur during processing, and notifies the user of any inconsistencies that are found and requesting corrections.
[0042] Finally, the server generates a document based on the completed information and sends the necessary notification to the user and relevant parties.
[0043] Specific example
[0044] For example, when a local government manages the payment of various taxes from its residents, this system can be used to digitize paper payment slips and automatically analyze tax classification and payment status. Users (e.g., city hall employees) can use this system to efficiently process a large amount of payment data and quickly issue payment notices. As described above, the present invention can reduce errors caused by manual work and expedite operations.
[0045] The following describes the processing flow.
[0046] Step 1:
[0047] The user places a paper invoice or payment slip into the scanner and presses the scan button. This action captures the information from the paper document as a digital image on the device.
[0048] Step 2:
[0049] The terminal sends the scanned image it has captured to the server. The server receives this digital image and prepares for the next processing step.
[0050] Step 3:
[0051] The server uses optical character recognition (OCR) technology on the received digital image to extract text data from the image data. This digitizes the text information that was printed on paper.
[0052] Step 4:
[0053] The server uses natural language processing technology to analyze and extract necessary processing information (e.g., amount, payment deadline, recipient information, etc.) from the extracted text data.
[0054] Step 5:
[0055] The server selects the optimal processing method based on the analysis results. Here, it utilizes a multimodal generative AI model to make decisions such as processing a large number of small payments in the most cost-effective way.
[0056] Step 6:
[0057] The server performs data integrity checks while processing is in progress. If errors or inconsistencies are detected, error handling is performed in real time, and the user is notified to request correction.
[0058] Step 7:
[0059] After the server verifies that all data is accurate, it proceeds with the actual payment processing. This includes transfer operations through payment systems and bank APIs.
[0060] Step 8:
[0061] Once the server completes the payment, it generates necessary documents, such as a payment notification, based on the results. The generated documents are sent to the user and relevant parties as needed, or made available for the user to access.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] The need to efficiently digitize paper-based information, quickly and accurately extract processing data, and determine the optimal processing method is paramount. However, traditional methods rely heavily on manual processes, leading to errors and processing delays. Furthermore, handling errors in real time and ensuring users receive necessary notifications is crucial. Additionally, a lack of automation technology for these processes has resulted in reduced operational efficiency.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes a device for converting information from paper documents into digital data, a device for extracting processing information, and a device for determining the optimal processing method. This enables improved operational efficiency through the digitization of information, rapid handling and notification of errors, and the proposal of efficient processing methods using a generative AI model.
[0067] "Paper media" refers to media in which information is recorded on physical paper, such as documents and printed materials.
[0068] "Digital data" refers to information that has been converted into a format that can be processed by computers and electronic devices.
[0069] "Processing information" refers to information extracted from digital data that is necessary for specific tasks or operations.
[0070] Optical character recognition (OCR) is a technology that automatically reads printed or handwritten text information and digitizes it using machines.
[0071] "Natural language processing technology" refers to technologies used to analyze, interpret, or generate human language using computers.
[0072] A "generative AI model" is a machine learning model that uses artificial intelligence to generate information, and in particular, it has the function of suggesting the most optimal processing method or solution.
[0073] Error handling is the process of detecting errors and inconsistencies that occur during processing and taking appropriate action to address them.
[0074] "Notification" refers to the act or means of communicating processing results or error information to users or relevant parties.
[0075] This invention is a system that efficiently digitizes information from paper documents and processes it optimally. The user places a paper document, such as an invoice, into the scanner and presses the scan button to digitally input the information. This input digital image is then transmitted to a server via a terminal.
[0076] The server first extracts text data from the received digital image using optical character recognition (OCR) software. This is done using, for example, Tesseract or a commercial API. This process converts the information from image format to text format.
[0077] The obtained text data is then analyzed using natural language processing (NLP) techniques. The server extracts necessary processing information, such as amounts and payment due dates, from this analysis. Examples of technologies used include Python's NLTK and SpaCy.
[0078] Next, the server uses a generative AI model to propose the optimal processing method. For example, it might use an OpenAI® generative model to suggest an efficient payment processing method. An example of a prompt used in this process might be, "Extract the amounts and payment due dates from multiple invoices and propose the optimal payment processing method."
[0079] Based on the information extracted and the proposed processing methods, the server handles errors in real time and requests corrections from the user as needed. Finally, the processing results are generated as text and sent to the user and relevant parties through the notification system.
[0080] For example, when local governments process tax payments from residents, paper payment slips are digitized, and a process is implemented to automatically analyze the payment status by tax category. In this way, the system achieves increased efficiency and reduced errors.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user places a paper document into the scanner and presses the scan button. The input is a paper document (e.g., an invoice), and the goal is to capture it as a digital image on the terminal. Specifically, the scanner converts the physical information of the paper into a digital signal, which the terminal then acquires as image data.
[0084] Step 2:
[0085] The terminal sends the acquired digital image data to the server. The data to be sent is image data generated by the scanner. The terminal first converts the data format appropriately and then sends it to the server via the network.
[0086] Step 3:
[0087] The server applies Optical Character Recognition (OCR) to the received image data to extract text data. The input is image data, and the output is text data. Specifically, the OCR software identifies characters from the image and extracts them as strings.
[0088] Step 4:
[0089] The server uses text data to analyze information using natural language processing (NLP). The input here is the text data obtained in step 3, and the output is structured information necessary for processing (e.g., amount, payment due date). The specific operation is a process of analyzing the text using an NLP library, filtering, and extracting information.
[0090] Step 5:
[0091] The server uses a generative AI model to determine the optimal processing method. The input is the structured information obtained in step 4, and the output is a suggestion of processing steps and methods. The server sends prompts to the AI model, for example, "Extract the amount and payment due date from the invoice and suggest the optimal payment processing method."
[0092] Step 6:
[0093] The server handles errors that may occur during processing in real time. Input is data related to any errors or inconsistencies during processing, and output is notifications to the user or instructions for error correction. Specifically, an error handling module on the server detects errors and takes appropriate action.
[0094] Step 7:
[0095] The server generates the final processing results as a document and sends notifications to the user and relevant parties. The input is the processing method and its result determined in step 5, and the output is document data in the form of a report or notice. Specifically, a document generation tool is used to distribute the information via email or other means of communication.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] The challenges include the inefficiency of managing information using paper documents and the lack of efficient processing of customer service that requires immediate attention in physical stores. In particular, tasks related to the use of discount coupons and loyalty cards require the digitalization of information and immediate decision-making. This leads to problems such as increased errors and time loss due to manual work.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information, and means for visualizing and suggesting information in real time through a portable visual device. This enables faster information processing in physical stores and improves the immediacy and accuracy of customer service.
[0101] "Paper media" refers to a form of media on which information is printed or written for recording purposes.
[0102] "Digital data" refers to a collection of information that has been converted into a format that can be processed electronically.
[0103] A "portable visual device" is a device that is portable by the user and used to display information visually.
[0104] "Optical character recognition technology" is a technology that converts printed or handwritten characters into digital data.
[0105] "Natural language processing technology" is a technology that uses computers to analyze and process human language.
[0106] "Real-time error handling" is a method of immediately detecting and correcting errors that occur during processing.
[0107] A "generative AI model" is an artificial intelligence model that generates appropriate predictions and suggestions based on input data.
[0108] "Visualization" is the process of visually displaying digital data and information, making it possible for humans to recognize it.
[0109] A "proposal" is the act of providing appropriate answers or means in advance for a specific purpose or condition.
[0110] In an embodiment of this invention, smart glasses are first used as a portable visual device. The user captures information written on paper using the camera function of the smart glasses. The captured information is converted into digital data by optical character recognition technology performed within the smart glasses.
[0111] The converted digital data is sent to the server via the terminal. The server uses natural language processing technology to extract necessary processing information from the digital data. The information obtained during this process is then used to select the optimal processing method using a generative AI model.
[0112] Furthermore, the server has real-time error handling capabilities, instantly detecting any errors that may occur during processing and instructing the user to correct them. This enables users to process information quickly and accurately.
[0113] As a concrete example, consider customer service operations in a physical store. When a customer presents a paper coupon, the user can instantly verify the coupon details through smart glasses and determine whether the discount is applicable. This procedure allows the store to improve customer satisfaction.
[0114] Based on the generated AI model, an example of a prompt message used to notify the user of processing recommendations would be the text, "Can I apply this coupon? We will give you a 50% discount on the product price," displayed in the user's field of view.
[0115] This invention makes it possible to reduce the effort required for managing information on paper, improve operational efficiency, and enhance the quality of service provided to customers.
[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0117] Step 1:
[0118] The user uses the camera function of smart glasses to capture information from paper documents. The input is visual information from the paper document, and the output is a digital image file. This process allows information printed on paper to be captured by an electronic device.
[0119] Step 2:
[0120] The terminal processes the captured image file using optical character recognition (OCR) technology to extract text data. The input is the digital image obtained in step 1, and the output is text data. OCR processing converts the characters in the image into digital data.
[0121] Step 3:
[0122] The terminal sends the extracted text data to the server. The input is text data, and the output is data transfer to the server. This process provides the server with text information for processing.
[0123] Step 4:
[0124] The server uses natural language processing (NLP) techniques to extract important processing information from text data. The input is text data, and the output is specific processing information (e.g., coupon details, discount rate, eligibility). Through NLP, the information necessary for digital transactions is organized.
[0125] Step 5:
[0126] The server uses a generative AI model to make optimal suggestions and decisions based on the extracted processing information. The input is processing information, and the output is the generated suggestions or action recommendations. In this step, the AI model constructs the suggestions and prepares them to be fed into the next step.
[0127] Step 6:
[0128] The server detects potential errors during processing in real time and notifies the user. Inputs are processing information and the current execution status, while outputs are error messages or instructions for correction. Error handling ensures that the user is immediately informed of any problems.
[0129] Step 7:
[0130] The server displays the final processing results and suggestions on the user's smart glasses. The input is the optimal suggestions and modifications. The output is the data displayed in the user's field of vision. The final display allows the user to take appropriate action immediately.
[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0132] This invention is a system designed to streamline payment and information processing operations for companies and local governments. In addition to digitizing conventional paper-based information, it recognizes user emotions and reflects them in business processes. It not only processes information based on paper documents but also incorporates an emotion engine to optimize the processing flow while judging the user's emotional state in real time.
[0133] System Implementation Example
[0134] This system is configured as follows: The user digitizes information from paper documents using a scanner and sends it to the server. The server uses optical character recognition (OCR) technology to convert the image data into text data, and then uses natural language processing (NLP) technology to extract the information necessary for processing.
[0135] Based on the extracted information, the server selects the optimal processing method through a multimodal AI model and handles any errors that may occur during processing in real time. During this process, the server analyzes the user's emotions using an emotion engine. This emotion information is used to adjust the tone and content of interactions and messages in processing selection and error handling.
[0136] Specific example
[0137] For example, if a user encounters a specific error while processing data related to a payment slip, the emotion engine recognizes frustration or anxiety based on the user's facial expressions and tone of voice. The server uses this emotion recognition information to generate a message containing more user-friendly language and gentler instructions during error handling, and notifies the user accordingly. By responding in accordance with emotions in this way, user stress can be reduced and work efficiency can be improved.
[0138] Finally, the completed processing results are generated as a document, including a personalized message based on the user's emotions as needed, and provided to the user. Thus, the present invention is a comprehensive system that not only improves operational efficiency but also enhances user satisfaction.
[0139] The following describes the processing flow.
[0140] Step 1:
[0141] The user places a paper invoice into the scanner and starts scanning. The scanner captures the image data into the terminal.
[0142] Step 2:
[0143] The terminal sends the scanned image data to the server. The server receives this digital image and prepares it for the next processing step.
[0144] Step 3:
[0145] The server applies optical character recognition (OCR) technology to the received image data to extract text data. In this process, the text information from paper documents is converted into a digital format.
[0146] Step 4:
[0147] The server uses natural language processing technology to analyze the text data and extract the information necessary for processing (e.g., amount, payment due date, recipient information).
[0148] Step 5:
[0149] The server uses a multimodal AI model to select the optimal processing method based on the extracted information. For example, it might consider a method for processing multiple payments in a single batch.
[0150] Step 6:
[0151] When an error occurs on the server during processing, the emotion engine is used to analyze the user's emotions. For example, facial recognition and voice analysis are used to determine whether the user is confused or distressed.
[0152] Step 7:
[0153] The server adjusts the tone and content of error messages based on emotional information and sends appropriate support or correction requests to the user. This allows users to resolve problems without stress.
[0154] Step 8:
[0155] The server generates a document based on the processing results and adds a personalized message based on emotions as needed. The generated document is then provided to the user.
[0156] Step 9:
[0157] The user reviews the provided documents and confirms that the processing was completed successfully. This entire process allows the user to complete payment transactions efficiently.
[0158] (Example 2)
[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0160] Conventional information processing systems can be inefficient in the process of digitizing paper-based information and in subsequent error handling. Furthermore, uniform responses that disregard user feelings hinder the improvement of the user experience. There is a need to resolve these issues and simultaneously improve both processing efficiency and user satisfaction.
[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0162] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for recognizing user emotions and adjusting interactions based on those emotions. This enables efficient information processing and user-friendly error handling.
[0163] "Paper media" refers to a means of representing information that is printed or written on physical paper.
[0164] "Digital data" refers to information that has been converted into a format that can be processed and stored by electronic devices such as computers.
[0165] "Processing information" refers to information that includes data and instructions necessary to perform a specific task or operation.
[0166] "Error handling" is the process of detecting errors and problems that occur during the information processing process and taking appropriate action.
[0167] "User emotions" refers to the emotional reactions and states that users exhibit while using the system.
[0168] "Interaction" refers to the exchange of information and communication that takes place between a system and a user.
[0169] This system efficiently converts information from paper documents into digital data and implements a process that optimizes processing based on that information. A distinctive feature of this system is its ability to recognize user emotions in real time and utilize that information to optimize processing.
[0170] First, the user uses a scanner to digitize the information on paper. This digital data is then processed using OCR (Optical Character Recognition) technology. The server uses this OCR technology to extract text data from the scanned image data. Typically, commercially available OCR software is used for this purpose.
[0171] Next, the server uses NLP (Natural Language Processing) technology to extract necessary processing information from the text data. This technology is used to identify specific keywords and instructions from text. Common natural language processing libraries can be applied to the software used.
[0172] Furthermore, the server determines the optimal processing method as needed through a generative AI model. This AI model learns from the dataset and provides flexible processing methods that adapt to the situation. In addition, if an error occurs, the emotion engine analyzes the user's emotions and assists in taking appropriate action. This emotion engine can read emotions from the user's facial expressions and voice.
[0173] A concrete example is the emotional response to error messages that occur when a user attempts to digitize their tax payment process. This involves displaying reassuring messages based on emotional information to support the user in continuing the process. For example, a prompt might read, "The payment slip could not be read. Please try again. If you have further problems, please contact support."
[0174] The overall flow of this system is designed to streamline and user-friendly information processing operations for companies and local governments.
[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0176] Step 1:
[0177] The user places a paper document in the scanner and begins the digitization process. The scanner acquires image data from the paper document and generates that data as an electronic image file. As output of this process, the image file is sent to the server.
[0178] Step 2:
[0179] The server applies Optical Character Recognition (OCR) technology to the received image file. The input is image data. In this step, the server identifies the text within the image and converts the visual data into text data. The output of this process is readable electronic text.
[0180] Step 3:
[0181] The server analyzes electronic text data using natural language processing (NLP) techniques. Text data is provided as input. The server scans the entire document and extracts the necessary processing information. Relevant keywords and phrases are retrieved as NLP output and stored in a database.
[0182] Step 4:
[0183] The server uses a generative AI model to select the optimal processing plan based on the extracted information. This model operates using a pre-trained algorithm to derive the most efficient processing steps from the input data. The output consists of specific processing steps and necessary configuration information.
[0184] Step 5:
[0185] The server monitors for errors in real time as processing progresses and performs error handling as needed. Simultaneously, it uses an emotion engine to analyze the input user's video and audio data and recognize their emotions. The output emotion data is reflected in the tone and content of the message.
[0186] Step 6:
[0187] The server generates the final processing results as a document, providing a customized report based on sentiment. In this step, all system processing and sentiment analysis outputs are aggregated to create a user-friendly report. The final document is sent to the user via a terminal.
[0188] (Application Example 2)
[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0190] The goal is to resolve issues that hinder the smooth progress of the payment process due to stress and confusion experienced by users when using electronic payment systems. In particular, emotional states can be a barrier to payment, so it is necessary to eliminate factors that disrupt a smooth user experience.
[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0192] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for analyzing the user's facial expressions and voice tone in real time and generating emotional information. This makes it possible to provide an appropriate interface and response that takes into account the user's emotional state, and to facilitate the payment process.
[0193] "Paper media" refers to physical documents and forms on which information is printed, and it represents the starting point of information before it is digitized.
[0194] "Digital data" refers to data in an electronic format converted from paper media, and in a form that can be processed by a computer system.
[0195] "Processing information" refers to specific information extracted from digital data that is required for business purposes or specific applications.
[0196] The "optimal processing method" is a method or process selected based on processing information to maximize operational efficiency.
[0197] "Real-time error handling" means immediately recognizing problems that arise during processing and responding appropriately.
[0198] "Analyzing the user's facial expressions and voice tone in real time" refers to using multimodal AI to evaluate and analyze the user's facial movements and voice characteristics on the spot.
[0199] "Emotional information" refers to data that represents a user's psychological or emotional state, derived from their analyzed facial expressions and voice tone.
[0200] "Adjusting interaction content" means dynamically changing communication methods, designs, and messages based on the user's emotional state.
[0201] "Generating and outputting as a document" refers to the process of compiling the final processing results into a report or record using text, charts, and other formats, and presenting it to the user or system.
[0202] The system for carrying out this invention is constructed by integrating multiple hardware and software components. The terminal is equipped with a scanner, camera, and microphone, and converts information from paper documents into digital data, collecting the user's facial expressions and voice in real time. The server receives this data and performs the following processing.
[0203] The server first uses optical character recognition (OCR) technology to convert scanned information into text data. Natural language processing (NLP) technology is then used to extract the necessary information from this text data. Next, emotional information is analyzed from the user's facial expressions and voice through a generative AI model and an emotion engine. This emotional information is used to adjust the content of the interaction with the user.
[0204] Specifically, when a user attempts to make an electronic payment, the system captures the user's facial expression with a camera and collects their voice tone with a microphone. This data is sent to a server and analyzed in real time. If the user appears confused, the application interface is modified to be more user-friendly, and a gentle guide message such as, "Are you okay? If you need assistance, we can call a staff member," is displayed. This reduces user stress.
[0205] To give a concrete example, imagine a user trying out a new QR code (registered trademark) payment method for the first time in a crowded cafe. If the generation AI model detects that the user is confused, the prompt message "If the user is frowning, generate a gentle, guiding message" will be used. This prompt allows the system to appropriately support the user and facilitate a smooth payment process.
[0206] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0207] Step 1:
[0208] The terminal scans the user's paper documents and saves them as digital image data. This image data becomes the input and is sent to the server.
[0209] Step 2:
[0210] The server performs optical character recognition (OCR) on the received image data and extracts text data from it. The input is image data, and the extracted text data is the output.
[0211] Step 3:
[0212] The server uses natural language processing (NLP) techniques based on text data to identify and extract the necessary processing information. The input is text data, and the output is the processed information.
[0213] Step 4:
[0214] The device captures the user's facial expressions with a camera and records their voice with a microphone. This raw data is sent to the server in real time. The input consists of the user's facial expressions and voice.
[0215] Step 5:
[0216] The server uses a generative AI model and an emotion engine to analyze the received facial and audio data. This generates the user's emotional information. The input is facial and audio data, and the output is emotional information.
[0217] Step 6:
[0218] The server dynamically modifies the user interface and uses prompts to generate appropriate messages based on the analyzed sentiment information. The input is sentiment information, and the output is a user-facing message generated based on the prompts.
[0219] Step 7:
[0220] The terminal displays emotion-based messages received from the server to the user and provides guidance as needed. The output is a message that the user sees directly, thereby creating a reassuring interaction.
[0221] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0222] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0223] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0224] [Second Embodiment]
[0225] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0226] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0227] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0228] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0229] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0230] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0231] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0232] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0233] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0234] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0235] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0236] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0237] This invention relates to a system that efficiently digitizes information from paper documents and processes it appropriately. This system converts information from paper documents into digital data using optical character recognition technology and extracts necessary processing information from the digital data using natural language processing technology. Based on this, the user, terminal, or server selects the optimal processing method and handles processing errors in real time.
[0238] System Implementation Example
[0239] The user places a paper invoice in the scanner and presses the scan button. The scanned information is sent to the server via the terminal.
[0240] The server applies optical character recognition (OCR) technology to the received scan data and extracts text data from the image.
[0241] The extracted text data is parsed on the server using natural language processing technology, and the information necessary for processing (e.g., amount, payment due date, recipient information, etc.) is extracted.
[0242] Based on this information, the server uses a multimodal generative AI model to determine the optimal processing method. For example, it can select a method for processing a large number of small payments in a batch.
[0243] Furthermore, the server has a real-time error handling function to prepare for errors that may occur during processing, and notifies the user of any inconsistencies that are found and requesting corrections.
[0244] Finally, the server generates a document based on the completed information and sends the necessary notification to the user and relevant parties.
[0245] Specific example
[0246] For example, when a local government manages the payment of various taxes from its residents, this system can be used to digitize paper payment slips and automatically analyze tax classification and payment status. Users (e.g., city hall employees) can use this system to efficiently process a large amount of payment data and quickly issue payment notices. As described above, the present invention can reduce errors caused by manual work and expedite operations.
[0247] The following describes the processing flow.
[0248] Step 1:
[0249] The user places a paper invoice or payment slip into the scanner and presses the scan button. This action captures the information from the paper document as a digital image on the device.
[0250] Step 2:
[0251] The terminal sends the scanned image it has captured to the server. The server receives this digital image and prepares for the next processing step.
[0252] Step 3:
[0253] The server uses optical character recognition (OCR) technology on the received digital image to extract text data from the image data. This digitizes the text information that was printed on paper.
[0254] Step 4:
[0255] The server uses natural language processing technology to analyze and extract necessary processing information (e.g., amount, payment deadline, recipient information, etc.) from the extracted text data.
[0256] Step 5:
[0257] The server selects the optimal processing method based on the analysis results. Here, it utilizes a multimodal generative AI model to make decisions such as processing a large number of small payments in the most cost-effective way.
[0258] Step 6:
[0259] The server performs data integrity checks while processing is in progress. If errors or inconsistencies are detected, error handling is performed in real time, and the user is notified to request correction.
[0260] Step 7:
[0261] After the server verifies that all data is accurate, it proceeds with the actual payment processing. This includes transfer operations through payment systems and bank APIs.
[0262] Step 8:
[0263] Once the server completes the payment, it generates necessary documents, such as a payment notification, based on the results. The generated documents are sent to the user and relevant parties as needed, or made available for the user to access.
[0264] (Example 1)
[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0266] The need to efficiently digitize paper-based information, quickly and accurately extract processing data, and determine the optimal processing method is paramount. However, traditional methods rely heavily on manual processes, leading to errors and processing delays. Furthermore, handling errors in real time and ensuring users receive necessary notifications is crucial. Additionally, a lack of automation technology for these processes has resulted in reduced operational efficiency.
[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0268] In this invention, the server includes a device for converting information from paper documents into digital data, a device for extracting processing information, and a device for determining the optimal processing method. This enables improved operational efficiency through the digitization of information, rapid handling and notification of errors, and the proposal of efficient processing methods using a generative AI model.
[0269] "Paper media" refers to media in which information is recorded on physical paper, such as documents and printed materials.
[0270] "Digital data" refers to information that has been converted into a format that can be processed by computers and electronic devices.
[0271] "Processing information" refers to information extracted from digital data that is necessary for specific tasks or operations.
[0272] Optical character recognition (OCR) is a technology that automatically reads printed or handwritten text information and digitizes it using machines.
[0273] "Natural language processing technology" refers to technologies used to analyze, interpret, or generate human language using computers.
[0274] A "generative AI model" is a machine learning model that uses artificial intelligence to generate information, and in particular, it has the function of suggesting the most optimal processing method or solution.
[0275] Error handling is the process of detecting errors and inconsistencies that occur during processing and taking appropriate action to address them.
[0276] "Notification" refers to the act or means of communicating processing results or error information to users or relevant parties.
[0277] This invention is a system that efficiently digitizes information from paper documents and processes it optimally. The user places a paper document, such as an invoice, into the scanner and presses the scan button to digitally input the information. This input digital image is then transmitted to a server via a terminal.
[0278] The server first extracts text data from the received digital image using optical character recognition (OCR) software. This is done using, for example, Tesseract or a commercial API. This process converts the information from image format to text format.
[0279] The obtained text data is then analyzed using natural language processing (NLP) techniques. The server extracts the necessary processing information, such as the amount and payment due date, through this analysis. Examples of the technologies used include NLTK and SpaCy in Python.
[0280] Subsequently, using a generative AI model, the server proposes an optimal processing method. For example, using a generative model from OpenAI, a method for efficient payment processing is presented. An example of the prompt text used in this case could be "Extract the amount and payment due date from multiple invoices and propose an optimal payment processing method."
[0281] Based on the information extracted and the processing method proposed in this way, the server handles errors in real time and requests corrections from the user as needed. Finally, the processing result is generated as text and sent to the user and relevant parties through the notification system.
[0282] As an example, when a local government processes tax payments from residents, a process of digitizing paper payment slips and automatically analyzing the payment status by tax category is carried out. In this way, the system achieves work efficiency improvement and error reduction.
[0283] The flow of the specific process in Example 1 will be described using FIG. 11.
[0284] Step 1:
[0285] The user sets the paper medium on the scanner and presses the scan button. The purpose is to have paper medium (e.g., invoice) as input and capture it as a digital image by the terminal. As a specific operation, the scanner converts the physical information of the paper into a digital signal, and the terminal acquires this as image data.
[0286] Step 2:
[0287] The terminal sends the acquired digital image data to the server. The data to be sent is image data generated by the scanner. The terminal first converts the data format appropriately and then sends it to the server via the network.
[0288] Step 3:
[0289] The server applies Optical Character Recognition (OCR) to the received image data to extract text data. The input is image data, and the output is text data. Specifically, the OCR software identifies characters from the image and extracts them as strings.
[0290] Step 4:
[0291] The server uses text data to analyze information using natural language processing (NLP). The input here is the text data obtained in step 3, and the output is structured information necessary for processing (e.g., amount, payment due date). The specific operation is a process of analyzing the text using an NLP library, filtering, and extracting information.
[0292] Step 5:
[0293] The server uses a generative AI model to determine the optimal processing method. The input is the structured information obtained in step 4, and the output is a suggestion of processing steps and methods. The server sends prompts to the AI model, for example, "Extract the amount and payment due date from the invoice and suggest the optimal payment processing method."
[0294] Step 6:
[0295] The server handles errors that may occur during processing in real time. Input is data related to any errors or inconsistencies during processing, and output is notifications to the user or instructions for error correction. Specifically, an error handling module on the server detects errors and takes appropriate action.
[0296] Step 7:
[0297] The server generates the final processing results as a document and sends notifications to the user and relevant parties. The input is the processing method and its result determined in step 5, and the output is document data in the form of a report or notice. Specifically, a document generation tool is used to distribute the information via email or other means of communication.
[0298] (Application Example 1)
[0299] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0300] The challenges include the inefficiency of managing information using paper documents and the lack of efficient processing of customer service that requires immediate attention in physical stores. In particular, tasks related to the use of discount coupons and loyalty cards require the digitalization of information and immediate decision-making. This leads to problems such as increased errors and time loss due to manual work.
[0301] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0302] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information, and means for visualizing and suggesting information in real time through a portable visual device. This enables faster information processing in physical stores and improves the immediacy and accuracy of customer service.
[0303] "Paper media" refers to a form of media on which information is printed or written for recording purposes.
[0304] "Digital data" refers to a collection of information that has been converted into a format that can be processed electronically.
[0305] A "portable visual device" is a device that can be carried by a user and displays information visually.
[0306] "Optical character recognition technology" is a technology that converts printed or handwritten characters into digital data.
[0307] "Natural language processing technology" is a technology that analyzes and processes human language using a computer.
[0308] "Real-time error handling" is a method of immediately detecting errors that occur during processing and performing corrections or responses.
[0309] A "generative AI model" is an artificial intelligence model that generates appropriate predictions or proposals based on input data.
[0310] "Visualization" is a process of visually displaying digital data or information to enable human recognition.
[0311] "Proposal" is an act of showing appropriate answers or means in advance for specific purposes or conditions.
[0312] In the form of implementing this invention, first, smart glasses are used as a portable visual device. The user captures the information described in the paper medium using the camera function of the smart glasses. The captured information is converted into digital data by the optical character recognition technology executed inside the smart glasses.
[0313] The converted digital data is transmitted to the server via the terminal. At the server, the necessary processing information is extracted from the digital data using natural language processing technology. The information obtained in this process is utilized to select an optimal processing method while using the generative AI model.
[0314] Furthermore, the server has real-time error handling capabilities, instantly detecting any errors that may occur during processing and instructing the user to correct them. This enables users to process information quickly and accurately.
[0315] As a concrete example, consider customer service operations in a physical store. When a customer presents a paper coupon, the user can instantly verify the coupon details through smart glasses and determine whether the discount is applicable. This procedure allows the store to improve customer satisfaction.
[0316] Based on the generated AI model, an example of a prompt message used to notify the user of processing recommendations would be the text, "Can I apply this coupon? We will give you a 50% discount on the product price," displayed in the user's field of view.
[0317] This invention makes it possible to reduce the effort required for managing information on paper, improve operational efficiency, and enhance the quality of service provided to customers.
[0318] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0319] Step 1:
[0320] The user uses the camera function of smart glasses to capture information from paper documents. The input is visual information from the paper document, and the output is a digital image file. This process allows information printed on paper to be captured by an electronic device.
[0321] Step 2:
[0322] The terminal processes the captured image file using optical character recognition (OCR) technology to extract text data. The input is the digital image obtained in step 1, and the output is text data. OCR processing converts the characters in the image into digital data.
[0323] Step 3:
[0324] The terminal sends the extracted text data to the server. The input is text data, and the output is data transfer to the server. This process provides the server with text information for processing.
[0325] Step 4:
[0326] The server uses natural language processing (NLP) techniques to extract important processing information from text data. The input is text data, and the output is specific processing information (e.g., coupon details, discount rate, eligibility). Through NLP, the information necessary for digital transactions is organized.
[0327] Step 5:
[0328] The server uses a generative AI model to make optimal suggestions and decisions based on the extracted processing information. The input is processing information, and the output is the generated suggestions or action recommendations. In this step, the AI model constructs the suggestions and prepares them to be fed into the next step.
[0329] Step 6:
[0330] The server detects potential errors during processing in real time and notifies the user. Inputs are processing information and the current execution status, while outputs are error messages or instructions for correction. Error handling ensures that the user is immediately informed of any problems.
[0331] Step 7:
[0332] The server displays the final processing results and suggestions on the user's smart glasses. The input is the optimal suggestions and modifications. The output is the data displayed in the user's field of vision. The final display allows the user to take appropriate action immediately.
[0333] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0334] This invention is a system designed to streamline payment and information processing operations for companies and local governments. In addition to digitizing conventional paper-based information, it recognizes user emotions and reflects them in business processes. It not only processes information based on paper documents but also incorporates an emotion engine to optimize the processing flow while judging the user's emotional state in real time.
[0335] System Implementation Example
[0336] This system is configured as follows: The user digitizes information from paper documents using a scanner and sends it to the server. The server uses optical character recognition (OCR) technology to convert the image data into text data, and then uses natural language processing (NLP) technology to extract the information necessary for processing.
[0337] Based on the extracted information, the server selects the optimal processing method through a multimodal AI model and handles any errors that may occur during processing in real time. During this process, the server analyzes the user's emotions using an emotion engine. This emotion information is used to adjust the tone and content of interactions and messages in processing selection and error handling.
[0338] Specific example
[0339] For example, if a user encounters a specific error while processing data related to a payment slip, the emotion engine recognizes frustration or anxiety based on the user's facial expressions and tone of voice. The server uses this emotion recognition information to generate a message containing more user-friendly language and gentler instructions during error handling, and notifies the user accordingly. By responding in accordance with emotions in this way, user stress can be reduced and work efficiency can be improved.
[0340] Finally, the completed processing results are generated as a document, including a personalized message based on the user's emotions as needed, and provided to the user. Thus, the present invention is a comprehensive system that not only improves operational efficiency but also enhances user satisfaction.
[0341] The following describes the processing flow.
[0342] Step 1:
[0343] The user places a paper invoice into the scanner and starts scanning. The scanner captures the image data into the terminal.
[0344] Step 2:
[0345] The terminal sends the scanned image data to the server. The server receives this digital image and prepares it for the next processing step.
[0346] Step 3:
[0347] The server applies optical character recognition (OCR) technology to the received image data to extract text data. In this process, the text information from paper documents is converted into a digital format.
[0348] Step 4:
[0349] The server uses natural language processing technology to analyze the text data and extract the information necessary for processing (e.g., amount, payment due date, recipient information).
[0350] Step 5:
[0351] The server uses a multimodal AI model to select the optimal processing method based on the extracted information. For example, it might consider a method for processing multiple payments in a single batch.
[0352] Step 6:
[0353] When an error occurs on the server during processing, the emotion engine is used to analyze the user's emotions. For example, facial recognition and voice analysis are used to determine whether the user is confused or distressed.
[0354] Step 7:
[0355] The server adjusts the tone and content of error messages based on emotional information and sends appropriate support or correction requests to the user. This allows users to resolve problems without stress.
[0356] Step 8:
[0357] The server generates a document based on the processing results and adds a personalized message based on emotions as needed. The generated document is then provided to the user.
[0358] Step 9:
[0359] The user reviews the provided documents and confirms that the processing was completed successfully. This entire process allows the user to complete payment transactions efficiently.
[0360] (Example 2)
[0361] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0362] Conventional information processing systems can be inefficient in the process of digitizing paper-based information and in subsequent error handling. Furthermore, uniform responses that disregard user feelings hinder the improvement of the user experience. There is a need to resolve these issues and simultaneously improve both processing efficiency and user satisfaction.
[0363] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0364] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for recognizing user emotions and adjusting interactions based on those emotions. This enables efficient information processing and user-friendly error handling.
[0365] "Paper media" refers to a means of representing information that is printed or written on physical paper.
[0366] "Digital data" refers to information that has been converted into a format that can be processed and stored by electronic devices such as computers.
[0367] "Processing information" refers to information that includes data and instructions necessary to perform a specific task or operation.
[0368] "Error handling" is the process of detecting errors and problems that occur during the information processing process and taking appropriate action.
[0369] "User emotions" refers to the emotional reactions and states that users exhibit while using the system.
[0370] "Interaction" refers to the exchange of information and communication that takes place between a system and a user.
[0371] This system efficiently converts information from paper documents into digital data and implements a process that optimizes processing based on that information. A distinctive feature of this system is its ability to recognize user emotions in real time and utilize that information to optimize processing.
[0372] First, the user uses a scanner to digitize the information on paper. This digital data is then processed using OCR (Optical Character Recognition) technology. The server uses this OCR technology to extract text data from the scanned image data. Typically, commercially available OCR software is used for this purpose.
[0373] Next, the server uses NLP (Natural Language Processing) technology to extract necessary processing information from the text data. This technology is used to identify specific keywords and instructions from text. Common natural language processing libraries can be applied to the software used.
[0374] Furthermore, the server determines the optimal processing method as needed through a generative AI model. This AI model learns from the dataset and provides flexible processing methods that adapt to the situation. In addition, if an error occurs, the emotion engine analyzes the user's emotions and assists in taking appropriate action. This emotion engine can read emotions from the user's facial expressions and voice.
[0375] A concrete example is the emotional response to error messages that occur when a user attempts to digitize their tax payment process. This involves displaying reassuring messages based on emotional information to support the user in continuing the process. For example, a prompt might read, "The payment slip could not be read. Please try again. If you have further problems, please contact support."
[0376] The overall flow of this system is designed to streamline and user-friendly information processing operations for companies and local governments.
[0377] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0378] Step 1:
[0379] The user places a paper document in the scanner and begins the digitization process. The scanner acquires image data from the paper document and generates that data as an electronic image file. As output of this process, the image file is sent to the server.
[0380] Step 2:
[0381] The server applies Optical Character Recognition (OCR) technology to the received image file. The input is image data. In this step, the server identifies the text within the image and converts the visual data into text data. The output of this process is readable electronic text.
[0382] Step 3:
[0383] The server analyzes electronic text data using natural language processing (NLP) techniques. Text data is provided as input. The server scans the entire document and extracts the necessary processing information. Relevant keywords and phrases are retrieved as NLP output and stored in a database.
[0384] Step 4:
[0385] The server uses a generative AI model to select the optimal processing plan based on the extracted information. This model operates using a pre-trained algorithm to derive the most efficient processing steps from the input data. The output consists of specific processing steps and necessary configuration information.
[0386] Step 5:
[0387] The server monitors for errors in real time as processing progresses and performs error handling as needed. Simultaneously, it uses an emotion engine to analyze the input user's video and audio data and recognize their emotions. The output emotion data is reflected in the tone and content of the message.
[0388] Step 6:
[0389] The server generates the final processing results as a document, providing a customized report based on sentiment. In this step, all system processing and sentiment analysis outputs are aggregated to create a user-friendly report. The final document is sent to the user via a terminal.
[0390] (Application Example 2)
[0391] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0392] The goal is to resolve issues that hinder the smooth progress of the payment process due to stress and confusion experienced by users when using electronic payment systems. In particular, emotional states can be a barrier to payment, so it is necessary to eliminate factors that disrupt a smooth user experience.
[0393] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0394] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for analyzing the user's facial expressions and voice tone in real time and generating emotional information. This makes it possible to provide an appropriate interface and response that takes into account the user's emotional state, and to facilitate the payment process.
[0395] "Paper media" refers to physical documents and forms on which information is printed, and it represents the starting point of information before it is digitized.
[0396] "Digital data" refers to data in an electronic format converted from paper media, and in a form that can be processed by a computer system.
[0397] "Processing information" refers to specific information extracted from digital data that is required for business purposes or specific applications.
[0398] The "optimal processing method" is a method or process selected based on processing information to maximize operational efficiency.
[0399] "Real-time error handling" means immediately recognizing problems that arise during processing and responding appropriately.
[0400] "Analyzing the user's facial expressions and voice tone in real time" refers to using multimodal AI to evaluate and analyze the user's facial movements and voice characteristics on the spot.
[0401] "Emotional information" refers to data that represents a user's psychological or emotional state, derived from their analyzed facial expressions and voice tone.
[0402] "Adjusting interaction content" means dynamically changing communication methods, designs, and messages based on the user's emotional state.
[0403] "Generating and outputting as a document" refers to the process of compiling the final processing results into a report or record using text, charts, and other formats, and presenting it to the user or system.
[0404] The system for carrying out this invention is constructed by integrating multiple hardware and software components. The terminal is equipped with a scanner, camera, and microphone, and converts information from paper documents into digital data, collecting the user's facial expressions and voice in real time. The server receives this data and performs the following processing.
[0405] The server first uses optical character recognition (OCR) technology to convert scanned information into text data. Natural language processing (NLP) technology is then used to extract the necessary information from this text data. Next, emotional information is analyzed from the user's facial expressions and voice through a generative AI model and an emotion engine. This emotional information is used to adjust the content of the interaction with the user.
[0406] Specifically, when a user attempts to make an electronic payment, the system captures the user's facial expression with a camera and collects their voice tone with a microphone. This data is sent to a server and analyzed in real time. If the user appears confused, the application interface is modified to be more user-friendly, and a gentle guide message such as, "Are you okay? If you need assistance, we can call a staff member," is displayed. This reduces user stress.
[0407] To give a concrete example, imagine a user trying a new QR code payment method for the first time in a crowded cafe. If the generation AI model detects that the user is confused, the prompt message "Generate a gentle guidance message if the user is frowning" will be used. This prompt allows the system to appropriately support the user and facilitate a smooth payment.
[0408] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0409] Step 1:
[0410] The terminal scans the user's paper documents and saves them as digital image data. This image data becomes the input and is sent to the server.
[0411] Step 2:
[0412] The server performs optical character recognition (OCR) on the received image data and extracts text data from it. The input is image data, and the extracted text data is the output.
[0413] Step 3:
[0414] The server uses natural language processing (NLP) techniques based on text data to identify and extract the necessary processing information. The input is text data, and the output is the processed information.
[0415] Step 4:
[0416] The device captures the user's facial expressions with a camera and records their voice with a microphone. This raw data is sent to the server in real time. The input consists of the user's facial expressions and voice.
[0417] Step 5:
[0418] The server uses a generative AI model and an emotion engine to analyze the received facial and audio data. This generates the user's emotional information. The input is facial and audio data, and the output is emotional information.
[0419] Step 6:
[0420] The server dynamically modifies the user interface and uses prompts to generate appropriate messages based on the analyzed sentiment information. The input is sentiment information, and the output is a user-facing message generated based on the prompts.
[0421] Step 7:
[0422] The terminal displays emotion-based messages received from the server to the user and provides guidance as needed. The output is a message that the user sees directly, thereby creating a reassuring interaction.
[0423] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0424] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0425] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0426] [Third Embodiment]
[0427] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0428] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0429] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0430] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0431] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0432] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0433] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0434] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0435] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0436] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0437] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0438] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0439] This invention relates to a system that efficiently digitizes information from paper documents and processes it appropriately. This system converts information from paper documents into digital data using optical character recognition technology and extracts necessary processing information from the digital data using natural language processing technology. Based on this, the user, terminal, or server selects the optimal processing method and handles processing errors in real time.
[0440] System Implementation Example
[0441] The user places a paper invoice in the scanner and presses the scan button. The scanned information is sent to the server via the terminal.
[0442] The server applies optical character recognition (OCR) technology to the received scan data and extracts text data from the image.
[0443] The extracted text data is parsed on the server using natural language processing technology, and the information necessary for processing (e.g., amount, payment due date, recipient information, etc.) is extracted.
[0444] Based on this information, the server uses a multimodal generative AI model to determine the optimal processing method. For example, it can select a method for processing a large number of small payments in a batch.
[0445] Furthermore, the server has a real-time error handling function to prepare for errors that may occur during processing, and notifies the user of any inconsistencies that are found and requesting corrections.
[0446] Finally, the server generates a document based on the completed information and sends the necessary notification to the user and relevant parties.
[0447] Specific example
[0448] For example, when a local government manages the payment of various taxes from its residents, this system can be used to digitize paper payment slips and automatically analyze tax classification and payment status. Users (e.g., city hall employees) can use this system to efficiently process a large amount of payment data and quickly issue payment notices. As described above, the present invention can reduce errors caused by manual work and expedite operations.
[0449] The following describes the processing flow.
[0450] Step 1:
[0451] The user places a paper invoice or payment slip into the scanner and presses the scan button. This action captures the information from the paper document as a digital image on the device.
[0452] Step 2:
[0453] The terminal sends the scanned image it has captured to the server. The server receives this digital image and prepares for the next processing step.
[0454] Step 3:
[0455] The server uses optical character recognition (OCR) technology on the received digital image to extract text data from the image data. This digitizes the text information that was printed on paper.
[0456] Step 4:
[0457] The server uses natural language processing technology to analyze and extract necessary processing information (e.g., amount, payment deadline, recipient information, etc.) from the extracted text data.
[0458] Step 5:
[0459] The server selects the optimal processing method based on the analysis results. Here, it utilizes a multimodal generative AI model to make decisions such as processing a large number of small payments in the most cost-effective way.
[0460] Step 6:
[0461] The server performs data integrity checks while processing is in progress. If errors or inconsistencies are detected, error handling is performed in real time, and the user is notified to request correction.
[0462] Step 7:
[0463] After the server verifies that all data is accurate, it proceeds with the actual payment processing. This includes transfer operations through payment systems and bank APIs.
[0464] Step 8:
[0465] Once the server completes the payment, it generates necessary documents, such as a payment notification, based on the results. The generated documents are sent to the user and relevant parties as needed, or made available for the user to access.
[0466] (Example 1)
[0467] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0468] The need to efficiently digitize paper-based information, quickly and accurately extract processing data, and determine the optimal processing method is paramount. However, traditional methods rely heavily on manual processes, leading to errors and processing delays. Furthermore, handling errors in real time and ensuring users receive necessary notifications is crucial. Additionally, a lack of automation technology for these processes has resulted in reduced operational efficiency.
[0469] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0470] In this invention, the server includes a device for converting information from paper documents into digital data, a device for extracting processing information, and a device for determining the optimal processing method. This enables improved operational efficiency through the digitization of information, rapid handling and notification of errors, and the proposal of efficient processing methods using a generative AI model.
[0471] "Paper media" refers to media in which information is recorded on physical paper, such as documents and printed materials.
[0472] "Digital data" refers to information that has been converted into a format that can be processed by computers and electronic devices.
[0473] "Processing information" refers to information extracted from digital data that is necessary for specific tasks or operations.
[0474] Optical character recognition (OCR) is a technology that automatically reads printed or handwritten text information and digitizes it using machines.
[0475] "Natural language processing technology" refers to technologies used to analyze, interpret, or generate human language using computers.
[0476] A "generative AI model" is a machine learning model that uses artificial intelligence to generate information, and in particular, it has the function of suggesting the most optimal processing method or solution.
[0477] Error handling is the process of detecting errors and inconsistencies that occur during processing and taking appropriate action to address them.
[0478] "Notification" refers to the act or means of communicating processing results or error information to users or relevant parties.
[0479] This invention is a system that efficiently digitizes information from paper documents and processes it optimally. The user places a paper document, such as an invoice, into the scanner and presses the scan button to digitally input the information. This input digital image is then transmitted to a server via a terminal.
[0480] The server first extracts text data from the received digital image using optical character recognition (OCR) software. This is done using, for example, Tesseract or a commercial API. This process converts the information from image format to text format.
[0481] The obtained text data is then analyzed using natural language processing (NLP) techniques. The server extracts necessary processing information, such as amounts and payment due dates, from this analysis. Examples of technologies used include Python's NLTK and SpaCy.
[0482] Next, the server uses a generative AI model to propose the optimal processing method. For example, it might use an OpenAI generative model to suggest an efficient payment processing method. An example of a prompt used in this process might be, "Extract the amounts and payment due dates from multiple invoices and propose the optimal payment processing method."
[0483] Based on the information extracted and the proposed processing methods, the server handles errors in real time and requests corrections from the user as needed. Finally, the processing results are generated as text and sent to the user and relevant parties through the notification system.
[0484] For example, when local governments process tax payments from residents, paper payment slips are digitized, and a process is implemented to automatically analyze the payment status by tax category. In this way, the system achieves increased efficiency and reduced errors.
[0485] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0486] Step 1:
[0487] The user places a paper document into the scanner and presses the scan button. The input is a paper document (e.g., an invoice), and the goal is to capture it as a digital image on the terminal. Specifically, the scanner converts the physical information of the paper into a digital signal, which the terminal then acquires as image data.
[0488] Step 2:
[0489] The terminal sends the acquired digital image data to the server. The data to be sent is image data generated by the scanner. The terminal first converts the data format appropriately and then sends it to the server via the network.
[0490] Step 3:
[0491] The server applies Optical Character Recognition (OCR) to the received image data to extract text data. The input is image data, and the output is text data. Specifically, the OCR software identifies characters from the image and extracts them as strings.
[0492] Step 4:
[0493] The server uses text data to analyze information using natural language processing (NLP). The input here is the text data obtained in step 3, and the output is structured information necessary for processing (e.g., amount, payment due date). The specific operation is a process of analyzing the text using an NLP library, filtering, and extracting information.
[0494] Step 5:
[0495] The server uses a generative AI model to determine the optimal processing method. The input is the structured information obtained in step 4, and the output is a suggestion of processing steps and methods. The server sends prompts to the AI model, for example, "Extract the amount and payment due date from the invoice and suggest the optimal payment processing method."
[0496] Step 6:
[0497] The server handles errors that may occur during processing in real time. Input is data related to any errors or inconsistencies during processing, and output is notifications to the user or instructions for error correction. Specifically, an error handling module on the server detects errors and takes appropriate action.
[0498] Step 7:
[0499] The server generates the final processing results as a document and sends notifications to the user and relevant parties. The input is the processing method and its result determined in step 5, and the output is document data in the form of a report or notice. Specifically, a document generation tool is used to distribute the information via email or other means of communication.
[0500] (Application Example 1)
[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] The challenges include the inefficiency of managing information using paper documents and the lack of efficient processing of customer service that requires immediate attention in physical stores. In particular, tasks related to the use of discount coupons and loyalty cards require the digitalization of information and immediate decision-making. This leads to problems such as increased errors and time loss due to manual work.
[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0504] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information, and means for visualizing and suggesting information in real time through a portable visual device. This enables faster information processing in physical stores and improves the immediacy and accuracy of customer service.
[0505] "Paper media" refers to a form of media on which information is printed or written for recording purposes.
[0506] "Digital data" refers to a collection of information that has been converted into a format that can be processed electronically.
[0507] A "portable visual device" is a device that is portable by the user and used to display information visually.
[0508] "Optical character recognition technology" is a technology that converts printed or handwritten characters into digital data.
[0509] "Natural language processing technology" is a technology that uses computers to analyze and process human language.
[0510] "Real-time error handling" is a method of immediately detecting and correcting errors that occur during processing.
[0511] A "generative AI model" is an artificial intelligence model that generates appropriate predictions and suggestions based on input data.
[0512] "Visualization" is the process of visually displaying digital data and information, making it possible for humans to recognize it.
[0513] A "proposal" is the act of providing appropriate answers or means in advance for a specific purpose or condition.
[0514] In an embodiment of this invention, smart glasses are first used as a portable visual device. The user captures information written on paper using the camera function of the smart glasses. The captured information is converted into digital data by optical character recognition technology performed within the smart glasses.
[0515] The converted digital data is sent to the server via the terminal. The server uses natural language processing technology to extract necessary processing information from the digital data. The information obtained during this process is then used to select the optimal processing method using a generative AI model.
[0516] Furthermore, the server has real-time error handling capabilities, instantly detecting any errors that may occur during processing and instructing the user to correct them. This enables users to process information quickly and accurately.
[0517] As a concrete example, consider customer service operations in a physical store. When a customer presents a paper coupon, the user can instantly verify the coupon details through smart glasses and determine whether the discount is applicable. This procedure allows the store to improve customer satisfaction.
[0518] Based on the generated AI model, an example of a prompt message used to notify the user of processing recommendations would be the text, "Can I apply this coupon? We will give you a 50% discount on the product price," displayed in the user's field of view.
[0519] This invention makes it possible to reduce the effort required for managing information on paper, improve operational efficiency, and enhance the quality of service provided to customers.
[0520] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0521] Step 1:
[0522] The user uses the camera function of smart glasses to capture information from paper documents. The input is visual information from the paper document, and the output is a digital image file. This process allows information printed on paper to be captured by an electronic device.
[0523] Step 2:
[0524] The terminal processes the captured image file using optical character recognition (OCR) technology to extract text data. The input is the digital image obtained in step 1, and the output is text data. OCR processing converts the characters in the image into digital data.
[0525] Step 3:
[0526] The terminal sends the extracted text data to the server. The input is text data, and the output is data transfer to the server. This process provides the server with text information for processing.
[0527] Step 4:
[0528] The server uses natural language processing (NLP) techniques to extract important processing information from text data. The input is text data, and the output is specific processing information (e.g., coupon details, discount rate, eligibility). Through NLP, the information necessary for digital transactions is organized.
[0529] Step 5:
[0530] The server uses a generative AI model to make optimal suggestions and decisions based on the extracted processing information. The input is processing information, and the output is the generated suggestions or action recommendations. In this step, the AI model constructs the suggestions and prepares them to be fed into the next step.
[0531] Step 6:
[0532] The server detects potential errors during processing in real time and notifies the user. Inputs are processing information and the current execution status, while outputs are error messages or instructions for correction. Error handling ensures that the user is immediately informed of any problems.
[0533] Step 7:
[0534] The server displays the final processing results and suggestions on the user's smart glasses. The input is the optimal suggestions and modifications. The output is the data displayed in the user's field of vision. The final display allows the user to take appropriate action immediately.
[0535] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0536] This invention is a system designed to streamline payment and information processing operations for companies and local governments. In addition to digitizing conventional paper-based information, it recognizes user emotions and reflects them in business processes. It not only processes information based on paper documents but also incorporates an emotion engine to optimize the processing flow while judging the user's emotional state in real time.
[0537] System Implementation Example
[0538] This system is configured as follows: The user digitizes information from paper documents using a scanner and sends it to the server. The server uses optical character recognition (OCR) technology to convert the image data into text data, and then uses natural language processing (NLP) technology to extract the information necessary for processing.
[0539] Based on the extracted information, the server selects the optimal processing method through a multimodal AI model and handles any errors that may occur during processing in real time. During this process, the server analyzes the user's emotions using an emotion engine. This emotion information is used to adjust the tone and content of interactions and messages in processing selection and error handling.
[0540] Specific example
[0541] For example, if a user encounters a specific error while processing data related to a payment slip, the emotion engine recognizes frustration or anxiety based on the user's facial expressions and tone of voice. The server uses this emotion recognition information to generate a message containing more user-friendly language and gentler instructions during error handling, and notifies the user accordingly. By responding in accordance with emotions in this way, user stress can be reduced and work efficiency can be improved.
[0542] Finally, the completed processing results are generated as a document, including a personalized message based on the user's emotions as needed, and provided to the user. Thus, the present invention is a comprehensive system that not only improves operational efficiency but also enhances user satisfaction.
[0543] The following describes the processing flow.
[0544] Step 1:
[0545] The user places a paper invoice into the scanner and starts scanning. The scanner captures the image data into the terminal.
[0546] Step 2:
[0547] The terminal sends the scanned image data to the server. The server receives this digital image and prepares it for the next processing step.
[0548] Step 3:
[0549] The server applies optical character recognition (OCR) technology to the received image data to extract text data. In this process, the text information from paper documents is converted into a digital format.
[0550] Step 4:
[0551] The server uses natural language processing technology to analyze the text data and extract the information necessary for processing (e.g., amount, payment due date, recipient information).
[0552] Step 5:
[0553] The server uses a multimodal AI model to select the optimal processing method based on the extracted information. For example, it might consider a method for processing multiple payments in a single batch.
[0554] Step 6:
[0555] When an error occurs on the server during processing, the emotion engine is used to analyze the user's emotions. For example, facial recognition and voice analysis are used to determine whether the user is confused or distressed.
[0556] Step 7:
[0557] The server adjusts the tone and content of error messages based on emotional information and sends appropriate support or correction requests to the user. This allows users to resolve problems without stress.
[0558] Step 8:
[0559] The server generates a document based on the processing results and adds a personalized message based on emotions as needed. The generated document is then provided to the user.
[0560] Step 9:
[0561] The user reviews the provided documents and confirms that the processing was completed successfully. This entire process allows the user to complete payment transactions efficiently.
[0562] (Example 2)
[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0564] Conventional information processing systems can be inefficient in the process of digitizing paper-based information and in subsequent error handling. Furthermore, uniform responses that disregard user feelings hinder the improvement of the user experience. There is a need to resolve these issues and simultaneously improve both processing efficiency and user satisfaction.
[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0566] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for recognizing user emotions and adjusting interactions based on those emotions. This enables efficient information processing and user-friendly error handling.
[0567] "Paper media" refers to a means of representing information that is printed or written on physical paper.
[0568] "Digital data" refers to information that has been converted into a format that can be processed and stored by electronic devices such as computers.
[0569] "Processing information" refers to information that includes data and instructions necessary to perform a specific task or operation.
[0570] "Error handling" is the process of detecting errors and problems that occur during the information processing process and taking appropriate action.
[0571] "User emotions" refers to the emotional reactions and states that users exhibit while using the system.
[0572] "Interaction" refers to the exchange of information and communication that takes place between a system and a user.
[0573] This system efficiently converts information from paper documents into digital data and implements a process that optimizes processing based on that information. A distinctive feature of this system is its ability to recognize user emotions in real time and utilize that information to optimize processing.
[0574] First, the user uses a scanner to digitize the information on paper. This digital data is then processed using OCR (Optical Character Recognition) technology. The server uses this OCR technology to extract text data from the scanned image data. Typically, commercially available OCR software is used for this purpose.
[0575] Next, the server uses NLP (Natural Language Processing) technology to extract necessary processing information from the text data. This technology is used to identify specific keywords and instructions from text. Common natural language processing libraries can be applied to the software used.
[0576] Furthermore, the server determines the optimal processing method as needed through a generative AI model. This AI model learns from the dataset and provides flexible processing methods that adapt to the situation. In addition, if an error occurs, the emotion engine analyzes the user's emotions and assists in taking appropriate action. This emotion engine can read emotions from the user's facial expressions and voice.
[0577] A concrete example is the emotional response to error messages that occur when a user attempts to digitize their tax payment process. This involves displaying reassuring messages based on emotional information to support the user in continuing the process. For example, a prompt might read, "The payment slip could not be read. Please try again. If you have further problems, please contact support."
[0578] The overall flow of this system is designed to streamline and user-friendly information processing operations for companies and local governments.
[0579] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0580] Step 1:
[0581] The user places a paper document in the scanner and begins the digitization process. The scanner acquires image data from the paper document and generates that data as an electronic image file. As output of this process, the image file is sent to the server.
[0582] Step 2:
[0583] The server applies Optical Character Recognition (OCR) technology to the received image file. The input is image data. In this step, the server identifies the text within the image and converts the visual data into text data. The output of this process is readable electronic text.
[0584] Step 3:
[0585] The server analyzes electronic text data using natural language processing (NLP) techniques. Text data is provided as input. The server scans the entire document and extracts the necessary processing information. Relevant keywords and phrases are retrieved as NLP output and stored in a database.
[0586] Step 4:
[0587] The server uses a generative AI model to select the optimal processing plan based on the extracted information. This model operates using a pre-trained algorithm to derive the most efficient processing steps from the input data. The output consists of specific processing steps and necessary configuration information.
[0588] Step 5:
[0589] The server monitors for errors in real time as processing progresses and performs error handling as needed. Simultaneously, it uses an emotion engine to analyze the input user's video and audio data and recognize their emotions. The output emotion data is reflected in the tone and content of the message.
[0590] Step 6:
[0591] The server generates the final processing results as a document, providing a customized report based on sentiment. In this step, all system processing and sentiment analysis outputs are aggregated to create a user-friendly report. The final document is sent to the user via a terminal.
[0592] (Application Example 2)
[0593] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0594] The goal is to resolve issues that hinder the smooth progress of the payment process due to stress and confusion experienced by users when using electronic payment systems. In particular, emotional states can be a barrier to payment, so it is necessary to eliminate factors that disrupt a smooth user experience.
[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0596] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for analyzing the user's facial expressions and voice tone in real time and generating emotional information. This makes it possible to provide an appropriate interface and response that takes into account the user's emotional state, and to facilitate the payment process.
[0597] "Paper media" refers to physical documents and forms on which information is printed, and it represents the starting point of information before it is digitized.
[0598] "Digital data" refers to data in an electronic format converted from paper media, and in a form that can be processed by a computer system.
[0599] "Processing information" refers to specific information extracted from digital data that is required for business purposes or specific applications.
[0600] The "optimal processing method" is a method or process selected based on processing information to maximize operational efficiency.
[0601] "Real-time error handling" means immediately recognizing problems that arise during processing and responding appropriately.
[0602] "Analyzing the user's facial expressions and voice tone in real time" refers to using multimodal AI to evaluate and analyze the user's facial movements and voice characteristics on the spot.
[0603] "Emotional information" refers to data that represents a user's psychological or emotional state, derived from their analyzed facial expressions and voice tone.
[0604] "Adjusting interaction content" means dynamically changing communication methods, designs, and messages based on the user's emotional state.
[0605] "Generating and outputting as a document" refers to the process of compiling the final processing results into a report or record using text, charts, and other formats, and presenting it to the user or system.
[0606] The system for carrying out this invention is constructed by integrating multiple hardware and software components. The terminal is equipped with a scanner, camera, and microphone, and converts information from paper documents into digital data, collecting the user's facial expressions and voice in real time. The server receives this data and performs the following processing.
[0607] The server first uses optical character recognition (OCR) technology to convert scanned information into text data. Natural language processing (NLP) technology is then used to extract the necessary information from this text data. Next, emotional information is analyzed from the user's facial expressions and voice through a generative AI model and an emotion engine. This emotional information is used to adjust the content of the interaction with the user.
[0608] Specifically, when a user attempts to make an electronic payment, the system captures the user's facial expression with a camera and collects their voice tone with a microphone. This data is sent to a server and analyzed in real time. If the user appears confused, the application interface is modified to be more user-friendly, and a gentle guide message such as, "Are you okay? If you need assistance, we can call a staff member," is displayed. This reduces user stress.
[0609] To give a concrete example, imagine a user trying a new QR code payment method for the first time in a crowded cafe. If the generation AI model detects that the user is confused, the prompt message "Generate a gentle guidance message if the user is frowning" will be used. This prompt allows the system to appropriately support the user and facilitate a smooth payment.
[0610] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0611] Step 1:
[0612] The terminal scans the user's paper documents and saves them as digital image data. This image data becomes the input and is sent to the server.
[0613] Step 2:
[0614] The server performs optical character recognition (OCR) on the received image data and extracts text data from it. The input is image data, and the extracted text data is the output.
[0615] Step 3:
[0616] The server uses natural language processing (NLP) techniques based on text data to identify and extract the necessary processing information. The input is text data, and the output is the processed information.
[0617] Step 4:
[0618] The device captures the user's facial expressions with a camera and records their voice with a microphone. This raw data is sent to the server in real time. The input consists of the user's facial expressions and voice.
[0619] Step 5:
[0620] The server uses a generative AI model and an emotion engine to analyze the received facial and audio data. This generates the user's emotional information. The input is facial and audio data, and the output is emotional information.
[0621] Step 6:
[0622] The server dynamically modifies the user interface and uses prompts to generate appropriate messages based on the analyzed sentiment information. The input is sentiment information, and the output is a user-facing message generated based on the prompts.
[0623] Step 7:
[0624] The terminal displays emotion-based messages received from the server to the user and provides guidance as needed. The output is a message that the user sees directly, thereby creating a reassuring interaction.
[0625] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0626] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0627] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0628] [Fourth Embodiment]
[0629] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0630] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0631] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0632] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0633] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0634] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0635] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0636] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0637] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0638] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0639] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0640] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0641] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0642] This invention relates to a system that efficiently digitizes information from paper documents and processes it appropriately. This system converts information from paper documents into digital data using optical character recognition technology and extracts necessary processing information from the digital data using natural language processing technology. Based on this, the user, terminal, or server selects the optimal processing method and handles processing errors in real time.
[0643] System Implementation Example
[0644] The user places a paper invoice in the scanner and presses the scan button. The scanned information is sent to the server via the terminal.
[0645] The server applies optical character recognition (OCR) technology to the received scan data and extracts text data from the image.
[0646] The extracted text data is parsed on the server using natural language processing technology, and the information necessary for processing (e.g., amount, payment due date, recipient information, etc.) is extracted.
[0647] Based on this information, the server uses a multimodal generative AI model to determine the optimal processing method. For example, it can select a method for processing a large number of small payments in a batch.
[0648] Furthermore, the server has a real-time error handling function to prepare for errors that may occur during processing, and notifies the user of any inconsistencies that are found and requesting corrections.
[0649] Finally, the server generates a document based on the completed information and sends the necessary notification to the user and relevant parties.
[0650] Specific example
[0651] For example, when a local government manages the payment of various taxes from its residents, this system can be used to digitize paper payment slips and automatically analyze tax classification and payment status. Users (e.g., city hall employees) can use this system to efficiently process a large amount of payment data and quickly issue payment notices. As described above, the present invention can reduce errors caused by manual work and expedite operations.
[0652] The following describes the processing flow.
[0653] Step 1:
[0654] The user places a paper invoice or payment slip into the scanner and presses the scan button. This action captures the information from the paper document as a digital image on the device.
[0655] Step 2:
[0656] The terminal sends the scanned image it has captured to the server. The server receives this digital image and prepares for the next processing step.
[0657] Step 3:
[0658] The server uses optical character recognition (OCR) technology on the received digital image to extract text data from the image data. This digitizes the text information that was printed on paper.
[0659] Step 4:
[0660] The server uses natural language processing technology to analyze and extract necessary processing information (e.g., amount, payment deadline, recipient information, etc.) from the extracted text data.
[0661] Step 5:
[0662] The server selects the optimal processing method based on the analysis results. Here, it utilizes a multimodal generative AI model to make decisions such as processing a large number of small payments in the most cost-effective way.
[0663] Step 6:
[0664] The server performs data integrity checks while processing is in progress. If errors or inconsistencies are detected, error handling is performed in real time, and the user is notified to request correction.
[0665] Step 7:
[0666] After the server verifies that all data is accurate, it proceeds with the actual payment processing. This includes transfer operations through payment systems and bank APIs.
[0667] Step 8:
[0668] Once the server completes the payment, it generates necessary documents, such as a payment notification, based on the results. The generated documents are sent to the user and relevant parties as needed, or made available for the user to access.
[0669] (Example 1)
[0670] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] The need to efficiently digitize paper-based information, quickly and accurately extract processing data, and determine the optimal processing method is paramount. However, traditional methods rely heavily on manual processes, leading to errors and processing delays. Furthermore, handling errors in real time and ensuring users receive necessary notifications is crucial. Additionally, a lack of automation technology for these processes has resulted in reduced operational efficiency.
[0672] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0673] In this invention, the server includes a device for converting information from paper documents into digital data, a device for extracting processing information, and a device for determining the optimal processing method. This enables improved operational efficiency through the digitization of information, rapid handling and notification of errors, and the proposal of efficient processing methods using a generative AI model.
[0674] "Paper media" refers to media in which information is recorded on physical paper, such as documents and printed materials.
[0675] "Digital data" refers to information that has been converted into a format that can be processed by computers and electronic devices.
[0676] "Processing information" refers to information extracted from digital data that is necessary for specific tasks or operations.
[0677] Optical character recognition (OCR) is a technology that automatically reads printed or handwritten text information and digitizes it using machines.
[0678] "Natural language processing technology" refers to technologies used to analyze, interpret, or generate human language using computers.
[0679] A "generative AI model" is a machine learning model that uses artificial intelligence to generate information, and in particular, it has the function of suggesting the most optimal processing method or solution.
[0680] Error handling is the process of detecting errors and inconsistencies that occur during processing and taking appropriate action to address them.
[0681] "Notification" refers to the act or means of communicating processing results or error information to users or relevant parties.
[0682] This invention is a system that efficiently digitizes information from paper documents and processes it optimally. The user places a paper document, such as an invoice, into the scanner and presses the scan button to digitally input the information. This input digital image is then transmitted to a server via a terminal.
[0683] The server first extracts text data from the received digital image using optical character recognition (OCR) software. This is done using, for example, Tesseract or a commercial API. This process converts the information from image format to text format.
[0684] The obtained text data is then analyzed using natural language processing (NLP) techniques. The server extracts necessary processing information, such as amounts and payment due dates, from this analysis. Examples of technologies used include Python's NLTK and SpaCy.
[0685] Next, the server uses a generative AI model to propose the optimal processing method. For example, it might use an OpenAI generative model to suggest an efficient payment processing method. An example of a prompt used in this process might be, "Extract the amounts and payment due dates from multiple invoices and propose the optimal payment processing method."
[0686] Based on the information extracted and the proposed processing methods, the server handles errors in real time and requests corrections from the user as needed. Finally, the processing results are generated as text and sent to the user and relevant parties through the notification system.
[0687] For example, when local governments process tax payments from residents, paper payment slips are digitized, and a process is implemented to automatically analyze the payment status by tax category. In this way, the system achieves increased efficiency and reduced errors.
[0688] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0689] Step 1:
[0690] The user places a paper document into the scanner and presses the scan button. The input is a paper document (e.g., an invoice), and the goal is to capture it as a digital image on the terminal. Specifically, the scanner converts the physical information of the paper into a digital signal, which the terminal then acquires as image data.
[0691] Step 2:
[0692] The terminal sends the acquired digital image data to the server. The data to be sent is image data generated by the scanner. The terminal first converts the data format appropriately and then sends it to the server via the network.
[0693] Step 3:
[0694] The server applies Optical Character Recognition (OCR) to the received image data to extract text data. The input is image data, and the output is text data. Specifically, the OCR software identifies characters from the image and extracts them as strings.
[0695] Step 4:
[0696] The server uses text data to analyze information using natural language processing (NLP). The input here is the text data obtained in step 3, and the output is structured information necessary for processing (e.g., amount, payment due date). The specific operation is a process of analyzing the text using an NLP library, filtering, and extracting information.
[0697] Step 5:
[0698] The server uses a generative AI model to determine the optimal processing method. The input is the structured information obtained in step 4, and the output is a suggestion of processing steps and methods. The server sends prompts to the AI model, for example, "Extract the amount and payment due date from the invoice and suggest the optimal payment processing method."
[0699] Step 6:
[0700] The server handles errors that may occur during processing in real time. Input is data related to any errors or inconsistencies during processing, and output is notifications to the user or instructions for error correction. Specifically, an error handling module on the server detects errors and takes appropriate action.
[0701] Step 7:
[0702] The server generates the final processing results as a document and sends notifications to the user and relevant parties. The input is the processing method and its result determined in step 5, and the output is document data in the form of a report or notice. Specifically, a document generation tool is used to distribute the information via email or other means of communication.
[0703] (Application Example 1)
[0704] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0705] The challenges include the inefficiency of managing information using paper documents and the lack of efficient processing of customer service that requires immediate attention in physical stores. In particular, tasks related to the use of discount coupons and loyalty cards require the digitalization of information and immediate decision-making. This leads to problems such as increased errors and time loss due to manual work.
[0706] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0707] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information, and means for visualizing and suggesting information in real time through a portable visual device. This enables faster information processing in physical stores and improves the immediacy and accuracy of customer service.
[0708] "Paper media" refers to a form of media on which information is printed or written for recording purposes.
[0709] "Digital data" refers to a collection of information that has been converted into a format that can be processed electronically.
[0710] A "portable visual device" is a device that is portable by the user and used to display information visually.
[0711] "Optical character recognition technology" is a technology that converts printed or handwritten characters into digital data.
[0712] "Natural language processing technology" is a technology that uses computers to analyze and process human language.
[0713] "Real-time error handling" is a method of immediately detecting and correcting errors that occur during processing.
[0714] A "generative AI model" is an artificial intelligence model that generates appropriate predictions and suggestions based on input data.
[0715] "Visualization" is the process of visually displaying digital data and information, making it possible for humans to recognize it.
[0716] A "proposal" is the act of providing appropriate answers or means in advance for a specific purpose or condition.
[0717] In an embodiment of this invention, smart glasses are first used as a portable visual device. The user captures information written on paper using the camera function of the smart glasses. The captured information is converted into digital data by optical character recognition technology performed within the smart glasses.
[0718] The converted digital data is sent to the server via the terminal. The server uses natural language processing technology to extract necessary processing information from the digital data. The information obtained during this process is then used to select the optimal processing method using a generative AI model.
[0719] Furthermore, the server has real-time error handling capabilities, instantly detecting any errors that may occur during processing and instructing the user to correct them. This enables users to process information quickly and accurately.
[0720] As a concrete example, consider customer service operations in a physical store. When a customer presents a paper coupon, the user can instantly verify the coupon details through smart glasses and determine whether the discount is applicable. This procedure allows the store to improve customer satisfaction.
[0721] Based on the generated AI model, an example of a prompt message used to notify the user of processing recommendations would be the text, "Can I apply this coupon? We will give you a 50% discount on the product price," displayed in the user's field of view.
[0722] This invention makes it possible to reduce the effort required for managing information on paper, improve operational efficiency, and enhance the quality of service provided to customers.
[0723] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0724] Step 1:
[0725] The user uses the camera function of smart glasses to capture information from paper documents. The input is visual information from the paper document, and the output is a digital image file. This process allows information printed on paper to be captured by an electronic device.
[0726] Step 2:
[0727] The terminal processes the captured image file using optical character recognition (OCR) technology to extract text data. The input is the digital image obtained in step 1, and the output is text data. OCR processing converts the characters in the image into digital data.
[0728] Step 3:
[0729] The terminal sends the extracted text data to the server. The input is text data, and the output is data transfer to the server. This process provides the server with text information for processing.
[0730] Step 4:
[0731] The server uses natural language processing (NLP) techniques to extract important processing information from text data. The input is text data, and the output is specific processing information (e.g., coupon details, discount rate, eligibility). Through NLP, the information necessary for digital transactions is organized.
[0732] Step 5:
[0733] The server uses a generative AI model to make optimal suggestions and decisions based on the extracted processing information. The input is processing information, and the output is the generated suggestions or action recommendations. In this step, the AI model constructs the suggestions and prepares them to be fed into the next step.
[0734] Step 6:
[0735] The server detects potential errors during processing in real time and notifies the user. Inputs are processing information and the current execution status, while outputs are error messages or instructions for correction. Error handling ensures that the user is immediately informed of any problems.
[0736] Step 7:
[0737] The server displays the final processing results and suggestions on the user's smart glasses. The input is the optimal suggestions and modifications. The output is the data displayed in the user's field of vision. The final display allows the user to take appropriate action immediately.
[0738] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0739] This invention is a system designed to streamline payment and information processing operations for companies and local governments. In addition to digitizing conventional paper-based information, it recognizes user emotions and reflects them in business processes. It not only processes information based on paper documents but also incorporates an emotion engine to optimize the processing flow while judging the user's emotional state in real time.
[0740] System Implementation Example
[0741] This system is configured as follows: The user digitizes information from paper documents using a scanner and sends it to the server. The server uses optical character recognition (OCR) technology to convert the image data into text data, and then uses natural language processing (NLP) technology to extract the information necessary for processing.
[0742] Based on the extracted information, the server selects the optimal processing method through a multimodal AI model and handles any errors that may occur during processing in real time. During this process, the server analyzes the user's emotions using an emotion engine. This emotion information is used to adjust the tone and content of interactions and messages in processing selection and error handling.
[0743] Specific example
[0744] For example, if a user encounters a specific error while processing data related to a payment slip, the emotion engine recognizes frustration or anxiety based on the user's facial expressions and tone of voice. The server uses this emotion recognition information to generate a message containing more user-friendly language and gentler instructions during error handling, and notifies the user accordingly. By responding in accordance with emotions in this way, user stress can be reduced and work efficiency can be improved.
[0745] Finally, the completed processing results are generated as a document, including a personalized message based on the user's emotions as needed, and provided to the user. Thus, the present invention is a comprehensive system that not only improves operational efficiency but also enhances user satisfaction.
[0746] The following describes the processing flow.
[0747] Step 1:
[0748] The user places a paper invoice into the scanner and starts scanning. The scanner captures the image data into the terminal.
[0749] Step 2:
[0750] The terminal sends the scanned image data to the server. The server receives this digital image and prepares it for the next processing step.
[0751] Step 3:
[0752] The server applies optical character recognition (OCR) technology to the received image data to extract text data. In this process, the text information from paper documents is converted into a digital format.
[0753] Step 4:
[0754] The server uses natural language processing technology to analyze the text data and extract the information necessary for processing (e.g., amount, payment due date, recipient information).
[0755] Step 5:
[0756] The server uses a multimodal AI model to select the optimal processing method based on the extracted information. For example, it might consider a method for processing multiple payments in a single batch.
[0757] Step 6:
[0758] When an error occurs on the server during processing, the emotion engine is used to analyze the user's emotions. For example, facial recognition and voice analysis are used to determine whether the user is confused or distressed.
[0759] Step 7:
[0760] The server adjusts the tone and content of error messages based on emotional information and sends appropriate support or correction requests to the user. This allows users to resolve problems without stress.
[0761] Step 8:
[0762] The server generates a document based on the processing results and adds a personalized message based on emotions as needed. The generated document is then provided to the user.
[0763] Step 9:
[0764] The user reviews the provided documents and confirms that the processing was completed successfully. This entire process allows the user to complete payment transactions efficiently.
[0765] (Example 2)
[0766] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0767] Conventional information processing systems can be inefficient in the process of digitizing paper-based information and in subsequent error handling. Furthermore, uniform responses that disregard user feelings hinder the improvement of the user experience. There is a need to resolve these issues and simultaneously improve both processing efficiency and user satisfaction.
[0768] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0769] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for recognizing user emotions and adjusting interactions based on those emotions. This enables efficient information processing and user-friendly error handling.
[0770] "Paper media" refers to a means of representing information that is printed or written on physical paper.
[0771] "Digital data" refers to information that has been converted into a format that can be processed and stored by electronic devices such as computers.
[0772] "Processing information" refers to information that includes data and instructions necessary to perform a specific task or operation.
[0773] "Error handling" is the process of detecting errors and problems that occur during the information processing process and taking appropriate action.
[0774] "User emotions" refers to the emotional reactions and states that users exhibit while using the system.
[0775] "Interaction" refers to the exchange of information and communication that takes place between a system and a user.
[0776] This system efficiently converts information from paper documents into digital data and implements a process that optimizes processing based on that information. A distinctive feature of this system is its ability to recognize user emotions in real time and utilize that information to optimize processing.
[0777] First, the user uses a scanner to digitize the information on paper. This digital data is then processed using OCR (Optical Character Recognition) technology. The server uses this OCR technology to extract text data from the scanned image data. Typically, commercially available OCR software is used for this purpose.
[0778] Next, the server uses NLP (Natural Language Processing) technology to extract necessary processing information from the text data. This technology is used to identify specific keywords and instructions from text. Common natural language processing libraries can be applied to the software used.
[0779] Furthermore, the server determines the optimal processing method as needed through a generative AI model. This AI model learns from the dataset and provides flexible processing methods that adapt to the situation. In addition, if an error occurs, the emotion engine analyzes the user's emotions and assists in taking appropriate action. This emotion engine can read emotions from the user's facial expressions and voice.
[0780] A concrete example is the emotional response to error messages that occur when a user attempts to digitize their tax payment process. This involves displaying reassuring messages based on emotional information to support the user in continuing the process. For example, a prompt might read, "The payment slip could not be read. Please try again. If you have further problems, please contact support."
[0781] The overall flow of this system is designed to streamline and user-friendly information processing operations for companies and local governments.
[0782] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0783] Step 1:
[0784] The user places a paper document in the scanner and begins the digitization process. The scanner acquires image data from the paper document and generates that data as an electronic image file. As output of this process, the image file is sent to the server.
[0785] Step 2:
[0786] The server applies Optical Character Recognition (OCR) technology to the received image file. The input is image data. In this step, the server identifies the text within the image and converts the visual data into text data. The output of this process is readable electronic text.
[0787] Step 3:
[0788] The server analyzes electronic text data using natural language processing (NLP) techniques. Text data is provided as input. The server scans the entire document and extracts the necessary processing information. Relevant keywords and phrases are retrieved as NLP output and stored in a database.
[0789] Step 4:
[0790] The server uses a generative AI model to select the optimal processing plan based on the extracted information. This model operates using a pre-trained algorithm to derive the most efficient processing steps from the input data. The output consists of specific processing steps and necessary configuration information.
[0791] Step 5:
[0792] The server monitors for errors in real time as processing progresses and performs error handling as needed. Simultaneously, it uses an emotion engine to analyze the input user's video and audio data and recognize their emotions. The output emotion data is reflected in the tone and content of the message.
[0793] Step 6:
[0794] The server generates the final processing results as a document, providing a customized report based on sentiment. In this step, all system processing and sentiment analysis outputs are aggregated to create a user-friendly report. The final document is sent to the user via a terminal.
[0795] (Application Example 2)
[0796] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0797] The goal is to resolve issues that hinder the smooth progress of the payment process due to stress and confusion experienced by users when using electronic payment systems. In particular, emotional states can be a barrier to payment, so it is necessary to eliminate factors that disrupt a smooth user experience.
[0798] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0799] In this invention, the server includes means for converting information from paper media into digital data, means for extracting necessary processing information from the digital data, and means for analyzing the user's facial expressions and voice tone in real time and generating emotional information. This makes it possible to provide an appropriate interface and response that takes into account the user's emotional state, and to facilitate the payment process.
[0800] "Paper media" refers to physical documents and forms on which information is printed, and it represents the starting point of information before it is digitized.
[0801] "Digital data" refers to data in an electronic format converted from paper media, and in a form that can be processed by a computer system.
[0802] "Processing information" refers to specific information extracted from digital data that is required for business purposes or specific applications.
[0803] The "optimal processing method" is a method or process selected based on processing information to maximize operational efficiency.
[0804] "Real-time error handling" means immediately recognizing problems that arise during processing and responding appropriately.
[0805] "Analyzing the user's facial expressions and voice tone in real time" refers to using multimodal AI to evaluate and analyze the user's facial movements and voice characteristics on the spot.
[0806] "Emotional information" refers to data that represents a user's psychological or emotional state, derived from their analyzed facial expressions and voice tone.
[0807] "Adjusting interaction content" means dynamically changing communication methods, designs, and messages based on the user's emotional state.
[0808] "Generating and outputting as a document" refers to the process of compiling the final processing results into a report or record using text, charts, and other formats, and presenting it to the user or system.
[0809] The system for carrying out this invention is constructed by integrating multiple hardware and software components. The terminal is equipped with a scanner, camera, and microphone, and converts information from paper documents into digital data, collecting the user's facial expressions and voice in real time. The server receives this data and performs the following processing.
[0810] The server first uses optical character recognition (OCR) technology to convert scanned information into text data. Natural language processing (NLP) technology is then used to extract the necessary information from this text data. Next, emotional information is analyzed from the user's facial expressions and voice through a generative AI model and an emotion engine. This emotional information is used to adjust the content of the interaction with the user.
[0811] Specifically, when a user attempts to make an electronic payment, the system captures the user's facial expression with a camera and collects their voice tone with a microphone. This data is sent to a server and analyzed in real time. If the user appears confused, the application interface is modified to be more user-friendly, and a gentle guide message such as, "Are you okay? If you need assistance, we can call a staff member," is displayed. This reduces user stress.
[0812] To give a concrete example, imagine a user trying a new QR code payment method for the first time in a crowded cafe. If the generation AI model detects that the user is confused, the prompt message "Generate a gentle guidance message if the user is frowning" will be used. This prompt allows the system to appropriately support the user and facilitate a smooth payment.
[0813] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0814] Step 1:
[0815] The terminal scans the user's paper documents and saves them as digital image data. This image data becomes the input and is sent to the server.
[0816] Step 2:
[0817] The server performs optical character recognition (OCR) on the received image data and extracts text data from it. The input is image data, and the extracted text data is the output.
[0818] Step 3:
[0819] The server uses natural language processing (NLP) techniques based on text data to identify and extract the necessary processing information. The input is text data, and the output is the processed information.
[0820] Step 4:
[0821] The device captures the user's facial expressions with a camera and records their voice with a microphone. This raw data is sent to the server in real time. The input consists of the user's facial expressions and voice.
[0822] Step 5:
[0823] The server uses a generative AI model and an emotion engine to analyze the received facial and audio data. This generates the user's emotional information. The input is facial and audio data, and the output is emotional information.
[0824] Step 6:
[0825] The server dynamically modifies the user interface and uses prompts to generate appropriate messages based on the analyzed sentiment information. The input is sentiment information, and the output is a user-facing message generated based on the prompts.
[0826] Step 7:
[0827] The terminal displays emotion-based messages received from the server to the user and provides guidance as needed. The output is a message that the user sees directly, thereby creating a reassuring interaction.
[0828] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0829] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0830] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0831] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0832] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0833] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0834] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0835] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0836] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0837] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0838] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0839] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0840] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0841] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0842] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0843] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0844] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0845] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0846] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0847] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0848] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0849] The following is further disclosed regarding the embodiments described above.
[0850] (Claim 1)
[0851] A means of converting information from paper documents into digital data,
[0852] Means for extracting necessary processing information from the aforementioned digital data,
[0853] A means for selecting the optimal processing method based on the processing information,
[0854] Means for handling errors in the aforementioned process in real time,
[0855] A means for generating and outputting the processing result as a document,
[0856] A system that includes this.
[0857] (Claim 2)
[0858] The system according to claim 1, which acquires information from paper media as text data using optical character recognition technology.
[0859] (Claim 3)
[0860] The system according to claim 1, which uses natural language processing technology to extract information necessary for processing from digital data.
[0861] "Example 1"
[0862] (Claim 1)
[0863] A device that converts information from paper documents into digital data,
[0864] A device for extracting processing information from the aforementioned digital data,
[0865] A device that determines the optimal processing method based on the processing information,
[0866] A device for controlling errors during the aforementioned processing in real time,
[0867] A device that creates and outputs the processing result as a document,
[0868] A device that acquires information from the aforementioned paper media as text data using optical character recognition,
[0869] A device that uses natural language processing technology to extract necessary information from digital data,
[0870] A device that proposes efficient processing methods using generative AI models,
[0871] A system that includes this.
[0872] (Claim 2)
[0873] The system according to claim 1, which uses a generative AI model to derive the optimal processing method.
[0874] (Claim 3)
[0875] The system according to claim 1, comprising a device that controls errors in real time and notifies the user.
[0876] "Application Example 1"
[0877] (Claim 1)
[0878] A means of converting information from paper documents into digital data,
[0879] Means for extracting necessary processing information from the aforementioned digital data,
[0880] A means for selecting the optimal processing method based on the processing information,
[0881] Means for handling errors in the aforementioned process in real time,
[0882] A means for generating and outputting the processing result as a document,
[0883] A means of capturing paper media information using a portable visual device and using it as a display means,
[0884] A means for visualizing and proposing information in real time through the aforementioned portable visual device,
[0885] A system that includes this.
[0886] (Claim 2)
[0887] The system according to claim 1, which acquires information from paper media as text data using optical character recognition technology.
[0888] (Claim 3)
[0889] The system according to claim 1, which uses natural language processing technology to extract information necessary for processing from digital data.
[0890] "Example 2 of combining an emotion engine"
[0891] (Claim 1)
[0892] A means of converting information from paper documents into digital data,
[0893] Means for extracting necessary processing information from the aforementioned digital data,
[0894] A means for selecting the optimal processing method based on the processing information,
[0895] Means for handling errors in the aforementioned process in real time,
[0896] A means for recognizing the user's emotions when the aforementioned error occurs,
[0897] A means for adjusting the selection of processing and the interactions and messages in error handling based on the user's emotions,
[0898] A means for generating and outputting the processing result as a document,
[0899] A system that includes this.
[0900] (Claim 2)
[0901] The system according to claim 1, which acquires information from paper media as text data using optical character recognition technology.
[0902] (Claim 3)
[0903] The system according to claim 1, which uses natural language processing technology to extract information necessary for processing from digital data.
[0904] "Application example 2 when combining with an emotional engine"
[0905] (Claim 1)
[0906] A means of converting information from paper documents into digital data,
[0907] Means for extracting necessary processing information from the aforementioned digital data,
[0908] A means for selecting the optimal processing method based on the processing information,
[0909] Means for handling errors in the aforementioned process in real time,
[0910] A means for analyzing the user's facial expressions and voice tone in real time and generating emotional information,
[0911] A means for adjusting the content of user interaction based on the aforementioned emotional information,
[0912] A means for generating and outputting the processing result as a document,
[0913] A system that includes this.
[0914] (Claim 2)
[0915] The system according to claim 1, which acquires information from paper media as text data using optical character recognition technology.
[0916] (Claim 3)
[0917] The system according to claim 1, which uses natural language processing technology to extract information necessary for processing from digital data. [Explanation of Symbols]
[0918] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of converting information from paper documents into digital data, Means for extracting necessary processing information from the aforementioned digital data, A means for selecting the optimal processing method based on the processing information, Means for handling errors in the aforementioned process in real time, A means for generating and outputting the processing result as a document, A system that includes this.
2. The system according to claim 1, which acquires information from paper media as text data by using optical character recognition technology.
3. The system according to claim 1, which uses natural language processing technology to extract information necessary for processing from digital data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A