System

The system automates data entry from documents and images through preprocessing, OCR, and validation, addressing inefficiencies and errors in manual processes to enhance productivity.

JP2026022445APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123962
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Manual data entry from documents and images is time-consuming and labor-intensive, leading to human error and reduced data accuracy, which negatively impacts business efficiency and information management.

Method used

A system that allows users to photograph documents or images using a terminal, which uploads them to a server for preprocessing, optical character recognition, data extraction, validation, and export to external systems, with user corrections and re-verification.

Benefits of technology

Improves the efficiency and accuracy of data processing from documents and images, enhancing overall business productivity by automating information extraction and verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022445000001_ABST
    Figure 2026022445000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to capture a document or an image using a terminal and upload it to a server; means for the server to perform pre-processing (skew correction, noise removal, resolution adjustment) on the uploaded image; means for using optical character recognition technology or artificial intelligence to extract text data from the pre-processed image; means for verifying the extracted data and notifying the user and asking for correction if there is any uncertainty; and means for re-verifying the corrected data and finally storing and exporting to the required external system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Currently, in many business processes, the manual data entry of information from documents and images is extremely time-consuming and labor-intensive, leading to human error. The inability to streamline this process negatively impacts the overall efficiency of the business. Furthermore, manual data entry reduces data accuracy, resulting in increased complexity in information management and increased difficulty in accessing information. The present invention aims to solve these problems and improve the speed and accuracy of data processing. [Means for solving the problem]

[0005] The present invention is a system that includes a means for a user to photograph a document or image using a terminal and upload it to a server, and a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image. It also includes a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image, a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties, a means for re-verifying the corrected data, and finally, a means for saving and exporting it to the necessary external systems. This improves the efficiency and accuracy of information extraction and data processing from documents and images, and is expected to improve overall business productivity.

[0006] "User" means a person or entity that operates the system and uploads and modifies documents and images.

[0007] A "terminal" is a device (such as a smartphone, tablet, or PC) that a user uses to take photos of documents and images and upload them to a server.

[0008] "Server" means the computing system that receives images uploaded by users and has the central functions of pre-processing, data extraction, validation, and data storage and export.

[0009] An "image" is digital data of a document or photograph that is taken or selected by a user using a terminal and uploaded to a server.

[0010] "Preprocessing" refers to processing that the server performs on uploaded images, and includes tilt correction, noise removal, resolution adjustment, and the like.

[0011] Optical character recognition (OCR) is a technology for extracting text data from images, converting letters and numbers written in documents and photographs into digital data.

[0012] "Artificial intelligence" is a technology in which computer systems process data through pattern recognition and machine learning, and is used to extract and analyze text data.

[0013] "Data validation" is the process of verifying that extracted data is accurate and complete, and that it meets specific criteria and formats.

[0014] A "notification" is a message or alert that notifies the user of specific information (e.g., a correction request or export completion).

[0015] "Export" is the process by which the server sends validated data to other external systems (e.g., accounting software, electronic medical record systems).

[0016] "Metadata" is supplementary information related to an image, and includes attribute information such as upload date and time, user ID, and image ID. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention is a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[0039] Program processing

[0040] 1. Acquiring and uploading images

[0041] Terminal

[0042] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request to send the photo or image to the server.

[0043] server

[0044] The server receives an image upload request and saves the image in storage. The saved image is assigned a unique identifier, and its metadata (upload date and time, user ID, image ID, etc.) is also saved in the database.

[0045] 2. Image Preprocessing

[0046] server

[0047] After capturing the image, the server begins pre-processing, which includes image deskewing, noise reduction, and resolution adjustment, making the characters and numbers in the image easier to recognize.

[0048] 3. Performing data extraction

[0049] server

[0050] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital data. The AI ​​model then further examines the extracted data based on contextual information to identify specific fields (such as names, dates, or numbers).

[0051] 4. Data verification and completion

[0052] server

[0053] The extracted data is automatically validated by the server, for example to ensure that date formats and numeric ranges are correct. If validation fails, the server notifies the user and asks them to correct the error.

[0054] User

[0055] The user receives the terminal notification, makes any necessary corrections, and then transmits the correction data to the server.

[0056] server

[0057] The server will then re-verify the corrected data, and once verification is complete, the data will be saved and go on to the next step.

[0058] 5. Saving and Exporting Data

[0059] server

[0060] The completed dataset is securely stored on the server. If necessary, the data can be exported to external applications or systems (e.g., accounting software, electronic medical record systems, etc.). If the export was successful, the server notifies the user.

[0061] Specific examples

[0062] Example 1: Processing receipts in the finance department

[0063] Terminal

[0064] The user takes a photo of the receipt using the terminal and uploads it to the server.

[0065] server

[0066] The server receives the image, performs deskewing, noise removal, and resolution adjustment. The pre-processed image is passed through an OCR engine, which extracts data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any ambiguities.

[0067] User

[0068] The user makes the corrections and sends the data to the server, which re-verifies the corrections and finally exports them to the accounting software and notifies the user of the completion.

[0069] Example 2: Medical record processing in a healthcare facility

[0070] Terminal

[0071] Medical staff scan handwritten medical records and upload them to a server.

[0072] server

[0073] The server receives the images and performs deskewing, noise reduction, and resolution adjustment. The preprocessed images are then run through an AI model to extract data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically verified, and medical staff are notified if there are any issues.

[0074] User

[0075] The medical staff makes the corrections and sends the data to the server, which re-verifies the corrected data, exports it to the electronic medical record system, and notifies the medical staff of the completion.

[0076] In this way, this system improves productivity throughout the entire business by efficiently extracting information from documents and images and automatically processing data.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] Terminal

[0080] A user launches an application on the device and uses the capture function to capture a document or image, or select an existing image. After the user completes the capture or selection, the device generates an upload request to send the image to the server.

[0081] Step 2:

[0082] server

[0083] The server receives an image upload request, assigns a unique identifier to the uploaded image, saves it in storage, and stores metadata such as the upload date and time, user ID, and image ID in a database.

[0084] Step 3:

[0085] server

[0086] The server then performs preprocessing on the stored image, which includes image deskewing, noise reduction, and resolution adjustment. Deskewing adjusts the text in the image so that it is level, noise reduction removes unwanted background elements, and resolution adjustment makes the details of the characters clearer.

[0087] Step 4:

[0088] server

[0089] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. Optionally, the AI ​​model analyzes the context and improves the accuracy of the extracted data.

[0090] Step 5:

[0091] server

[0092] Specific fields (e.g., names, dates, numbers, etc.) are identified from the extracted text data and saved as structured data (e.g., JSON, CSV format) for subsequent validation and export processes.

[0093] Step 6:

[0094] server

[0095] The extracted data is automatically validated, including checking date formats, numeric ranges, required fields, etc. If any errors are found during the validation process, the server notifies the user and asks them to correct them.

[0096] Step 7:

[0097] Terminal

[0098] The user receives the notification and modifies the data on the terminal screen. Once the modification is complete, the user generates a request to send the modified data to the server.

[0099] Step 8:

[0100] server

[0101] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be accurate, it is saved as the final data. Once the correction process is complete and the verification is successful, the server proceeds to the next step.

[0102] Step 9:

[0103] server

[0104] Securely store the completed dataset in a database. Once the storage process is complete, generate API requests to export the data to any external systems required (e.g., accounting software, electronic medical record systems).

[0105] Step 10:

[0106] server

[0107] If the export is successful, the server records the result in the database and sends the user a notification of completion, including confirmation of the destination system and the contents of the data.

[0108] This detailed process at each step allows for efficient image-based data extraction and automated data processing, improving overall business productivity.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] In modern business processes and healthcare facilities, manual information extraction and validation from documents and images is time-consuming and error-prone. This reduces operational efficiency and reduces reliability in critical data processing. Furthermore, there is a need to export extracted data quickly and accurately to external systems. To address these challenges, an automated system is needed.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (such as deskewing, noise removal, and resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for re-verifying the corrected data and finally saving and exporting it to required external systems; a means for identifying specific fields (such as names, dates, and numbers) of the extracted data and saving them as structured data; and a means for examining the data based on a generative AI model and generating automated notifications using prompts. This enables the automation of information extraction and verification from documents and images, improving business efficiency and the reliability of data processing.

[0114] A "user" is an entity that uses the system to capture and upload documents and images.

[0115] A "terminal" is a device used by a user, which has the function of taking pictures of documents and images and uploading them to a server.

[0116] A "server" is a computer system that pre-processes images uploaded by users and performs functions such as data extraction, validation, storage and export.

[0117] "Upload" refers to the act of sending image or document data from a terminal to a server.

[0118] "Preprocessing" refers to processing such as tilt correction, noise removal, and resolution adjustment on the image to facilitate subsequent data extraction.

[0119] "Tilt correction" is a process of adjusting the orientation of an image so that the characters and figures in the image are easier to recognize.

[0120] "Noise reduction" is a process that removes unnecessary elements from an image to make it clearer.

[0121] "Resolution adjustment" is a process of changing the image resolution to an appropriate level.

[0122] Optical character recognition (OCR) is a technology that recognizes text within an image and converts it into digital data.

[0123] "Artificial intelligence (AI)" is the technology that enables computers to learn, reason, and solve problems like humans.

[0124] "Data validation" is the process of checking whether extracted data is accurate and valid.

[0125] "Notification" is a message or signal that notifies the user when there is a validation result or uncertain data.

[0126] "Correction" refers to the act of the user correcting data based on a notification from the server.

[0127] "Storage" is the act of securely recording verified data in storage.

[0128] "Export" is the act of sending stored data to an external system.

[0129] A "specific field" is a data item that indicates specific information such as a name, date, or number.

[0130] "Structured data" is data that is organized according to a predefined data format.

[0131] A "generative AI model" is an artificial intelligence algorithm that is trained to perform a specific task.

[0132] A "prompt" is a sentence of text that instructs a generative AI model on a task based on specific input.

[0133] This invention is a system in which a user uses a terminal to take a photo of a document or image and upload it to a server, which then extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[0134] Hardware and software used

[0135] Terminal

[0136] The device used by the user is a device with a photographic function, such as a smartphone, tablet, or digital camera. The device is responsible for sending the photographed or selected images to the server. The application running on the device should preferably have an internet connection for data transmission and an image editing function.

[0137] server

[0138] The server performs the functions of data reception, image pre-processing, data extraction, validation, storage, and export, using the following software:

[0139] Image preprocessing library (e.g. OpenCV)

[0140] Optical Character Recognition (OCR) engine (e.g., Tesseract OCR)

[0141] Artificial Intelligence (AI) models (e.g., Google Cloud Vision API)

[0142] Database system (e.g. MySQL)

[0143] Notification systems (e.g., Firebase Cloud Messaging)

[0144] Data processing and calculation in the program processing flow

[0145] 1. Acquiring and uploading images

[0146] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request and sends the image data and metadata to the server, including information such as the user ID and image ID.

[0147] 2. Image Preprocessing

[0148] The server performs deskew, noise reduction, and resolution adjustment on the images received, making it easier to recognize characters and numbers. The OpenCV library is used for these preprocessing steps.

[0149] 3. Data extraction

[0150] The preprocessed images are then fed into an optical character recognition (OCR) or artificial intelligence (AI) model, using Tesseract OCR to extract text data from the image and Google Cloud Vision API to identify specific fields (such as names, dates, or amounts).

[0151] 4. Data verification and completion

[0152] The extracted data is validated by the server. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. If there is any problem with the validation, the server notifies the user and asks them to make corrections. After the user makes the necessary corrections, the corrected data is sent back to the server, which then validates the data again.

[0153] 5. Saving and Exporting Data

[0154] The complete dataset is stored on the server and can be exported to external applications or systems as needed, with the user notified if the export was successful.

[0155] Specific examples

[0156] Example 1: Processing receipts in the finance department

[0157] The user takes a photo of the receipt using a smartphone app and uploads it to the server. The server then performs image deskewing, noise reduction, and resolution adjustment. The preprocessed image is then input into the Tesseract OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any uncertainties. If the user corrects any errors and submits it to the server again, the server revalidates the data, exports it to the accounting software, and notifies the user of completion.

[0158] Example 2: Medical record processing in a healthcare facility

[0159] Medical staff scan handwritten medical records and upload them to the server, which then performs image deskewing, noise removal, and resolution adjustment. The preprocessed images are then input into the Google Cloud Vision API, which extracts data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically validated, and if there are any problems, the medical staff is notified. After the medical staff corrects any errors and resubmits the data to the server, the server revalidates the data, exports it to the electronic medical record system, and notifies the medical staff of its completion.

[0160] Prompt Sentence Examples

[0161] Below are some example prompts using generative AI models:

[0162] "Input the following image into an OCR engine and extract data such as the amount, transaction date, and store name. After extraction, verify this data and, if necessary, notify the user."

[0163] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0164] Step 1:

[0165] User

[0166] The user uses a device to take a photo of a document or image, or to select an existing image. To do this, the user uses an application on the device. The data input here is the image taken or selected by the user, and the image data is output. In a specific example, the user might take a photo of a receipt using a camera app on their smartphone.

[0167] Step 2:

[0168] Terminal

[0169] The device internally checks the captured or selected image data and generates an upload request to send to the server. The request includes metadata such as the user ID and image ID. The input is the image data captured or selected by the user and the metadata, and the output is an upload request sent to the server. Here, the data is sent using an Internet connection.

[0170] Step 3:

[0171] server

[0172] The server receives an image upload request and saves the image data and metadata in storage. When saving, it assigns a unique identifier to the image and records the metadata in a database. The input is the image data and metadata sent from the device, and the output is the image data saved in storage and the metadata saved in the database. For example, the image data is saved in a dedicated folder, and the metadata is recorded in a database table.

[0173] Step 4:

[0174] server

[0175] The server retrieves the image from storage and begins preprocessing. In preprocessing, the image is tilted and rotated to the correct orientation. Noise reduction is also performed to remove unnecessary parts. Furthermore, the image resolution is adjusted to make text and numbers easier to read. The input is the image data retrieved from storage, and the output is the preprocessed image data. Here, the OpenCV library is used to preprocess the image.

[0176] Step 5:

[0177] server

[0178] The server inputs the preprocessed image data into an optical character recognition (OCR) engine or artificial intelligence (AI) model. Tesseract OCR is used to recognize text in the image and convert it into digital data. Google Cloud Vision API is then used to identify specific fields (such as names, dates, or amounts) based on contextual information from the extracted text. The input is the preprocessed image data, and the output is the extracted text data.

[0179] Step 6:

[0180] server

[0181] The server validates the extracted data. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. The input is the extracted text data, and the output is the validation result. If there is any validation failure, the server notifies the user and asks them to correct it. This notification is sent using Firebase Cloud Messaging or similar.

[0182] Step 7:

[0183] User

[0184] The user receives the notification on the device and makes the necessary corrections. Here, the input is the notification from the server, and the output is the corrected data. For example, the user manually corrects misreadings made by OCR on the device screen.

[0185] Step 8:

[0186] Terminal

[0187] The user modifies the data and sends it back to the server. The input is the modified data, and the output is the upload request sent again.

[0188] Step 9:

[0189] server

[0190] The server re-verifies the modified data. If the re-verification is successful, the data proceeds to the next processing step. The input is the modified data submitted by the user, and the output is the re-verified data.

[0191] Step 10:

[0192] server

[0193] The validated dataset is securely stored and exported to external applications or systems as needed. For example, sending data to accounting software or electronic medical record systems. If the export is successful, the server sends a notification to the user. The input is the validated data, and the output is the export result to the external system and a notification to the user.

[0194] (Application example 1)

[0195] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0196] In logistics centers, the receiving and inventory management of items is often done manually, which can lead to problems such as data entry errors and the time required for confirmation work. For this reason, there is a need for a system that can manage item information efficiently and accurately.

[0197] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0198] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user if there is any uncertainty and requesting correction; a means for re-verifying the corrected data and finally saving and exporting it to the necessary external system; and a means for automatically recognizing item information at the logistics center and providing a function for managing data on the terminal, thereby enabling efficient and accurate management of item information at the logistics center.

[0199] A "terminal" is a device that a user uses to take documents and images and upload the data to a server.

[0200] A "server" is a computer system that receives image data uploaded by users and performs pre-processing, data extraction, validation, and export.

[0201] "Preprocessing" refers to the process of correcting tilt, removing noise, and adjusting resolution of uploaded images.

[0202] "Optical character recognition (OCR) technology" is a technology that identifies characters in an image and converts them into digital text data.

[0203] "Artificial intelligence (AI)" is a technology that gives computer systems the ability to understand meaning and context from images and text, and to extract, classify, and verify data.

[0204] "Verification" is the process of checking whether the extracted data is accurate and asking the user to correct any uncertainties.

[0205] "Export" is the process of outputting the final validated data to the required external systems or applications.

[0206] A "logistics center" is a facility that handles logistics operations such as receiving, storing, and shipping items.

[0207] "Item information" refers to data such as the identification information, quantity, and date of receipt of products handled at the logistics center.

[0208] The "database" is an information management system for organizing and storing extracted item information.

[0209] An "inventory management system" is a software system that manages the inventory status of items and records information on incoming and outgoing goods.

[0210] The present invention is a system for improving the efficiency and accuracy of item management in a logistics center. A specific embodiment of this system will be described below.

[0211] System Configuration

[0212] This system consists of a terminal used by users, a server that performs processing, and a database that stores data. The main software used is OpenCV for image preprocessing, Tesseract as an optical character recognition (OCR) engine, and requests for communication.

[0213] Program processing overview

[0214] 1. Acquiring and uploading images

[0215] The device provides a means for users to take photos of items and upload them to the server. The images are sent to the server and assigned a unique identifier. Image metadata (upload date and time, user ID, image ID, etc.) is also generated and stored in a database.

[0216] 2. Image Preprocessing

[0217] The server receives the uploaded image and performs deskewing, noise reduction, and resolution adjustment, which makes the text and numbers in the image more readable.

[0218] 3. Data Extraction

[0219] The server inputs the preprocessed image into an OCR engine (Tesseract) to recognize text in the image and convert it into digital data. Additionally, it uses an artificial intelligence (AI) model to extract item information (item ID, quantity, date of receipt, etc.).

[0220] 4. Data verification and completion

[0221] The extracted data is automatically validated by the server. For example, it checks whether the date format and numeric range are correct. If there are any errors during validation, the user is notified and can make corrections. If the user makes corrections and submits the data to the server again, the server will validate the data again.

[0222] 5. Saving and Exporting Data

[0223] The final verified data is stored in a database and, if required, exported to an inventory management system.

[0224] Specific examples

[0225] Example 1: Taking a photo of an item and exporting the data

[0226] Staff at the logistics center use a terminal to take a photo of an item and upload it to the server. The server receives the image, preprocesses it, and then uses an OCR engine to extract item information. The extracted data is verified and exported to the inventory management system. If there is an error in the data, a notification is sent to the user, who can then make the necessary corrections.

[0227] Prompt Sentence Examples

[0228] "Please take a photo of this item. The system will automatically extract the item information and update it in your inventory."

[0229] In this way, the system of the present invention automates item management in a logistics center, improving work efficiency and accuracy.

[0230] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0231] Step 1:

[0232] A user takes a photo of an item using a device, which captures the item image, generates a unique identifier, and prepares the captured image and its metadata (upload date and time, user ID, image ID, etc.).

[0233] Step 2:

[0234] The device uploads the captured images and metadata to a server, which receives the data and stores them in a database using a unique identifier assigned to the image.

[0235] Step 3:

[0236] The server retrieves the received image from the database and begins pre-processing it, i.e., correcting the image's distortion (image rotation), removing noise (image filtering), and adjusting the resolution (resizing or rescaling). This pre-processing process makes the features in the image more visible.

[0237] Step 4:

[0238] The preprocessed image is input into the server's OCR engine (Tesseract) and OCR processing is performed. During this process, characters within the image are extracted and converted into digital text data. The extracted text data includes the item ID, quantity, and inventory date.

[0239] Step 5:

[0240] The server uses an AI model to more precisely identify item information from the text data extracted by OCR, particularly identifying specific fields (item ID, quantity, inventory date, etc.) and saving them as structured data, utilizing machine learning and natural language processing techniques.

[0241] Step 6:

[0242] The server validates the extracted and identified data, specifically ensuring that date formats and numeric ranges are correct. If validation fails, the server notifies the user, including the specific corrections required.

[0243] Step 7:

[0244] The user receives a notification and makes the necessary corrections on the device, after which the corrected data is sent back to the server.

[0245] Step 8:

[0246] The server will again verify the corrected data resubmitted by the user, and if the verification is successful, the data will be finally saved in the database.

[0247] Step 9:

[0248] The server exports the saved data to an external system (e.g., inventory management system) as needed. Once the export is complete, the server notifies the user.

[0249] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0250] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[0251] Program processing

[0252] 1. Acquiring and uploading images

[0253] Terminal

[0254] The user opens the application on the device and uses the camera function to take a document or image. At the same time, the device captures the user's face, and the emotion engine analyzes the user's emotional state. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[0255] server

[0256] The server receives the upload request and stores the image and emotion data. The image is assigned a unique identifier, and metadata such as the upload date and time and user ID are also stored in the database.

[0257] 2. Image Preprocessing

[0258] server

[0259] The server then performs preprocessing on the stored image, which includes deskewing, noise reduction, and resolution adjustment, making the letters and numbers in the image more recognizable.

[0260] 3. Performing data extraction

[0261] server

[0262] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. The AI ​​model then analyzes the context to further refine the extracted data.

[0263] 4. Data verification and completion

[0264] server

[0265] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, an emotion engine evaluates the user's emotional state. If a specific emotional state (e.g., stress or confusion) is detected, the server sends the user an improved notification.

[0266] User

[0267] The user receives the notification on the device and can modify the data. Once the modification is complete, a request is generated to send the modified data to the server.

[0268] server

[0269] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved as the final data. The verification process and the content of the correction request may be adjusted based on the emotion data.

[0270] 5. Saving and Exporting Data

[0271] server

[0272] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). If the export is successful, the server sends a completion notification to the user and records the export results in the database.

[0273] Specific examples

[0274] Example 1: Processing receipts in the finance department

[0275] Terminal

[0276] The user takes a photo of the receipt using the device and uploads it to the server, while the emotion engine simultaneously analyzes the user's emotional state.

[0277] server

[0278] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification tailored based on the emotion data is sent to the user.

[0279] User

[0280] The user receives a notification, makes the necessary corrections, and sends the corrected data to the server, which then validates it again. Finally, the data is exported to the accounting software, and the user is notified of the completion.

[0281] Example 2: Medical record processing in a healthcare facility

[0282] Terminal

[0283] Medical staff scan handwritten medical records and upload them to the server, and at the same time, the emotion engine analyzes the emotional state of the medical staff.

[0284] server

[0285] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is then passed through an AI model to extract data such as patient name, diagnosis, and prescribed medication. The extracted data is automatically verified, and if there are any issues, a tailored notification based on the emotion data is sent to medical staff.

[0286] User

[0287] The medical staff receives the notification, makes the necessary corrections, and sends the data to the server, which re-verifies the corrected data, exports the data to the electronic medical record system, and notifies the medical staff of the completion.

[0288] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[0289] The processing flow will be explained below.

[0290] Step 1:

[0291] Terminal

[0292] The user launches the application on the device and uses the camera function to take a document or image. At the same time, the device camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., joy, anger, sadness, etc.) in real time. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[0293] Step 2:

[0294] server

[0295] The server receives the upload request and saves the submitted image and emotion data in storage. A unique identifier is assigned to the image, and metadata such as the upload date and time, user ID, and emotion data are recorded in the database.

[0296] Step 3:

[0297] server

[0298] The server begins pre-processing the stored image, straightening the image to make the text easier to read, applying a noise reduction filter to remove unnecessary elements in the image, and adjusting the resolution to make the details of the text and numbers clearer.

[0299] Step 4:

[0300] server

[0301] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital text data. The AI ​​model analyzes the context and improves the accuracy of the extracted data.

[0302] Step 5:

[0303] server

[0304] The text data extracted by the OCR engine or AI model is parsed to identify specific fields (e.g., names, dates, numbers, etc.) and stored as structured data for subsequent validation and export processes.

[0305] Step 6:

[0306] server

[0307] The extracted data is automatically validated, for example, checking whether the date format is correct or whether the numbers are within a reasonable range. If any flaws are detected during the validation process, the server will also evaluate the user's emotional state and send a notification with an appropriate tone if a specific emotional state (e.g., stress, confusion) is detected through the emotion engine.

[0308] Step 7:

[0309] Terminal

[0310] The user receives the notification on the terminal and makes the necessary modifications through the data modification interface. After completing the modifications, the user generates a request to send the data to the server.

[0311] Step 8:

[0312] server

[0313] The server receives the corrected data sent by the user and performs automatic verification again. If the corrected data is confirmed to be correct, it is saved as the final data set.

[0314] Step 9:

[0315] server

[0316] The completed dataset is securely stored. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems). If the export is successful, the server records the result in a database and sends a completion notification to the user.

[0317] Specific examples

[0318] Example 1: Processing receipts in the finance department

[0319] Step 1:

[0320] Terminal

[0321] The user takes a photo of the receipt using the device, and at the same time, the emotion engine analyzes the user's emotional state in real time.

[0322] Step 2:

[0323] server

[0324] The server receives the image and emotion data, stores them in storage, generates a unique identifier and metadata, and records them in a database.

[0325] Step 3:

[0326] server

[0327] Preprocessing such as tilt correction, noise removal, and resolution adjustment is performed on the received image.

[0328] Step 4:

[0329] server

[0330] The preprocessed image is input into an OCR engine to extract data such as the amount, transaction date, and store name.

[0331] Step 5:

[0332] server

[0333] The extracted data is saved as structured data and passed to subsequent validation processing.

[0334] Step 6:

[0335] server

[0336] The extracted data is automatically verified, and if there are any deficiencies, an appropriate notification is sent based on the user's emotional state.

[0337] Step 7:

[0338] Terminal

[0339] The user receives a notification on the device, makes the necessary corrections, and sends the corrected data to the server.

[0340] Step 8:

[0341] server

[0342] The corrected data is verified again, and if it is correct, it is saved as the final data.

[0343] Step 9:

[0344] server

[0345] The final data is exported to accounting software and the results are notified to the user.

[0346] Example 2: Medical record processing in a healthcare facility

[0347] Step 1:

[0348] Terminal

[0349] Medical staff use the device to scan handwritten medical records, while the emotion engine simultaneously analyzes the emotional state of the medical staff in real time.

[0350] Step 2:

[0351] server

[0352] The server receives and stores the image and emotion data, assigns a unique identifier, and generates and records metadata.

[0353] Step 3:

[0354] server

[0355] The saved image is then tilted, noise removed, and resolution adjusted.

[0356] Step 4:

[0357] server

[0358] The pre-processed images are fed into an AI model to extract data such as patient name, diagnosis, and prescribed medications.

[0359] Step 5:

[0360] server

[0361] The extracted data is saved as structured data and passed to subsequent validation processing.

[0362] Step 6:

[0363] server

[0364] The extracted data is automatically verified, and if there are any problems, notifications are sent that are tailored to the emotional state of the medical staff.

[0365] Step 7:

[0366] Terminal

[0367] Medical staff are notified and make any necessary corrections, which are then sent to the server.

[0368] Step 8:

[0369] server

[0370] The corrected data is verified again, and if appropriate, is saved as the final data.

[0371] Step 9:

[0372] server

[0373] The final data is exported to the electronic medical record system and the results are notified to medical staff.

[0374] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[0375] Example 2

[0376] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0377] In conventional systems, when extracting information from documents and images, data processing is performed without taking the user's emotional state into consideration, which means that errors or uncertainties in the extracted data are not properly communicated to the user, resulting in reduced work efficiency. Furthermore, when users feel stressed or confused, it becomes even more difficult to make corrections.

[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0379] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for the server to analyze the user's emotional state simultaneously with image capture and transmit the image to the server; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for adjusting the notification content based on the user's emotional state; and a means for re-verifying the corrected data and finally saving and exporting it to a required external system. This enables appropriate notifications that take the user's emotional state into consideration, thereby improving the reliability of data and reducing the burden on the user.

[0380] A "terminal" is an electronic device that a user uses to capture documents and images and upload them to a server.

[0381] "Server" refers to a computer system that receives and stores images and data sent from the terminals, and performs preprocessing, data extraction, validation, and export.

[0382] "Preprocessing" refers to image processing operations such as deskewing, noise removal, and resolution adjustment that are performed on uploaded images.

[0383] Optical character recognition (OCR) is a technology that analyzes letters and numbers in an image and converts them into digital text data.

[0384] Artificial intelligence (AI) is a technology that analyzes image and text data, automatically recognizing specific patterns and contexts, and extracting and verifying data.

[0385] "Emotional state" refers to the user's psychological and emotional state (e.g., Happy, Angry, Sad, etc.) obtained by analyzing the user's facial image.

[0386] "Validation" is the process of verifying whether the extracted data is correct, checking whether it conforms to a specific format or range.

[0387] A "request for correction" means that if there are any uncertainties as a result of the verification, the user is notified of the details and asked to correct the data.

[0388] "Adjusting notification content" means appropriately changing the wording and method of notification based on the user's emotional state.

[0389] "Export" refers to sending or outputting data that has undergone final verification and correction to the required external system.

[0390] A "unique identifier" is unique identification information assigned to each image to distinguish it from other images.

[0391] "Metadata" is information mainly about image data (e.g., upload date and time, user ID, etc.), and is used by the server to manage image data.

[0392] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[0393] Acquiring and uploading images

[0394] Terminal

[0395] A user opens an application on their device and uses the camera function to take a document or image. At this time, the device's camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., Happy, Angry, Sad, etc.). After the user selects an image or completes the capture, the device generates an HTTP request to send the document or image and emotion information to the server. The transmitted data includes Base64-encoded image data and emotion data.

[0396] server

[0397] The server receives HTTP requests from devices and stores images and emotion data. Images are assigned a unique identifier, and metadata such as upload date and time, user ID, etc. are also stored in a database.

[0398] Image preprocessing

[0399] server

[0400] The server then begins preprocessing the saved image. This includes correcting the image's deskew (for example, using a Hough transform), applying a noise reduction filter, and adjusting the resolution. These processes make the characters and symbols in the image easier to recognize. The preprocessed image is then stored in a temporary folder.

[0401] Running the Data Extraction

[0402] server

[0403] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the letters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. The AI ​​model highlights specific keywords and phrases, improving the accuracy of the extracted data.

[0404] Data validation and completion

[0405] server

[0406] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. During the validation process, the emotion engine reassess the user's emotional state and adjusts the notification content if a specific emotional state (e.g., stress or confusion) is detected. Based on the validation results, the server sends a notification to the user requesting corrections.

[0407] User

[0408] The user receives a notification on their device, checks the content, and corrects the data if necessary, generating an HTTP request to send the corrected data to the server.

[0409] server

[0410] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved in the database as the final data. The verification process and the content of the correction request are adjusted appropriately based on the emotion data.

[0411] Saving and exporting data

[0412] server

[0413] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). The data is sent using protocols such as RESTful APIs or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database.

[0414] Examples of concrete examples and prompts

[0415] Example 1: Receipt processing in the finance department

[0416] The user takes a photo of the receipt using their device and uploads it to the server. At the same time, the emotion engine analyzes the user's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification adjusted based on the emotion data is sent to the user. The user receives the notification and makes any necessary corrections. The corrected data is then sent to the server, which then performs another verification. Finally, the data is exported to the accounting software, and a completion notification is sent to the user.

[0417] Example prompt (for the finance department):

[0418] Please take a picture of your receipt and upload it.

[0419] "Analyzing emotions..."

[0420] "Image preprocessing complete, optical character recognition begins."

[0421] "Validating extracted data..."

[0422] "There is a discrepancy in the amount verification. Please check again."

[0423] Example 2: Medical record processing in a healthcare facility

[0424] Medical staff scan handwritten medical records and upload them to the server. At the same time, an emotion engine analyzes the medical staff's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an AI model to extract data such as the patient's name, diagnosis, and prescribed medications. The extracted data is automatically verified, and if there are any problems, a notification with adjustments based on the emotion data is sent to the medical staff. The medical staff receives the notification, makes any necessary corrections, and sends the data to the server. The server re-verifies the corrected data, exports it to the electronic medical record system, and sends a completion notification to the medical staff.

[0425] Example prompts (for medical facilities)

[0426] "Scan and upload your medical records."

[0427] "Analyzing emotions..."

[0428] "Image preprocessing complete, data extraction begins."

[0429] "Validating extracted data..."

[0430] "There is a discrepancy in the diagnostic results confirmation. Please check again."

[0431] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0432] Step 1:

[0433] Device-based image capture and emotion analysis

[0434] The user opens the device application and uses the camera function to take a photo of a document or image. At this time, the device camera also captures the user's face, and the emotion engine analyzes the user's emotional state. Emotional data (e.g., Happy, Angry, Sad, etc.) is generated as a result of the analysis. Specifically, the device application calls the camera module and captures the image and a photo of the user's face.

[0435] Input: Documents and images captured by the camera, and the user's facial image

[0436] Output: Photographed documents, image data, analyzed emotional data

[0437] Step 2:

[0438] Uploading images and emotion data to the server

[0439] After the user selects or captures an image, the device sends the document, image, and emotion information to the server as an HTTP request. The sent data includes Base64-encoded image data and emotion data. Specifically, the device application generates an HTTP request and sends the data to the server's API endpoint.

[0440] Input: Photographed documents, image data, analyzed emotional data

[0441] Output: HTTP request sent to the server

[0442] Step 3:

[0443] Receiving and storing data by the server

[0444] The server receives the HTTP request sent from the device and saves the image and emotion data. A unique identifier is assigned to the image data, and metadata such as the upload date and time and user ID are also saved in the database. Specifically, the server analyzes the request and creates an entry to save in the database.

[0445] Input: HTTP request (image data, emotion data, metadata)

[0446] Output: Image data and metadata stored in a database

[0447] Step 4:

[0448] Image preprocessing by the server

[0449] The server begins preprocessing the saved image. This includes correcting the image's tilt using a Hough transform, applying a noise reduction filter, and adjusting the resolution. This makes it easier to recognize characters and numbers in the image. The preprocessed image is stored in a temporary folder. Specific operations involve the use of an image processing library (e.g., OpenCV).

[0450] Input: Saved image data

[0451] Output: Preprocessed image data

[0452] Step 5:

[0453] Server performs data extraction

[0454] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the characters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. Specific operations involve the use of APIs for the OCR engine and the AI ​​model.

[0455] Input: Preprocessed image data

[0456] Output: Extracted text data

[0457] Step 6:

[0458] Server-based data validation and completion

[0459] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, the emotion engine reassess the user's emotional state, and if a specific emotional state (e.g., stress or confusion) is detected, the server adjusts the notification content and sends a notification to the user requesting corrections. Specific operations include the execution of the validation algorithm and the emotion analysis engine.

[0460] Input: Extracted text data, emotion data

[0461] Output: Notification content, correction request notification

[0462] Step 7:

[0463] User Modifications

[0464] The user receives the notification on the device, checks the content, and corrects the data if necessary. An HTTP request is generated to send the corrected data to the server. Specifically, a user interface is provided, and the data is sent again after the user makes the corrections.

[0465] Input: Notification of correction request, data to be corrected

[0466] Output: Corrected data sent to the server

[0467] Step 8:

[0468] Server verifies and saves modified data

[0469] The server receives the corrected data from the user and performs a re-verification. If the corrected data is confirmed to be correct, it is saved as the final data in the database. The verification process and the content of the correction request are adjusted appropriately based on the emotion data. Specifically, the re-verification algorithm is executed and the data is saved in the database.

[0470] Input: Modified text data

[0471] Output: Final saved data

[0472] Step 9:

[0473] Saving and exporting data

[0474] The server securely stores the completed dataset in a database. If necessary, the data is exported to an external system (e.g., accounting software, electronic medical record system, etc.). The data is sent using protocols such as RESTful API or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database. The specific operation is the execution of the export module.

[0475] Input: Final saved data

[0476] Output: Data exported to external systems, completion notification

[0477] (Application example 2)

[0478] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0479] Although the product quality inspection process in modern factories is highly automated, it still requires a lot of manual work and human intervention, which can reduce efficiency. In particular, when inspectors are stressed or confused, the likelihood of errors or inaccurate data increases. Therefore, there is a need not only to automate the quality inspection process, but also to develop a system that can monitor the emotional state of inspectors and respond appropriately.

[0480] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0481] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting correction if there is any uncertainty; a means for re-verifying the corrected data and finally saving it or exporting it to a required external system; a means for photographing and uploading product images and having an emotion engine for analyzing the emotional state of the inspector; and a means for generating and sending a notification according to the inspector's state based on the received emotion information, thereby improving the efficiency and accuracy of the product quality inspection process.

[0482] "Document or Image" refers to visual data captured or saved by a user using a device.

[0483] "Server" means a computer on a network that processes, stores, and analyzes data uploaded by users.

[0484] "Preprocessing" refers to initial data processing such as tilt correction, noise removal, and resolution adjustment performed on uploaded images.

[0485] "Optical character recognition technology" refers to technology that converts character information in an image into digital text.

[0486] "Artificial intelligence" refers to computer programs that analyze and learn from large amounts of data and perform various tasks automatically.

[0487] "Validation" refers to the process of verifying the accuracy of extracted data.

[0488] "Correction" refers to a user making corrections to data that is determined to be uncertain during the validation process.

[0489] "Export" refers to sending processed data to an external system or application.

[0490] "Emotion engine" refers to a system that analyzes a user's emotional state from facial images.

[0491] "Emotional Information" refers to data on the user's emotional state analyzed by the emotion engine.

[0492] "Generating a notification" refers to the process of creating and sending a message to a user based on specific information.

[0493] The system embodying this invention combines multiple advanced technologies to improve the efficiency of product quality inspections by factory robots and monitor the status of inspectors. Specific hardware, software, and system configurations are shown below.

[0494] Hardware and software used

[0495] Hardware

[0496] Device: The device (smartphone, tablet, etc.) where a user takes or uploads documents or images.

[0497] Robot: Industrial robot for quality inspection.

[0498] Camera: A high-resolution camera that takes images of the product.

[0499] Device with emotion engine: A camera for analyzing the emotional state of users and inspectors.

[0500] software

[0501] OCR engine: For example, use Tesseract OCR.

[0502] AI Model: Custom model using TensorFlow or PyTorch.

[0503] Image processing: Uses the OpenCV library.

[0504] Database: Use MySQL or PostgreSQL.

[0505] Sentiment analysis engine: Uses Amazon Rekognition and Microsoft Azure Face API.

[0506] System Overview

[0507] To inspect the quality of products, the factory robot first uses a high-resolution camera to take a picture of the product and uploads the image data to a server. At the same time, the robot uses an electronic device to take a picture of the inspector's face, which is then used by an emotion analysis engine to analyze the inspector's emotional state.

[0508] The server performs preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image. The preprocessed image is converted into text data through an OCR engine, and the data undergoes contextual analysis using an AI model. The extracted data is automatically verified, and notifications are sent to the user if necessary.

[0509] Based on the emotional information analyzed by the emotion analysis engine, the server generates and sends notifications according to the inspector's emotional state. If the inspector is in a specific emotional state, such as stress or confusion, the server will send a notification urging them to improve their processing or reconfirm their actions.

[0510] Finally, the validated data set is stored in a database and can be exported to external quality control systems if required, improving the efficiency and accuracy of the product quality inspection process.

[0511] Examples of concrete examples and prompts

[0512] Specific examples

[0513] Consider a scenario in which a factory robot is performing an inspection. The robot takes a picture of the product and uploads it to a server. At the same time, it takes a facial image of the inspector to obtain emotional information. The server analyzes the image, extracts and verifies the necessary information, and if the inspector is feeling stressed, a notification is sent to encourage process improvement based on that state.

[0514] Prompt Sentence Examples

[0515] We will create an application that automates the product quality inspection process performed by factory robots. A high-resolution camera takes images of the product and an OCR engine (Tesseract) extracts information. The inspector's face is captured and an emotion analysis engine (Amazon Rekognition) evaluates their emotional state and sends notifications as needed. The data is further analyzed by an AI model (TensorFlow) and exported to a quality control system.

[0516] Hardware used: industrial robots, high-resolution cameras, emotion analysis devices

[0517] Software used: Tesseract OCR, TensorFlow, OpenCV, MySQL, Amazon Rekognition

[0518] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0519] Step 1:

[0520] A user takes a picture of a product using a terminal, and the terminal simultaneously captures a facial image of the inspector. The input is the captured product image and the facial image of the inspector, and the output is a request sent from the terminal to the server. With this operation, the terminal prepares to send the image data and facial image data to the server.

[0521] Step 2:

[0522] The server receives a request from the device and stores the image data and facial image data. The input is the image data and facial image data sent from the device, and the output is the image data and facial image data stored in the database, as well as metadata (e.g., the date and time of the photo, a unique identifier, etc.). With this operation, the server manages the received data.

[0523] Step 3:

[0524] The server starts preprocessing on the stored product image. The input is the stored product image, and the output is the preprocessed product image. Preprocessing includes deskewing, noise removal, and resolution adjustment. Through this operation, the server improves the image quality so that important information in the image can be more easily extracted.

[0525] Step 4:

[0526] The server sends the preprocessed image to the OCR engine to perform optical character recognition. The input is the preprocessed image and the output is the extracted text data. In this operation, the server converts the character and numeric information from the product image into digital text.

[0527] Step 5:

[0528] The server inputs the text data extracted by the OCR engine into the AI ​​model and performs contextual analysis. The input is the text data extracted by OCR, and the output is the analyzed data. In this process, the server adds contextual information to improve the accuracy of the data.

[0529] Step 6:

[0530] The server automatically validates the parsed data. The input is the text data parsed by the AI ​​model, and the output is the validation result. Validation includes checking the date format and numeric range. With this operation, the server verifies whether the data is accurate.

[0531] Step 7:

[0532] The facial image of the inspector is sent to the emotion analysis engine to analyze the emotional state. The input is the facial image of the inspector, and the output is emotion information. In this operation, the server evaluates the emotional state of the inspector (e.g., stress, confusion).

[0533] Step 8:

[0534] The server generates a notification based on the emotion information and sends it to the user. The input is the result of the validation and emotion analysis, and the output is a notification to the user. In this operation, the server provides the necessary notification to the user and prompts them to improve the process.

[0535] Step 9:

[0536] The user receives the notification on the terminal and makes the necessary corrections. The input is the notification from the server, and the output is the corrected data. This operation allows the user to smoothly correct data.

[0537] Step 10:

[0538] The server again verifies the modified data from the user and saves the final data. The input is the modified data, and the output is the final saved data. With this operation, the server confirms the final data set.

[0539] Step 11:

[0540] The server exports the final dataset to an external quality control system as needed and sends a completion notification to the user. The inputs are the final data and information from the external system, and the outputs are the export results and a completion notification. Through this operation, the server achieves data integration.

[0541] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0542] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0543] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0544] [Second embodiment]

[0545] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0546] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0547] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0548] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0549] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0550] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0551] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0552] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0553] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0554] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0555] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0556] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0557] The present invention is a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[0558] Program processing

[0559] 1. Acquiring and uploading images

[0560] Terminal

[0561] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request to send the photo or image to the server.

[0562] server

[0563] The server receives an image upload request and saves the image in storage. The saved image is assigned a unique identifier, and its metadata (upload date and time, user ID, image ID, etc.) is also saved in the database.

[0564] 2. Image Preprocessing

[0565] server

[0566] After capturing the image, the server begins pre-processing, which includes image deskewing, noise reduction, and resolution adjustment, making the characters and numbers in the image easier to recognize.

[0567] 3. Performing data extraction

[0568] server

[0569] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital data. The AI ​​model then further examines the extracted data based on contextual information to identify specific fields (such as names, dates, or numbers).

[0570] 4. Data verification and completion

[0571] server

[0572] The extracted data is automatically validated by the server, for example to ensure that date formats and numeric ranges are correct. If validation fails, the server notifies the user and asks them to correct the error.

[0573] User

[0574] The user receives the terminal notification, makes any necessary corrections, and then transmits the correction data to the server.

[0575] server

[0576] The server will then re-verify the corrected data, and once verification is complete, the data will be saved and go on to the next step.

[0577] 5. Saving and Exporting Data

[0578] server

[0579] The completed dataset is securely stored on the server. If necessary, the data can be exported to external applications or systems (e.g., accounting software, electronic medical record systems, etc.). If the export was successful, the server notifies the user.

[0580] Specific examples

[0581] Example 1: Processing receipts in the finance department

[0582] Terminal

[0583] The user takes a photo of the receipt using the terminal and uploads it to the server.

[0584] server

[0585] The server receives the image, performs deskewing, noise removal, and resolution adjustment. The pre-processed image is passed through an OCR engine, which extracts data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any ambiguities.

[0586] User

[0587] The user makes the corrections and sends the data to the server, which re-verifies the corrections and finally exports them to the accounting software and notifies the user of the completion.

[0588] Example 2: Medical record processing in a healthcare facility

[0589] Terminal

[0590] Medical staff scan handwritten medical records and upload them to a server.

[0591] server

[0592] The server receives the images and performs deskewing, noise reduction, and resolution adjustment. The preprocessed images are then run through an AI model to extract data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically verified, and medical staff are notified if there are any issues.

[0593] User

[0594] The medical staff makes the corrections and sends the data to the server, which re-verifies the corrected data, exports it to the electronic medical record system, and notifies the medical staff of the completion.

[0595] In this way, this system improves productivity throughout the entire business by efficiently extracting information from documents and images and automatically processing data.

[0596] The processing flow will be explained below.

[0597] Step 1:

[0598] Terminal

[0599] A user launches an application on the device and uses the capture function to capture a document or image, or select an existing image. After the user completes the capture or selection, the device generates an upload request to send the image to the server.

[0600] Step 2:

[0601] server

[0602] The server receives an image upload request, assigns a unique identifier to the uploaded image, saves it in storage, and stores metadata such as the upload date and time, user ID, and image ID in a database.

[0603] Step 3:

[0604] server

[0605] The server then performs preprocessing on the stored image, which includes image deskewing, noise reduction, and resolution adjustment. Deskewing adjusts the text in the image so that it is level, noise reduction removes unwanted background elements, and resolution adjustment makes the details of the characters clearer.

[0606] Step 4:

[0607] server

[0608] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. Optionally, the AI ​​model analyzes the context and improves the accuracy of the extracted data.

[0609] Step 5:

[0610] server

[0611] Specific fields (e.g., names, dates, numbers, etc.) are identified from the extracted text data and saved as structured data (e.g., JSON, CSV format) for subsequent validation and export processes.

[0612] Step 6:

[0613] server

[0614] The extracted data is automatically validated, including checking date formats, numeric ranges, required fields, etc. If any errors are found during the validation process, the server notifies the user and asks them to correct them.

[0615] Step 7:

[0616] Terminal

[0617] The user receives the notification and modifies the data on the terminal screen. Once the modification is complete, the user generates a request to send the modified data to the server.

[0618] Step 8:

[0619] server

[0620] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be accurate, it is saved as the final data. Once the correction process is complete and the verification is successful, the server proceeds to the next step.

[0621] Step 9:

[0622] server

[0623] Securely store the completed dataset in a database. Once the storage process is complete, generate API requests to export the data to any external systems required (e.g., accounting software, electronic medical record systems).

[0624] Step 10:

[0625] server

[0626] If the export is successful, the server records the result in the database and sends the user a notification of completion, including confirmation of the destination system and the contents of the data.

[0627] This detailed process at each step allows for efficient image-based data extraction and automated data processing, improving overall business productivity.

[0628] Example 1

[0629] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0630] In modern business processes and healthcare facilities, manual information extraction and validation from documents and images is time-consuming and error-prone. This reduces operational efficiency and reduces reliability in critical data processing. Furthermore, there is a need to export extracted data quickly and accurately to external systems. To address these challenges, an automated system is needed.

[0631] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0632] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (such as deskewing, noise removal, and resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for re-verifying the corrected data and finally saving and exporting it to required external systems; a means for identifying specific fields (such as names, dates, and numbers) of the extracted data and saving them as structured data; and a means for examining the data based on a generative AI model and generating automated notifications using prompts. This enables the automation of information extraction and verification from documents and images, improving business efficiency and the reliability of data processing.

[0633] A "user" is an entity that uses the system to capture and upload documents and images.

[0634] A "terminal" is a device used by a user, which has the function of taking pictures of documents and images and uploading them to a server.

[0635] A "server" is a computer system that pre-processes images uploaded by users and performs functions such as data extraction, validation, storage and export.

[0636] "Upload" refers to the act of sending image or document data from a terminal to a server.

[0637] "Preprocessing" refers to processing such as tilt correction, noise removal, and resolution adjustment on the image to facilitate subsequent data extraction.

[0638] "Tilt correction" is a process of adjusting the orientation of an image so that the characters and figures in the image are easier to recognize.

[0639] "Noise reduction" is a process that removes unnecessary elements from an image to make it clearer.

[0640] "Resolution adjustment" is a process of changing the image resolution to an appropriate level.

[0641] Optical character recognition (OCR) is a technology that recognizes text within an image and converts it into digital data.

[0642] "Artificial intelligence (AI)" is the technology that enables computers to learn, reason, and solve problems like humans.

[0643] "Data validation" is the process of checking whether extracted data is accurate and valid.

[0644] "Notification" is a message or signal that notifies the user when there is a validation result or uncertain data.

[0645] "Correction" refers to the act of the user correcting data based on a notification from the server.

[0646] "Storage" is the act of securely recording verified data in storage.

[0647] "Export" is the act of sending stored data to an external system.

[0648] A "specific field" is a data item that indicates specific information such as a name, date, or number.

[0649] "Structured data" is data that is organized according to a predefined data format.

[0650] A "generative AI model" is an artificial intelligence algorithm that is trained to perform a specific task.

[0651] A "prompt" is a sentence of text that instructs a generative AI model on a task based on specific input.

[0652] This invention is a system in which a user uses a terminal to take a photo of a document or image and upload it to a server, which then extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[0653] Hardware and software used

[0654] Terminal

[0655] The device used by the user is a device with a photographic function, such as a smartphone, tablet, or digital camera. The device is responsible for sending the photographed or selected images to the server. The application running on the device should preferably have an internet connection for data transmission and an image editing function.

[0656] server

[0657] The server performs the functions of data reception, image pre-processing, data extraction, validation, storage, and export, using the following software:

[0658] Image preprocessing library (e.g. OpenCV)

[0659] Optical Character Recognition (OCR) engine (e.g., Tesseract OCR)

[0660] Artificial Intelligence (AI) models (e.g., Google Cloud Vision API)

[0661] Database system (e.g. MySQL)

[0662] Notification systems (e.g., Firebase Cloud Messaging)

[0663] Data processing and calculation in the program processing flow

[0664] 1. Acquiring and uploading images

[0665] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request and sends the image data and metadata to the server, including information such as the user ID and image ID.

[0666] 2. Image Preprocessing

[0667] The server performs deskew, noise reduction, and resolution adjustment on the images received, making it easier to recognize characters and numbers. The OpenCV library is used for these preprocessing steps.

[0668] 3. Data extraction

[0669] The preprocessed images are then fed into an optical character recognition (OCR) or artificial intelligence (AI) model, using Tesseract OCR to extract text data from the image and Google Cloud Vision API to identify specific fields (such as names, dates, or amounts).

[0670] 4. Data verification and completion

[0671] The extracted data is validated by the server. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. If there is any problem with the validation, the server notifies the user and asks them to make corrections. After the user makes the necessary corrections, the corrected data is sent back to the server, which then validates the data again.

[0672] 5. Saving and Exporting Data

[0673] The complete dataset is stored on the server and can be exported to external applications or systems as needed, with the user notified if the export was successful.

[0674] Specific examples

[0675] Example 1: Processing receipts in the finance department

[0676] The user takes a photo of the receipt using a smartphone app and uploads it to the server. The server then performs image deskewing, noise reduction, and resolution adjustment. The preprocessed image is then input into the Tesseract OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any uncertainties. If the user corrects any errors and submits it to the server again, the server revalidates the data, exports it to the accounting software, and notifies the user of completion.

[0677] Example 2: Medical record processing in a healthcare facility

[0678] Medical staff scan handwritten medical records and upload them to the server, which then performs image deskewing, noise removal, and resolution adjustment. The preprocessed images are then input into the Google Cloud Vision API, which extracts data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically validated, and if there are any problems, the medical staff is notified. After the medical staff corrects any errors and resubmits the data to the server, the server revalidates the data, exports it to the electronic medical record system, and notifies the medical staff of its completion.

[0679] Prompt Sentence Examples

[0680] Below are some example prompts using generative AI models:

[0681] "Input the following image into an OCR engine and extract data such as the amount, transaction date, and store name. After extraction, verify this data and, if necessary, notify the user."

[0682] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0683] Step 1:

[0684] User

[0685] The user uses a device to take a photo of a document or image, or to select an existing image. To do this, the user uses an application on the device. The data input here is the image taken or selected by the user, and the image data is output. In a specific example, the user might take a photo of a receipt using a camera app on their smartphone.

[0686] Step 2:

[0687] Terminal

[0688] The device internally checks the captured or selected image data and generates an upload request to send to the server. The request includes metadata such as the user ID and image ID. The input is the image data captured or selected by the user and the metadata, and the output is an upload request sent to the server. Here, the data is sent using an Internet connection.

[0689] Step 3:

[0690] server

[0691] The server receives an image upload request and saves the image data and metadata in storage. When saving, it assigns a unique identifier to the image and records the metadata in a database. The input is the image data and metadata sent from the device, and the output is the image data saved in storage and the metadata saved in the database. For example, the image data is saved in a dedicated folder, and the metadata is recorded in a database table.

[0692] Step 4:

[0693] server

[0694] The server retrieves the image from storage and begins preprocessing. In preprocessing, the image is tilted and rotated to the correct orientation. Noise reduction is also performed to remove unnecessary parts. Furthermore, the image resolution is adjusted to make text and numbers easier to read. The input is the image data retrieved from storage, and the output is the preprocessed image data. Here, the OpenCV library is used to preprocess the image.

[0695] Step 5:

[0696] server

[0697] The server inputs the preprocessed image data into an optical character recognition (OCR) engine or artificial intelligence (AI) model. Tesseract OCR is used to recognize text in the image and convert it into digital data. Google Cloud Vision API is then used to identify specific fields (such as names, dates, or amounts) based on contextual information from the extracted text. The input is the preprocessed image data, and the output is the extracted text data.

[0698] Step 6:

[0699] server

[0700] The server validates the extracted data. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. The input is the extracted text data, and the output is the validation result. If there is any validation failure, the server notifies the user and asks them to correct it. This notification is sent using Firebase Cloud Messaging or similar.

[0701] Step 7:

[0702] User

[0703] The user receives the notification on the device and makes the necessary corrections. Here, the input is the notification from the server, and the output is the corrected data. For example, the user manually corrects misreadings made by OCR on the device screen.

[0704] Step 8:

[0705] Terminal

[0706] The user modifies the data and sends it back to the server. The input is the modified data, and the output is the upload request sent again.

[0707] Step 9:

[0708] server

[0709] The server re-verifies the modified data. If the re-verification is successful, the data proceeds to the next processing step. The input is the modified data submitted by the user, and the output is the re-verified data.

[0710] Step 10:

[0711] server

[0712] The validated dataset is securely stored and exported to external applications or systems as needed. For example, sending data to accounting software or electronic medical record systems. If the export is successful, the server sends a notification to the user. The input is the validated data, and the output is the export result to the external system and a notification to the user.

[0713] (Application example 1)

[0714] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0715] In logistics centers, the receiving and inventory management of items is often done manually, which can lead to problems such as data entry errors and the time required for confirmation work. For this reason, there is a need for a system that can manage item information efficiently and accurately.

[0716] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0717] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user if there is any uncertainty and requesting correction; a means for re-verifying the corrected data and finally saving and exporting it to the necessary external system; and a means for automatically recognizing item information at the logistics center and providing a function for managing data on the terminal, thereby enabling efficient and accurate management of item information at the logistics center.

[0718] A "terminal" is a device that a user uses to take documents and images and upload the data to a server.

[0719] A "server" is a computer system that receives image data uploaded by users and performs pre-processing, data extraction, validation, and export.

[0720] "Preprocessing" refers to the process of correcting tilt, removing noise, and adjusting resolution of uploaded images.

[0721] "Optical character recognition (OCR) technology" is a technology that identifies characters in an image and converts them into digital text data.

[0722] "Artificial intelligence (AI)" is a technology that gives computer systems the ability to understand meaning and context from images and text, and to extract, classify, and verify data.

[0723] "Verification" is the process of checking whether the extracted data is accurate and asking the user to correct any uncertainties.

[0724] "Export" is the process of outputting the final validated data to the required external systems or applications.

[0725] A "logistics center" is a facility that handles logistics operations such as receiving, storing, and shipping items.

[0726] "Item information" refers to data such as the identification information, quantity, and date of receipt of products handled at the logistics center.

[0727] The "database" is an information management system for organizing and storing extracted item information.

[0728] An "inventory management system" is a software system that manages the inventory status of items and records information on incoming and outgoing goods.

[0729] The present invention is a system for improving the efficiency and accuracy of item management in a logistics center. A specific embodiment of this system will be described below.

[0730] System Configuration

[0731] This system consists of a terminal used by users, a server that performs processing, and a database that stores data. The main software used is OpenCV for image preprocessing, Tesseract as an optical character recognition (OCR) engine, and requests for communication.

[0732] Program processing overview

[0733] 1. Acquiring and uploading images

[0734] The device provides a means for users to take photos of items and upload them to the server. The images are sent to the server and assigned a unique identifier. Image metadata (upload date and time, user ID, image ID, etc.) is also generated and stored in a database.

[0735] 2. Image Preprocessing

[0736] The server receives the uploaded image and performs deskewing, noise reduction, and resolution adjustment, which makes the text and numbers in the image more readable.

[0737] 3. Data Extraction

[0738] The server inputs the preprocessed image into an OCR engine (Tesseract) to recognize text in the image and convert it into digital data. Additionally, it uses an artificial intelligence (AI) model to extract item information (item ID, quantity, date of receipt, etc.).

[0739] 4. Data verification and completion

[0740] The extracted data is automatically validated by the server. For example, it checks whether the date format and numeric range are correct. If there are any errors during validation, the user is notified and can make corrections. If the user makes corrections and submits the data to the server again, the server will validate the data again.

[0741] 5. Saving and Exporting Data

[0742] The final verified data is stored in a database and, if required, exported to an inventory management system.

[0743] Specific examples

[0744] Example 1: Taking a photo of an item and exporting the data

[0745] Staff at the logistics center use a terminal to take a photo of an item and upload it to the server. The server receives the image, preprocesses it, and then uses an OCR engine to extract item information. The extracted data is verified and exported to the inventory management system. If there is an error in the data, a notification is sent to the user, who can then make the necessary corrections.

[0746] Prompt Sentence Examples

[0747] "Please take a photo of this item. The system will automatically extract the item information and update it in your inventory."

[0748] In this way, the system of the present invention automates item management in a logistics center, improving work efficiency and accuracy.

[0749] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0750] Step 1:

[0751] A user takes a photo of an item using a device, which captures the item image, generates a unique identifier, and prepares the captured image and its metadata (upload date and time, user ID, image ID, etc.).

[0752] Step 2:

[0753] The device uploads the captured images and metadata to a server, which receives the data and stores them in a database using a unique identifier assigned to the image.

[0754] Step 3:

[0755] The server retrieves the received image from the database and begins pre-processing it, i.e., correcting the image's distortion (image rotation), removing noise (image filtering), and adjusting the resolution (resizing or rescaling). This pre-processing process makes the features in the image more visible.

[0756] Step 4:

[0757] The preprocessed image is input into the server's OCR engine (Tesseract) and OCR processing is performed. During this process, characters within the image are extracted and converted into digital text data. The extracted text data includes the item ID, quantity, and inventory date.

[0758] Step 5:

[0759] The server uses an AI model to more precisely identify item information from the text data extracted by OCR, particularly identifying specific fields (item ID, quantity, inventory date, etc.) and saving them as structured data, utilizing machine learning and natural language processing techniques.

[0760] Step 6:

[0761] The server validates the extracted and identified data, specifically ensuring that date formats and numeric ranges are correct. If validation fails, the server notifies the user, including the specific corrections required.

[0762] Step 7:

[0763] The user receives a notification and makes the necessary corrections on the device, after which the corrected data is sent back to the server.

[0764] Step 8:

[0765] The server will again verify the corrected data resubmitted by the user, and if the verification is successful, the data will be finally saved in the database.

[0766] Step 9:

[0767] The server exports the saved data to an external system (e.g., inventory management system) as needed. Once the export is complete, the server notifies the user.

[0768] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0769] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[0770] Program processing

[0771] 1. Acquiring and uploading images

[0772] Terminal

[0773] The user opens the application on the device and uses the camera function to take a document or image. At the same time, the device captures the user's face, and the emotion engine analyzes the user's emotional state. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[0774] server

[0775] The server receives the upload request and stores the image and emotion data. The image is assigned a unique identifier, and metadata such as the upload date and time and user ID are also stored in the database.

[0776] 2. Image Preprocessing

[0777] server

[0778] The server then performs preprocessing on the stored image, which includes deskewing, noise reduction, and resolution adjustment, making the letters and numbers in the image more recognizable.

[0779] 3. Performing data extraction

[0780] server

[0781] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. The AI ​​model then analyzes the context to further refine the extracted data.

[0782] 4. Data verification and completion

[0783] server

[0784] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, an emotion engine evaluates the user's emotional state. If a specific emotional state (e.g., stress or confusion) is detected, the server sends the user an improved notification.

[0785] User

[0786] The user receives the notification on the device and can modify the data. Once the modification is complete, a request is generated to send the modified data to the server.

[0787] server

[0788] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved as the final data. The verification process and the content of the correction request may be adjusted based on the emotion data.

[0789] 5. Saving and Exporting Data

[0790] server

[0791] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). If the export is successful, the server sends a completion notification to the user and records the export results in the database.

[0792] Specific examples

[0793] Example 1: Processing receipts in the finance department

[0794] Terminal

[0795] The user takes a photo of the receipt using the device and uploads it to the server, while the emotion engine simultaneously analyzes the user's emotional state.

[0796] server

[0797] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification tailored based on the emotion data is sent to the user.

[0798] User

[0799] The user receives a notification, makes the necessary corrections, and sends the corrected data to the server, which then validates it again. Finally, the data is exported to the accounting software, and the user is notified of the completion.

[0800] Example 2: Medical record processing in a healthcare facility

[0801] Terminal

[0802] Medical staff scan handwritten medical records and upload them to the server, and at the same time, the emotion engine analyzes the emotional state of the medical staff.

[0803] server

[0804] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is then passed through an AI model to extract data such as patient name, diagnosis, and prescribed medication. The extracted data is automatically verified, and if there are any issues, a tailored notification based on the emotion data is sent to medical staff.

[0805] User

[0806] The medical staff receives the notification, makes the necessary corrections, and sends the data to the server, which re-verifies the corrected data, exports the data to the electronic medical record system, and notifies the medical staff of the completion.

[0807] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[0808] The processing flow will be explained below.

[0809] Step 1:

[0810] Terminal

[0811] The user launches the application on the device and uses the camera function to take a document or image. At the same time, the device camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., joy, anger, sadness, etc.) in real time. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[0812] Step 2:

[0813] server

[0814] The server receives the upload request and saves the submitted image and emotion data in storage. A unique identifier is assigned to the image, and metadata such as the upload date and time, user ID, and emotion data are recorded in the database.

[0815] Step 3:

[0816] server

[0817] The server begins pre-processing the stored image, straightening the image to make the text easier to read, applying a noise reduction filter to remove unnecessary elements in the image, and adjusting the resolution to make the details of the text and numbers clearer.

[0818] Step 4:

[0819] server

[0820] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital text data. The AI ​​model analyzes the context and improves the accuracy of the extracted data.

[0821] Step 5:

[0822] server

[0823] The text data extracted by the OCR engine or AI model is parsed to identify specific fields (e.g., names, dates, numbers, etc.) and stored as structured data for subsequent validation and export processes.

[0824] Step 6:

[0825] server

[0826] The extracted data is automatically validated, for example, checking whether the date format is correct or whether the numbers are within a reasonable range. If any flaws are detected during the validation process, the server will also evaluate the user's emotional state and send a notification with an appropriate tone if a specific emotional state (e.g., stress, confusion) is detected through the emotion engine.

[0827] Step 7:

[0828] Terminal

[0829] The user receives the notification on the terminal and makes the necessary modifications through the data modification interface. After completing the modifications, the user generates a request to send the data to the server.

[0830] Step 8:

[0831] server

[0832] The server receives the corrected data sent by the user and performs automatic verification again. If the corrected data is confirmed to be correct, it is saved as the final data set.

[0833] Step 9:

[0834] server

[0835] The completed dataset is securely stored. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems). If the export is successful, the server records the result in a database and sends a completion notification to the user.

[0836] Specific examples

[0837] Example 1: Processing receipts in the finance department

[0838] Step 1:

[0839] Terminal

[0840] The user takes a photo of the receipt using the device, and at the same time, the emotion engine analyzes the user's emotional state in real time.

[0841] Step 2:

[0842] server

[0843] The server receives the image and emotion data, stores them in storage, generates a unique identifier and metadata, and records them in a database.

[0844] Step 3:

[0845] server

[0846] Preprocessing such as tilt correction, noise removal, and resolution adjustment is performed on the received image.

[0847] Step 4:

[0848] server

[0849] The preprocessed image is input into an OCR engine to extract data such as the amount, transaction date, and store name.

[0850] Step 5:

[0851] server

[0852] The extracted data is saved as structured data and passed to subsequent validation processing.

[0853] Step 6:

[0854] server

[0855] The extracted data is automatically verified, and if there are any deficiencies, an appropriate notification is sent based on the user's emotional state.

[0856] Step 7:

[0857] Terminal

[0858] The user receives a notification on the device, makes the necessary corrections, and sends the corrected data to the server.

[0859] Step 8:

[0860] server

[0861] The corrected data is verified again, and if it is correct, it is saved as the final data.

[0862] Step 9:

[0863] server

[0864] The final data is exported to accounting software and the results are notified to the user.

[0865] Example 2: Medical record processing in a healthcare facility

[0866] Step 1:

[0867] Terminal

[0868] Medical staff use the device to scan handwritten medical records, while the emotion engine simultaneously analyzes the emotional state of the medical staff in real time.

[0869] Step 2:

[0870] server

[0871] The server receives and stores the image and emotion data, assigns a unique identifier, and generates and records metadata.

[0872] Step 3:

[0873] server

[0874] The saved image is then tilted, noise removed, and resolution adjusted.

[0875] Step 4:

[0876] server

[0877] The pre-processed images are fed into an AI model to extract data such as patient name, diagnosis, and prescribed medications.

[0878] Step 5:

[0879] server

[0880] The extracted data is saved as structured data and passed to subsequent validation processing.

[0881] Step 6:

[0882] server

[0883] The extracted data is automatically verified, and if there are any problems, notifications are sent that are tailored to the emotional state of the medical staff.

[0884] Step 7:

[0885] Terminal

[0886] Medical staff are notified and make any necessary corrections, which are then sent to the server.

[0887] Step 8:

[0888] server

[0889] The corrected data is verified again, and if appropriate, is saved as the final data.

[0890] Step 9:

[0891] server

[0892] The final data is exported to the electronic medical record system and the results are notified to medical staff.

[0893] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[0894] Example 2

[0895] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0896] In conventional systems, when extracting information from documents and images, data processing is performed without taking the user's emotional state into consideration, which means that errors or uncertainties in the extracted data are not properly communicated to the user, resulting in reduced work efficiency. Furthermore, when users feel stressed or confused, it becomes even more difficult to make corrections.

[0897] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0898] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for the server to analyze the user's emotional state simultaneously with image capture and transmit the image to the server; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for adjusting the notification content based on the user's emotional state; and a means for re-verifying the corrected data and finally saving and exporting it to a required external system. This enables appropriate notifications that take the user's emotional state into consideration, thereby improving the reliability of data and reducing the burden on the user.

[0899] A "terminal" is an electronic device that a user uses to capture documents and images and upload them to a server.

[0900] "Server" refers to a computer system that receives and stores images and data sent from the terminals, and performs preprocessing, data extraction, validation, and export.

[0901] "Preprocessing" refers to image processing operations such as deskewing, noise removal, and resolution adjustment that are performed on uploaded images.

[0902] Optical character recognition (OCR) is a technology that analyzes letters and numbers in an image and converts them into digital text data.

[0903] Artificial intelligence (AI) is a technology that analyzes image and text data, automatically recognizing specific patterns and contexts, and extracting and verifying data.

[0904] "Emotional state" refers to the user's psychological and emotional state (e.g., Happy, Angry, Sad, etc.) obtained by analyzing the user's facial image.

[0905] "Validation" is the process of verifying whether the extracted data is correct, checking whether it conforms to a specific format or range.

[0906] A "request for correction" means that if there are any uncertainties as a result of the verification, the user is notified of the details and asked to correct the data.

[0907] "Adjusting notification content" means appropriately changing the wording and method of notification based on the user's emotional state.

[0908] "Export" refers to sending or outputting data that has undergone final verification and correction to the required external system.

[0909] A "unique identifier" is unique identification information assigned to each image to distinguish it from other images.

[0910] "Metadata" is information mainly about image data (e.g., upload date and time, user ID, etc.), and is used by the server to manage image data.

[0911] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[0912] Acquiring and uploading images

[0913] Terminal

[0914] A user opens an application on their device and uses the camera function to take a document or image. At this time, the device's camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., Happy, Angry, Sad, etc.). After the user selects an image or completes the capture, the device generates an HTTP request to send the document or image and emotion information to the server. The transmitted data includes Base64-encoded image data and emotion data.

[0915] server

[0916] The server receives HTTP requests from devices and stores images and emotion data. Images are assigned a unique identifier, and metadata such as upload date and time, user ID, etc. are also stored in a database.

[0917] Image preprocessing

[0918] server

[0919] The server then begins preprocessing the saved image. This includes correcting the image's deskew (for example, using a Hough transform), applying a noise reduction filter, and adjusting the resolution. These processes make the characters and symbols in the image easier to recognize. The preprocessed image is then stored in a temporary folder.

[0920] Running the Data Extraction

[0921] server

[0922] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the letters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. The AI ​​model highlights specific keywords and phrases, improving the accuracy of the extracted data.

[0923] Data validation and completion

[0924] server

[0925] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. During the validation process, the emotion engine reassess the user's emotional state and adjusts the notification content if a specific emotional state (e.g., stress or confusion) is detected. Based on the validation results, the server sends a notification to the user requesting corrections.

[0926] User

[0927] The user receives a notification on their device, checks the content, and corrects the data if necessary, generating an HTTP request to send the corrected data to the server.

[0928] server

[0929] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved in the database as the final data. The verification process and the content of the correction request are adjusted appropriately based on the emotion data.

[0930] Saving and exporting data

[0931] server

[0932] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). The data is sent using protocols such as RESTful APIs or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database.

[0933] Examples of concrete examples and prompts

[0934] Example 1: Receipt processing in the finance department

[0935] The user takes a photo of the receipt using their device and uploads it to the server. At the same time, the emotion engine analyzes the user's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification adjusted based on the emotion data is sent to the user. The user receives the notification and makes any necessary corrections. The corrected data is then sent to the server, which then performs another verification. Finally, the data is exported to the accounting software, and a completion notification is sent to the user.

[0936] Example prompt (for the finance department):

[0937] Please take a picture of your receipt and upload it.

[0938] "Analyzing emotions..."

[0939] "Image preprocessing complete, optical character recognition begins."

[0940] "Validating extracted data..."

[0941] "There is a discrepancy in the amount verification. Please check again."

[0942] Example 2: Medical record processing in a healthcare facility

[0943] Medical staff scan handwritten medical records and upload them to the server. At the same time, an emotion engine analyzes the medical staff's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an AI model to extract data such as the patient's name, diagnosis, and prescribed medications. The extracted data is automatically verified, and if there are any problems, a notification with adjustments based on the emotion data is sent to the medical staff. The medical staff receives the notification, makes any necessary corrections, and sends the data to the server. The server re-verifies the corrected data, exports it to the electronic medical record system, and sends a completion notification to the medical staff.

[0944] Example prompts (for medical facilities)

[0945] "Scan and upload your medical records."

[0946] "Analyzing emotions..."

[0947] "Image preprocessing complete, data extraction begins."

[0948] "Validating extracted data..."

[0949] "There is a discrepancy in the diagnostic results confirmation. Please check again."

[0950] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0951] Step 1:

[0952] Device-based image capture and emotion analysis

[0953] The user opens the device application and uses the camera function to take a photo of a document or image. At this time, the device camera also captures the user's face, and the emotion engine analyzes the user's emotional state. Emotional data (e.g., Happy, Angry, Sad, etc.) is generated as a result of the analysis. Specifically, the device application calls the camera module and captures the image and a photo of the user's face.

[0954] Input: Documents and images captured by the camera, and the user's facial image

[0955] Output: Photographed documents, image data, analyzed emotional data

[0956] Step 2:

[0957] Uploading images and emotion data to the server

[0958] After the user selects or captures an image, the device sends the document, image, and emotion information to the server as an HTTP request. The sent data includes Base64-encoded image data and emotion data. Specifically, the device application generates an HTTP request and sends the data to the server's API endpoint.

[0959] Input: Photographed documents, image data, analyzed emotional data

[0960] Output: HTTP request sent to the server

[0961] Step 3:

[0962] Receiving and storing data by the server

[0963] The server receives the HTTP request sent from the device and saves the image and emotion data. A unique identifier is assigned to the image data, and metadata such as the upload date and time and user ID are also saved in the database. Specifically, the server analyzes the request and creates an entry to save in the database.

[0964] Input: HTTP request (image data, emotion data, metadata)

[0965] Output: Image data and metadata stored in a database

[0966] Step 4:

[0967] Image preprocessing by the server

[0968] The server begins preprocessing the saved image. This includes correcting the image's tilt using a Hough transform, applying a noise reduction filter, and adjusting the resolution. This makes it easier to recognize characters and numbers in the image. The preprocessed image is stored in a temporary folder. Specific operations involve the use of an image processing library (e.g., OpenCV).

[0969] Input: Saved image data

[0970] Output: Preprocessed image data

[0971] Step 5:

[0972] Server performs data extraction

[0973] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the characters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. Specific operations involve the use of APIs for the OCR engine and the AI ​​model.

[0974] Input: Preprocessed image data

[0975] Output: Extracted text data

[0976] Step 6:

[0977] Server-based data validation and completion

[0978] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, the emotion engine reassess the user's emotional state, and if a specific emotional state (e.g., stress or confusion) is detected, the server adjusts the notification content and sends a notification to the user requesting corrections. Specific operations include the execution of the validation algorithm and the emotion analysis engine.

[0979] Input: Extracted text data, emotion data

[0980] Output: Notification content, correction request notification

[0981] Step 7:

[0982] User Modifications

[0983] The user receives the notification on the device, checks the content, and corrects the data if necessary. An HTTP request is generated to send the corrected data to the server. Specifically, a user interface is provided, and the data is sent again after the user makes the corrections.

[0984] Input: Notification of correction request, data to be corrected

[0985] Output: Corrected data sent to the server

[0986] Step 8:

[0987] Server verifies and saves modified data

[0988] The server receives the corrected data from the user and performs a re-verification. If the corrected data is confirmed to be correct, it is saved as the final data in the database. The verification process and the content of the correction request are adjusted appropriately based on the emotion data. Specifically, the re-verification algorithm is executed and the data is saved in the database.

[0989] Input: Modified text data

[0990] Output: Final saved data

[0991] Step 9:

[0992] Saving and exporting data

[0993] The server securely stores the completed dataset in a database. If necessary, the data is exported to an external system (e.g., accounting software, electronic medical record system, etc.). The data is sent using protocols such as RESTful API or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database. The specific operation is the execution of the export module.

[0994] Input: Final saved data

[0995] Output: Data exported to external systems, completion notification

[0996] (Application example 2)

[0997] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0998] Although the product quality inspection process in modern factories is highly automated, it still requires a lot of manual work and human intervention, which can reduce efficiency. In particular, when inspectors are stressed or confused, the likelihood of errors or inaccurate data increases. Therefore, there is a need not only to automate the quality inspection process, but also to develop a system that can monitor the emotional state of inspectors and respond appropriately.

[0999] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1000] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting correction if there is any uncertainty; a means for re-verifying the corrected data and finally saving it or exporting it to a required external system; a means for photographing and uploading product images and having an emotion engine for analyzing the emotional state of the inspector; and a means for generating and sending a notification according to the inspector's state based on the received emotion information, thereby improving the efficiency and accuracy of the product quality inspection process.

[1001] "Document or Image" refers to visual data captured or saved by a user using a device.

[1002] "Server" means a computer on a network that processes, stores, and analyzes data uploaded by users.

[1003] "Preprocessing" refers to initial data processing such as tilt correction, noise removal, and resolution adjustment performed on uploaded images.

[1004] "Optical character recognition technology" refers to technology that converts character information in an image into digital text.

[1005] "Artificial intelligence" refers to computer programs that analyze and learn from large amounts of data and perform various tasks automatically.

[1006] "Validation" refers to the process of verifying the accuracy of extracted data.

[1007] "Correction" refers to a user making corrections to data that is determined to be uncertain during the validation process.

[1008] "Export" refers to sending processed data to an external system or application.

[1009] "Emotion engine" refers to a system that analyzes a user's emotional state from facial images.

[1010] "Emotional Information" refers to data on the user's emotional state analyzed by the emotion engine.

[1011] "Generating a notification" refers to the process of creating and sending a message to a user based on specific information.

[1012] The system embodying this invention combines multiple advanced technologies to improve the efficiency of product quality inspections by factory robots and monitor the status of inspectors. Specific hardware, software, and system configurations are shown below.

[1013] Hardware and software used

[1014] Hardware

[1015] Device: The device (smartphone, tablet, etc.) where a user takes or uploads documents or images.

[1016] Robot: Industrial robot for quality inspection.

[1017] Camera: A high-resolution camera that takes images of the product.

[1018] Device with emotion engine: A camera for analyzing the emotional state of users and inspectors.

[1019] software

[1020] OCR engine: For example, use Tesseract OCR.

[1021] AI Model: Custom model using TensorFlow or PyTorch.

[1022] Image processing: Uses the OpenCV library.

[1023] Database: Use MySQL or PostgreSQL.

[1024] Sentiment analysis engine: Uses Amazon Rekognition and Microsoft Azure Face API.

[1025] System Overview

[1026] To inspect the quality of products, the factory robot first uses a high-resolution camera to take a picture of the product and uploads the image data to a server. At the same time, the robot uses an electronic device to take a picture of the inspector's face, which is then used by an emotion analysis engine to analyze the inspector's emotional state.

[1027] The server performs preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image. The preprocessed image is converted into text data through an OCR engine, and the data undergoes contextual analysis using an AI model. The extracted data is automatically verified, and notifications are sent to the user if necessary.

[1028] Based on the emotional information analyzed by the emotion analysis engine, the server generates and sends notifications according to the inspector's emotional state. If the inspector is in a specific emotional state, such as stress or confusion, the server will send a notification urging them to improve their processing or reconfirm their actions.

[1029] Finally, the validated data set is stored in a database and can be exported to external quality control systems if required, improving the efficiency and accuracy of the product quality inspection process.

[1030] Examples of concrete examples and prompts

[1031] Specific examples

[1032] Consider a scenario in which a factory robot is performing an inspection. The robot takes a picture of the product and uploads it to a server. At the same time, it takes a facial image of the inspector to obtain emotional information. The server analyzes the image, extracts and verifies the necessary information, and if the inspector is feeling stressed, a notification is sent to encourage process improvement based on that state.

[1033] Prompt Sentence Examples

[1034] We will create an application that automates the product quality inspection process performed by factory robots. A high-resolution camera takes images of the product and an OCR engine (Tesseract) extracts information. The inspector's face is captured and an emotion analysis engine (Amazon Rekognition) evaluates their emotional state and sends notifications as needed. The data is further analyzed by an AI model (TensorFlow) and exported to a quality control system.

[1035] Hardware used: industrial robots, high-resolution cameras, emotion analysis devices

[1036] Software used: Tesseract OCR, TensorFlow, OpenCV, MySQL, Amazon Rekognition

[1037] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1038] Step 1:

[1039] A user takes a picture of a product using a terminal, and the terminal simultaneously captures a facial image of the inspector. The input is the captured product image and the facial image of the inspector, and the output is a request sent from the terminal to the server. With this operation, the terminal prepares to send the image data and facial image data to the server.

[1040] Step 2:

[1041] The server receives a request from the device and stores the image data and facial image data. The input is the image data and facial image data sent from the device, and the output is the image data and facial image data stored in the database, as well as metadata (e.g., the date and time of the photo, a unique identifier, etc.). With this operation, the server manages the received data.

[1042] Step 3:

[1043] The server starts preprocessing on the stored product image. The input is the stored product image, and the output is the preprocessed product image. Preprocessing includes deskewing, noise removal, and resolution adjustment. Through this operation, the server improves the image quality so that important information in the image can be more easily extracted.

[1044] Step 4:

[1045] The server sends the preprocessed image to the OCR engine to perform optical character recognition. The input is the preprocessed image and the output is the extracted text data. In this operation, the server converts the character and numeric information from the product image into digital text.

[1046] Step 5:

[1047] The server inputs the text data extracted by the OCR engine into the AI ​​model and performs contextual analysis. The input is the text data extracted by OCR, and the output is the analyzed data. In this process, the server adds contextual information to improve the accuracy of the data.

[1048] Step 6:

[1049] The server automatically validates the parsed data. The input is the text data parsed by the AI ​​model, and the output is the validation result. Validation includes checking the date format and numeric range. With this operation, the server verifies whether the data is accurate.

[1050] Step 7:

[1051] The facial image of the inspector is sent to the emotion analysis engine to analyze the emotional state. The input is the facial image of the inspector, and the output is emotion information. In this operation, the server evaluates the emotional state of the inspector (e.g., stress, confusion).

[1052] Step 8:

[1053] The server generates a notification based on the emotion information and sends it to the user. The input is the result of the validation and emotion analysis, and the output is a notification to the user. In this operation, the server provides the necessary notification to the user and prompts them to improve the process.

[1054] Step 9:

[1055] The user receives the notification on the terminal and makes the necessary corrections. The input is the notification from the server, and the output is the corrected data. This operation allows the user to smoothly correct data.

[1056] Step 10:

[1057] The server again verifies the modified data from the user and saves the final data. The input is the modified data, and the output is the final saved data. With this operation, the server confirms the final data set.

[1058] Step 11:

[1059] The server exports the final dataset to an external quality control system as needed and sends a completion notification to the user. The inputs are the final data and information from the external system, and the outputs are the export results and a completion notification. Through this operation, the server achieves data integration.

[1060] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1061] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1062] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1063] [Third embodiment]

[1064] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1065] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1066] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1067] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1068] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1069] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1070] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1071] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1072] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1073] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1074] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1075] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1076] The present invention is a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[1077] Program processing

[1078] 1. Acquiring and uploading images

[1079] Terminal

[1080] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request to send the photo or image to the server.

[1081] server

[1082] The server receives an image upload request and saves the image in storage. The saved image is assigned a unique identifier, and its metadata (upload date and time, user ID, image ID, etc.) is also saved in the database.

[1083] 2. Image Preprocessing

[1084] server

[1085] After capturing the image, the server begins pre-processing, which includes image deskewing, noise reduction, and resolution adjustment, making the characters and numbers in the image easier to recognize.

[1086] 3. Performing data extraction

[1087] server

[1088] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital data. The AI ​​model then further examines the extracted data based on contextual information to identify specific fields (such as names, dates, or numbers).

[1089] 4. Data verification and completion

[1090] server

[1091] The extracted data is automatically validated by the server, for example to ensure that date formats and numeric ranges are correct. If validation fails, the server notifies the user and asks them to correct the error.

[1092] User

[1093] The user receives the terminal notification, makes any necessary corrections, and then transmits the correction data to the server.

[1094] server

[1095] The server will then re-verify the corrected data, and once verification is complete, the data will be saved and go on to the next step.

[1096] 5. Saving and Exporting Data

[1097] server

[1098] The completed dataset is securely stored on the server. If necessary, the data can be exported to external applications or systems (e.g., accounting software, electronic medical record systems, etc.). If the export was successful, the server notifies the user.

[1099] Specific examples

[1100] Example 1: Processing receipts in the finance department

[1101] Terminal

[1102] The user takes a photo of the receipt using the terminal and uploads it to the server.

[1103] server

[1104] The server receives the image, performs deskewing, noise removal, and resolution adjustment. The pre-processed image is passed through an OCR engine, which extracts data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any ambiguities.

[1105] User

[1106] The user makes the corrections and sends the data to the server, which re-verifies the corrections and finally exports them to the accounting software and notifies the user of the completion.

[1107] Example 2: Medical record processing in a healthcare facility

[1108] Terminal

[1109] Medical staff scan handwritten medical records and upload them to a server.

[1110] server

[1111] The server receives the images and performs deskewing, noise reduction, and resolution adjustment. The preprocessed images are then run through an AI model to extract data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically verified, and medical staff are notified if there are any issues.

[1112] User

[1113] The medical staff makes the corrections and sends the data to the server, which re-verifies the corrected data, exports it to the electronic medical record system, and notifies the medical staff of the completion.

[1114] In this way, this system improves productivity throughout the entire business by efficiently extracting information from documents and images and automatically processing data.

[1115] The processing flow will be explained below.

[1116] Step 1:

[1117] Terminal

[1118] A user launches an application on the device and uses the capture function to capture a document or image, or select an existing image. After the user completes the capture or selection, the device generates an upload request to send the image to the server.

[1119] Step 2:

[1120] server

[1121] The server receives an image upload request, assigns a unique identifier to the uploaded image, saves it in storage, and stores metadata such as the upload date and time, user ID, and image ID in a database.

[1122] Step 3:

[1123] server

[1124] The server then performs preprocessing on the stored image, which includes image deskewing, noise reduction, and resolution adjustment. Deskewing adjusts the text in the image so that it is level, noise reduction removes unwanted background elements, and resolution adjustment makes the details of the characters clearer.

[1125] Step 4:

[1126] server

[1127] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. Optionally, the AI ​​model analyzes the context and improves the accuracy of the extracted data.

[1128] Step 5:

[1129] server

[1130] Specific fields (e.g., names, dates, numbers, etc.) are identified from the extracted text data and saved as structured data (e.g., JSON, CSV format) for subsequent validation and export processes.

[1131] Step 6:

[1132] server

[1133] The extracted data is automatically validated, including checking date formats, numeric ranges, required fields, etc. If any errors are found during the validation process, the server notifies the user and asks them to correct them.

[1134] Step 7:

[1135] Terminal

[1136] The user receives the notification and modifies the data on the terminal screen. Once the modification is complete, the user generates a request to send the modified data to the server.

[1137] Step 8:

[1138] server

[1139] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be accurate, it is saved as the final data. Once the correction process is complete and the verification is successful, the server proceeds to the next step.

[1140] Step 9:

[1141] server

[1142] Securely store the completed dataset in a database. Once the storage process is complete, generate API requests to export the data to any external systems required (e.g., accounting software, electronic medical record systems).

[1143] Step 10:

[1144] server

[1145] If the export is successful, the server records the result in the database and sends the user a notification of completion, including confirmation of the destination system and the contents of the data.

[1146] This detailed process at each step allows for efficient image-based data extraction and automated data processing, improving overall business productivity.

[1147] Example 1

[1148] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1149] In modern business processes and healthcare facilities, manual information extraction and validation from documents and images is time-consuming and error-prone. This reduces operational efficiency and reduces reliability in critical data processing. Furthermore, there is a need to export extracted data quickly and accurately to external systems. To address these challenges, an automated system is needed.

[1150] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1151] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (such as deskewing, noise removal, and resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for re-verifying the corrected data and finally saving and exporting it to required external systems; a means for identifying specific fields (such as names, dates, and numbers) of the extracted data and saving them as structured data; and a means for examining the data based on a generative AI model and generating automated notifications using prompts. This enables the automation of information extraction and verification from documents and images, improving business efficiency and the reliability of data processing.

[1152] A "user" is an entity that uses the system to capture and upload documents and images.

[1153] A "terminal" is a device used by a user, which has the function of taking pictures of documents and images and uploading them to a server.

[1154] A "server" is a computer system that pre-processes images uploaded by users and performs functions such as data extraction, validation, storage and export.

[1155] "Upload" refers to the act of sending image or document data from a terminal to a server.

[1156] "Preprocessing" refers to processing such as tilt correction, noise removal, and resolution adjustment on the image to facilitate subsequent data extraction.

[1157] "Tilt correction" is a process of adjusting the orientation of an image so that the characters and figures in the image are easier to recognize.

[1158] "Noise reduction" is a process that removes unnecessary elements from an image to make it clearer.

[1159] "Resolution adjustment" is a process of changing the image resolution to an appropriate level.

[1160] Optical character recognition (OCR) is a technology that recognizes text within an image and converts it into digital data.

[1161] "Artificial intelligence (AI)" is the technology that enables computers to learn, reason, and solve problems like humans.

[1162] "Data validation" is the process of checking whether extracted data is accurate and valid.

[1163] "Notification" is a message or signal that notifies the user when there is a validation result or uncertain data.

[1164] "Correction" refers to the act of the user correcting data based on a notification from the server.

[1165] "Storage" is the act of securely recording verified data in storage.

[1166] "Export" is the act of sending stored data to an external system.

[1167] A "specific field" is a data item that indicates specific information such as a name, date, or number.

[1168] "Structured data" is data that is organized according to a predefined data format.

[1169] A "generative AI model" is an artificial intelligence algorithm that is trained to perform a specific task.

[1170] A "prompt" is a sentence of text that instructs a generative AI model on a task based on specific input.

[1171] This invention is a system in which a user uses a terminal to take a photo of a document or image and upload it to a server, which then extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[1172] Hardware and software used

[1173] Terminal

[1174] The device used by the user is a device with a photographic function, such as a smartphone, tablet, or digital camera. The device is responsible for sending the photographed or selected images to the server. The application running on the device should preferably have an internet connection for data transmission and an image editing function.

[1175] server

[1176] The server performs the functions of data reception, image pre-processing, data extraction, validation, storage, and export, using the following software:

[1177] Image preprocessing library (e.g. OpenCV)

[1178] Optical Character Recognition (OCR) engine (e.g., Tesseract OCR)

[1179] Artificial Intelligence (AI) models (e.g., Google Cloud Vision API)

[1180] Database system (e.g. MySQL)

[1181] Notification systems (e.g., Firebase Cloud Messaging)

[1182] Data processing and calculation in the program processing flow

[1183] 1. Acquiring and uploading images

[1184] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request and sends the image data and metadata to the server, including information such as the user ID and image ID.

[1185] 2. Image Preprocessing

[1186] The server performs deskew, noise reduction, and resolution adjustment on the images received, making it easier to recognize characters and numbers. The OpenCV library is used for these preprocessing steps.

[1187] 3. Data extraction

[1188] The preprocessed images are then fed into an optical character recognition (OCR) or artificial intelligence (AI) model, using Tesseract OCR to extract text data from the image and Google Cloud Vision API to identify specific fields (such as names, dates, or amounts).

[1189] 4. Data verification and completion

[1190] The extracted data is validated by the server. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. If there is any problem with the validation, the server notifies the user and asks them to make corrections. After the user makes the necessary corrections, the corrected data is sent back to the server, which then validates the data again.

[1191] 5. Saving and Exporting Data

[1192] The complete dataset is stored on the server and can be exported to external applications or systems as needed, with the user notified if the export was successful.

[1193] Specific examples

[1194] Example 1: Processing receipts in the finance department

[1195] The user takes a photo of the receipt using a smartphone app and uploads it to the server. The server then performs image deskewing, noise reduction, and resolution adjustment. The preprocessed image is then input into the Tesseract OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any uncertainties. If the user corrects any errors and submits it to the server again, the server revalidates the data, exports it to the accounting software, and notifies the user of completion.

[1196] Example 2: Medical record processing in a healthcare facility

[1197] Medical staff scan handwritten medical records and upload them to the server, which then performs image deskewing, noise removal, and resolution adjustment. The preprocessed images are then input into the Google Cloud Vision API, which extracts data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically validated, and if there are any problems, the medical staff is notified. After the medical staff corrects any errors and resubmits the data to the server, the server revalidates the data, exports it to the electronic medical record system, and notifies the medical staff of its completion.

[1198] Prompt Sentence Examples

[1199] Below are some example prompts using generative AI models:

[1200] "Input the following image into an OCR engine and extract data such as the amount, transaction date, and store name. After extraction, verify this data and, if necessary, notify the user."

[1201] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1202] Step 1:

[1203] User

[1204] The user uses a device to take a photo of a document or image, or to select an existing image. To do this, the user uses an application on the device. The data input here is the image taken or selected by the user, and the image data is output. In a specific example, the user might take a photo of a receipt using a camera app on their smartphone.

[1205] Step 2:

[1206] Terminal

[1207] The device internally checks the captured or selected image data and generates an upload request to send to the server. The request includes metadata such as the user ID and image ID. The input is the image data captured or selected by the user and the metadata, and the output is an upload request sent to the server. Here, the data is sent using an Internet connection.

[1208] Step 3:

[1209] server

[1210] The server receives an image upload request and saves the image data and metadata in storage. When saving, it assigns a unique identifier to the image and records the metadata in a database. The input is the image data and metadata sent from the device, and the output is the image data saved in storage and the metadata saved in the database. For example, the image data is saved in a dedicated folder, and the metadata is recorded in a database table.

[1211] Step 4:

[1212] server

[1213] The server retrieves the image from storage and begins preprocessing. In preprocessing, the image is tilted and rotated to the correct orientation. Noise reduction is also performed to remove unnecessary parts. Furthermore, the image resolution is adjusted to make text and numbers easier to read. The input is the image data retrieved from storage, and the output is the preprocessed image data. Here, the OpenCV library is used to preprocess the image.

[1214] Step 5:

[1215] server

[1216] The server inputs the preprocessed image data into an optical character recognition (OCR) engine or artificial intelligence (AI) model. Tesseract OCR is used to recognize text in the image and convert it into digital data. Google Cloud Vision API is then used to identify specific fields (such as names, dates, or amounts) based on contextual information from the extracted text. The input is the preprocessed image data, and the output is the extracted text data.

[1217] Step 6:

[1218] server

[1219] The server validates the extracted data. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. The input is the extracted text data, and the output is the validation result. If there is any validation failure, the server notifies the user and asks them to correct it. This notification is sent using Firebase Cloud Messaging or similar.

[1220] Step 7:

[1221] User

[1222] The user receives the notification on the device and makes the necessary corrections. Here, the input is the notification from the server, and the output is the corrected data. For example, the user manually corrects misreadings made by OCR on the device screen.

[1223] Step 8:

[1224] Terminal

[1225] The user modifies the data and sends it back to the server. The input is the modified data, and the output is the upload request sent again.

[1226] Step 9:

[1227] server

[1228] The server re-verifies the modified data. If the re-verification is successful, the data proceeds to the next processing step. The input is the modified data submitted by the user, and the output is the re-verified data.

[1229] Step 10:

[1230] server

[1231] The validated dataset is securely stored and exported to external applications or systems as needed. For example, sending data to accounting software or electronic medical record systems. If the export is successful, the server sends a notification to the user. The input is the validated data, and the output is the export result to the external system and a notification to the user.

[1232] (Application example 1)

[1233] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1234] In logistics centers, the receiving and inventory management of items is often done manually, which can lead to problems such as data entry errors and the time required for confirmation work. For this reason, there is a need for a system that can manage item information efficiently and accurately.

[1235] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1236] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user if there is any uncertainty and requesting correction; a means for re-verifying the corrected data and finally saving and exporting it to the necessary external system; and a means for automatically recognizing item information at the logistics center and providing a function for managing data on the terminal, thereby enabling efficient and accurate management of item information at the logistics center.

[1237] A "terminal" is a device that a user uses to take documents and images and upload the data to a server.

[1238] A "server" is a computer system that receives image data uploaded by users and performs pre-processing, data extraction, validation, and export.

[1239] "Preprocessing" refers to the process of correcting tilt, removing noise, and adjusting resolution of uploaded images.

[1240] "Optical character recognition (OCR) technology" is a technology that identifies characters in an image and converts them into digital text data.

[1241] "Artificial intelligence (AI)" is a technology that gives computer systems the ability to understand meaning and context from images and text, and to extract, classify, and verify data.

[1242] "Verification" is the process of checking whether the extracted data is accurate and asking the user to correct any uncertainties.

[1243] "Export" is the process of outputting the final validated data to the required external systems or applications.

[1244] A "logistics center" is a facility that handles logistics operations such as receiving, storing, and shipping items.

[1245] "Item information" refers to data such as the identification information, quantity, and date of receipt of products handled at the logistics center.

[1246] The "database" is an information management system for organizing and storing extracted item information.

[1247] An "inventory management system" is a software system that manages the inventory status of items and records information on incoming and outgoing goods.

[1248] The present invention is a system for improving the efficiency and accuracy of item management in a logistics center. A specific embodiment of this system will be described below.

[1249] System Configuration

[1250] This system consists of a terminal used by users, a server that performs processing, and a database that stores data. The main software used is OpenCV for image preprocessing, Tesseract as an optical character recognition (OCR) engine, and requests for communication.

[1251] Program processing overview

[1252] 1. Acquiring and uploading images

[1253] The device provides a means for users to take photos of items and upload them to the server. The images are sent to the server and assigned a unique identifier. Image metadata (upload date and time, user ID, image ID, etc.) is also generated and stored in a database.

[1254] 2. Image Preprocessing

[1255] The server receives the uploaded image and performs deskewing, noise reduction, and resolution adjustment, which makes the text and numbers in the image more readable.

[1256] 3. Data Extraction

[1257] The server inputs the preprocessed image into an OCR engine (Tesseract) to recognize text in the image and convert it into digital data. Additionally, it uses an artificial intelligence (AI) model to extract item information (item ID, quantity, date of receipt, etc.).

[1258] 4. Data verification and completion

[1259] The extracted data is automatically validated by the server. For example, it checks whether the date format and numeric range are correct. If there are any errors during validation, the user is notified and can make corrections. If the user makes corrections and submits the data to the server again, the server will validate the data again.

[1260] 5. Saving and Exporting Data

[1261] The final verified data is stored in a database and, if required, exported to an inventory management system.

[1262] Specific examples

[1263] Example 1: Taking a photo of an item and exporting the data

[1264] Staff at the logistics center use a terminal to take a photo of an item and upload it to the server. The server receives the image, preprocesses it, and then uses an OCR engine to extract item information. The extracted data is verified and exported to the inventory management system. If there is an error in the data, a notification is sent to the user, who can then make the necessary corrections.

[1265] Prompt Sentence Examples

[1266] "Please take a photo of this item. The system will automatically extract the item information and update it in your inventory."

[1267] In this way, the system of the present invention automates item management in a logistics center, improving work efficiency and accuracy.

[1268] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1269] Step 1:

[1270] A user takes a photo of an item using a device, which captures the item image, generates a unique identifier, and prepares the captured image and its metadata (upload date and time, user ID, image ID, etc.).

[1271] Step 2:

[1272] The device uploads the captured images and metadata to a server, which receives the data and stores them in a database using a unique identifier assigned to the image.

[1273] Step 3:

[1274] The server retrieves the received image from the database and begins pre-processing it, i.e., correcting the image's distortion (image rotation), removing noise (image filtering), and adjusting the resolution (resizing or rescaling). This pre-processing process makes the features in the image more visible.

[1275] Step 4:

[1276] The preprocessed image is input into the server's OCR engine (Tesseract) and OCR processing is performed. During this process, characters within the image are extracted and converted into digital text data. The extracted text data includes the item ID, quantity, and inventory date.

[1277] Step 5:

[1278] The server uses an AI model to more precisely identify item information from the text data extracted by OCR, particularly identifying specific fields (item ID, quantity, inventory date, etc.) and saving them as structured data, utilizing machine learning and natural language processing techniques.

[1279] Step 6:

[1280] The server validates the extracted and identified data, specifically ensuring that date formats and numeric ranges are correct. If validation fails, the server notifies the user, including the specific corrections required.

[1281] Step 7:

[1282] The user receives a notification and makes the necessary corrections on the device, after which the corrected data is sent back to the server.

[1283] Step 8:

[1284] The server will again verify the corrected data resubmitted by the user, and if the verification is successful, the data will be finally saved in the database.

[1285] Step 9:

[1286] The server exports the saved data to an external system (e.g., inventory management system) as needed. Once the export is complete, the server notifies the user.

[1287] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1288] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[1289] Program processing

[1290] 1. Acquiring and uploading images

[1291] Terminal

[1292] The user opens the application on the device and uses the camera function to take a document or image. At the same time, the device captures the user's face, and the emotion engine analyzes the user's emotional state. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[1293] server

[1294] The server receives the upload request and stores the image and emotion data. The image is assigned a unique identifier, and metadata such as the upload date and time and user ID are also stored in the database.

[1295] 2. Image Preprocessing

[1296] server

[1297] The server then performs preprocessing on the stored image, which includes deskewing, noise reduction, and resolution adjustment, making the letters and numbers in the image more recognizable.

[1298] 3. Performing data extraction

[1299] server

[1300] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. The AI ​​model then analyzes the context to further refine the extracted data.

[1301] 4. Data verification and completion

[1302] server

[1303] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, an emotion engine evaluates the user's emotional state. If a specific emotional state (e.g., stress or confusion) is detected, the server sends the user an improved notification.

[1304] User

[1305] The user receives the notification on the device and can modify the data. Once the modification is complete, a request is generated to send the modified data to the server.

[1306] server

[1307] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved as the final data. The verification process and the content of the correction request may be adjusted based on the emotion data.

[1308] 5. Saving and Exporting Data

[1309] server

[1310] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). If the export is successful, the server sends a completion notification to the user and records the export results in the database.

[1311] Specific examples

[1312] Example 1: Processing receipts in the finance department

[1313] Terminal

[1314] The user takes a photo of the receipt using the device and uploads it to the server, while the emotion engine simultaneously analyzes the user's emotional state.

[1315] server

[1316] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification tailored based on the emotion data is sent to the user.

[1317] User

[1318] The user receives a notification, makes the necessary corrections, and sends the corrected data to the server, which then validates it again. Finally, the data is exported to the accounting software, and the user is notified of the completion.

[1319] Example 2: Medical record processing in a healthcare facility

[1320] Terminal

[1321] Medical staff scan handwritten medical records and upload them to the server, and at the same time, the emotion engine analyzes the emotional state of the medical staff.

[1322] server

[1323] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is then passed through an AI model to extract data such as patient name, diagnosis, and prescribed medication. The extracted data is automatically verified, and if there are any issues, a tailored notification based on the emotion data is sent to medical staff.

[1324] User

[1325] The medical staff receives the notification, makes the necessary corrections, and sends the data to the server, which re-verifies the corrected data, exports the data to the electronic medical record system, and notifies the medical staff of the completion.

[1326] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[1327] The processing flow will be explained below.

[1328] Step 1:

[1329] Terminal

[1330] The user launches the application on the device and uses the camera function to take a document or image. At the same time, the device camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., joy, anger, sadness, etc.) in real time. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[1331] Step 2:

[1332] server

[1333] The server receives the upload request and saves the submitted image and emotion data in storage. A unique identifier is assigned to the image, and metadata such as the upload date and time, user ID, and emotion data are recorded in the database.

[1334] Step 3:

[1335] server

[1336] The server begins pre-processing the stored image, straightening the image to make the text easier to read, applying a noise reduction filter to remove unnecessary elements in the image, and adjusting the resolution to make the details of the text and numbers clearer.

[1337] Step 4:

[1338] server

[1339] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital text data. The AI ​​model analyzes the context and improves the accuracy of the extracted data.

[1340] Step 5:

[1341] server

[1342] The text data extracted by the OCR engine or AI model is parsed to identify specific fields (e.g., names, dates, numbers, etc.) and stored as structured data for subsequent validation and export processes.

[1343] Step 6:

[1344] server

[1345] The extracted data is automatically validated, for example, checking whether the date format is correct or whether the numbers are within a reasonable range. If any flaws are detected during the validation process, the server will also evaluate the user's emotional state and send a notification with an appropriate tone if a specific emotional state (e.g., stress, confusion) is detected through the emotion engine.

[1346] Step 7:

[1347] Terminal

[1348] The user receives the notification on the terminal and makes the necessary modifications through the data modification interface. After completing the modifications, the user generates a request to send the data to the server.

[1349] Step 8:

[1350] server

[1351] The server receives the corrected data sent by the user and performs automatic verification again. If the corrected data is confirmed to be correct, it is saved as the final data set.

[1352] Step 9:

[1353] server

[1354] The completed dataset is securely stored. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems). If the export is successful, the server records the result in a database and sends a completion notification to the user.

[1355] Specific examples

[1356] Example 1: Processing receipts in the finance department

[1357] Step 1:

[1358] Terminal

[1359] The user takes a photo of the receipt using the device, and at the same time, the emotion engine analyzes the user's emotional state in real time.

[1360] Step 2:

[1361] server

[1362] The server receives the image and emotion data, stores them in storage, generates a unique identifier and metadata, and records them in a database.

[1363] Step 3:

[1364] server

[1365] Preprocessing such as tilt correction, noise removal, and resolution adjustment is performed on the received image.

[1366] Step 4:

[1367] server

[1368] The preprocessed image is input into an OCR engine to extract data such as the amount, transaction date, and store name.

[1369] Step 5:

[1370] server

[1371] The extracted data is saved as structured data and passed to subsequent validation processing.

[1372] Step 6:

[1373] server

[1374] The extracted data is automatically verified, and if there are any deficiencies, an appropriate notification is sent based on the user's emotional state.

[1375] Step 7:

[1376] Terminal

[1377] The user receives a notification on the device, makes the necessary corrections, and sends the corrected data to the server.

[1378] Step 8:

[1379] server

[1380] The corrected data is verified again, and if it is correct, it is saved as the final data.

[1381] Step 9:

[1382] server

[1383] The final data is exported to accounting software and the results are notified to the user.

[1384] Example 2: Medical record processing in a healthcare facility

[1385] Step 1:

[1386] Terminal

[1387] Medical staff use the device to scan handwritten medical records, while the emotion engine simultaneously analyzes the emotional state of the medical staff in real time.

[1388] Step 2:

[1389] server

[1390] The server receives and stores the image and emotion data, assigns a unique identifier, and generates and records metadata.

[1391] Step 3:

[1392] server

[1393] The saved image is then tilted, noise removed, and resolution adjusted.

[1394] Step 4:

[1395] server

[1396] The pre-processed images are fed into an AI model to extract data such as patient name, diagnosis, and prescribed medications.

[1397] Step 5:

[1398] server

[1399] The extracted data is saved as structured data and passed to subsequent validation processing.

[1400] Step 6:

[1401] server

[1402] The extracted data is automatically verified, and if there are any problems, notifications are sent that are tailored to the emotional state of the medical staff.

[1403] Step 7:

[1404] Terminal

[1405] Medical staff are notified and make any necessary corrections, which are then sent to the server.

[1406] Step 8:

[1407] server

[1408] The corrected data is verified again, and if appropriate, is saved as the final data.

[1409] Step 9:

[1410] server

[1411] The final data is exported to the electronic medical record system and the results are notified to medical staff.

[1412] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[1413] Example 2

[1414] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1415] In conventional systems, when extracting information from documents and images, data processing is performed without taking the user's emotional state into consideration, which means that errors or uncertainties in the extracted data are not properly communicated to the user, resulting in reduced work efficiency. Furthermore, when users feel stressed or confused, it becomes even more difficult to make corrections.

[1416] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1417] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for the server to analyze the user's emotional state simultaneously with image capture and transmit the image to the server; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for adjusting the notification content based on the user's emotional state; and a means for re-verifying the corrected data and finally saving and exporting it to a required external system. This enables appropriate notifications that take the user's emotional state into consideration, thereby improving the reliability of data and reducing the burden on the user.

[1418] A "terminal" is an electronic device that a user uses to capture documents and images and upload them to a server.

[1419] "Server" refers to a computer system that receives and stores images and data sent from the terminals, and performs preprocessing, data extraction, validation, and export.

[1420] "Preprocessing" refers to image processing operations such as deskewing, noise removal, and resolution adjustment that are performed on uploaded images.

[1421] Optical character recognition (OCR) is a technology that analyzes letters and numbers in an image and converts them into digital text data.

[1422] Artificial intelligence (AI) is a technology that analyzes image and text data, automatically recognizing specific patterns and contexts, and extracting and verifying data.

[1423] "Emotional state" refers to the user's psychological and emotional state (e.g., Happy, Angry, Sad, etc.) obtained by analyzing the user's facial image.

[1424] "Validation" is the process of verifying whether the extracted data is correct, checking whether it conforms to a specific format or range.

[1425] A "request for correction" means that if there are any uncertainties as a result of the verification, the user is notified of the details and asked to correct the data.

[1426] "Adjusting notification content" means appropriately changing the wording and method of notification based on the user's emotional state.

[1427] "Export" refers to sending or outputting data that has undergone final verification and correction to the required external system.

[1428] A "unique identifier" is unique identification information assigned to each image to distinguish it from other images.

[1429] "Metadata" is information mainly about image data (e.g., upload date and time, user ID, etc.), and is used by the server to manage image data.

[1430] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[1431] Acquiring and uploading images

[1432] Terminal

[1433] A user opens an application on their device and uses the camera function to take a document or image. At this time, the device's camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., Happy, Angry, Sad, etc.). After the user selects an image or completes the capture, the device generates an HTTP request to send the document or image and emotion information to the server. The transmitted data includes Base64-encoded image data and emotion data.

[1434] server

[1435] The server receives HTTP requests from devices and stores images and emotion data. Images are assigned a unique identifier, and metadata such as upload date and time, user ID, etc. are also stored in a database.

[1436] Image preprocessing

[1437] server

[1438] The server then begins preprocessing the saved image. This includes correcting the image's deskew (for example, using a Hough transform), applying a noise reduction filter, and adjusting the resolution. These processes make the characters and symbols in the image easier to recognize. The preprocessed image is then stored in a temporary folder.

[1439] Running the Data Extraction

[1440] server

[1441] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the letters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. The AI ​​model highlights specific keywords and phrases, improving the accuracy of the extracted data.

[1442] Data validation and completion

[1443] server

[1444] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. During the validation process, the emotion engine reassess the user's emotional state and adjusts the notification content if a specific emotional state (e.g., stress or confusion) is detected. Based on the validation results, the server sends a notification to the user requesting corrections.

[1445] User

[1446] The user receives a notification on their device, checks the content, and corrects the data if necessary, generating an HTTP request to send the corrected data to the server.

[1447] server

[1448] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved in the database as the final data. The verification process and the content of the correction request are adjusted appropriately based on the emotion data.

[1449] Saving and exporting data

[1450] server

[1451] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). The data is sent using protocols such as RESTful APIs or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database.

[1452] Examples of concrete examples and prompts

[1453] Example 1: Receipt processing in the finance department

[1454] The user takes a photo of the receipt using their device and uploads it to the server. At the same time, the emotion engine analyzes the user's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification adjusted based on the emotion data is sent to the user. The user receives the notification and makes any necessary corrections. The corrected data is then sent to the server, which then performs another verification. Finally, the data is exported to the accounting software, and a completion notification is sent to the user.

[1455] Example prompt (for the finance department):

[1456] Please take a picture of your receipt and upload it.

[1457] "Analyzing emotions..."

[1458] "Image preprocessing complete, optical character recognition begins."

[1459] "Validating extracted data..."

[1460] "There is a discrepancy in the amount verification. Please check again."

[1461] Example 2: Medical record processing in a healthcare facility

[1462] Medical staff scan handwritten medical records and upload them to the server. At the same time, an emotion engine analyzes the medical staff's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an AI model to extract data such as the patient's name, diagnosis, and prescribed medications. The extracted data is automatically verified, and if there are any problems, a notification with adjustments based on the emotion data is sent to the medical staff. The medical staff receives the notification, makes any necessary corrections, and sends the data to the server. The server re-verifies the corrected data, exports it to the electronic medical record system, and sends a completion notification to the medical staff.

[1463] Example prompts (for medical facilities)

[1464] "Scan and upload your medical records."

[1465] "Analyzing emotions..."

[1466] "Image preprocessing complete, data extraction begins."

[1467] "Validating extracted data..."

[1468] "There is a discrepancy in the diagnostic results confirmation. Please check again."

[1469] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1470] Step 1:

[1471] Device-based image capture and emotion analysis

[1472] The user opens the device application and uses the camera function to take a photo of a document or image. At this time, the device camera also captures the user's face, and the emotion engine analyzes the user's emotional state. Emotional data (e.g., Happy, Angry, Sad, etc.) is generated as a result of the analysis. Specifically, the device application calls the camera module and captures the image and a photo of the user's face.

[1473] Input: Documents and images captured by the camera, and the user's facial image

[1474] Output: Photographed documents, image data, analyzed emotional data

[1475] Step 2:

[1476] Uploading images and emotion data to the server

[1477] After the user selects or captures an image, the device sends the document, image, and emotion information to the server as an HTTP request. The sent data includes Base64-encoded image data and emotion data. Specifically, the device application generates an HTTP request and sends the data to the server's API endpoint.

[1478] Input: Photographed documents, image data, analyzed emotional data

[1479] Output: HTTP request sent to the server

[1480] Step 3:

[1481] Receiving and storing data by the server

[1482] The server receives the HTTP request sent from the device and saves the image and emotion data. A unique identifier is assigned to the image data, and metadata such as the upload date and time and user ID are also saved in the database. Specifically, the server analyzes the request and creates an entry to save in the database.

[1483] Input: HTTP request (image data, emotion data, metadata)

[1484] Output: Image data and metadata stored in a database

[1485] Step 4:

[1486] Image preprocessing by the server

[1487] The server begins preprocessing the saved image. This includes correcting the image's tilt using a Hough transform, applying a noise reduction filter, and adjusting the resolution. This makes it easier to recognize characters and numbers in the image. The preprocessed image is stored in a temporary folder. Specific operations involve the use of an image processing library (e.g., OpenCV).

[1488] Input: Saved image data

[1489] Output: Preprocessed image data

[1490] Step 5:

[1491] Server performs data extraction

[1492] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the characters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. Specific operations involve the use of APIs for the OCR engine and the AI ​​model.

[1493] Input: Preprocessed image data

[1494] Output: Extracted text data

[1495] Step 6:

[1496] Server-based data validation and completion

[1497] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, the emotion engine reassess the user's emotional state, and if a specific emotional state (e.g., stress or confusion) is detected, the server adjusts the notification content and sends a notification to the user requesting corrections. Specific operations include the execution of the validation algorithm and the emotion analysis engine.

[1498] Input: Extracted text data, emotion data

[1499] Output: Notification content, correction request notification

[1500] Step 7:

[1501] User Modifications

[1502] The user receives the notification on the device, checks the content, and corrects the data if necessary. An HTTP request is generated to send the corrected data to the server. Specifically, a user interface is provided, and the data is sent again after the user makes the corrections.

[1503] Input: Notification of correction request, data to be corrected

[1504] Output: Corrected data sent to the server

[1505] Step 8:

[1506] Server verifies and saves modified data

[1507] The server receives the corrected data from the user and performs a re-verification. If the corrected data is confirmed to be correct, it is saved as the final data in the database. The verification process and the content of the correction request are adjusted appropriately based on the emotion data. Specifically, the re-verification algorithm is executed and the data is saved in the database.

[1508] Input: Modified text data

[1509] Output: Final saved data

[1510] Step 9:

[1511] Saving and exporting data

[1512] The server securely stores the completed dataset in a database. If necessary, the data is exported to an external system (e.g., accounting software, electronic medical record system, etc.). The data is sent using protocols such as RESTful API or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database. The specific operation is the execution of the export module.

[1513] Input: Final saved data

[1514] Output: Data exported to external systems, completion notification

[1515] (Application example 2)

[1516] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1517] Although the product quality inspection process in modern factories is highly automated, it still requires a lot of manual work and human intervention, which can reduce efficiency. In particular, when inspectors are stressed or confused, the likelihood of errors or inaccurate data increases. Therefore, there is a need not only to automate the quality inspection process, but also to develop a system that can monitor the emotional state of inspectors and respond appropriately.

[1518] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1519] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting correction if there is any uncertainty; a means for re-verifying the corrected data and finally saving it or exporting it to a required external system; a means for photographing and uploading product images and having an emotion engine for analyzing the emotional state of the inspector; and a means for generating and sending a notification according to the inspector's state based on the received emotion information, thereby improving the efficiency and accuracy of the product quality inspection process.

[1520] "Document or Image" refers to visual data captured or saved by a user using a device.

[1521] "Server" means a computer on a network that processes, stores, and analyzes data uploaded by users.

[1522] "Preprocessing" refers to initial data processing such as tilt correction, noise removal, and resolution adjustment performed on uploaded images.

[1523] "Optical character recognition technology" refers to technology that converts character information in an image into digital text.

[1524] "Artificial intelligence" refers to computer programs that analyze and learn from large amounts of data and perform various tasks automatically.

[1525] "Validation" refers to the process of verifying the accuracy of extracted data.

[1526] "Correction" refers to a user making corrections to data that is determined to be uncertain during the validation process.

[1527] "Export" refers to sending processed data to an external system or application.

[1528] "Emotion engine" refers to a system that analyzes a user's emotional state from facial images.

[1529] "Emotional Information" refers to data on the user's emotional state analyzed by the emotion engine.

[1530] "Generating a notification" refers to the process of creating and sending a message to a user based on specific information.

[1531] The system embodying this invention combines multiple advanced technologies to improve the efficiency of product quality inspections by factory robots and monitor the status of inspectors. Specific hardware, software, and system configurations are shown below.

[1532] Hardware and software used

[1533] Hardware

[1534] Device: The device (smartphone, tablet, etc.) where a user takes or uploads documents or images.

[1535] Robot: Industrial robot for quality inspection.

[1536] Camera: A high-resolution camera that takes images of the product.

[1537] Device with emotion engine: A camera for analyzing the emotional state of users and inspectors.

[1538] software

[1539] OCR engine: For example, use Tesseract OCR.

[1540] AI Model: Custom model using TensorFlow or PyTorch.

[1541] Image processing: Uses the OpenCV library.

[1542] Database: Use MySQL or PostgreSQL.

[1543] Sentiment analysis engine: Uses Amazon Rekognition and Microsoft Azure Face API.

[1544] System Overview

[1545] To inspect the quality of products, the factory robot first uses a high-resolution camera to take a picture of the product and uploads the image data to a server. At the same time, the robot uses an electronic device to take a picture of the inspector's face, which is then used by an emotion analysis engine to analyze the inspector's emotional state.

[1546] The server performs preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image. The preprocessed image is converted into text data through an OCR engine, and the data undergoes contextual analysis using an AI model. The extracted data is automatically verified, and notifications are sent to the user if necessary.

[1547] Based on the emotional information analyzed by the emotion analysis engine, the server generates and sends notifications according to the inspector's emotional state. If the inspector is in a specific emotional state, such as stress or confusion, the server will send a notification urging them to improve their processing or reconfirm their actions.

[1548] Finally, the validated data set is stored in a database and can be exported to external quality control systems if required, improving the efficiency and accuracy of the product quality inspection process.

[1549] Examples of concrete examples and prompts

[1550] Specific examples

[1551] Consider a scenario in which a factory robot is performing an inspection. The robot takes a picture of the product and uploads it to a server. At the same time, it takes a facial image of the inspector to obtain emotional information. The server analyzes the image, extracts and verifies the necessary information, and if the inspector is feeling stressed, a notification is sent to encourage process improvement based on that state.

[1552] Prompt Sentence Examples

[1553] We will create an application that automates the product quality inspection process performed by factory robots. A high-resolution camera takes images of the product and an OCR engine (Tesseract) extracts information. The inspector's face is captured and an emotion analysis engine (Amazon Rekognition) evaluates their emotional state and sends notifications as needed. The data is further analyzed by an AI model (TensorFlow) and exported to a quality control system.

[1554] Hardware used: industrial robots, high-resolution cameras, emotion analysis devices

[1555] Software used: Tesseract OCR, TensorFlow, OpenCV, MySQL, Amazon Rekognition

[1556] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1557] Step 1:

[1558] A user takes a picture of a product using a terminal, and the terminal simultaneously captures a facial image of the inspector. The input is the captured product image and the facial image of the inspector, and the output is a request sent from the terminal to the server. With this operation, the terminal prepares to send the image data and facial image data to the server.

[1559] Step 2:

[1560] The server receives a request from the device and stores the image data and facial image data. The input is the image data and facial image data sent from the device, and the output is the image data and facial image data stored in the database, as well as metadata (e.g., the date and time of the photo, a unique identifier, etc.). With this operation, the server manages the received data.

[1561] Step 3:

[1562] The server starts preprocessing on the stored product image. The input is the stored product image, and the output is the preprocessed product image. Preprocessing includes deskewing, noise removal, and resolution adjustment. Through this operation, the server improves the image quality so that important information in the image can be more easily extracted.

[1563] Step 4:

[1564] The server sends the preprocessed image to the OCR engine to perform optical character recognition. The input is the preprocessed image and the output is the extracted text data. In this operation, the server converts the character and numeric information from the product image into digital text.

[1565] Step 5:

[1566] The server inputs the text data extracted by the OCR engine into the AI ​​model and performs contextual analysis. The input is the text data extracted by OCR, and the output is the analyzed data. In this process, the server adds contextual information to improve the accuracy of the data.

[1567] Step 6:

[1568] The server automatically validates the parsed data. The input is the text data parsed by the AI ​​model, and the output is the validation result. Validation includes checking the date format and numeric range. With this operation, the server verifies whether the data is accurate.

[1569] Step 7:

[1570] The facial image of the inspector is sent to the emotion analysis engine to analyze the emotional state. The input is the facial image of the inspector, and the output is emotion information. In this operation, the server evaluates the emotional state of the inspector (e.g., stress, confusion).

[1571] Step 8:

[1572] The server generates a notification based on the emotion information and sends it to the user. The input is the result of the validation and emotion analysis, and the output is a notification to the user. In this operation, the server provides the necessary notification to the user and prompts them to improve the process.

[1573] Step 9:

[1574] The user receives the notification on the terminal and makes the necessary corrections. The input is the notification from the server, and the output is the corrected data. This operation allows the user to smoothly correct data.

[1575] Step 10:

[1576] The server again verifies the modified data from the user and saves the final data. The input is the modified data, and the output is the final saved data. With this operation, the server confirms the final data set.

[1577] Step 11:

[1578] The server exports the final dataset to an external quality control system as needed and sends a completion notification to the user. The inputs are the final data and information from the external system, and the outputs are the export results and a completion notification. Through this operation, the server achieves data integration.

[1579] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1580] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1581] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1582] [Fourth embodiment]

[1583] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1584] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1585] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1586] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1587] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1588] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1589] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1590] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1591] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1592] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1593] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1594] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1595] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1596] The present invention is a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[1597] Program processing

[1598] 1. Acquiring and uploading images

[1599] Terminal

[1600] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request to send the photo or image to the server.

[1601] server

[1602] The server receives an image upload request and saves the image in storage. The saved image is assigned a unique identifier, and its metadata (upload date and time, user ID, image ID, etc.) is also saved in the database.

[1603] 2. Image Preprocessing

[1604] server

[1605] After capturing the image, the server begins pre-processing, which includes image deskewing, noise reduction, and resolution adjustment, making the characters and numbers in the image easier to recognize.

[1606] 3. Performing data extraction

[1607] server

[1608] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital data. The AI ​​model then further examines the extracted data based on contextual information to identify specific fields (such as names, dates, or numbers).

[1609] 4. Data verification and completion

[1610] server

[1611] The extracted data is automatically validated by the server, for example to ensure that date formats and numeric ranges are correct. If validation fails, the server notifies the user and asks them to correct the error.

[1612] User

[1613] The user receives the terminal notification, makes any necessary corrections, and then transmits the correction data to the server.

[1614] server

[1615] The server will then re-verify the corrected data, and once verification is complete, the data will be saved and go on to the next step.

[1616] 5. Saving and Exporting Data

[1617] server

[1618] The completed dataset is securely stored on the server. If necessary, the data can be exported to external applications or systems (e.g., accounting software, electronic medical record systems, etc.). If the export was successful, the server notifies the user.

[1619] Specific examples

[1620] Example 1: Processing receipts in the finance department

[1621] Terminal

[1622] The user takes a photo of the receipt using the terminal and uploads it to the server.

[1623] server

[1624] The server receives the image, performs deskewing, noise removal, and resolution adjustment. The pre-processed image is passed through an OCR engine, which extracts data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any ambiguities.

[1625] User

[1626] The user makes the corrections and sends the data to the server, which re-verifies the corrections and finally exports them to the accounting software and notifies the user of the completion.

[1627] Example 2: Medical record processing in a healthcare facility

[1628] Terminal

[1629] Medical staff scan handwritten medical records and upload them to a server.

[1630] server

[1631] The server receives the images and performs deskewing, noise reduction, and resolution adjustment. The preprocessed images are then run through an AI model to extract data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically verified, and medical staff are notified if there are any issues.

[1632] User

[1633] The medical staff makes the corrections and sends the data to the server, which re-verifies the corrected data, exports it to the electronic medical record system, and notifies the medical staff of the completion.

[1634] In this way, this system improves productivity throughout the entire business by efficiently extracting information from documents and images and automatically processing data.

[1635] The processing flow will be explained below.

[1636] Step 1:

[1637] Terminal

[1638] A user launches an application on the device and uses the capture function to capture a document or image, or select an existing image. After the user completes the capture or selection, the device generates an upload request to send the image to the server.

[1639] Step 2:

[1640] server

[1641] The server receives an image upload request, assigns a unique identifier to the uploaded image, saves it in storage, and stores metadata such as the upload date and time, user ID, and image ID in a database.

[1642] Step 3:

[1643] server

[1644] The server then performs preprocessing on the stored image, which includes image deskewing, noise reduction, and resolution adjustment. Deskewing adjusts the text in the image so that it is level, noise reduction removes unwanted background elements, and resolution adjustment makes the details of the characters clearer.

[1645] Step 4:

[1646] server

[1647] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. Optionally, the AI ​​model analyzes the context and improves the accuracy of the extracted data.

[1648] Step 5:

[1649] server

[1650] Specific fields (e.g., names, dates, numbers, etc.) are identified from the extracted text data and saved as structured data (e.g., JSON, CSV format) for subsequent validation and export processes.

[1651] Step 6:

[1652] server

[1653] The extracted data is automatically validated, including checking date formats, numeric ranges, required fields, etc. If any errors are found during the validation process, the server notifies the user and asks them to correct them.

[1654] Step 7:

[1655] Terminal

[1656] The user receives the notification and modifies the data on the terminal screen. Once the modification is complete, the user generates a request to send the modified data to the server.

[1657] Step 8:

[1658] server

[1659] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be accurate, it is saved as the final data. Once the correction process is complete and the verification is successful, the server proceeds to the next step.

[1660] Step 9:

[1661] server

[1662] Securely store the completed dataset in a database. Once the storage process is complete, generate API requests to export the data to any external systems required (e.g., accounting software, electronic medical record systems).

[1663] Step 10:

[1664] server

[1665] If the export is successful, the server records the result in the database and sends the user a notification of completion, including confirmation of the destination system and the contents of the data.

[1666] This detailed process at each step allows for efficient image-based data extraction and automated data processing, improving overall business productivity.

[1667] Example 1

[1668] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1669] In modern business processes and healthcare facilities, manual information extraction and validation from documents and images is time-consuming and error-prone. This reduces operational efficiency and reduces reliability in critical data processing. Furthermore, there is a need to export extracted data quickly and accurately to external systems. To address these challenges, an automated system is needed.

[1670] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1671] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (such as deskewing, noise removal, and resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for re-verifying the corrected data and finally saving and exporting it to required external systems; a means for identifying specific fields (such as names, dates, and numbers) of the extracted data and saving them as structured data; and a means for examining the data based on a generative AI model and generating automated notifications using prompts. This enables the automation of information extraction and verification from documents and images, improving business efficiency and the reliability of data processing.

[1672] A "user" is an entity that uses the system to capture and upload documents and images.

[1673] A "terminal" is a device used by a user, which has the function of taking pictures of documents and images and uploading them to a server.

[1674] A "server" is a computer system that pre-processes images uploaded by users and performs functions such as data extraction, validation, storage and export.

[1675] "Upload" refers to the act of sending image or document data from a terminal to a server.

[1676] "Preprocessing" refers to processing such as tilt correction, noise removal, and resolution adjustment on the image to facilitate subsequent data extraction.

[1677] "Tilt correction" is a process of adjusting the orientation of an image so that the characters and figures in the image are easier to recognize.

[1678] "Noise reduction" is a process that removes unnecessary elements from an image to make it clearer.

[1679] "Resolution adjustment" is a process of changing the image resolution to an appropriate level.

[1680] Optical character recognition (OCR) is a technology that recognizes text within an image and converts it into digital data.

[1681] "Artificial intelligence (AI)" is the technology that enables computers to learn, reason, and solve problems like humans.

[1682] "Data validation" is the process of checking whether extracted data is accurate and valid.

[1683] "Notification" is a message or signal that notifies the user when there is a validation result or uncertain data.

[1684] "Correction" refers to the act of the user correcting data based on a notification from the server.

[1685] "Storage" is the act of securely recording verified data in storage.

[1686] "Export" is the act of sending stored data to an external system.

[1687] A "specific field" is a data item that indicates specific information such as a name, date, or number.

[1688] "Structured data" is data that is organized according to a predefined data format.

[1689] A "generative AI model" is an artificial intelligence algorithm that is trained to perform a specific task.

[1690] A "prompt" is a sentence of text that instructs a generative AI model on a task based on specific input.

[1691] This invention is a system in which a user uses a terminal to take a photo of a document or image and upload it to a server, which then extracts, verifies, and exports information from the image. Specific embodiments of this system are described below.

[1692] Hardware and software used

[1693] Terminal

[1694] The device used by the user is a device with a photographic function, such as a smartphone, tablet, or digital camera. The device is responsible for sending the photographed or selected images to the server. The application running on the device should preferably have an internet connection for data transmission and an image editing function.

[1695] server

[1696] The server performs the functions of data reception, image pre-processing, data extraction, validation, storage, and export, using the following software:

[1697] Image preprocessing library (e.g. OpenCV)

[1698] Optical Character Recognition (OCR) engine (e.g., Tesseract OCR)

[1699] Artificial Intelligence (AI) models (e.g., Google Cloud Vision API)

[1700] Database system (e.g. MySQL)

[1701] Notification systems (e.g., Firebase Cloud Messaging)

[1702] Data processing and calculation in the program processing flow

[1703] 1. Acquiring and uploading images

[1704] A user opens a device application and takes a photo of a document or image, or selects an existing image. The device generates an upload request and sends the image data and metadata to the server, including information such as the user ID and image ID.

[1705] 2. Image Preprocessing

[1706] The server performs deskew, noise reduction, and resolution adjustment on the images received, making it easier to recognize characters and numbers. The OpenCV library is used for these preprocessing steps.

[1707] 3. Data extraction

[1708] The preprocessed images are then fed into an optical character recognition (OCR) or artificial intelligence (AI) model, using Tesseract OCR to extract text data from the image and Google Cloud Vision API to identify specific fields (such as names, dates, or amounts).

[1709] 4. Data verification and completion

[1710] The extracted data is validated by the server. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. If there is any problem with the validation, the server notifies the user and asks them to make corrections. After the user makes the necessary corrections, the corrected data is sent back to the server, which then validates the data again.

[1711] 5. Saving and Exporting Data

[1712] The complete dataset is stored on the server and can be exported to external applications or systems as needed, with the user notified if the export was successful.

[1713] Specific examples

[1714] Example 1: Processing receipts in the finance department

[1715] The user takes a photo of the receipt using a smartphone app and uploads it to the server. The server then performs image deskewing, noise reduction, and resolution adjustment. The preprocessed image is then input into the Tesseract OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically validated, and the user is notified if there are any uncertainties. If the user corrects any errors and submits it to the server again, the server revalidates the data, exports it to the accounting software, and notifies the user of completion.

[1716] Example 2: Medical record processing in a healthcare facility

[1717] Medical staff scan handwritten medical records and upload them to the server, which then performs image deskewing, noise removal, and resolution adjustment. The preprocessed images are then input into the Google Cloud Vision API, which extracts data such as patient name, diagnosis, and prescribed medications. The extracted data is automatically validated, and if there are any problems, the medical staff is notified. After the medical staff corrects any errors and resubmits the data to the server, the server revalidates the data, exports it to the electronic medical record system, and notifies the medical staff of its completion.

[1718] Prompt Sentence Examples

[1719] Below are some example prompts using generative AI models:

[1720] "Input the following image into an OCR engine and extract data such as the amount, transaction date, and store name. After extraction, verify this data and, if necessary, notify the user."

[1721] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1722] Step 1:

[1723] User

[1724] The user uses a device to take a photo of a document or image, or to select an existing image. To do this, the user uses an application on the device. The data input here is the image taken or selected by the user, and the image data is output. In a specific example, the user might take a photo of a receipt using a camera app on their smartphone.

[1725] Step 2:

[1726] Terminal

[1727] The device internally checks the captured or selected image data and generates an upload request to send to the server. The request includes metadata such as the user ID and image ID. The input is the image data captured or selected by the user and the metadata, and the output is an upload request sent to the server. Here, the data is sent using an Internet connection.

[1728] Step 3:

[1729] server

[1730] The server receives an image upload request and saves the image data and metadata in storage. When saving, it assigns a unique identifier to the image and records the metadata in a database. The input is the image data and metadata sent from the device, and the output is the image data saved in storage and the metadata saved in the database. For example, the image data is saved in a dedicated folder, and the metadata is recorded in a database table.

[1731] Step 4:

[1732] server

[1733] The server retrieves the image from storage and begins preprocessing. In preprocessing, the image is tilted and rotated to the correct orientation. Noise reduction is also performed to remove unnecessary parts. Furthermore, the image resolution is adjusted to make text and numbers easier to read. The input is the image data retrieved from storage, and the output is the preprocessed image data. Here, the OpenCV library is used to preprocess the image.

[1734] Step 5:

[1735] server

[1736] The server inputs the preprocessed image data into an optical character recognition (OCR) engine or artificial intelligence (AI) model. Tesseract OCR is used to recognize text in the image and convert it into digital data. Google Cloud Vision API is then used to identify specific fields (such as names, dates, or amounts) based on contextual information from the extracted text. The input is the preprocessed image data, and the output is the extracted text data.

[1737] Step 6:

[1738] server

[1739] The server validates the extracted data. For example, it checks whether the date format is correct or whether the amount is within a reasonable range. The input is the extracted text data, and the output is the validation result. If there is any validation failure, the server notifies the user and asks them to correct it. This notification is sent using Firebase Cloud Messaging or similar.

[1740] Step 7:

[1741] User

[1742] The user receives the notification on the device and makes the necessary corrections. Here, the input is the notification from the server, and the output is the corrected data. For example, the user manually corrects misreadings made by OCR on the device screen.

[1743] Step 8:

[1744] Terminal

[1745] The user modifies the data and sends it back to the server. The input is the modified data, and the output is the upload request sent again.

[1746] Step 9:

[1747] server

[1748] The server re-verifies the modified data. If the re-verification is successful, the data proceeds to the next processing step. The input is the modified data submitted by the user, and the output is the re-verified data.

[1749] Step 10:

[1750] server

[1751] The validated dataset is securely stored and exported to external applications or systems as needed. For example, sending data to accounting software or electronic medical record systems. If the export is successful, the server sends a notification to the user. The input is the validated data, and the output is the export result to the external system and a notification to the user.

[1752] (Application example 1)

[1753] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1754] In logistics centers, the receiving and inventory management of items is often done manually, which can lead to problems such as data entry errors and the time required for confirmation work. For this reason, there is a need for a system that can manage item information efficiently and accurately.

[1755] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1756] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user if there is any uncertainty and requesting correction; a means for re-verifying the corrected data and finally saving and exporting it to the necessary external system; and a means for automatically recognizing item information at the logistics center and providing a function for managing data on the terminal, thereby enabling efficient and accurate management of item information at the logistics center.

[1757] A "terminal" is a device that a user uses to take documents and images and upload the data to a server.

[1758] A "server" is a computer system that receives image data uploaded by users and performs pre-processing, data extraction, validation, and export.

[1759] "Preprocessing" refers to the process of correcting tilt, removing noise, and adjusting resolution of uploaded images.

[1760] "Optical character recognition (OCR) technology" is a technology that identifies characters in an image and converts them into digital text data.

[1761] "Artificial intelligence (AI)" is a technology that gives computer systems the ability to understand meaning and context from images and text, and to extract, classify, and verify data.

[1762] "Verification" is the process of checking whether the extracted data is accurate and asking the user to correct any uncertainties.

[1763] "Export" is the process of outputting the final validated data to the required external systems or applications.

[1764] A "logistics center" is a facility that handles logistics operations such as receiving, storing, and shipping items.

[1765] "Item information" refers to data such as the identification information, quantity, and date of receipt of products handled at the logistics center.

[1766] The "database" is an information management system for organizing and storing extracted item information.

[1767] An "inventory management system" is a software system that manages the inventory status of items and records information on incoming and outgoing goods.

[1768] The present invention is a system for improving the efficiency and accuracy of item management in a logistics center. A specific embodiment of this system will be described below.

[1769] System Configuration

[1770] This system consists of a terminal used by users, a server that performs processing, and a database that stores data. The main software used is OpenCV for image preprocessing, Tesseract as an optical character recognition (OCR) engine, and requests for communication.

[1771] Program processing overview

[1772] 1. Acquiring and uploading images

[1773] The device provides a means for users to take photos of items and upload them to the server. The images are sent to the server and assigned a unique identifier. Image metadata (upload date and time, user ID, image ID, etc.) is also generated and stored in a database.

[1774] 2. Image Preprocessing

[1775] The server receives the uploaded image and performs deskewing, noise reduction, and resolution adjustment, which makes the text and numbers in the image more readable.

[1776] 3. Data Extraction

[1777] The server inputs the preprocessed image into an OCR engine (Tesseract) to recognize text in the image and convert it into digital data. Additionally, it uses an artificial intelligence (AI) model to extract item information (item ID, quantity, date of receipt, etc.).

[1778] 4. Data verification and completion

[1779] The extracted data is automatically validated by the server. For example, it checks whether the date format and numeric range are correct. If there are any errors during validation, the user is notified and can make corrections. If the user makes corrections and submits the data to the server again, the server will validate the data again.

[1780] 5. Saving and Exporting Data

[1781] The final verified data is stored in a database and, if required, exported to an inventory management system.

[1782] Specific examples

[1783] Example 1: Taking a photo of an item and exporting the data

[1784] Staff at the logistics center use a terminal to take a photo of an item and upload it to the server. The server receives the image, preprocesses it, and then uses an OCR engine to extract item information. The extracted data is verified and exported to the inventory management system. If there is an error in the data, a notification is sent to the user, who can then make the necessary corrections.

[1785] Prompt Sentence Examples

[1786] "Please take a photo of this item. The system will automatically extract the item information and update it in your inventory."

[1787] In this way, the system of the present invention automates item management in a logistics center, improving work efficiency and accuracy.

[1788] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1789] Step 1:

[1790] A user takes a photo of an item using a device, which captures the item image, generates a unique identifier, and prepares the captured image and its metadata (upload date and time, user ID, image ID, etc.).

[1791] Step 2:

[1792] The device uploads the captured images and metadata to a server, which receives the data and stores them in a database using a unique identifier assigned to the image.

[1793] Step 3:

[1794] The server retrieves the received image from the database and begins pre-processing it, i.e., correcting the image's distortion (image rotation), removing noise (image filtering), and adjusting the resolution (resizing or rescaling). This pre-processing process makes the features in the image more visible.

[1795] Step 4:

[1796] The preprocessed image is input into the server's OCR engine (Tesseract) and OCR processing is performed. During this process, characters within the image are extracted and converted into digital text data. The extracted text data includes the item ID, quantity, and inventory date.

[1797] Step 5:

[1798] The server uses an AI model to more precisely identify item information from the text data extracted by OCR, particularly identifying specific fields (item ID, quantity, inventory date, etc.) and saving them as structured data, utilizing machine learning and natural language processing techniques.

[1799] Step 6:

[1800] The server validates the extracted and identified data, specifically ensuring that date formats and numeric ranges are correct. If validation fails, the server notifies the user, including the specific corrections required.

[1801] Step 7:

[1802] The user receives a notification and makes the necessary corrections on the device, after which the corrected data is sent back to the server.

[1803] Step 8:

[1804] The server will again verify the corrected data resubmitted by the user, and if the verification is successful, the data will be finally saved in the database.

[1805] Step 9:

[1806] The server exports the saved data to an external system (e.g., inventory management system) as needed. Once the export is complete, the server notifies the user.

[1807] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1808] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[1809] Program processing

[1810] 1. Acquiring and uploading images

[1811] Terminal

[1812] The user opens the application on the device and uses the camera function to take a document or image. At the same time, the device captures the user's face, and the emotion engine analyzes the user's emotional state. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[1813] server

[1814] The server receives the upload request and stores the image and emotion data. The image is assigned a unique identifier, and metadata such as the upload date and time and user ID are also stored in the database.

[1815] 2. Image Preprocessing

[1816] server

[1817] The server then performs preprocessing on the stored image, which includes deskewing, noise reduction, and resolution adjustment, making the letters and numbers in the image more recognizable.

[1818] 3. Performing data extraction

[1819] server

[1820] The server then feeds the preprocessed image into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine analyzes the letters and numbers in the image and converts them into digital text data. The AI ​​model then analyzes the context to further refine the extracted data.

[1821] 4. Data verification and completion

[1822] server

[1823] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, an emotion engine evaluates the user's emotional state. If a specific emotional state (e.g., stress or confusion) is detected, the server sends the user an improved notification.

[1824] User

[1825] The user receives the notification on the device and can modify the data. Once the modification is complete, a request is generated to send the modified data to the server.

[1826] server

[1827] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved as the final data. The verification process and the content of the correction request may be adjusted based on the emotion data.

[1828] 5. Saving and Exporting Data

[1829] server

[1830] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). If the export is successful, the server sends a completion notification to the user and records the export results in the database.

[1831] Specific examples

[1832] Example 1: Processing receipts in the finance department

[1833] Terminal

[1834] The user takes a photo of the receipt using the device and uploads it to the server, while the emotion engine simultaneously analyzes the user's emotional state.

[1835] server

[1836] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification tailored based on the emotion data is sent to the user.

[1837] User

[1838] The user receives a notification, makes the necessary corrections, and sends the corrected data to the server, which then validates it again. Finally, the data is exported to the accounting software, and the user is notified of the completion.

[1839] Example 2: Medical record processing in a healthcare facility

[1840] Terminal

[1841] Medical staff scan handwritten medical records and upload them to the server, and at the same time, the emotion engine analyzes the emotional state of the medical staff.

[1842] server

[1843] The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is then passed through an AI model to extract data such as patient name, diagnosis, and prescribed medication. The extracted data is automatically verified, and if there are any issues, a tailored notification based on the emotion data is sent to medical staff.

[1844] User

[1845] The medical staff receives the notification, makes the necessary corrections, and sends the data to the server, which re-verifies the corrected data, exports the data to the electronic medical record system, and notifies the medical staff of the completion.

[1846] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[1847] The processing flow will be explained below.

[1848] Step 1:

[1849] Terminal

[1850] The user launches the application on the device and uses the camera function to take a document or image. At the same time, the device camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., joy, anger, sadness, etc.) in real time. After the user selects or takes an image, the device generates a request to send the document or image and emotion information to the server.

[1851] Step 2:

[1852] server

[1853] The server receives the upload request and saves the submitted image and emotion data in storage. A unique identifier is assigned to the image, and metadata such as the upload date and time, user ID, and emotion data are recorded in the database.

[1854] Step 3:

[1855] server

[1856] The server begins pre-processing the stored image, straightening the image to make the text easier to read, applying a noise reduction filter to remove unnecessary elements in the image, and adjusting the resolution to make the details of the text and numbers clearer.

[1857] Step 4:

[1858] server

[1859] The preprocessed image is then fed into an optical character recognition (OCR) engine or artificial intelligence (AI) model. The OCR engine recognizes the text in the image and converts it into digital text data. The AI ​​model analyzes the context and improves the accuracy of the extracted data.

[1860] Step 5:

[1861] server

[1862] The text data extracted by the OCR engine or AI model is parsed to identify specific fields (e.g., names, dates, numbers, etc.) and stored as structured data for subsequent validation and export processes.

[1863] Step 6:

[1864] server

[1865] The extracted data is automatically validated, for example, checking whether the date format is correct or whether the numbers are within a reasonable range. If any flaws are detected during the validation process, the server will also evaluate the user's emotional state and send a notification with an appropriate tone if a specific emotional state (e.g., stress, confusion) is detected through the emotion engine.

[1866] Step 7:

[1867] Terminal

[1868] The user receives the notification on the terminal and makes the necessary modifications through the data modification interface. After completing the modifications, the user generates a request to send the data to the server.

[1869] Step 8:

[1870] server

[1871] The server receives the corrected data sent by the user and performs automatic verification again. If the corrected data is confirmed to be correct, it is saved as the final data set.

[1872] Step 9:

[1873] server

[1874] The completed dataset is securely stored. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems). If the export is successful, the server records the result in a database and sends a completion notification to the user.

[1875] Specific examples

[1876] Example 1: Processing receipts in the finance department

[1877] Step 1:

[1878] Terminal

[1879] The user takes a photo of the receipt using the device, and at the same time, the emotion engine analyzes the user's emotional state in real time.

[1880] Step 2:

[1881] server

[1882] The server receives the image and emotion data, stores them in storage, generates a unique identifier and metadata, and records them in a database.

[1883] Step 3:

[1884] server

[1885] Preprocessing such as tilt correction, noise removal, and resolution adjustment is performed on the received image.

[1886] Step 4:

[1887] server

[1888] The preprocessed image is input into an OCR engine to extract data such as the amount, transaction date, and store name.

[1889] Step 5:

[1890] server

[1891] The extracted data is saved as structured data and passed to subsequent validation processing.

[1892] Step 6:

[1893] server

[1894] The extracted data is automatically verified, and if there are any deficiencies, an appropriate notification is sent based on the user's emotional state.

[1895] Step 7:

[1896] Terminal

[1897] The user receives a notification on the device, makes the necessary corrections, and sends the corrected data to the server.

[1898] Step 8:

[1899] server

[1900] The corrected data is verified again, and if it is correct, it is saved as the final data.

[1901] Step 9:

[1902] server

[1903] The final data is exported to accounting software and the results are notified to the user.

[1904] Example 2: Medical record processing in a healthcare facility

[1905] Step 1:

[1906] Terminal

[1907] Medical staff use the device to scan handwritten medical records, while the emotion engine simultaneously analyzes the emotional state of the medical staff in real time.

[1908] Step 2:

[1909] server

[1910] The server receives and stores the image and emotion data, assigns a unique identifier, and generates and records metadata.

[1911] Step 3:

[1912] server

[1913] The saved image is then tilted, noise removed, and resolution adjusted.

[1914] Step 4:

[1915] server

[1916] The pre-processed images are fed into an AI model to extract data such as patient name, diagnosis, and prescribed medications.

[1917] Step 5:

[1918] server

[1919] The extracted data is saved as structured data and passed to subsequent validation processing.

[1920] Step 6:

[1921] server

[1922] The extracted data is automatically verified, and if there are any problems, notifications are sent that are tailored to the emotional state of the medical staff.

[1923] Step 7:

[1924] Terminal

[1925] Medical staff are notified and make any necessary corrections, which are then sent to the server.

[1926] Step 8:

[1927] server

[1928] The corrected data is verified again, and if appropriate, is saved as the final data.

[1929] Step 9:

[1930] server

[1931] The final data is exported to the electronic medical record system and the results are notified to medical staff.

[1932] In this way, this system adds an emotion engine to information extraction from documents and images and automatic data processing, improving the user experience and further increasing overall business productivity.

[1933] Example 2

[1934] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1935] In conventional systems, when extracting information from documents and images, data processing is performed without taking the user's emotional state into consideration, which means that errors or uncertainties in the extracted data are not properly communicated to the user, resulting in reduced work efficiency. Furthermore, when users feel stressed or confused, it becomes even more difficult to make corrections.

[1936] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1937] In this invention, the server includes: a means for a user to photograph a document or image using a terminal and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means for the server to analyze the user's emotional state simultaneously with image capture and transmit the image to the server; a means for using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting corrections if there are any uncertainties; a means for adjusting the notification content based on the user's emotional state; and a means for re-verifying the corrected data and finally saving and exporting it to a required external system. This enables appropriate notifications that take the user's emotional state into consideration, thereby improving the reliability of data and reducing the burden on the user.

[1938] A "terminal" is an electronic device that a user uses to capture documents and images and upload them to a server.

[1939] "Server" refers to a computer system that receives and stores images and data sent from the terminals, and performs preprocessing, data extraction, validation, and export.

[1940] "Preprocessing" refers to image processing operations such as deskewing, noise removal, and resolution adjustment that are performed on uploaded images.

[1941] Optical character recognition (OCR) is a technology that analyzes letters and numbers in an image and converts them into digital text data.

[1942] Artificial intelligence (AI) is a technology that analyzes image and text data, automatically recognizing specific patterns and contexts, and extracting and verifying data.

[1943] "Emotional state" refers to the user's psychological and emotional state (e.g., Happy, Angry, Sad, etc.) obtained by analyzing the user's facial image.

[1944] "Validation" is the process of verifying whether the extracted data is correct, checking whether it conforms to a specific format or range.

[1945] A "request for correction" means that if there are any uncertainties as a result of the verification, the user is notified of the details and asked to correct the data.

[1946] "Adjusting notification content" means appropriately changing the wording and method of notification based on the user's emotional state.

[1947] "Export" refers to sending or outputting data that has undergone final verification and correction to the required external system.

[1948] A "unique identifier" is unique identification information assigned to each image to distinguish it from other images.

[1949] "Metadata" is information mainly about image data (e.g., upload date and time, user ID, etc.), and is used by the server to manage image data.

[1950] The present invention combines a system in which a user uses a terminal to upload a document or image to a server, and the server extracts, verifies, and exports information from the image, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system and program processing are described below.

[1951] Acquiring and uploading images

[1952] Terminal

[1953] A user opens an application on their device and uses the camera function to take a document or image. At this time, the device's camera captures the user's face, and the emotion engine analyzes the user's emotional state (e.g., Happy, Angry, Sad, etc.). After the user selects an image or completes the capture, the device generates an HTTP request to send the document or image and emotion information to the server. The transmitted data includes Base64-encoded image data and emotion data.

[1954] server

[1955] The server receives HTTP requests from devices and stores images and emotion data. Images are assigned a unique identifier, and metadata such as upload date and time, user ID, etc. are also stored in a database.

[1956] Image preprocessing

[1957] server

[1958] The server then begins preprocessing the saved image. This includes correcting the image's deskew (for example, using a Hough transform), applying a noise reduction filter, and adjusting the resolution. These processes make the characters and symbols in the image easier to recognize. The preprocessed image is then stored in a temporary folder.

[1959] Running the Data Extraction

[1960] server

[1961] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the letters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. The AI ​​model highlights specific keywords and phrases, improving the accuracy of the extracted data.

[1962] Data validation and completion

[1963] server

[1964] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. During the validation process, the emotion engine reassess the user's emotional state and adjusts the notification content if a specific emotional state (e.g., stress or confusion) is detected. Based on the validation results, the server sends a notification to the user requesting corrections.

[1965] User

[1966] The user receives a notification on their device, checks the content, and corrects the data if necessary, generating an HTTP request to send the corrected data to the server.

[1967] server

[1968] The server receives the corrected data from the user and verifies it again. If the corrected data is confirmed to be correct, it is saved in the database as the final data. The verification process and the content of the correction request are adjusted appropriately based on the emotion data.

[1969] Saving and exporting data

[1970] server

[1971] The completed dataset is securely stored in a database. If necessary, the data is exported to external systems (e.g., accounting software, electronic medical record systems, etc.). The data is sent using protocols such as RESTful APIs or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database.

[1972] Examples of concrete examples and prompts

[1973] Example 1: Receipt processing in the finance department

[1974] The user takes a photo of the receipt using their device and uploads it to the server. At the same time, the emotion engine analyzes the user's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an OCR engine to extract data such as the amount, transaction date, and store name. The extracted data is automatically verified, and if there are any uncertainties, a notification adjusted based on the emotion data is sent to the user. The user receives the notification and makes any necessary corrections. The corrected data is then sent to the server, which then performs another verification. Finally, the data is exported to the accounting software, and a completion notification is sent to the user.

[1975] Example prompt (for the finance department):

[1976] Please take a picture of your receipt and upload it.

[1977] "Analyzing emotions..."

[1978] "Image preprocessing complete, optical character recognition begins."

[1979] "Validating extracted data..."

[1980] "There is a discrepancy in the amount verification. Please check again."

[1981] Example 2: Medical record processing in a healthcare facility

[1982] Medical staff scan handwritten medical records and upload them to the server. At the same time, an emotion engine analyzes the medical staff's emotional state. The server receives the image and emotion data and performs preprocessing such as tilt correction, noise removal, and resolution adjustment. The preprocessed image is passed through an AI model to extract data such as the patient's name, diagnosis, and prescribed medications. The extracted data is automatically verified, and if there are any problems, a notification with adjustments based on the emotion data is sent to the medical staff. The medical staff receives the notification, makes any necessary corrections, and sends the data to the server. The server re-verifies the corrected data, exports it to the electronic medical record system, and sends a completion notification to the medical staff.

[1983] Example prompts (for medical facilities)

[1984] "Scan and upload your medical records."

[1985] "Analyzing emotions..."

[1986] "Image preprocessing complete, data extraction begins."

[1987] "Validating extracted data..."

[1988] "There is a discrepancy in the diagnostic results confirmation. Please check again."

[1989] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1990] Step 1:

[1991] Device-based image capture and emotion analysis

[1992] The user opens the device application and uses the camera function to take a photo of a document or image. At this time, the device camera also captures the user's face, and the emotion engine analyzes the user's emotional state. Emotional data (e.g., Happy, Angry, Sad, etc.) is generated as a result of the analysis. Specifically, the device application calls the camera module and captures the image and a photo of the user's face.

[1993] Input: Documents and images captured by the camera, and the user's facial image

[1994] Output: Photographed documents, image data, analyzed emotional data

[1995] Step 2:

[1996] Uploading images and emotion data to the server

[1997] After the user selects or captures an image, the device sends the document, image, and emotion information to the server as an HTTP request. The sent data includes Base64-encoded image data and emotion data. Specifically, the device application generates an HTTP request and sends the data to the server's API endpoint.

[1998] Input: Photographed documents, image data, analyzed emotional data

[1999] Output: HTTP request sent to the server

[2000] Step 3:

[2001] Receiving and storing data by the server

[2002] The server receives the HTTP request sent from the device and saves the image and emotion data. A unique identifier is assigned to the image data, and metadata such as the upload date and time and user ID are also saved in the database. Specifically, the server analyzes the request and creates an entry to save in the database.

[2003] Input: HTTP request (image data, emotion data, metadata)

[2004] Output: Image data and metadata stored in a database

[2005] Step 4:

[2006] Image preprocessing by the server

[2007] The server begins preprocessing the saved image. This includes correcting the image's tilt using a Hough transform, applying a noise reduction filter, and adjusting the resolution. This makes it easier to recognize characters and numbers in the image. The preprocessed image is stored in a temporary folder. Specific operations involve the use of an image processing library (e.g., OpenCV).

[2008] Input: Saved image data

[2009] Output: Preprocessed image data

[2010] Step 5:

[2011] Server performs data extraction

[2012] The preprocessed image is input into an optical character recognition (OCR) engine (e.g., Tesseract OCR). The OCR engine analyzes the characters and numbers in the image and converts them into text data. The converted text data is then passed to an AI model (e.g., BERT, GPT, etc.) to analyze the context and improve the accuracy of the data. Specific operations involve the use of APIs for the OCR engine and the AI ​​model.

[2013] Input: Preprocessed image data

[2014] Output: Extracted text data

[2015] Step 6:

[2016] Server-based data validation and completion

[2017] The extracted data is automatically validated, for example, to ensure that date formats and numeric ranges are correct. Based on the validation results, the emotion engine reassess the user's emotional state, and if a specific emotional state (e.g., stress or confusion) is detected, the server adjusts the notification content and sends a notification to the user requesting corrections. Specific operations include the execution of the validation algorithm and the emotion analysis engine.

[2018] Input: Extracted text data, emotion data

[2019] Output: Notification content, correction request notification

[2020] Step 7:

[2021] User Modifications

[2022] The user receives the notification on the device, checks the content, and corrects the data if necessary. An HTTP request is generated to send the corrected data to the server. Specifically, a user interface is provided, and the data is sent again after the user makes the corrections.

[2023] Input: Notification of correction request, data to be corrected

[2024] Output: Corrected data sent to the server

[2025] Step 8:

[2026] Server verifies and saves modified data

[2027] The server receives the corrected data from the user and performs a re-verification. If the corrected data is confirmed to be correct, it is saved as the final data in the database. The verification process and the content of the correction request are adjusted appropriately based on the emotion data. Specifically, the re-verification algorithm is executed and the data is saved in the database.

[2028] Input: Modified text data

[2029] Output: Final saved data

[2030] Step 9:

[2031] Saving and exporting data

[2032] The server securely stores the completed dataset in a database. If necessary, the data is exported to an external system (e.g., accounting software, electronic medical record system, etc.). The data is sent using protocols such as RESTful API or SOAP. If the export is successful, a completion notification is sent to the user and the export results are recorded in the database. The specific operation is the execution of the export module.

[2033] Input: Final saved data

[2034] Output: Data exported to external systems, completion notification

[2035] (Application example 2)

[2036] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2037] Although the product quality inspection process in modern factories is highly automated, it still requires a lot of manual work and human intervention, which can reduce efficiency. In particular, when inspectors are stressed or confused, the likelihood of errors or inaccurate data increases. Therefore, there is a need not only to automate the quality inspection process, but also to develop a system that can monitor the emotional state of inspectors and respond appropriately.

[2038] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2039] In this invention, the server includes: a means for a user to use a terminal to photograph a document or image and upload it to the server; a means for the server to perform preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image; a means using optical character recognition technology or artificial intelligence to extract text data from the preprocessed image; a means for verifying the extracted data and notifying the user and requesting correction if there is any uncertainty; a means for re-verifying the corrected data and finally saving it or exporting it to a required external system; a means for photographing and uploading product images and having an emotion engine for analyzing the emotional state of the inspector; and a means for generating and sending a notification according to the inspector's state based on the received emotion information, thereby improving the efficiency and accuracy of the product quality inspection process.

[2040] "Document or Image" refers to visual data captured or saved by a user using a device.

[2041] "Server" means a computer on a network that processes, stores, and analyzes data uploaded by users.

[2042] "Preprocessing" refers to initial data processing such as tilt correction, noise removal, and resolution adjustment performed on uploaded images.

[2043] "Optical character recognition technology" refers to technology that converts character information in an image into digital text.

[2044] "Artificial intelligence" refers to computer programs that analyze and learn from large amounts of data and perform various tasks automatically.

[2045] "Validation" refers to the process of verifying the accuracy of extracted data.

[2046] "Correction" refers to a user making corrections to data that is determined to be uncertain during the validation process.

[2047] "Export" refers to sending processed data to an external system or application.

[2048] "Emotion engine" refers to a system that analyzes a user's emotional state from facial images.

[2049] "Emotional Information" refers to data on the user's emotional state analyzed by the emotion engine.

[2050] "Generating a notification" refers to the process of creating and sending a message to a user based on specific information.

[2051] The system embodying this invention combines multiple advanced technologies to improve the efficiency of product quality inspections by factory robots and monitor the status of inspectors. Specific hardware, software, and system configurations are shown below.

[2052] Hardware and software used

[2053] Hardware

[2054] Device: The device (smartphone, tablet, etc.) where a user takes or uploads documents or images.

[2055] Robot: Industrial robot for quality inspection.

[2056] Camera: A high-resolution camera that takes images of the product.

[2057] Device with emotion engine: A camera for analyzing the emotional state of users and inspectors.

[2058] software

[2059] OCR engine: For example, use Tesseract OCR.

[2060] AI Model: Custom model using TensorFlow or PyTorch.

[2061] Image processing: Uses the OpenCV library.

[2062] Database: Use MySQL or PostgreSQL.

[2063] Sentiment analysis engine: Uses Amazon Rekognition and Microsoft Azure Face API.

[2064] System Overview

[2065] To inspect the quality of products, the factory robot first uses a high-resolution camera to take a picture of the product and uploads the image data to a server. At the same time, the robot uses an electronic device to take a picture of the inspector's face, which is then used by an emotion analysis engine to analyze the inspector's emotional state.

[2066] The server performs preprocessing (skew correction, noise removal, resolution adjustment) on the uploaded image. The preprocessed image is converted into text data through an OCR engine, and the data undergoes contextual analysis using an AI model. The extracted data is automatically verified, and notifications are sent to the user if necessary.

[2067] Based on the emotional information analyzed by the emotion analysis engine, the server generates and sends notifications according to the inspector's emotional state. If the inspector is in a specific emotional state, such as stress or confusion, the server will send a notification urging them to improve their processing or reconfirm their actions.

[2068] Finally, the validated data set is stored in a database and can be exported to external quality control systems if required, improving the efficiency and accuracy of the product quality inspection process.

[2069] Examples of concrete examples and prompts

[2070] Specific examples

[2071] Consider a scenario in which a factory robot is performing an inspection. The robot takes a picture of the product and uploads it to a server. At the same time, it takes a facial image of the inspector to obtain emotional information. The server analyzes the image, extracts and verifies the necessary information, and if the inspector is feeling stressed, a notification is sent to encourage process improvement based on that state.

[2072] Prompt Sentence Examples

[2073] We will create an application that automates the product quality inspection process performed by factory robots. A high-resolution camera takes images of the product and an OCR engine (Tesseract) extracts information. The inspector's face is captured and an emotion analysis engine (Amazon Rekognition) evaluates their emotional state and sends notifications as needed. The data is further analyzed by an AI model (TensorFlow) and exported to a quality control system.

[2074] Hardware used: industrial robots, high-resolution cameras, emotion analysis devices

[2075] Software used: Tesseract OCR, TensorFlow, OpenCV, MySQL, Amazon Rekognition

[2076] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2077] Step 1:

[2078] A user takes a picture of a product using a terminal, and the terminal simultaneously captures a facial image of the inspector. The input is the captured product image and the facial image of the inspector, and the output is a request sent from the terminal to the server. With this operation, the terminal prepares to send the image data and facial image data to the server.

[2079] Step 2:

[2080] The server receives a request from the device and stores the image data and facial im...

Claims

1. A means for a user to take a photograph of a document or image using a terminal and upload it to a server; A means for the server to perform pre-processing (tilt correction, noise removal, resolution adjustment) on the uploaded images; means for using optical character recognition techniques or artificial intelligence to extract text data from the preprocessed images; A means for validating the extracted data and informing the user of any uncertainties and requesting corrections; A means to re-validate the corrected data and finally store and export it to the required external systems; A system including:

2. A means for assigning a unique identifier to an image taken by a user when the image is uploaded to a server by the terminal; means for the server to generate and store metadata for the uploaded images; The system of claim 1 , comprising:

3. A means for the server to identify specific fields (such as names, dates, and numbers) from the extracted text data and store them as structured data; a means of checking whether the automatically validated data meets certain conditions; The system of claim 1 , comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A