system

The system enhances OCR accuracy by using generative AI to supplement missing or misrecognized characters in document images, addressing the limitations of conventional OCR technologies and improving recognition across diverse document types.

JP2026062113APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional optical character recognition (OCR) technologies struggle with low accuracy, particularly for semi-standard and non-standard documents, and fail to handle complex layouts or different formats, limiting their utility and scalability.

Method used

A system that includes uploading document images to a server for OCR processing, followed by feeding the extracted text data and original images into a generative AI model for training, which supplements missing or misrecognized characters to enhance recognition accuracy.

Benefits of technology

Improves OCR accuracy significantly, enabling high-precision character recognition and expanding the applicability of OCR systems to various document formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062113000001_ABST
    Figure 2026062113000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for uploading document images to a server, A means for performing optical character recognition processing on a document image and extracting text data, A method for feeding extracted text data and the original document image into a generative artificial intelligence model for training, A method for supplementing text data using generative artificial intelligence models, A means of providing the user with completed text data, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional optical character recognition (OCR) technology, the reading accuracy for semi-standard and non-standard documents is not sufficient, and there may be many misrecognized or unrecognized characters. As a result, the utility value of OCR is limited, and there has been a problem that it is difficult to expand the introduction scale. In particular, for documents with complex layouts or different formats, high-precision character recognition is required, but the conventional technology has a problem that it cannot cope with this.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides the following means. First, it includes means for a user to upload a document image to a server. Next, it includes means for the server to perform OCR processing on the document image and extract text data. Subsequently, it includes means for feeding the extracted text data and the original document image into a generative artificial intelligence (AI) model for training. The generative artificial intelligence model includes means for supplementing the text data based on the learned patterns and structures. Finally, by providing means for providing the supplemented text data to the user, high-precision character recognition is achieved, and the value of OCR is improved.

[0006] A "document image" is an image file that contains text and graphics, and image formats include JPEG, PNG, and PDF.

[0007] A "server" is a computer device that interacts with other computers and devices via a network to process data and provide services.

[0008] Optical Character Recognition (OCR) processing refers to the technology and process of analyzing characters in an image and converting them into digital text.

[0009] "Text data" refers to digitized character information, specifically characters and words extracted through OCR (Optical Character Recognition) processing.

[0010] A "generative artificial intelligence (AI) model" is a machine learning model that can generate new data based on input data, and is particularly capable of learning patterns and structures in text and images.

[0011] "Learning" is the process by which a generative AI model analyzes input data and acquires new knowledge and patterns based on that analysis.

[0012] "Complementation" means adding missing information to existing information to make it complete.

[0013] "User" refers to anyone who uses a system or service, including, for example, someone who uploads a document image that requires OCR processing. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit, or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit, or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] <{0000104}>In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] System Overview

[0036] This system utilizes generative artificial intelligence (AI) to perform supplementary processing with the aim of improving the accuracy of optical character recognition (OCR). Users upload their document images to the server, where OCR processing is performed. The generative AI then supplements the images, and the results are provided to the user.

[0037] Processing flow and specific actions

[0038] 1. Upload document images

[0039] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0040] 2. Execute OCR processing

[0041] The server performs OCR processing on the received document image. It uses an OCR engine to analyze the image, recognize the characters it contains, and extract the text data. This process utilizes common OCR libraries and services (such as Tesseract or Google® Cloud Vision API).

[0042] 3. Learning using generative AI

[0043] The server feeds the OCR results and the original document image into the generative AI for training. The generative AI receives the OCR results (text data) and the original image data as input and learns patterns and structures based on this data. In this step, advanced generative AI models such as GPT-4® or DALL-E are used.

[0044] 4. Completion process

[0045] The server uses generative AI to supplement text portions that are missing or misrecognized during OCR processing. Based on the information learned by the generative AI model, it generates text data to improve overall accuracy.

[0046] 5. Output of completion results

[0047] The server provides the user with the completed text data. This completed text data is returned to the user via API responses or the user interface. As a result, the user can obtain highly accurate character recognition results.

[0048] Specific example

[0049] The following scenario will be described as a specific example.

[0050] The user uploads a JPEG file of the contract scanned on their computer to the server. The server first extracts the text data from the contract using an OCR engine. At this point, some characters may be misrecognized, so the server inputs the OCR results and the original contract image into a generative AI to train the generative AI model. After the generative AI model has trained, it completes the misrecognized parts and finally generates highly accurate text data. The server provides this result to the user, allowing the user to obtain accurate digital text data of the contract. This system significantly improves the accuracy of OCR for semi-standard and non-standard documents, increasing the value of OCR.

[0051] The following describes the processing flow.

[0052] Step 1:

[0053] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0054] Step 2:

[0055] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0056] Step 3:

[0057] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0058] python

[0059] ocr_result = ocr_engine.process(image_file)

[0060] Step 4:

[0061] The server feeds the extracted text data and the original document image into a generative AI. The generative AI learns from this data to understand the patterns and structure of the document.

[0062] python

[0063] gen_ai_model.learn(ocr_result, image_file)

[0064] Step 5:

[0065] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0066] python

[0067] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0068] Step 6:

[0069] The server provides the user with the completed text data. The final result is returned to the user, allowing them to obtain accurate character recognition results.

[0070] python

[0071] send_response_to_user(enhanced_result)

[0072] Specific example

[0073] Consider a scenario where a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI to train the generative AI model (Step 4). Based on the analyzed data, the generative AI fills in any misrecognized parts and generates the final, highly accurate text data (Step 5). Finally, this completed text data is sent back to the user, who obtains the accurate text data (Step 6).

[0074] (Example 1)

[0075] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0076] Conventional optical character recognition (OCR) systems have suffered from problems such as misrecognition and low recognition accuracy. Furthermore, recognition accuracy deteriorated even further for documents with inconsistent formats or handwritten characters, making it difficult for users to obtain useful digital data. This limited their use in many applications that require accurate character data.

[0077] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0078] In this invention, the server includes means for uploading a document image to an electronic device, means for performing optical character recognition processing on the document image and extracting character data, means for feeding the extracted character data and the original document image into a generative artificial intelligence model for training, means for supplementing the character data using the generative artificial intelligence model, and means for providing the supplemented character data to the user. This improves the accuracy of OCR processing and makes it possible to provide highly accurate character data by supplementing missing or misrecognized parts.

[0079] A "document image" is a digital image that holds information in a visual format, including text and graphics.

[0080] "Electronic devices" refer to devices used for processing and communicating digital data, such as computers, smartphones, and tablets.

[0081] "Uploading" refers to the act of a user sending data from their electronic device to a server.

[0082] "Optical character recognition processing" is a technology that analyzes characters in an image and converts them into text data.

[0083] "Character data" refers to text-based information extracted through optical character recognition (OCR) processing.

[0084] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to learn data patterns and perform generation and completion.

[0085] "Learning" refers to the process by which a generative artificial intelligence model analyzes patterns and structures from input data and improves its processing capabilities based on that information.

[0086] "Complementation" refers to the process of supplementing partially missing or misidentified information with accurate data to generate complete data.

[0087] "Providing" refers to the act of returning data generated by the server as a result of processing to the user.

[0088] This section describes specific embodiments for carrying out this invention. To understand the detailed processing of the invention, the following description will specify the names of the hardware and software and show the processing flow and operation.

[0089] First, the user uploads the document image to an electronic device. Specifically, the user selects the image file (e.g., JPEG, PNG, PDF) using a web browser or a dedicated application and clicks the "Upload" button. Electronic devices used here include personal computers, smartphones, and tablets.

[0090] Next, the server receives the uploaded document image and performs optical character recognition (OCR) processing. This process uses an OCR engine (e.g., Tesseract or Google Cloud Vision API). The OCR engine analyzes the received image and extracts character data. The text data obtained at this stage is an ideal digital representation of the characters contained in the original document.

[0091] Subsequently, the server feeds the extracted character data and the original document image into a generative artificial intelligence model for training. This generative AI model uses a high-performance model (e.g., GPT-4 or DALL-E). Based on the OCR results and the original image data, the generative AI model learns the patterns and structures of the character data, enabling it to fill in inaccuracies and missing parts.

[0092] Next, the server uses a generative artificial intelligence model to supplement the missing or misrecognized character data from the OCR process. Based on the information learned by the generative AI model, it generates highly accurate character data, resulting in accurate digital text data overall.

[0093] Ultimately, the server provides the completed text data to the user through a user interface or API response. The user can view the completed text data in a web browser and download it if necessary.

[0094] As a concrete example, consider a scenario where a user uploads a JPEG file of a contract scanned on their PC to a server. The server uses an OCR engine to extract text data from the contract, but some characters may be misrecognized. Therefore, the server inputs the OCR results and the original contract image into a generative artificial intelligence model, corrects the misrecognized parts, generates final high-precision text data, and provides it to the user. This concrete example demonstrates how the present invention improves the accuracy of inaccurate character recognition results.

[0095] Examples of prompt messages are shown below.

[0096] "Based on the character recognition results after OCR processing, please complete the content using the scanned image of the contract. The original image is a scanned image of the entire contract page. Please correct any misrecognized text by replacing it with the correct characters, taking into consideration the possibility of errors, and output high-precision text data."

[0097] The above describes the embodiments for carrying out this invention.

[0098] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0099] Step 1: Upload document images

[0100] The user opens a web browser or dedicated application from their electronic device.

[0101] The user selects an image file (e.g., JPEG, PNG, PDF) and clicks the "Upload" button.

[0102] Input: Document image selected by the user.

[0103] Output: Document image file sent to the server.

[0104] Specific operation: The user selects a JPEG file of the contract on their computer and clicks the "Upload" button in the web browser. The image file is sent to the server via an HTTP POST request.

[0105] Step 2: Execute OCR processing

[0106] The server receives the uploaded document image.

[0107] The server calls an OCR engine (such as Tesseract or Google Cloud Vision API) to analyze the image and extract text data.

[0108] Input: Document image file stored on the server.

[0109] Output: Character data generated by OCR processing.

[0110] Specific operation: The server receives a JPEG file, inputs it into Tesseract, and extracts the text from the image as text data. For example, the word "contract" will be recognized as "contract," but some characters may be recognized as "misspellings."

[0111] Step 3: Learning with Generative AI

[0112] The server inputs the OCR result text data and the original document image into a generative artificial intelligence model.

[0113] The server trains a generative AI (such as GPT-4 or DALL-E).

[0114] Input: OCR result text data and original document image.

[0115] Output: Training results from a generative artificial intelligence model.

[0116] Specific operation: The server inputs the OCR result text data and the original contract image into the GPT-4 model. The generative AI analyzes the OCR result text data and the original image to learn relevant patterns and context.

[0117] Step 4: Completion process

[0118] The server uses generative AI to identify missing or misrecognized character data during OCR processing.

[0119] The server uses generative AI to fill in any gaps or misrecognitions.

[0120] Input: Training results from a generative artificial intelligence model.

[0121] Output: High-precision text data generated by a generative artificial intelligence model.

[0122] Specific operation: The generative AI model reviews the OCR results and corrects the misrecognized character parts using an algorithm for error correction. Specifically, "misspellings" are corrected to "correct characters."

[0123] Step 5: Output of completion results

[0124] The server generates the completed text data.

[0125] The server returns the completed text data to the user as a user interface or API response.

[0126] Input: High-accuracy text data augmented by a generative AI model.

[0127] Output: The final, high-precision text data provided to the user.

[0128] Specific operation: The server generates the final, high-precision text data and sends it back to the user via the user interface or API response. The user can then view and download the completed text data in their browser.

[0129] (Application Example 1)

[0130] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0131] Conventional optical character recognition (OCR) systems have problems with the accuracy of text data extracted from document images, resulting in high misrecognition rates, especially with blurry images and documents in different formats. Furthermore, while high-precision data entry is required in on-site settings such as logistics centers, manual data entry is time-consuming and labor-intensive, and prone to errors. In addition, existing systems lack sufficient post-image analysis processing, making efficient operation difficult.

[0132] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0133] In this invention, the server includes means for uploading document images to the server, means for performing optical character recognition (OCR) processing on the document images and extracting text data, means for feeding the extracted text data and the original document images into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for providing the supplemented text data to the user, and means for performing OCR processing and generative AI supplementation processing on document images input via scanning with a smartphone at a logistics center. This enables highly accurate text data extraction and supplementation of document images at a logistics center.

[0134] A "document image" is data that saves the contents of a paper or electronic document in image format.

[0135] A "server" is a computer system that provides information and services to clients via a network in response to their requests.

[0136] Optical Character Recognition (OCR) processing is a process that analyzes characters from images such as scans and photographs and extracts them as digital text data.

[0137] "Text data" refers to data stored digitally as string information.

[0138] A "generative artificial intelligence (AI) model" is an artificial intelligence system that learns patterns and structures based on large amounts of data, and then generates or completes new data.

[0139] "Methods for training" refer to methods of inputting specific data into a generative artificial intelligence model and training that model to understand the patterns and structures of the data.

[0140] "Means of supplementation" refer to methods of generating accurate data by supplementing or correcting insufficient or misidentified data.

[0141] "Means of providing text data to users" refers to methods of making processed text data publicly available and accessible to users.

[0142] A "logistics center" is a facility used for storing, distributing, and receiving goods.

[0143] A "smartphone" is a portable information terminal that, in addition to making phone calls, is capable of internet communication and operating applications.

[0144] "Scanning" is the act of digitizing an image into electronic data.

[0145] The system for realizing this invention involves uploading document images to a server using a smartphone at a logistics center, and then performing optical character recognition (OCR) processing and supplementary processing using a generative artificial intelligence (AI) model on those images.

[0146] System Configuration

[0147] hardware

[0148] Server: A server with high-performance computing capabilities. This server performs OCR processing and auto-completion processing.

[0149] Smartphone: A portable information terminal that includes a camera function. Employees at the logistics center use it to scan document images.

[0150] software

[0151] OCR engine: Software used to perform character recognition, such as Tesseract or Google Cloud Vision API.

[0152] Generative AI models: Advanced generative AI models such as GPT-4. Used for interpolation processing.

[0153] Smartphone application: A dedicated application for users to scan images and upload them to a server.

[0154] Processing flow

[0155] 1. Upload document images

[0156] Users scan document images (such as shipping labels or delivery slips) using their smartphones and upload them to the server. The process is easily managed through a dedicated application.

[0157] 2. Execute OCR processing

[0158] The server performs OCR processing on uploaded document images. It uses an OCR engine to analyze the images and extract text data. Typically, this is done using Tesseract or the Google Cloud Vision API.

[0159] 3. Completion processing by generative AI

[0160] The server feeds the OCR results and the original document image into a generative AI model for training. A model such as GPT-4 is used for this training. Based on the patterns and structures learned by the generative AI model, it compensates for misrecognized parts of the OCR results to generate highly accurate text data.

[0161] 4. Providing completion results

[0162] The server provides the augmented text data to users via a smartphone application. This allows employees at the logistics center to proceed with their work based on highly accurate data.

[0163] Specific example

[0164] Logistics center employees use the "SmartLogistics OCR" app on their smartphones to scan shipping labels. The scanned images are uploaded to a server where OCR processing is performed. Because OCR results may contain misrecognitions, the server uses a generative AI model based on GPT-4 to fill in any missing information and generate accurate text data. This text data is then provided to employees through the app.

[0165] Example of a prompt

[0166] "The OCR results are below. Please make corrections."

[0167] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0168] 2. Date: October 15, 2023

[0169] 3. Contents: 10 laptop computers

[0170] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0172] Step 1:

[0173] Users scan document images (e.g., shipping labels and delivery slips) within the logistics center using a dedicated smartphone application. These scanned image files (e.g., in JPEG format) are uploaded to a server via the dedicated application. The input is image files from the smartphone, and the output is the document images stored on the server.

[0174] Step 2:

[0175] The server performs optical character recognition (OCR) processing on uploaded document images. The server launches an OCR engine (such as Tesseract or Google Cloud Vision API), analyzes the image data, and extracts text data from it. The input is a document image, and the output is text data. However, misrecognition may occur due to unclear areas or unusual formats.

[0176] Step 3:

[0177] The server feeds the extracted text data and the original document image into a generative artificial intelligence (AI) model for training. For example, an advanced generative AI model such as GPT-4 is used. The input consists of the OCR result text data and the original document image, and the output is the patterns and structures that the AI ​​model has learned. This process generates data to supplement misrecognitions and missing parts of the OCR result.

[0178] Step 4:

[0179] The server performs text data completion processing using a generative AI model. Based on learned patterns and structures, the AI ​​model corrects misrecognized parts of the OCR results and generates accurate text data. The input consists of training data and OCR results, and the output is highly accurate, completed text data.

[0180] Step 5:

[0181] The server provides the user with completed text data. This completed text data is returned to the user via API responses or the user interface of a smartphone application. The input is completed text data, and the output is highly accurate text data accessible to the user. The user can then continue their work based on this accurate text data.

[0182] Specific example

[0183] Example of a prompt:

[0184] "The OCR results are below. Please make corrections."

[0185] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0186] 2. Date: October 15, 2023

[0187] 3. Contents: 10 laptop computers

[0188] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0189] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0190] System Overview

[0191] This system aims to improve the accuracy of optical character recognition (OCR) by combining generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions. Users upload document images to the server, where OCR processing is performed. The system then uses generative AI to complete the images, and finally, the emotion engine adjusts the completion results based on the user's emotions before providing them to the user.

[0192] Processing flow and specific actions

[0193] Upload document images

[0194] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0195] Execution of OCR processing

[0196] The server saves the received document image and then performs OCR processing. The OCR engine (for example, Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0197] Learning by generative AI

[0198] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand document patterns and structures.

[0199] Completion process

[0200] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0201] Adjustment by the emotion engine

[0202] The server adjusts the completed text data based on the user's emotions. The emotion engine recognizes emotions from the user's facial expressions, voice, or input text, and the generative AI model adjusts the completion results according to that emotional information.

[0203] Output of completion results

[0204] The server provides the user with adjusted, supplementary text data. The final result is returned to the user, allowing them to obtain accurate and emotionally sensitive character recognition results.

[0205] Specific example

[0206] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer. The server receives it and saves the image. Next, the server analyzes the image using an OCR engine and extracts the initial text data. Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model. Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data.

[0207] The emotion engine analyzes the generated text data to identify information (facial expressions, voice, input characters, etc.) necessary to recognize the user's emotions. For example, if the user is satisfied, the basic completion result is provided as is; if the user is dissatisfied, the generative AI adjusts the result to provide an even more accurate one. Finally, this completed, high-accuracy text data is provided to the user, allowing them to obtain highly accurate and emotion-sensitive character recognition results.

[0208] This invention not only improves the accuracy of OCR but also enables the provision of services that take user emotions into consideration, thereby enhancing the user experience and significantly expanding the value and scale of deployment of OCR systems.

[0209] The following describes the processing flow.

[0210] Step 1:

[0211] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0212] Step 2:

[0213] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0214] Step 3:

[0215] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the initial text data.

[0216] python

[0217] ocr_result = ocr_engine.process(image_file)

[0218] Step 4:

[0219] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand the patterns and structure of documents.

[0220] python

[0221] gen_ai_model.learn(ocr_result, image_file)

[0222] Step 5:

[0223] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0224] python

[0225] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0226] Step 6:

[0227] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expressions, voice, or input text to determine whether the user is satisfied, irritated, anxious, etc.

[0228] python

[0229] user_emotion = emotion_engine.analyze(user_input)

[0230] Step 7:

[0231] The server uses a generative AI model to adjust the supplementary text data based on the user's emotions. For example, if the user expresses dissatisfaction, the generative AI adjusts the data to produce more accurate supplementary results.

[0232] python

[0233] adjusted_result = gen_ai_model.adjust_based_on_emotion(enhanced_result, user_emotion)

[0234] Step 8:

[0235] The server provides the user with adjusted, supplementary text data. The final result is returned to the user via an API response or user interface, allowing the user to obtain accurate and emotionally sensitive character recognition results.

[0236] python

[0237] send_response_to_user(adjusted_result)

[0238] Specific example

[0239] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts the initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model (Step 4). Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data (Step 5).

[0240] For this generated text data, the emotion engine analyzes the user's facial expressions, voice, or input text to recognize the user's emotions (Step 6). For example, if the user is expressing dissatisfaction, the generative AI adjusts the results to produce even more accurate results (Step 7). Finally, this adjusted supplementary text data is provided to the user, allowing them to obtain accurate and emotion-sensitive character recognition results (Step 8). This improves the overall user experience by enabling flexible responses that respond to the user's emotions, in addition to improving the accuracy of OCR.

[0241] (Example 2)

[0242] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0243] Optical character recognition (OCR) technology is used in many fields, but it has challenges in recognition accuracy, particularly with misrecognition and missed recognition. Furthermore, recognition results can affect user satisfaction, and there is a need for methods that reflect user emotions regarding the recognition results. Therefore, high-precision OCR processing and the provision of results that take user emotions into consideration are necessary.

[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0245] In this invention, the server includes means for uploading a document image to an information processing device, means for performing optical character recognition processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence model for training, means for supplementing the text data using the generative artificial intelligence model, means for adjusting the supplemented text data based on the user's emotions, and means for providing the adjusted text data to the user. This enables highly accurate OCR processing and the provision of results that take into account the user's emotions.

[0246] A "document image" is image data that is the target of optical character recognition processing, and is generally an image format in which text is embedded (e.g., JPEG, PNG, PDF).

[0247] "Information processing equipment" refers to all devices that perform data input, processing, storage, output, etc., and in this context mainly includes servers and client terminals.

[0248] Optical Character Recognition (OCR) processing is a technology that analyzes characters contained in an image and extracts them as text data.

[0249] "Text data" refers to data representing character information extracted through optical character recognition processing.

[0250] A "generative artificial intelligence model" is an artificial intelligence model that performs data completion or generation based on training data, and in this context, it is used to complete misrecognized or unrecognized parts of OCR results.

[0251] "User" refers to an individual or organization that utilizes optical character recognition (OCR) processing services.

[0252] "Means of adjusting based on emotions" refers to a function that analyzes the user's emotional information and has a generative artificial intelligence model make adjustments in accordance with those emotions.

[0253] "Means of complementarity" refers to a function that complements OCR results based on data learned by a generative artificial intelligence model, thereby generating highly accurate text data.

[0254] This invention relates to a system that combines generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions, with the aim of improving the accuracy of optical character recognition (OCR). Specific embodiments of this system are described below.

[0255] System Overview

[0256] The user uploads a document image to an information processing device (server), where OCR processing is performed. This is then supplemented by a generative AI model, and finally, an emotion engine is used to adjust the supplemented result based on the user's emotions before providing it. Specifically, the process proceeds in the following steps:

[0257] Hardware and software

[0258] This system uses the following hardware and software:

[0259] 1. User's device:

[0260] Web browser or dedicated application

[0261] 2. Server:

[0262] Data storage function

[0263] OCR engine (e.g., Tesseract, Google Cloud Vision API)

[0264] Generative AI models

[0265] Emotional Engine

[0266] Data processing and data calculation

[0267] 1. The user uploads a document image from their device to the server.

[0268] 2. The server saves the received document image and extracts the text data using an OCR engine.

[0269] 3. The server inputs the extracted text data and the original document image into a generative AI model for training.

[0270] 4. Generative AI models are used to supplement the OCR results and generate high-accuracy text data.

[0271] 5. The emotion engine analyzes the user's emotional information, and the generative AI model adjusts the completion results based on that information.

[0272] 6. The server provides users with pre-calibrated, high-precision text data.

[0273] Specific example

[0274] Next, we will show a specific example of a user uploading a JPEG file of a contract to the server.

[0275] 1. Upload document images:

[0276] The user uses the browser on their own computer to access the upload page of the contract JPEG file and clicks the "Upload" button. The image file is sent to the server.

[0277] 2. Execution of OCR processing:

[0278] The server saves the received image file and then executes OCR processing using Tesseract or Google Cloud Vision API to extract the initial text data.

[0279] 3. Learning by generative AI:

[0280] The server inputs the extracted text data and the original document image into the generative AI program to train it to understand the pattern and structure of the document.

[0281] 4. Completion processing:

[0282] The generative AI model completes the misrecognized or unrecognized parts to generate high-precision text data.

[0283] 5. Adjustment by emotion engine:

[0284] If the user is dissatisfied, the server further adjusts the completion result based on the user's emotion information (e.g., feedback by input).

[0285] 6. Output of completion result:

[0286] The server provides the adjusted high-precision text data to the user, and the user checks the result in the browser.

[0287] Specific examples of prompt sentences

[0288] The following are specific examples of prompt sentences for training this system in the generative AI model:

[0289] Image file: Contract_2023.jpeg

[0290] Extracted text: "This agreement is effective as of March 1, 2023…"

[0291] User sentiment: "Satisfied"

[0292] AI-generated output: "This agreement was concluded on March 1, 2023…"

[0293] Adjusted and supplemented result: "This agreement was entered into on March 1, 2023, and is effective as of that date…"

[0294] Image file: receipt_2023.pdf

[0295] Extracted text: "Item Quantity Price"

[0296] User sentiment: "Dissatisfied"

[0297] AI-generated completion result: "Item Quantity Price Total Amount"

[0298] Adjusted and completed result: "Item Quantity Price Total Amount Promotion Applied"

[0299] This enables highly accurate OCR processing and the provision of results that take user emotions into consideration.

[0300] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0301] Step 1:

[0302] The user selects a document image on their device and uploads it to the information processing device (server).

[0303] Input: User-selected document image (e.g., JPEG, PNG, PDF file)

[0304] Output: Document image file stored in the server

[0305] Specific operations:

[0306] The user uses the web browser on their own computer to click the file selection button and select the document image to be uploaded. Next, by clicking the "Upload" button, the image file is sent to the server and stored in the designated storage of the server.

[0307] Step 2:

[0308] The server performs OCR processing on the stored document image.

[0309] Input: Stored document image file

[0310] Output: Extracted text data

[0311] Specific operations:

[0312] The server uses an OCR engine (e.g., Tesseract or Google Cloud Vision API) to analyze the stored document image file. The OCR engine recognizes the characters in the image and extracts the text data. The extracted text data is stored within the server.

[0313] Step 3:

[0314] The server inputs the extracted text data and the original document image into a generative AI model for training.

[0315] Input: Extracted text data, original document image

[0316] Output: Trained model by the generative AI model

[0317] Specific operations:

[0318] The server inputs the text data obtained through OCR processing and the original document image into a generative AI model. The generative AI model learns and understands the document's patterns and structure based on this data. The trained model is then used for subsequent completion processing.

[0319] Step 4:

[0320] The server uses a generative AI model to supplement the OCR results.

[0321] Input: Pre-trained generative AI model, initial OCR text data

[0322] Output: Interpolated high-precision text data

[0323] Specific actions:

[0324] The server uses a pre-trained generative AI model to perform completion processing on the initial OCR text data. Specifically, misrecognized or unrecognized portions are identified, and these portions are completed by the generative AI model. The completed, high-accuracy text data is stored on the server.

[0325] Step 5:

[0326] The server uses an emotion engine to adjust the completion results based on the user's emotions.

[0327] Input: Completed, high-precision text data, user sentiment information (facial expressions, voice, input characters)

[0328] Output: Text data adjusted based on emotions

[0329] Specific actions:

[0330] Users provide emotional information in real time through their camera and microphone, or input emotional feedback. The server uses an emotion engine to analyze the user's emotional information, has a generative AI model make adjustments based on the emotions, and further refines the completion results.

[0331] Step 6:

[0332] The server then provides the user with the final, adjusted, high-precision text data.

[0333] Input: Text data adjusted based on emotions

[0334] Output: Final high-precision text data provided to the user.

[0335] Specific actions:

[0336] The server sends the adjusted, high-precision text data to the user. The user then downloads or views the results again via a web browser or dedicated application and performs the necessary processing.

[0337] (Application Example 2)

[0338] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0339] Conventional optical character recognition (OCR) systems have suffered from frequent misrecognition in extracting text data from document images, resulting in low accuracy. Furthermore, they often disregarded the user experience, leading to user dissatisfaction. This invention aims not only to improve the accuracy of OCR but also to enhance the user experience by considering user feelings.

[0340] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0341] In this invention, the server includes means for uploading a document image to the server, means for performing optical character recognition (OCR) processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for adjusting the supplementation results with an emotion engine that analyzes the user's emotional information, and means for providing the supplemented text data to the user. This makes it possible to improve OCR accuracy and provide text data that takes the user's emotions into consideration.

[0342] A "document image" is an image file containing text information provided by the user.

[0343] A "server" is a computer system that provides services to client terminals over a network.

[0344] Optical Character Recognition (OCR) processing is a technology that recognizes characters and numbers in an image and extracts them as text data.

[0345] "Text data" refers to data containing characters and numbers extracted from a document image.

[0346] A "generative artificial intelligence (AI) model" is an artificial intelligence that has the ability to learn from large amounts of data and generate new data.

[0347] An "emotion engine" is software that recognizes a user's emotions and adjusts the processing results based on those emotions.

[0348] "Completion" is the process of filling in omissions and errors in initial text data to make it closer to a complete form.

[0349] "Means of providing to the user" refers to the method of displaying or transmitting the final processing result to the user.

[0350] This invention provides a system that performs high-precision OCR processing on document images while also taking user emotions into consideration. It is basically implemented by the following means.

[0351] Hardware and software

[0352] Hardware: Primarily uses servers, client terminals, and smartphones.

[0353] The server is a central computer system that performs OCR processing, executes generative artificial intelligence (AI) models, and manages data, and is connected to users by applications running on client terminals and smartphones.

[0354] Software: Use the following software.

[0355] OCR engine: Uses tools such as pytesseract and Google Cloud Vision API to perform character recognition from document images.

[0356] Generative artificial intelligence models: Text data is completed using the Transformers library (e.g., GPT-3(registered trademark)).

[0357] Emotion Engine: Analyzes user emotions using TextBlob.

[0358] Processing flow

[0359] 1. Uploading Document Images: Users upload document images (e.g., JPEG, PNG, PDF) to the server. Uploads are performed via a dedicated application on a smartphone or client device.

[0360] 2. OCR processing: The server saves the received document image and extracts the text data using an OCR engine (e.g., pytesseract). The extracted text data is saved as an intermediate result.

[0361] 3. Completion by Generative AI: The server feeds the extracted text data and the original document image into a generative AI model (e.g., GPT-3) for training. The generative AI model uses this data to complete the text and generate high-accuracy text data.

[0362] 4. Adjustment by the emotion engine: The server analyzes the user's emotion information using the emotion engine (TextBlob). If the emotion information is "satisfied," the completion result is used as is; if it is "dissatisfied," the text is further modified and completed.

[0363] 5. Provision to the user: The server ultimately provides the user with the completed and adjusted text data. The user obtains highly accurate and emotion-sensitive character recognition results through a dedicated application.

[0364] Specific example

[0365] For example, suppose a user uploads an image of a warranty certificate for a virtual store, and OCR processing is required. In this case, the user takes a picture of the warranty certificate within the virtual store app and sends it to the server. On the server, the OCR engine analyzes the image and extracts initial text data. Then, a generative AI model (e.g., GPT-3) completes this text data, and an emotion engine (e.g., TextBlob) analyzes the user's input text to recognize emotions.

[0366] For example, if a user enters "The contents of this warranty are inaccurate," the emotion engine recognizes the dissatisfaction, and the generative AI model corrects and supplements the text data again. Finally, the completed and adjusted text is returned to the user. An example of the prompt text in this case is as follows:

[0367] "The information on the warranty is inaccurate. Please see below for details:"

[0368] This allows users to obtain highly accurate character recognition results and makes it easier to manage warranty certificates.

[0369] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0370] Step 1:

[0371] The user uploads a document image.

[0372] Input: Document image (JPEG, PNG, PDF, etc.)

[0373] Specific operation: The user launches a dedicated application, takes or selects a document image, and sends it to the server. Clicking the upload button transfers the image to the server.

[0374] Output: Uploaded document image file

[0375] Step 2:

[0376] The server saves the document image and performs OCR processing.

[0377] Input: Uploaded document image file

[0378] Specific operation: The server saves the received image file and runs an OCR engine (e.g., pytesseract) to extract text data from the image. It then performs grayscale conversion, noise reduction, and character recognition on the image.

[0379] Output: Extracted text data

[0380] Step 3:

[0381] The server loads the extracted text data and the original document images into a generative AI model for training.

[0382] Input: Text data, document images

[0383] Specific operation: The server inputs text data and images into a generative artificial intelligence model (e.g., GPT-3) to train it on document patterns and structures. The model then learns to correct misrecognitions based on the input data.

[0384] Output: Generated high-precision text data

[0385] Step 4:

[0386] The server analyzes the user's emotional information and adjusts the completion results accordingly.

[0387] Input: Generated text data, user sentiment information

[0388] Specific operation: The server receives additional input from the user (e.g., "This part is inaccurate") and uses a sentiment engine (e.g., TextBlob) to analyze the user's sentiment. It then further completes or adjusts the text data based on the sentiment.

[0389] Output: Adjusted text data

[0390] Step 5:

[0391] The server provides the user with the final, completed, and adjusted text data.

[0392] Input: Adjusted text data

[0393] Specific operation: The server saves the final adjustment results as a file and sends it to the user through a dedicated application. The user can then view the text data within the application.

[0394] Output: High-precision, emotion-sensitive text data acquired by the user.

[0395] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0396] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0397] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0398] [Second Embodiment]

[0399] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0400] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0401] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0402] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0403] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0404] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0405] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0406] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0407] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0408] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0409] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0410] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0411] System Overview

[0412] This system utilizes generative artificial intelligence (AI) to perform supplementary processing with the aim of improving the accuracy of optical character recognition (OCR). Users upload their document images to the server, where OCR processing is performed. The generative AI then supplements the images, and the results are provided to the user.

[0413] Processing flow and specific actions

[0414] 1. Upload document images

[0415] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0416] 2. Execute OCR processing

[0417] The server performs OCR processing on the received document image. It uses an OCR engine to analyze the image, recognize the characters it contains, and extract text data. This process utilizes common OCR libraries and services (such as Tesseract or Google Cloud Vision API).

[0418] 3. Learning using generative AI

[0419] The server feeds the OCR results and the original document image into a generative AI for training. The generative AI receives the OCR results (text data) and the original image data as input and learns patterns and structures based on this data. This step uses advanced generative AI models such as GPT-4 or DALL-E.

[0420] 4. Completion process

[0421] The server uses generative AI to supplement text portions that are missing or misrecognized during OCR processing. Based on the information learned by the generative AI model, it generates text data to improve overall accuracy.

[0422] 5. Output of completion results

[0423] The server provides the user with the completed text data. This completed text data is returned to the user via API responses or the user interface. As a result, the user can obtain highly accurate character recognition results.

[0424] Specific example

[0425] The following scenario will be described as a specific example.

[0426] The user uploads a JPEG file of the contract scanned on their computer to the server. The server first extracts the text data from the contract using an OCR engine. At this point, some characters may be misrecognized, so the server inputs the OCR results and the original contract image into a generative AI to train the generative AI model. After the generative AI model has trained, it completes the misrecognized parts and finally generates highly accurate text data. The server provides this result to the user, allowing the user to obtain accurate digital text data of the contract. This system significantly improves the accuracy of OCR for semi-standard and non-standard documents, increasing the value of OCR.

[0427] The following describes the processing flow.

[0428] Step 1:

[0429] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0430] Step 2:

[0431] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0432] Step 3:

[0433] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0434] python

[0435] ocr_result = ocr_engine.process(image_file)

[0436] Step 4:

[0437] The server feeds the extracted text data and the original document image into a generative AI. The generative AI learns from this data to understand the patterns and structure of the document.

[0438] python

[0439] gen_ai_model.learn(ocr_result, image_file)

[0440] Step 5:

[0441] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0442] python

[0443] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0444] Step 6:

[0445] The server provides the user with the completed text data. The final result is returned to the user, allowing them to obtain accurate character recognition results.

[0446] python

[0447] send_response_to_user(enhanced_result)

[0448] Specific example

[0449] Consider a scenario where a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI to train the generative AI model (Step 4). Based on the analyzed data, the generative AI fills in any misrecognized parts and generates the final, highly accurate text data (Step 5). Finally, this completed text data is sent back to the user, who obtains the accurate text data (Step 6).

[0450] (Example 1)

[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0452] Conventional optical character recognition (OCR) systems have suffered from problems such as misrecognition and low recognition accuracy. Furthermore, recognition accuracy deteriorated even further for documents with inconsistent formats or handwritten characters, making it difficult for users to obtain useful digital data. This limited their use in many applications that require accurate character data.

[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0454] In this invention, the server includes means for uploading a document image to an electronic device, means for performing optical character recognition processing on the document image and extracting character data, means for feeding the extracted character data and the original document image into a generative artificial intelligence model for training, means for supplementing the character data using the generative artificial intelligence model, and means for providing the supplemented character data to the user. This improves the accuracy of OCR processing and makes it possible to provide highly accurate character data by supplementing missing or misrecognized parts.

[0455] A "document image" is a digital image that holds information in a visual format, including text and graphics.

[0456] "Electronic devices" refer to devices used for processing and communicating digital data, such as computers, smartphones, and tablets.

[0457] "Uploading" refers to the act of a user sending data from their electronic device to a server.

[0458] "Optical character recognition processing" is a technology that analyzes characters in an image and converts them into text data.

[0459] "Character data" refers to text-based information extracted through optical character recognition (OCR) processing.

[0460] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to learn data patterns and perform generation and completion.

[0461] "Learning" refers to the process by which a generative artificial intelligence model analyzes patterns and structures from input data and improves its processing capabilities based on that information.

[0462] "Complementation" refers to the process of supplementing partially missing or misidentified information with accurate data to generate complete data.

[0463] "Providing" refers to the act of returning data generated by the server as a result of processing to the user.

[0464] This section describes specific embodiments for carrying out this invention. To understand the detailed processing of the invention, the following description will specify the names of the hardware and software and show the processing flow and operation.

[0465] First, the user uploads the document image to an electronic device. Specifically, the user selects the image file (e.g., JPEG, PNG, PDF) using a web browser or a dedicated application and clicks the "Upload" button. Electronic devices used here include personal computers, smartphones, and tablets.

[0466] Next, the server receives the uploaded document image and performs optical character recognition (OCR) processing. This process uses an OCR engine (e.g., Tesseract or Google Cloud Vision API). The OCR engine analyzes the received image and extracts character data. The text data obtained at this stage is an ideal digital representation of the characters contained in the original document.

[0467] Subsequently, the server feeds the extracted character data and the original document image into a generative artificial intelligence model for training. This generative AI model uses a high-performance model (e.g., GPT-4 or DALL-E). Based on the OCR results and the original image data, the generative AI model learns the patterns and structures of the character data, enabling it to fill in inaccuracies and missing parts.

[0468] Next, the server uses a generative artificial intelligence model to supplement the missing or misrecognized character data from the OCR process. Based on the information learned by the generative AI model, it generates highly accurate character data, resulting in accurate digital text data overall.

[0469] Ultimately, the server provides the completed text data to the user through a user interface or API response. The user can view the completed text data in a web browser and download it if necessary.

[0470] As a concrete example, consider a scenario where a user uploads a JPEG file of a contract scanned on their PC to a server. The server uses an OCR engine to extract text data from the contract, but some characters may be misrecognized. Therefore, the server inputs the OCR results and the original contract image into a generative artificial intelligence model, corrects the misrecognized parts, generates final high-precision text data, and provides it to the user. This concrete example demonstrates how the present invention improves the accuracy of inaccurate character recognition results.

[0471] Examples of prompt messages are shown below.

[0472] "Based on the character recognition results after OCR processing, please complete the content using the scanned image of the contract. The original image is a scanned image of the entire contract page. Please correct any misrecognized text by replacing it with the correct characters, taking into consideration the possibility of errors, and output high-precision text data."

[0473] The above describes the embodiments for carrying out this invention.

[0474] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0475] Step 1: Upload document images

[0476] The user opens a web browser or dedicated application from their electronic device.

[0477] The user selects an image file (e.g., JPEG, PNG, PDF) and clicks the "Upload" button.

[0478] Input: Document image selected by the user.

[0479] Output: Document image file sent to the server.

[0480] Specific operation: The user selects a JPEG file of the contract on their computer and clicks the "Upload" button in the web browser. The image file is sent to the server via an HTTP POST request.

[0481] Step 2: Execute OCR processing

[0482] The server receives the uploaded document image.

[0483] The server calls an OCR engine (such as Tesseract or Google Cloud Vision API) to analyze the image and extract text data.

[0484] Input: Document image file stored on the server.

[0485] Output: Character data generated by OCR processing.

[0486] Specific operation: The server receives a JPEG file, inputs it into Tesseract, and extracts the text from the image as text data. For example, the word "contract" will be recognized as "contract," but some characters may be recognized as "misspellings."

[0487] Step 3: Learning with Generative AI

[0488] The server inputs the OCR result text data and the original document image into a generative artificial intelligence model.

[0489] The server trains a generative AI (such as GPT-4 or DALL-E).

[0490] Input: OCR result text data and original document image.

[0491] Output: Training results from a generative artificial intelligence model.

[0492] Specific operation: The server inputs the OCR result text data and the original contract image into the GPT-4 model. The generative AI analyzes the OCR result text data and the original image to learn relevant patterns and context.

[0493] Step 4: Completion process

[0494] The server uses generative AI to identify missing or misrecognized character data during OCR processing.

[0495] The server uses generative AI to fill in any gaps or misrecognitions.

[0496] Input: Training results from a generative artificial intelligence model.

[0497] Output: High-precision text data generated by a generative artificial intelligence model.

[0498] Specific operation: The generative AI model reviews the OCR results and corrects the misrecognized character parts using an algorithm for error correction. Specifically, "misspellings" are corrected to "correct characters."

[0499] Step 5: Output of completion results

[0500] The server generates the completed text data.

[0501] The server returns the completed text data to the user as a user interface or API response.

[0502] Input: High-accuracy text data augmented by a generative AI model.

[0503] Output: The final, high-precision text data provided to the user.

[0504] Specific operation: The server generates the final, high-precision text data and sends it back to the user via the user interface or API response. The user can then view and download the completed text data in their browser.

[0505] (Application Example 1)

[0506] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0507] Conventional optical character recognition (OCR) systems have problems with the accuracy of text data extracted from document images, resulting in high misrecognition rates, especially with blurry images and documents in different formats. Furthermore, while high-precision data entry is required in on-site settings such as logistics centers, manual data entry is time-consuming and labor-intensive, and prone to errors. In addition, existing systems lack sufficient post-image analysis processing, making efficient operation difficult.

[0508] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0509] In this invention, the server includes means for uploading document images to the server, means for performing optical character recognition (OCR) processing on the document images and extracting text data, means for feeding the extracted text data and the original document images into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for providing the supplemented text data to the user, and means for performing OCR processing and generative AI supplementation processing on document images input via scanning with a smartphone at a logistics center. This enables highly accurate text data extraction and supplementation of document images at a logistics center.

[0510] A "document image" is data that saves the contents of a paper or electronic document in image format.

[0511] A "server" is a computer system that provides information and services to clients via a network in response to their requests.

[0512] Optical Character Recognition (OCR) processing is a process that analyzes characters from images such as scans and photographs and extracts them as digital text data.

[0513] "Text data" refers to data stored digitally as string information.

[0514] A "generative artificial intelligence (AI) model" is an artificial intelligence system that learns patterns and structures based on large amounts of data, and then generates or completes new data.

[0515] "Methods for training" refer to methods of inputting specific data into a generative artificial intelligence model and training that model to understand the patterns and structures of the data.

[0516] "Means of supplementation" refer to methods of generating accurate data by supplementing or correcting insufficient or misidentified data.

[0517] "Means of providing text data to users" refers to methods of making processed text data publicly available and accessible to users.

[0518] A "logistics center" is a facility used for storing, distributing, and receiving goods.

[0519] A "smartphone" is a portable information terminal that, in addition to making phone calls, is capable of internet communication and operating applications.

[0520] "Scanning" is the act of digitizing an image into electronic data.

[0521] The system for realizing this invention involves uploading document images to a server using a smartphone at a logistics center, and then performing optical character recognition (OCR) processing and supplementary processing using a generative artificial intelligence (AI) model on those images.

[0522] System Configuration

[0523] hardware

[0524] Server: A server with high-performance computing capabilities. This server performs OCR processing and auto-completion processing.

[0525] Smartphone: A portable information terminal that includes a camera function. Employees at the logistics center use it to scan document images.

[0526] software

[0527] OCR engine: Software used to perform character recognition, such as Tesseract or Google Cloud Vision API.

[0528] Generative AI models: Advanced generative AI models such as GPT-4. Used for interpolation processing.

[0529] Smartphone application: A dedicated application for users to scan images and upload them to a server.

[0530] Processing flow

[0531] 1. Upload document images

[0532] Users scan document images (such as shipping labels or delivery slips) using their smartphones and upload them to the server. The process is easily managed through a dedicated application.

[0533] 2. Execute OCR processing

[0534] The server performs OCR processing on uploaded document images. It uses an OCR engine to analyze the images and extract text data. Typically, this is done using Tesseract or the Google Cloud Vision API.

[0535] 3. Completion processing by generative AI

[0536] The server feeds the OCR results and the original document image into a generative AI model for training. A model such as GPT-4 is used for this training. Based on the patterns and structures learned by the generative AI model, it compensates for misrecognized parts of the OCR results to generate highly accurate text data.

[0537] 4. Providing completion results

[0538] The server provides the augmented text data to users via a smartphone application. This allows employees at the logistics center to proceed with their work based on highly accurate data.

[0539] Specific example

[0540] Logistics center employees use the "SmartLogistics OCR" app on their smartphones to scan shipping labels. The scanned images are uploaded to a server where OCR processing is performed. Because OCR results may contain misrecognitions, the server uses a generative AI model based on GPT-4 to fill in any missing information and generate accurate text data. This text data is then provided to employees through the app.

[0541] Example of a prompt

[0542] "The OCR results are below. Please make corrections."

[0543] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0544] 2. Date: October 15, 2023

[0545] 3. Contents: 10 laptop computers

[0546] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0548] Step 1:

[0549] Users scan document images (e.g., shipping labels and delivery slips) within the logistics center using a dedicated smartphone application. These scanned image files (e.g., in JPEG format) are uploaded to a server via the dedicated application. The input is image files from the smartphone, and the output is the document images stored on the server.

[0550] Step 2:

[0551] The server performs optical character recognition (OCR) processing on uploaded document images. The server launches an OCR engine (such as Tesseract or Google Cloud Vision API), analyzes the image data, and extracts text data from it. The input is a document image, and the output is text data. However, misrecognition may occur due to unclear areas or unusual formats.

[0552] Step 3:

[0553] The server feeds the extracted text data and the original document image into a generative artificial intelligence (AI) model for training. For example, an advanced generative AI model such as GPT-4 is used. The input consists of the OCR result text data and the original document image, and the output is the patterns and structures that the AI ​​model has learned. This process generates data to supplement misrecognitions and missing parts of the OCR result.

[0554] Step 4:

[0555] The server performs text data completion processing using a generative AI model. Based on learned patterns and structures, the AI ​​model corrects misrecognized parts of the OCR results and generates accurate text data. The input consists of training data and OCR results, and the output is highly accurate, completed text data.

[0556] Step 5:

[0557] The server provides the user with completed text data. This completed text data is returned to the user via API responses or the user interface of a smartphone application. The input is completed text data, and the output is highly accurate text data accessible to the user. The user can then continue their work based on this accurate text data.

[0558] Specific example

[0559] Example of a prompt:

[0560] "The OCR results are below. Please make corrections."

[0561] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0562] 2. Date: October 15, 2023

[0563] 3. Contents: 10 laptop computers

[0564] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0565] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0566] System Overview

[0567] This system aims to improve the accuracy of optical character recognition (OCR) by combining generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions. Users upload document images to the server, where OCR processing is performed. The system then uses generative AI to complete the images, and finally, the emotion engine adjusts the completion results based on the user's emotions before providing them to the user.

[0568] Processing flow and specific actions

[0569] Upload document images

[0570] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0571] Execution of OCR processing

[0572] The server saves the received document image and then performs OCR processing. The OCR engine (for example, Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0573] Learning by generative AI

[0574] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand document patterns and structures.

[0575] Completion process

[0576] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0577] Adjustment by the emotion engine

[0578] The server adjusts the completed text data based on the user's emotions. The emotion engine recognizes emotions from the user's facial expressions, voice, or input text, and the generative AI model adjusts the completion results according to that emotional information.

[0579] Output of completion results

[0580] The server provides the user with adjusted, supplementary text data. The final result is returned to the user, allowing them to obtain accurate and emotionally sensitive character recognition results.

[0581] Specific example

[0582] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer. The server receives it and saves the image. Next, the server analyzes the image using an OCR engine and extracts the initial text data. Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model. Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data.

[0583] The emotion engine analyzes the generated text data to identify information (facial expressions, voice, input characters, etc.) necessary to recognize the user's emotions. For example, if the user is satisfied, the basic completion result is provided as is; if the user is dissatisfied, the generative AI adjusts the result to provide an even more accurate one. Finally, this completed, high-accuracy text data is provided to the user, allowing them to obtain highly accurate and emotion-sensitive character recognition results.

[0584] This invention not only improves the accuracy of OCR but also enables the provision of services that take user emotions into consideration, thereby enhancing the user experience and significantly expanding the value and scale of deployment of OCR systems.

[0585] The following describes the processing flow.

[0586] Step 1:

[0587] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0588] Step 2:

[0589] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0590] Step 3:

[0591] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the initial text data.

[0592] python

[0593] ocr_result = ocr_engine.process(image_file)

[0594] Step 4:

[0595] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand the patterns and structure of documents.

[0596] python

[0597] gen_ai_model.learn(ocr_result, image_file)

[0598] Step 5:

[0599] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0600] python

[0601] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0602] Step 6:

[0603] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expressions, voice, or input text to determine whether the user is satisfied, irritated, anxious, etc.

[0604] python

[0605] user_emotion = emotion_engine.analyze(user_input)

[0606] Step 7:

[0607] The server uses a generative AI model to adjust the supplementary text data based on the user's emotions. For example, if the user expresses dissatisfaction, the generative AI adjusts the data to produce more accurate supplementary results.

[0608] python

[0609] adjusted_result = gen_ai_model.adjust_based_on_emotion(enhanced_result, user_emotion)

[0610] Step 8:

[0611] The server provides the user with adjusted, supplementary text data. The final result is returned to the user via an API response or user interface, allowing the user to obtain accurate and emotionally sensitive character recognition results.

[0612] python

[0613] send_response_to_user(adjusted_result)

[0614] Specific example

[0615] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts the initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model (Step 4). Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data (Step 5).

[0616] For this generated text data, the emotion engine analyzes the user's facial expressions, voice, or input text to recognize the user's emotions (Step 6). For example, if the user is expressing dissatisfaction, the generative AI adjusts the results to produce even more accurate results (Step 7). Finally, this adjusted supplementary text data is provided to the user, allowing them to obtain accurate and emotion-sensitive character recognition results (Step 8). This improves the overall user experience by enabling flexible responses that respond to the user's emotions, in addition to improving the accuracy of OCR.

[0617] (Example 2)

[0618] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0619] Optical character recognition (OCR) technology is used in many fields, but it has challenges in recognition accuracy, particularly with misrecognition and missed recognition. Furthermore, recognition results can affect user satisfaction, and there is a need for methods that reflect user emotions regarding the recognition results. Therefore, high-precision OCR processing and the provision of results that take user emotions into consideration are necessary.

[0620] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0621] In this invention, the server includes means for uploading a document image to an information processing device, means for performing optical character recognition processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence model for training, means for supplementing the text data using the generative artificial intelligence model, means for adjusting the supplemented text data based on the user's emotions, and means for providing the adjusted text data to the user. This enables highly accurate OCR processing and the provision of results that take into account the user's emotions.

[0622] A "document image" is image data that is the target of optical character recognition processing, and is generally an image format in which text is embedded (e.g., JPEG, PNG, PDF).

[0623] "Information processing equipment" refers to all devices that perform data input, processing, storage, output, etc., and in this context mainly includes servers and client terminals.

[0624] Optical Character Recognition (OCR) processing is a technology that analyzes characters contained in an image and extracts them as text data.

[0625] "Text data" refers to data representing character information extracted through optical character recognition processing.

[0626] A "generative artificial intelligence model" is an artificial intelligence model that performs data completion or generation based on training data, and in this context, it is used to complete misrecognized or unrecognized parts of OCR results.

[0627] "User" refers to an individual or organization that utilizes optical character recognition (OCR) processing services.

[0628] "Means of adjusting based on emotions" refers to a function that analyzes the user's emotional information and has a generative artificial intelligence model make adjustments in accordance with those emotions.

[0629] "Means of complementarity" refers to a function that complements OCR results based on data learned by a generative artificial intelligence model, thereby generating highly accurate text data.

[0630] This invention relates to a system that combines generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions, with the aim of improving the accuracy of optical character recognition (OCR). Specific embodiments of this system are described below.

[0631] System Overview

[0632] The user uploads a document image to an information processing device (server), where OCR processing is performed. This is then supplemented by a generative AI model, and finally, an emotion engine is used to adjust the supplemented result based on the user's emotions before providing it. Specifically, the process proceeds in the following steps:

[0633] Hardware and software

[0634] This system uses the following hardware and software:

[0635] 1. User's device:

[0636] Web browser or dedicated application

[0637] 2. Server:

[0638] Data storage function

[0639] OCR engine (e.g., Tesseract, Google Cloud Vision API)

[0640] Generative AI models

[0641] Emotional Engine

[0642] Data processing and data calculation

[0643] 1. The user uploads a document image from their device to the server.

[0644] 2. The server saves the received document image and extracts the text data using an OCR engine.

[0645] 3. The server inputs the extracted text data and the original document image into a generative AI model for training.

[0646] 4. Generative AI models are used to supplement the OCR results and generate high-accuracy text data.

[0647] 5. The emotion engine analyzes the user's emotional information, and the generative AI model adjusts the completion results based on that information.

[0648] 6. The server provides users with pre-calibrated, high-precision text data.

[0649] Specific example

[0650] Next, we will show a specific example of a user uploading a JPEG file of a contract to the server.

[0651] 1. Upload document images:

[0652] The user accesses the upload page for the JPEG file of the contract using a browser on their computer and clicks the "Upload" button. The image file is then sent to the server.

[0653] 2. Perform OCR processing:

[0654] The server saves the received image file and then performs OCR processing using Tesseract or the Google Cloud Vision API to extract the initial text data.

[0655] 3. Learning using generative AI:

[0656] The server inputs the extracted text data and the original document image into a generative AI program, which then learns to understand the patterns and structure of the document.

[0657] 4. Completion process:

[0658] The generative AI model fills in the gaps in misrecognition and unrecognized parts, generating highly accurate text data.

[0659] 5. Adjustment by the Emotion Engine:

[0660] If a user is dissatisfied, the server will further adjust the completion results based on the user's emotional information (e.g., feedback from input).

[0661] 6. Output of completion results:

[0662] The server provides the user with adjusted, high-precision text data, and the user checks the results in their browser.

[0663] Examples of prompt statements

[0664] The following is a concrete example of a prompt statement for training an AI model to generate this system:

[0665] Image file: Contract_2023.jpeg

[0666] Extracted text: "This agreement is effective as of March 1, 2023…"

[0667] User sentiment: "Satisfied"

[0668] AI-generated output: "This agreement was concluded on March 1, 2023…"

[0669] Adjusted and supplemented result: "This agreement was entered into on March 1, 2023, and is effective as of that date…"

[0670] Image file: receipt_2023.pdf

[0671] Extracted text: "Item Quantity Price"

[0672] User sentiment: "Dissatisfied"

[0673] AI-generated completion result: "Item Quantity Price Total Amount"

[0674] Adjusted and completed result: "Item Quantity Price Total Amount Promotion Applied"

[0675] This enables highly accurate OCR processing and the provision of results that take user emotions into consideration.

[0676] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0677] Step 1:

[0678] The user selects a document image on their device and uploads it to the information processing device (server).

[0679] Input: User-selected document image (e.g., JPEG, PNG, PDF file)

[0680] Output: Document image files saved on the server

[0681] Specific actions:

[0682] The user uses a web browser on their computer to click the file selection button and choose the document image they want to upload. Then, by clicking the "Upload" button, the image file is sent to the server and saved to the server's designated storage.

[0683] Step 2:

[0684] The server performs OCR processing on the saved document images.

[0685] Input: Saved document image file

[0686] Output: Extracted text data

[0687] Specific actions:

[0688] The server uses an OCR engine (e.g., Tesseract or Google Cloud Vision API) to analyze the stored document image files. The OCR engine recognizes the characters in the image and extracts the text data. The extracted text data is stored on the server.

[0689] Step 3:

[0690] The server inputs the extracted text data and the original document image into a generative AI model for training.

[0691] Input: Extracted text data, original document image

[0692] Output: Trained model using a generative AI model

[0693] Specific actions:

[0694] The server inputs the text data obtained through OCR processing and the original document image into a generative AI model. The generative AI model learns and understands the document's patterns and structure based on this data. The trained model is then used for subsequent completion processing.

[0695] Step 4:

[0696] The server uses a generative AI model to supplement the OCR results.

[0697] Input: Pre-trained generative AI model, initial OCR text data

[0698] Output: Interpolated high-precision text data

[0699] Specific actions:

[0700] The server uses a pre-trained generative AI model to perform completion processing on the initial OCR text data. Specifically, misrecognized or unrecognized portions are identified, and these portions are completed by the generative AI model. The completed, high-accuracy text data is stored on the server.

[0701] Step 5:

[0702] The server uses an emotion engine to adjust the completion results based on the user's emotions.

[0703] Input: Completed, high-precision text data, user sentiment information (facial expressions, voice, input characters)

[0704] Output: Text data adjusted based on emotions

[0705] Specific actions:

[0706] Users provide emotional information in real time through their camera and microphone, or input emotional feedback. The server uses an emotion engine to analyze the user's emotional information, has a generative AI model make adjustments based on the emotions, and further refines the completion results.

[0707] Step 6:

[0708] The server then provides the user with the final, adjusted, high-precision text data.

[0709] Input: Text data adjusted based on emotions

[0710] Output: Final high-precision text data provided to the user.

[0711] Specific actions:

[0712] The server sends the adjusted, high-precision text data to the user. The user then downloads or views the results again via a web browser or dedicated application and performs the necessary processing.

[0713] (Application Example 2)

[0714] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0715] Conventional optical character recognition (OCR) systems have suffered from frequent misrecognition in extracting text data from document images, resulting in low accuracy. Furthermore, they often disregarded the user experience, leading to user dissatisfaction. This invention aims not only to improve the accuracy of OCR but also to enhance the user experience by considering user feelings.

[0716] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0717] In this invention, the server includes means for uploading a document image to the server, means for performing optical character recognition (OCR) processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for adjusting the supplementation results with an emotion engine that analyzes the user's emotional information, and means for providing the supplemented text data to the user. This makes it possible to improve OCR accuracy and provide text data that takes the user's emotions into consideration.

[0718] A "document image" is an image file containing text information provided by the user.

[0719] A "server" is a computer system that provides services to client terminals over a network.

[0720] Optical Character Recognition (OCR) processing is a technology that recognizes characters and numbers in an image and extracts them as text data.

[0721] "Text data" refers to data containing characters and numbers extracted from a document image.

[0722] A "generative artificial intelligence (AI) model" is an artificial intelligence that has the ability to learn from large amounts of data and generate new data.

[0723] An "emotion engine" is software that recognizes a user's emotions and adjusts the processing results based on those emotions.

[0724] "Completion" is the process of filling in omissions and errors in initial text data to make it closer to a complete form.

[0725] "Means of providing to the user" refers to the method of displaying or transmitting the final processing result to the user.

[0726] This invention provides a system that performs high-precision OCR processing on document images while also taking user emotions into consideration. It is basically implemented by the following means.

[0727] Hardware and software

[0728] Hardware: Primarily uses servers, client terminals, and smartphones.

[0729] The server is a central computer system that performs OCR processing, executes generative artificial intelligence (AI) models, and manages data, and is connected to users by applications running on client terminals and smartphones.

[0730] Software: Use the following software.

[0731] OCR engine: Uses tools such as pytesseract and Google Cloud Vision API to perform character recognition from document images.

[0732] Generative artificial intelligence models: Text data is completed using the Transformers library (e.g., GPT-3).

[0733] Emotion Engine: Analyzes user emotions using TextBlob.

[0734] Processing flow

[0735] 1. Uploading Document Images: Users upload document images (e.g., JPEG, PNG, PDF) to the server. Uploads are performed via a dedicated application on a smartphone or client device.

[0736] 2. OCR processing: The server saves the received document image and extracts the text data using an OCR engine (e.g., pytesseract). The extracted text data is saved as an intermediate result.

[0737] 3. Completion by Generative AI: The server feeds the extracted text data and the original document image into a generative AI model (e.g., GPT-3) for training. The generative AI model uses this data to complete the text and generate high-accuracy text data.

[0738] 4. Adjustment by the emotion engine: The server analyzes the user's emotion information using the emotion engine (TextBlob). If the emotion information is "satisfied," the completion result is used as is; if it is "dissatisfied," the text is further modified and completed.

[0739] 5. Provision to the user: The server ultimately provides the user with the completed and adjusted text data. The user obtains highly accurate and emotion-sensitive character recognition results through a dedicated application.

[0740] Specific example

[0741] For example, suppose a user uploads an image of a warranty certificate for a virtual store, and OCR processing is required. In this case, the user takes a picture of the warranty certificate within the virtual store app and sends it to the server. On the server, the OCR engine analyzes the image and extracts initial text data. Then, a generative AI model (e.g., GPT-3) completes this text data, and an emotion engine (e.g., TextBlob) analyzes the user's input text to recognize emotions.

[0742] For example, if a user enters "The contents of this warranty are inaccurate," the emotion engine recognizes the dissatisfaction, and the generative AI model corrects and supplements the text data again. Finally, the completed and adjusted text is returned to the user. An example of the prompt text in this case is as follows:

[0743] "The information on the warranty is inaccurate. Please see below for details:"

[0744] This allows users to obtain highly accurate character recognition results and makes it easier to manage warranty certificates.

[0745] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0746] Step 1:

[0747] The user uploads a document image.

[0748] Input: Document image (JPEG, PNG, PDF, etc.)

[0749] Specific operation: The user launches a dedicated application, takes or selects a document image, and sends it to the server. Clicking the upload button transfers the image to the server.

[0750] Output: Uploaded document image file

[0751] Step 2:

[0752] The server saves the document image and performs OCR processing.

[0753] Input: Uploaded document image file

[0754] Specific operation: The server saves the received image file and runs an OCR engine (e.g., pytesseract) to extract text data from the image. It then performs grayscale conversion, noise reduction, and character recognition on the image.

[0755] Output: Extracted text data

[0756] Step 3:

[0757] The server loads the extracted text data and the original document images into a generative AI model for training.

[0758] Input: Text data, document images

[0759] Specific operation: The server inputs text data and images into a generative artificial intelligence model (e.g., GPT-3) to train it on document patterns and structures. The model then learns to correct misrecognitions based on the input data.

[0760] Output: Generated high-precision text data

[0761] Step 4:

[0762] The server analyzes the user's emotional information and adjusts the completion results accordingly.

[0763] Input: Generated text data, user sentiment information

[0764] Specific operation: The server receives additional input from the user (e.g., "This part is inaccurate") and uses a sentiment engine (e.g., TextBlob) to analyze the user's sentiment. It then further completes or adjusts the text data based on the sentiment.

[0765] Output: Adjusted text data

[0766] Step 5:

[0767] The server provides the user with the final, completed, and adjusted text data.

[0768] Input: Adjusted text data

[0769] Specific operation: The server saves the final adjustment results as a file and sends it to the user through a dedicated application. The user can then view the text data within the application.

[0770] Output: High-precision, emotion-sensitive text data acquired by the user.

[0771] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0772] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0773] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0774] [Third Embodiment]

[0775] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0776] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0777] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0778] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0779] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0780] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0781] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0782] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0783] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0784] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0785] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0786] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0787] System Overview

[0788] This system utilizes generative artificial intelligence (AI) to perform supplementary processing with the aim of improving the accuracy of optical character recognition (OCR). Users upload their document images to the server, where OCR processing is performed. The generative AI then supplements the images, and the results are provided to the user.

[0789] Processing flow and specific actions

[0790] 1. Upload document images

[0791] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0792] 2. Execute OCR processing

[0793] The server performs OCR processing on the received document image. It uses an OCR engine to analyze the image, recognize the characters it contains, and extract text data. This process utilizes common OCR libraries and services (such as Tesseract or Google Cloud Vision API).

[0794] 3. Learning using generative AI

[0795] The server feeds the OCR results and the original document image into a generative AI for training. The generative AI receives the OCR results (text data) and the original image data as input and learns patterns and structures based on this data. This step uses advanced generative AI models such as GPT-4 or DALL-E.

[0796] 4. Completion process

[0797] The server uses generative AI to supplement text portions that are missing or misrecognized during OCR processing. Based on the information learned by the generative AI model, it generates text data to improve overall accuracy.

[0798] 5. Output of completion results

[0799] The server provides the user with the completed text data. This completed text data is returned to the user via API responses or the user interface. As a result, the user can obtain highly accurate character recognition results.

[0800] Specific example

[0801] The following scenario will be described as a specific example.

[0802] The user uploads a JPEG file of the contract scanned on their computer to the server. The server first extracts the text data from the contract using an OCR engine. At this point, some characters may be misrecognized, so the server inputs the OCR results and the original contract image into a generative AI to train the generative AI model. After the generative AI model has trained, it completes the misrecognized parts and finally generates highly accurate text data. The server provides this result to the user, allowing the user to obtain accurate digital text data of the contract. This system significantly improves the accuracy of OCR for semi-standard and non-standard documents, increasing the value of OCR.

[0803] The following describes the processing flow.

[0804] Step 1:

[0805] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0806] Step 2:

[0807] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0808] Step 3:

[0809] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0810] python

[0811] ocr_result = ocr_engine.process(image_file)

[0812] Step 4:

[0813] The server feeds the extracted text data and the original document image into a generative AI. The generative AI learns from this data to understand the patterns and structure of the document.

[0814] python

[0815] gen_ai_model.learn(ocr_result, image_file)

[0816] Step 5:

[0817] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0818] python

[0819] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0820] Step 6:

[0821] The server provides the user with the completed text data. The final result is returned to the user, allowing them to obtain accurate character recognition results.

[0822] python

[0823] send_response_to_user(enhanced_result)

[0824] Specific example

[0825] Consider a scenario where a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI to train the generative AI model (Step 4). Based on the analyzed data, the generative AI fills in any misrecognized parts and generates the final, highly accurate text data (Step 5). Finally, this completed text data is sent back to the user, who obtains the accurate text data (Step 6).

[0826] (Example 1)

[0827] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0828] Conventional optical character recognition (OCR) systems have suffered from problems such as misrecognition and low recognition accuracy. Furthermore, recognition accuracy deteriorated even further for documents with inconsistent formats or handwritten characters, making it difficult for users to obtain useful digital data. This limited their use in many applications that require accurate character data.

[0829] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0830] In this invention, the server includes means for uploading a document image to an electronic device, means for performing optical character recognition processing on the document image and extracting character data, means for feeding the extracted character data and the original document image into a generative artificial intelligence model for training, means for supplementing the character data using the generative artificial intelligence model, and means for providing the supplemented character data to the user. This improves the accuracy of OCR processing and makes it possible to provide highly accurate character data by supplementing missing or misrecognized parts.

[0831] A "document image" is a digital image that holds information in a visual format, including text and graphics.

[0832] "Electronic devices" refer to devices used for processing and communicating digital data, such as computers, smartphones, and tablets.

[0833] "Uploading" refers to the act of a user sending data from their electronic device to a server.

[0834] "Optical character recognition processing" is a technology that analyzes characters in an image and converts them into text data.

[0835] "Character data" refers to text-based information extracted through optical character recognition (OCR) processing.

[0836] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to learn data patterns and perform generation and completion.

[0837] "Learning" refers to the process by which a generative artificial intelligence model analyzes patterns and structures from input data and improves its processing capabilities based on that information.

[0838] "Complementation" refers to the process of supplementing partially missing or misidentified information with accurate data to generate complete data.

[0839] "Providing" refers to the act of returning data generated by the server as a result of processing to the user.

[0840] This section describes specific embodiments for carrying out this invention. To understand the detailed processing of the invention, the following description will specify the names of the hardware and software and show the processing flow and operation.

[0841] First, the user uploads the document image to an electronic device. Specifically, the user selects the image file (e.g., JPEG, PNG, PDF) using a web browser or a dedicated application and clicks the "Upload" button. Electronic devices used here include personal computers, smartphones, and tablets.

[0842] Next, the server receives the uploaded document image and performs optical character recognition (OCR) processing. This process uses an OCR engine (e.g., Tesseract or Google Cloud Vision API). The OCR engine analyzes the received image and extracts character data. The text data obtained at this stage is an ideal digital representation of the characters contained in the original document.

[0843] Subsequently, the server feeds the extracted character data and the original document image into a generative artificial intelligence model for training. This generative AI model uses a high-performance model (e.g., GPT-4 or DALL-E). Based on the OCR results and the original image data, the generative AI model learns the patterns and structures of the character data, enabling it to fill in inaccuracies and missing parts.

[0844] Next, the server uses a generative artificial intelligence model to supplement the missing or misrecognized character data from the OCR process. Based on the information learned by the generative AI model, it generates highly accurate character data, resulting in accurate digital text data overall.

[0845] Ultimately, the server provides the completed text data to the user through a user interface or API response. The user can view the completed text data in a web browser and download it if necessary.

[0846] As a concrete example, consider a scenario where a user uploads a JPEG file of a contract scanned on their PC to a server. The server uses an OCR engine to extract text data from the contract, but some characters may be misrecognized. Therefore, the server inputs the OCR results and the original contract image into a generative artificial intelligence model, corrects the misrecognized parts, generates final high-precision text data, and provides it to the user. This concrete example demonstrates how the present invention improves the accuracy of inaccurate character recognition results.

[0847] Examples of prompt messages are shown below.

[0848] "Based on the character recognition results after OCR processing, please complete the content using the scanned image of the contract. The original image is a scanned image of the entire contract page. Please correct any misrecognized text by replacing it with the correct characters, taking into consideration the possibility of errors, and output high-precision text data."

[0849] The above describes the embodiments for carrying out this invention.

[0850] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0851] Step 1: Upload document images

[0852] The user opens a web browser or dedicated application from their electronic device.

[0853] The user selects an image file (e.g., JPEG, PNG, PDF) and clicks the "Upload" button.

[0854] Input: Document image selected by the user.

[0855] Output: Document image file sent to the server.

[0856] Specific operation: The user selects a JPEG file of the contract on their computer and clicks the "Upload" button in the web browser. The image file is sent to the server via an HTTP POST request.

[0857] Step 2: Execute OCR processing

[0858] The server receives the uploaded document image.

[0859] The server calls an OCR engine (such as Tesseract or Google Cloud Vision API) to analyze the image and extract text data.

[0860] Input: Document image file stored on the server.

[0861] Output: Character data generated by OCR processing.

[0862] Specific operation: The server receives a JPEG file, inputs it into Tesseract, and extracts the text from the image as text data. For example, the word "contract" will be recognized as "contract," but some characters may be recognized as "misspellings."

[0863] Step 3: Learning with Generative AI

[0864] The server inputs the OCR result text data and the original document image into a generative artificial intelligence model.

[0865] The server trains a generative AI (such as GPT-4 or DALL-E).

[0866] Input: OCR result text data and original document image.

[0867] Output: Training results from a generative artificial intelligence model.

[0868] Specific operation: The server inputs the OCR result text data and the original contract image into the GPT-4 model. The generative AI analyzes the OCR result text data and the original image to learn relevant patterns and context.

[0869] Step 4: Completion process

[0870] The server uses generative AI to identify missing or misrecognized character data during OCR processing.

[0871] The server uses generative AI to fill in any gaps or misrecognitions.

[0872] Input: Training results from a generative artificial intelligence model.

[0873] Output: High-precision text data generated by a generative artificial intelligence model.

[0874] Specific operation: The generative AI model reviews the OCR results and corrects the misrecognized character parts using an algorithm for error correction. Specifically, "misspellings" are corrected to "correct characters."

[0875] Step 5: Output of completion results

[0876] The server generates the completed text data.

[0877] The server returns the completed text data to the user as a user interface or API response.

[0878] Input: High-accuracy text data augmented by a generative AI model.

[0879] Output: The final, high-precision text data provided to the user.

[0880] Specific operation: The server generates the final, high-precision text data and sends it back to the user via the user interface or API response. The user can then view and download the completed text data in their browser.

[0881] (Application Example 1)

[0882] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0883] Conventional optical character recognition (OCR) systems have problems with the accuracy of text data extracted from document images, resulting in high misrecognition rates, especially with blurry images and documents in different formats. Furthermore, while high-precision data entry is required in on-site settings such as logistics centers, manual data entry is time-consuming and labor-intensive, and prone to errors. In addition, existing systems lack sufficient post-image analysis processing, making efficient operation difficult.

[0884] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0885] In this invention, the server includes means for uploading document images to the server, means for performing optical character recognition (OCR) processing on the document images and extracting text data, means for feeding the extracted text data and the original document images into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for providing the supplemented text data to the user, and means for performing OCR processing and generative AI supplementation processing on document images input via scanning with a smartphone at a logistics center. This enables highly accurate text data extraction and supplementation of document images at a logistics center.

[0886] A "document image" is data that saves the contents of a paper or electronic document in image format.

[0887] A "server" is a computer system that provides information and services to clients via a network in response to their requests.

[0888] Optical Character Recognition (OCR) processing is a process that analyzes characters from images such as scans and photographs and extracts them as digital text data.

[0889] "Text data" refers to data stored digitally as string information.

[0890] A "generative artificial intelligence (AI) model" is an artificial intelligence system that learns patterns and structures based on large amounts of data, and then generates or completes new data.

[0891] "Methods for training" refer to methods of inputting specific data into a generative artificial intelligence model and training that model to understand the patterns and structures of the data.

[0892] "Means of supplementation" refer to methods of generating accurate data by supplementing or correcting insufficient or misidentified data.

[0893] "Means of providing text data to users" refers to methods of making processed text data publicly available and accessible to users.

[0894] A "logistics center" is a facility used for storing, distributing, and receiving goods.

[0895] A "smartphone" is a portable information terminal that, in addition to making phone calls, is capable of internet communication and operating applications.

[0896] "Scanning" is the act of digitizing an image into electronic data.

[0897] The system for realizing this invention involves uploading document images to a server using a smartphone at a logistics center, and then performing optical character recognition (OCR) processing and supplementary processing using a generative artificial intelligence (AI) model on those images.

[0898] System Configuration

[0899] hardware

[0900] Server: A server with high-performance computing capabilities. This server performs OCR processing and auto-completion processing.

[0901] Smartphone: A portable information terminal that includes a camera function. Employees at the logistics center use it to scan document images.

[0902] software

[0903] OCR engine: Software used to perform character recognition, such as Tesseract or Google Cloud Vision API.

[0904] Generative AI models: Advanced generative AI models such as GPT-4. Used for interpolation processing.

[0905] Smartphone application: A dedicated application for users to scan images and upload them to a server.

[0906] Processing flow

[0907] 1. Upload document images

[0908] Users scan document images (such as shipping labels or delivery slips) using their smartphones and upload them to the server. The process is easily managed through a dedicated application.

[0909] 2. Execute OCR processing

[0910] The server performs OCR processing on uploaded document images. It uses an OCR engine to analyze the images and extract text data. Typically, this is done using Tesseract or the Google Cloud Vision API.

[0911] 3. Completion processing by generative AI

[0912] The server feeds the OCR results and the original document image into a generative AI model for training. A model such as GPT-4 is used for this training. Based on the patterns and structures learned by the generative AI model, it compensates for misrecognized parts of the OCR results to generate highly accurate text data.

[0913] 4. Providing completion results

[0914] The server provides the augmented text data to users via a smartphone application. This allows employees at the logistics center to proceed with their work based on highly accurate data.

[0915] Specific example

[0916] Logistics center employees use the "SmartLogistics OCR" app on their smartphones to scan shipping labels. The scanned images are uploaded to a server where OCR processing is performed. Because OCR results may contain misrecognitions, the server uses a generative AI model based on GPT-4 to fill in any missing information and generate accurate text data. This text data is then provided to employees through the app.

[0917] Example of a prompt

[0918] "The OCR results are below. Please make corrections."

[0919] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0920] 2. Date: October 15, 2023

[0921] 3. Contents: 10 laptop computers

[0922] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0923] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0924] Step 1:

[0925] Users scan document images (e.g., shipping labels and delivery slips) within the logistics center using a dedicated smartphone application. These scanned image files (e.g., in JPEG format) are uploaded to a server via the dedicated application. The input is image files from the smartphone, and the output is the document images stored on the server.

[0926] Step 2:

[0927] The server performs optical character recognition (OCR) processing on uploaded document images. The server launches an OCR engine (such as Tesseract or Google Cloud Vision API), analyzes the image data, and extracts text data from it. The input is a document image, and the output is text data. However, misrecognition may occur due to unclear areas or unusual formats.

[0928] Step 3:

[0929] The server feeds the extracted text data and the original document image into a generative artificial intelligence (AI) model for training. For example, an advanced generative AI model such as GPT-4 is used. The input consists of the OCR result text data and the original document image, and the output is the patterns and structures that the AI ​​model has learned. This process generates data to supplement misrecognitions and missing parts of the OCR result.

[0930] Step 4:

[0931] The server performs text data completion processing using a generative AI model. Based on learned patterns and structures, the AI ​​model corrects misrecognized parts of the OCR results and generates accurate text data. The input consists of training data and OCR results, and the output is highly accurate, completed text data.

[0932] Step 5:

[0933] The server provides the user with completed text data. This completed text data is returned to the user via API responses or the user interface of a smartphone application. The input is completed text data, and the output is highly accurate text data accessible to the user. The user can then continue their work based on this accurate text data.

[0934] Specific example

[0935] Example of a prompt:

[0936] "The OCR results are below. Please make corrections."

[0937] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[0938] 2. Date: October 15, 2023

[0939] 3. Contents: 10 laptop computers

[0940] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[0941] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0942] System Overview

[0943] This system aims to improve the accuracy of optical character recognition (OCR) by combining generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions. Users upload document images to the server, where OCR processing is performed. The system then uses generative AI to complete the images, and finally, the emotion engine adjusts the completion results based on the user's emotions before providing them to the user.

[0944] Processing flow and specific actions

[0945] Upload document images

[0946] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0947] Execution of OCR processing

[0948] The server saves the received document image and then performs OCR processing. The OCR engine (for example, Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[0949] Learning by generative AI

[0950] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand document patterns and structures.

[0951] Completion process

[0952] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0953] Adjustment by the emotion engine

[0954] The server adjusts the completed text data based on the user's emotions. The emotion engine recognizes emotions from the user's facial expressions, voice, or input text, and the generative AI model adjusts the completion results according to that emotional information.

[0955] Output of completion results

[0956] The server provides the user with adjusted, supplementary text data. The final result is returned to the user, allowing them to obtain accurate and emotionally sensitive character recognition results.

[0957] Specific example

[0958] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer. The server receives it and saves the image. Next, the server analyzes the image using an OCR engine and extracts the initial text data. Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model. Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data.

[0959] The emotion engine analyzes the generated text data to identify information (facial expressions, voice, input characters, etc.) necessary to recognize the user's emotions. For example, if the user is satisfied, the basic completion result is provided as is; if the user is dissatisfied, the generative AI adjusts the result to provide an even more accurate one. Finally, this completed, high-accuracy text data is provided to the user, allowing them to obtain highly accurate and emotion-sensitive character recognition results.

[0960] This invention not only improves the accuracy of OCR but also enables the provision of services that take user emotions into consideration, thereby enhancing the user experience and significantly expanding the value and scale of deployment of OCR systems.

[0961] The following describes the processing flow.

[0962] Step 1:

[0963] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[0964] Step 2:

[0965] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[0966] Step 3:

[0967] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the initial text data.

[0968] python

[0969] ocr_result = ocr_engine.process(image_file)

[0970] Step 4:

[0971] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand the patterns and structure of documents.

[0972] python

[0973] gen_ai_model.learn(ocr_result, image_file)

[0974] Step 5:

[0975] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[0976] python

[0977] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[0978] Step 6:

[0979] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expressions, voice, or input text to determine whether the user is satisfied, irritated, anxious, etc.

[0980] python

[0981] user_emotion = emotion_engine.analyze(user_input)

[0982] Step 7:

[0983] The server uses a generative AI model to adjust the supplementary text data based on the user's emotions. For example, if the user expresses dissatisfaction, the generative AI adjusts the data to produce more accurate supplementary results.

[0984] python

[0985] adjusted_result = gen_ai_model.adjust_based_on_emotion(enhanced_result, user_emotion)

[0986] Step 8:

[0987] The server provides the user with adjusted, supplementary text data. The final result is returned to the user via an API response or user interface, allowing the user to obtain accurate and emotionally sensitive character recognition results.

[0988] python

[0989] send_response_to_user(adjusted_result)

[0990] Specific example

[0991] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts the initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model (Step 4). Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data (Step 5).

[0992] For this generated text data, the emotion engine analyzes the user's facial expressions, voice, or input text to recognize the user's emotions (Step 6). For example, if the user is expressing dissatisfaction, the generative AI adjusts the results to produce even more accurate results (Step 7). Finally, this adjusted supplementary text data is provided to the user, allowing them to obtain accurate and emotion-sensitive character recognition results (Step 8). This improves the overall user experience by enabling flexible responses that respond to the user's emotions, in addition to improving the accuracy of OCR.

[0993] (Example 2)

[0994] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0995] Optical character recognition (OCR) technology is used in many fields, but it has challenges in recognition accuracy, particularly with misrecognition and missed recognition. Furthermore, recognition results can affect user satisfaction, and there is a need for methods that reflect user emotions regarding the recognition results. Therefore, high-precision OCR processing and the provision of results that take user emotions into consideration are necessary.

[0996] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0997] In this invention, the server includes means for uploading a document image to an information processing device, means for performing optical character recognition processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence model for training, means for supplementing the text data using the generative artificial intelligence model, means for adjusting the supplemented text data based on the user's emotions, and means for providing the adjusted text data to the user. This enables highly accurate OCR processing and the provision of results that take into account the user's emotions.

[0998] A "document image" is image data that is the target of optical character recognition processing, and is generally an image format in which text is embedded (e.g., JPEG, PNG, PDF).

[0999] "Information processing equipment" refers to all devices that perform data input, processing, storage, output, etc., and in this context mainly includes servers and client terminals.

[1000] Optical Character Recognition (OCR) processing is a technology that analyzes characters contained in an image and extracts them as text data.

[1001] "Text data" refers to data representing character information extracted through optical character recognition processing.

[1002] A "generative artificial intelligence model" is an artificial intelligence model that performs data completion or generation based on training data, and in this context, it is used to complete misrecognized or unrecognized parts of OCR results.

[1003] "User" refers to an individual or organization that utilizes optical character recognition (OCR) processing services.

[1004] "Means of adjusting based on emotions" refers to a function that analyzes the user's emotional information and has a generative artificial intelligence model make adjustments in accordance with those emotions.

[1005] "Means of complementarity" refers to a function that complements OCR results based on data learned by a generative artificial intelligence model, thereby generating highly accurate text data.

[1006] This invention relates to a system that combines generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions, with the aim of improving the accuracy of optical character recognition (OCR). Specific embodiments of this system are described below.

[1007] System Overview

[1008] The user uploads a document image to an information processing device (server), where OCR processing is performed. This is then supplemented by a generative AI model, and finally, an emotion engine is used to adjust the supplemented result based on the user's emotions before providing it. Specifically, the process proceeds in the following steps:

[1009] Hardware and software

[1010] This system uses the following hardware and software:

[1011] 1. User's device:

[1012] Web browser or dedicated application

[1013] 2. Server:

[1014] Data storage function

[1015] OCR engine (e.g., Tesseract, Google Cloud Vision API)

[1016] Generative AI models

[1017] Emotional Engine

[1018] Data processing and data calculation

[1019] 1. The user uploads a document image from their device to the server.

[1020] 2. The server saves the received document image and extracts the text data using an OCR engine.

[1021] 3. The server inputs the extracted text data and the original document image into a generative AI model for training.

[1022] 4. Generative AI models are used to supplement the OCR results and generate high-accuracy text data.

[1023] 5. The emotion engine analyzes the user's emotional information, and the generative AI model adjusts the completion results based on that information.

[1024] 6. The server provides users with pre-calibrated, high-precision text data.

[1025] Specific example

[1026] Next, we will show a specific example of a user uploading a JPEG file of a contract to the server.

[1027] 1. Upload document images:

[1028] The user accesses the upload page for the JPEG file of the contract using a browser on their computer and clicks the "Upload" button. The image file is then sent to the server.

[1029] 2. Perform OCR processing:

[1030] The server saves the received image file and then performs OCR processing using Tesseract or the Google Cloud Vision API to extract the initial text data.

[1031] 3. Learning using generative AI:

[1032] The server inputs the extracted text data and the original document image into a generative AI program, which then learns to understand the patterns and structure of the document.

[1033] 4. Completion process:

[1034] The generative AI model fills in the gaps in misrecognition and unrecognized parts, generating highly accurate text data.

[1035] 5. Adjustment by the Emotion Engine:

[1036] If a user is dissatisfied, the server will further adjust the completion results based on the user's emotional information (e.g., feedback from input).

[1037] 6. Output of completion results:

[1038] The server provides the user with adjusted, high-precision text data, and the user checks the results in their browser.

[1039] Examples of prompt statements

[1040] The following is a concrete example of a prompt statement for training an AI model to generate this system:

[1041] Image file: Contract_2023.jpeg

[1042] Extracted text: "This agreement is effective as of March 1, 2023…"

[1043] User sentiment: "Satisfied"

[1044] AI-generated output: "This agreement was concluded on March 1, 2023…"

[1045] Adjusted and supplemented result: "This agreement was entered into on March 1, 2023, and is effective as of that date…"

[1046] Image file: receipt_2023.pdf

[1047] Extracted text: "Item Quantity Price"

[1048] User sentiment: "Dissatisfied"

[1049] AI-generated completion result: "Item Quantity Price Total Amount"

[1050] Adjusted and completed result: "Item Quantity Price Total Amount Promotion Applied"

[1051] This enables highly accurate OCR processing and the provision of results that take user emotions into consideration.

[1052] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1053] Step 1:

[1054] The user selects a document image on their device and uploads it to the information processing device (server).

[1055] Input: User-selected document image (e.g., JPEG, PNG, PDF file)

[1056] Output: Document image files saved on the server

[1057] Specific actions:

[1058] The user uses a web browser on their computer to click the file selection button and choose the document image they want to upload. Then, by clicking the "Upload" button, the image file is sent to the server and saved to the server's designated storage.

[1059] Step 2:

[1060] The server performs OCR processing on the saved document images.

[1061] Input: Saved document image file

[1062] Output: Extracted text data

[1063] Specific actions:

[1064] The server uses an OCR engine (e.g., Tesseract or Google Cloud Vision API) to analyze the stored document image files. The OCR engine recognizes the characters in the image and extracts the text data. The extracted text data is stored on the server.

[1065] Step 3:

[1066] The server inputs the extracted text data and the original document image into a generative AI model for training.

[1067] Input: Extracted text data, original document image

[1068] Output: Trained model using a generative AI model

[1069] Specific actions:

[1070] The server inputs the text data obtained through OCR processing and the original document image into a generative AI model. The generative AI model learns and understands the document's patterns and structure based on this data. The trained model is then used for subsequent completion processing.

[1071] Step 4:

[1072] The server uses a generative AI model to supplement the OCR results.

[1073] Input: Pre-trained generative AI model, initial OCR text data

[1074] Output: Interpolated high-precision text data

[1075] Specific actions:

[1076] The server uses a pre-trained generative AI model to perform completion processing on the initial OCR text data. Specifically, misrecognized or unrecognized portions are identified, and these portions are completed by the generative AI model. The completed, high-accuracy text data is stored on the server.

[1077] Step 5:

[1078] The server uses an emotion engine to adjust the completion results based on the user's emotions.

[1079] Input: Completed, high-precision text data, user sentiment information (facial expressions, voice, input characters)

[1080] Output: Text data adjusted based on emotions

[1081] Specific actions:

[1082] Users provide emotional information in real time through their camera and microphone, or input emotional feedback. The server uses an emotion engine to analyze the user's emotional information, has a generative AI model make adjustments based on the emotions, and further refines the completion results.

[1083] Step 6:

[1084] The server then provides the user with the final, adjusted, high-precision text data.

[1085] Input: Text data adjusted based on emotions

[1086] Output: Final high-precision text data provided to the user.

[1087] Specific actions:

[1088] The server sends the adjusted, high-precision text data to the user. The user then downloads or views the results again via a web browser or dedicated application and performs the necessary processing.

[1089] (Application Example 2)

[1090] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1091] Conventional optical character recognition (OCR) systems have suffered from frequent misrecognition in extracting text data from document images, resulting in low accuracy. Furthermore, they often disregarded the user experience, leading to user dissatisfaction. This invention aims not only to improve the accuracy of OCR but also to enhance the user experience by considering user feelings.

[1092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1093] In this invention, the server includes means for uploading a document image to the server, means for performing optical character recognition (OCR) processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for adjusting the supplementation results with an emotion engine that analyzes the user's emotional information, and means for providing the supplemented text data to the user. This makes it possible to improve OCR accuracy and provide text data that takes the user's emotions into consideration.

[1094] A "document image" is an image file containing text information provided by the user.

[1095] A "server" is a computer system that provides services to client terminals over a network.

[1096] Optical Character Recognition (OCR) processing is a technology that recognizes characters and numbers in an image and extracts them as text data.

[1097] "Text data" refers to data containing characters and numbers extracted from a document image.

[1098] A "generative artificial intelligence (AI) model" is an artificial intelligence that has the ability to learn from large amounts of data and generate new data.

[1099] An "emotion engine" is software that recognizes a user's emotions and adjusts the processing results based on those emotions.

[1100] "Completion" is the process of filling in omissions and errors in initial text data to make it closer to a complete form.

[1101] "Means of providing to the user" refers to the method of displaying or transmitting the final processing result to the user.

[1102] This invention provides a system that performs high-precision OCR processing on document images while also taking user emotions into consideration. It is basically implemented by the following means.

[1103] Hardware and software

[1104] Hardware: Primarily uses servers, client terminals, and smartphones.

[1105] The server is a central computer system that performs OCR processing, executes generative artificial intelligence (AI) models, and manages data, and is connected to users by applications running on client terminals and smartphones.

[1106] Software: Use the following software.

[1107] OCR engine: Uses tools such as pytesseract and Google Cloud Vision API to perform character recognition from document images.

[1108] Generative artificial intelligence models: Text data is completed using the Transformers library (e.g., GPT-3).

[1109] Emotion Engine: Analyzes user emotions using TextBlob.

[1110] Processing flow

[1111] 1. Uploading Document Images: Users upload document images (e.g., JPEG, PNG, PDF) to the server. Uploads are performed via a dedicated application on a smartphone or client device.

[1112] 2. OCR processing: The server saves the received document image and extracts the text data using an OCR engine (e.g., pytesseract). The extracted text data is saved as an intermediate result.

[1113] 3. Completion by Generative AI: The server feeds the extracted text data and the original document image into a generative AI model (e.g., GPT-3) for training. The generative AI model uses this data to complete the text and generate high-accuracy text data.

[1114] 4. Adjustment by the emotion engine: The server analyzes the user's emotion information using the emotion engine (TextBlob). If the emotion information is "satisfied," the completion result is used as is; if it is "dissatisfied," the text is further modified and completed.

[1115] 5. Provision to the user: The server ultimately provides the user with the completed and adjusted text data. The user obtains highly accurate and emotion-sensitive character recognition results through a dedicated application.

[1116] Specific example

[1117] For example, suppose a user uploads an image of a warranty certificate for a virtual store, and OCR processing is required. In this case, the user takes a picture of the warranty certificate within the virtual store app and sends it to the server. On the server, the OCR engine analyzes the image and extracts initial text data. Then, a generative AI model (e.g., GPT-3) completes this text data, and an emotion engine (e.g., TextBlob) analyzes the user's input text to recognize emotions.

[1118] For example, if a user enters "The contents of this warranty are inaccurate," the emotion engine recognizes the dissatisfaction, and the generative AI model corrects and supplements the text data again. Finally, the completed and adjusted text is returned to the user. An example of the prompt text in this case is as follows:

[1119] "The information on the warranty is inaccurate. Please see below for details:"

[1120] This allows users to obtain highly accurate character recognition results and makes it easier to manage warranty certificates.

[1121] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1122] Step 1:

[1123] The user uploads a document image.

[1124] Input: Document image (JPEG, PNG, PDF, etc.)

[1125] Specific operation: The user launches a dedicated application, takes or selects a document image, and sends it to the server. Clicking the upload button transfers the image to the server.

[1126] Output: Uploaded document image file

[1127] Step 2:

[1128] The server saves the document image and performs OCR processing.

[1129] Input: Uploaded document image file

[1130] Specific operation: The server saves the received image file and runs an OCR engine (e.g., pytesseract) to extract text data from the image. It then performs grayscale conversion, noise reduction, and character recognition on the image.

[1131] Output: Extracted text data

[1132] Step 3:

[1133] The server loads the extracted text data and the original document images into a generative AI model for training.

[1134] Input: Text data, document images

[1135] Specific operation: The server inputs text data and images into a generative artificial intelligence model (e.g., GPT-3) to train it on document patterns and structures. The model then learns to correct misrecognitions based on the input data.

[1136] Output: Generated high-precision text data

[1137] Step 4:

[1138] The server analyzes the user's emotional information and adjusts the completion results accordingly.

[1139] Input: Generated text data, user sentiment information

[1140] Specific operation: The server receives additional input from the user (e.g., "This part is inaccurate") and uses a sentiment engine (e.g., TextBlob) to analyze the user's sentiment. It then further completes or adjusts the text data based on the sentiment.

[1141] Output: Adjusted text data

[1142] Step 5:

[1143] The server provides the user with the final, completed, and adjusted text data.

[1144] Input: Adjusted text data

[1145] Specific operation: The server saves the final adjustment results as a file and sends it to the user through a dedicated application. The user can then view the text data within the application.

[1146] Output: High-precision, emotion-sensitive text data acquired by the user.

[1147] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1148] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1149] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1150] [Fourth Embodiment]

[1151] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1152] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1153] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1154] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1155] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1156] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1157] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1158] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1159] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1160] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1161] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1162] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1163] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1164] System Overview

[1165] This system utilizes generative artificial intelligence (AI) to perform supplementary processing with the aim of improving the accuracy of optical character recognition (OCR). Users upload their document images to the server, where OCR processing is performed. The generative AI then supplements the images, and the results are provided to the user.

[1166] Processing flow and specific actions

[1167] 1. Upload document images

[1168] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[1169] 2. Execute OCR processing

[1170] The server performs OCR processing on the received document image. It uses an OCR engine to analyze the image, recognize the characters it contains, and extract text data. This process utilizes common OCR libraries and services (such as Tesseract or Google Cloud Vision API).

[1171] 3. Learning using generative AI

[1172] The server feeds the OCR results and the original document image into a generative AI for training. The generative AI receives the OCR results (text data) and the original image data as input and learns patterns and structures based on this data. This step uses advanced generative AI models such as GPT-4 or DALL-E.

[1173] 4. Completion process

[1174] The server uses generative AI to supplement text portions that are missing or misrecognized during OCR processing. Based on the information learned by the generative AI model, it generates text data to improve overall accuracy.

[1175] 5. Output of completion results

[1176] The server provides the user with the completed text data. This completed text data is returned to the user via API responses or the user interface. As a result, the user can obtain highly accurate character recognition results.

[1177] Specific example

[1178] The following scenario will be described as a specific example.

[1179] The user uploads a JPEG file of the contract scanned on their computer to the server. The server first extracts the text data from the contract using an OCR engine. At this point, some characters may be misrecognized, so the server inputs the OCR results and the original contract image into a generative AI to train the generative AI model. After the generative AI model has trained, it completes the misrecognized parts and finally generates highly accurate text data. The server provides this result to the user, allowing the user to obtain accurate digital text data of the contract. This system significantly improves the accuracy of OCR for semi-standard and non-standard documents, increasing the value of OCR.

[1180] The following describes the processing flow.

[1181] Step 1:

[1182] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[1183] Step 2:

[1184] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[1185] Step 3:

[1186] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[1187] python

[1188] ocr_result = ocr_engine.process(image_file)

[1189] Step 4:

[1190] The server feeds the extracted text data and the original document image into a generative AI. The generative AI learns from this data to understand the patterns and structure of the document.

[1191] python

[1192] gen_ai_model.learn(ocr_result, image_file)

[1193] Step 5:

[1194] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[1195] python

[1196] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[1197] Step 6:

[1198] The server provides the user with the completed text data. The final result is returned to the user, allowing them to obtain accurate character recognition results.

[1199] python

[1200] send_response_to_user(enhanced_result)

[1201] Specific example

[1202] Consider a scenario where a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI to train the generative AI model (Step 4). Based on the analyzed data, the generative AI fills in any misrecognized parts and generates the final, highly accurate text data (Step 5). Finally, this completed text data is sent back to the user, who obtains the accurate text data (Step 6).

[1203] (Example 1)

[1204] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1205] Conventional optical character recognition (OCR) systems have suffered from problems such as misrecognition and low recognition accuracy. Furthermore, recognition accuracy deteriorated even further for documents with inconsistent formats or handwritten characters, making it difficult for users to obtain useful digital data. This limited their use in many applications that require accurate character data.

[1206] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1207] In this invention, the server includes means for uploading a document image to an electronic device, means for performing optical character recognition processing on the document image and extracting character data, means for feeding the extracted character data and the original document image into a generative artificial intelligence model for training, means for supplementing the character data using the generative artificial intelligence model, and means for providing the supplemented character data to the user. This improves the accuracy of OCR processing and makes it possible to provide highly accurate character data by supplementing missing or misrecognized parts.

[1208] A "document image" is a digital image that holds information in a visual format, including text and graphics.

[1209] "Electronic devices" refer to devices used for processing and communicating digital data, such as computers, smartphones, and tablets.

[1210] "Uploading" refers to the act of a user sending data from their electronic device to a server.

[1211] "Optical character recognition processing" is a technology that analyzes characters in an image and converts them into text data.

[1212] "Character data" refers to text-based information extracted through optical character recognition (OCR) processing.

[1213] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to learn data patterns and perform generation and completion.

[1214] "Learning" refers to the process by which a generative artificial intelligence model analyzes patterns and structures from input data and improves its processing capabilities based on that information.

[1215] "Complementation" refers to the process of supplementing partially missing or misidentified information with accurate data to generate complete data.

[1216] "Providing" refers to the act of returning data generated by the server as a result of processing to the user.

[1217] This section describes specific embodiments for carrying out this invention. To understand the detailed processing of the invention, the following description will specify the names of the hardware and software and show the processing flow and operation.

[1218] First, the user uploads the document image to an electronic device. Specifically, the user selects the image file (e.g., JPEG, PNG, PDF) using a web browser or a dedicated application and clicks the "Upload" button. Electronic devices used here include personal computers, smartphones, and tablets.

[1219] Next, the server receives the uploaded document image and performs optical character recognition (OCR) processing. This process uses an OCR engine (e.g., Tesseract or Google Cloud Vision API). The OCR engine analyzes the received image and extracts character data. The text data obtained at this stage is an ideal digital representation of the characters contained in the original document.

[1220] Subsequently, the server feeds the extracted character data and the original document image into a generative artificial intelligence model for training. This generative AI model uses a high-performance model (e.g., GPT-4 or DALL-E). Based on the OCR results and the original image data, the generative AI model learns the patterns and structures of the character data, enabling it to fill in inaccuracies and missing parts.

[1221] Next, the server uses a generative artificial intelligence model to supplement the missing or misrecognized character data from the OCR process. Based on the information learned by the generative AI model, it generates highly accurate character data, resulting in accurate digital text data overall.

[1222] Ultimately, the server provides the completed text data to the user through a user interface or API response. The user can view the completed text data in a web browser and download it if necessary.

[1223] As a concrete example, consider a scenario where a user uploads a JPEG file of a contract scanned on their PC to a server. The server uses an OCR engine to extract text data from the contract, but some characters may be misrecognized. Therefore, the server inputs the OCR results and the original contract image into a generative artificial intelligence model, corrects the misrecognized parts, generates final high-precision text data, and provides it to the user. This concrete example demonstrates how the present invention improves the accuracy of inaccurate character recognition results.

[1224] Examples of prompt messages are shown below.

[1225] "Based on the character recognition results after OCR processing, please complete the content using the scanned image of the contract. The original image is a scanned image of the entire contract page. Please correct any misrecognized text by replacing it with the correct characters, taking into consideration the possibility of errors, and output high-precision text data."

[1226] The above describes the embodiments for carrying out this invention.

[1227] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1228] Step 1: Upload document images

[1229] The user opens a web browser or dedicated application from their electronic device.

[1230] The user selects an image file (e.g., JPEG, PNG, PDF) and clicks the "Upload" button.

[1231] Input: Document image selected by the user.

[1232] Output: Document image file sent to the server.

[1233] Specific operation: The user selects a JPEG file of the contract on their computer and clicks the "Upload" button in the web browser. The image file is sent to the server via an HTTP POST request.

[1234] Step 2: Execute OCR processing

[1235] The server receives the uploaded document image.

[1236] The server calls an OCR engine (such as Tesseract or Google Cloud Vision API) to analyze the image and extract text data.

[1237] Input: Document image file stored on the server.

[1238] Output: Character data generated by OCR processing.

[1239] Specific operation: The server receives a JPEG file, inputs it into Tesseract, and extracts the text from the image as text data. For example, the word "contract" will be recognized as "contract," but some characters may be recognized as "misspellings."

[1240] Step 3: Learning with Generative AI

[1241] The server inputs the OCR result text data and the original document image into a generative artificial intelligence model.

[1242] The server trains a generative AI (such as GPT-4 or DALL-E).

[1243] Input: OCR result text data and original document image.

[1244] Output: Training results from a generative artificial intelligence model.

[1245] Specific operation: The server inputs the OCR result text data and the original contract image into the GPT-4 model. The generative AI analyzes the OCR result text data and the original image to learn relevant patterns and context.

[1246] Step 4: Completion process

[1247] The server uses generative AI to identify missing or misrecognized character data during OCR processing.

[1248] The server uses generative AI to fill in any gaps or misrecognitions.

[1249] Input: Training results from a generative artificial intelligence model.

[1250] Output: High-precision text data generated by a generative artificial intelligence model.

[1251] Specific operation: The generative AI model reviews the OCR results and corrects the misrecognized character parts using an algorithm for error correction. Specifically, "misspellings" are corrected to "correct characters."

[1252] Step 5: Output of completion results

[1253] The server generates the completed text data.

[1254] The server returns the completed text data to the user as a user interface or API response.

[1255] Input: High-accuracy text data augmented by a generative AI model.

[1256] Output: The final, high-precision text data provided to the user.

[1257] Specific operation: The server generates the final, high-precision text data and sends it back to the user via the user interface or API response. The user can then view and download the completed text data in their browser.

[1258] (Application Example 1)

[1259] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1260] Conventional optical character recognition (OCR) systems have problems with the accuracy of text data extracted from document images, resulting in high misrecognition rates, especially with blurry images and documents in different formats. Furthermore, while high-precision data entry is required in on-site settings such as logistics centers, manual data entry is time-consuming and labor-intensive, and prone to errors. In addition, existing systems lack sufficient post-image analysis processing, making efficient operation difficult.

[1261] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1262] In this invention, the server includes means for uploading document images to the server, means for performing optical character recognition (OCR) processing on the document images and extracting text data, means for feeding the extracted text data and the original document images into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for providing the supplemented text data to the user, and means for performing OCR processing and generative AI supplementation processing on document images input via scanning with a smartphone at a logistics center. This enables highly accurate text data extraction and supplementation of document images at a logistics center.

[1263] A "document image" is data that saves the contents of a paper or electronic document in image format.

[1264] A "server" is a computer system that provides information and services to clients via a network in response to their requests.

[1265] Optical Character Recognition (OCR) processing is a process that analyzes characters from images such as scans and photographs and extracts them as digital text data.

[1266] "Text data" refers to data stored digitally as string information.

[1267] A "generative artificial intelligence (AI) model" is an artificial intelligence system that learns patterns and structures based on large amounts of data, and then generates or completes new data.

[1268] "Methods for training" refer to methods of inputting specific data into a generative artificial intelligence model and training that model to understand the patterns and structures of the data.

[1269] "Means of supplementation" refer to methods of generating accurate data by supplementing or correcting insufficient or misidentified data.

[1270] "Means of providing text data to users" refers to methods of making processed text data publicly available and accessible to users.

[1271] A "logistics center" is a facility used for storing, distributing, and receiving goods.

[1272] A "smartphone" is a portable information terminal that, in addition to making phone calls, is capable of internet communication and operating applications.

[1273] "Scanning" is the act of digitizing an image into electronic data.

[1274] The system for realizing this invention involves uploading document images to a server using a smartphone at a logistics center, and then performing optical character recognition (OCR) processing and supplementary processing using a generative artificial intelligence (AI) model on those images.

[1275] System Configuration

[1276] hardware

[1277] Server: A server with high-performance computing capabilities. This server performs OCR processing and auto-completion processing.

[1278] Smartphone: A portable information terminal that includes a camera function. Employees at the logistics center use it to scan document images.

[1279] software

[1280] OCR engine: Software used to perform character recognition, such as Tesseract or Google Cloud Vision API.

[1281] Generative AI models: Advanced generative AI models such as GPT-4. Used for interpolation processing.

[1282] Smartphone application: A dedicated application for users to scan images and upload them to a server.

[1283] Processing flow

[1284] 1. Upload document images

[1285] Users scan document images (such as shipping labels or delivery slips) using their smartphones and upload them to the server. The process is easily managed through a dedicated application.

[1286] 2. Execute OCR processing

[1287] The server performs OCR processing on uploaded document images. It uses an OCR engine to analyze the images and extract text data. Typically, this is done using Tesseract or the Google Cloud Vision API.

[1288] 3. Completion processing by generative AI

[1289] The server feeds the OCR results and the original document image into a generative AI model for training. A model such as GPT-4 is used for this training. Based on the patterns and structures learned by the generative AI model, it compensates for misrecognized parts of the OCR results to generate highly accurate text data.

[1290] 4. Providing completion results

[1291] The server provides the augmented text data to users via a smartphone application. This allows employees at the logistics center to proceed with their work based on highly accurate data.

[1292] Specific example

[1293] Logistics center employees use the "SmartLogistics OCR" app on their smartphones to scan shipping labels. The scanned images are uploaded to a server where OCR processing is performed. Because OCR results may contain misrecognitions, the server uses a generative AI model based on GPT-4 to fill in any missing information and generate accurate text data. This text data is then provided to employees through the app.

[1294] Example of a prompt

[1295] "The OCR results are below. Please make corrections."

[1296] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[1297] 2. Date: October 15, 2023

[1298] 3. Contents: 10 laptop computers

[1299] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[1300] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1301] Step 1:

[1302] Users scan document images (e.g., shipping labels and delivery slips) within the logistics center using a dedicated smartphone application. These scanned image files (e.g., in JPEG format) are uploaded to a server via the dedicated application. The input is image files from the smartphone, and the output is the document images stored on the server.

[1303] Step 2:

[1304] The server performs optical character recognition (OCR) processing on uploaded document images. The server launches an OCR engine (such as Tesseract or Google Cloud Vision API), analyzes the image data, and extracts text data from it. The input is a document image, and the output is text data. However, misrecognition may occur due to unclear areas or unusual formats.

[1305] Step 3:

[1306] The server feeds the extracted text data and the original document image into a generative artificial intelligence (AI) model for training. For example, an advanced generative AI model such as GPT-4 is used. The input consists of the OCR result text data and the original document image, and the output is the patterns and structures that the AI ​​model has learned. This process generates data to supplement misrecognitions and missing parts of the OCR result.

[1307] Step 4:

[1308] The server performs text data completion processing using a generative AI model. Based on learned patterns and structures, the AI ​​model corrects misrecognized parts of the OCR results and generates accurate text data. The input consists of training data and OCR results, and the output is highly accurate, completed text data.

[1309] Step 5:

[1310] The server provides the user with completed text data. This completed text data is returned to the user via API responses or the user interface of a smartphone application. The input is completed text data, and the output is highly accurate text data accessible to the user. The user can then continue their work based on this accurate text data.

[1311] Specific example

[1312] Example of a prompt:

[1313] "The OCR results are below. Please make corrections."

[1314] 1. Delivery address: Roppongi, Minato-ku, Tokyo

[1315] 2. Date: October 15, 2023

[1316] 3. Contents: 10 laptop computers

[1317] The original image was unclear and was mistakenly identified as "10 laptops." Please correct it to reflect the correct content.

[1318] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1319] System Overview

[1320] This system aims to improve the accuracy of optical character recognition (OCR) by combining generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions. Users upload document images to the server, where OCR processing is performed. The system then uses generative AI to complete the images, and finally, the emotion engine adjusts the completion results based on the user's emotions before providing them to the user.

[1321] Processing flow and specific actions

[1322] Upload document images

[1323] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[1324] Execution of OCR processing

[1325] The server saves the received document image and then performs OCR processing. The OCR engine (for example, Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the text data.

[1326] Learning by generative AI

[1327] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand document patterns and structures.

[1328] Completion process

[1329] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[1330] Adjustment by the emotion engine

[1331] The server adjusts the completed text data based on the user's emotions. The emotion engine recognizes emotions from the user's facial expressions, voice, or input text, and the generative AI model adjusts the completion results according to that emotional information.

[1332] Output of completion results

[1333] The server provides the user with adjusted, supplementary text data. The final result is returned to the user, allowing them to obtain accurate and emotionally sensitive character recognition results.

[1334] Specific example

[1335] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer. The server receives it and saves the image. Next, the server analyzes the image using an OCR engine and extracts the initial text data. Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model. Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data.

[1336] The emotion engine analyzes the generated text data to identify information (facial expressions, voice, input characters, etc.) necessary to recognize the user's emotions. For example, if the user is satisfied, the basic completion result is provided as is; if the user is dissatisfied, the generative AI adjusts the result to provide an even more accurate one. Finally, this completed, high-accuracy text data is provided to the user, allowing them to obtain highly accurate and emotion-sensitive character recognition results.

[1337] This invention not only improves the accuracy of OCR but also enables the provision of services that take user emotions into consideration, thereby enhancing the user experience and significantly expanding the value and scale of deployment of OCR systems.

[1338] The following describes the processing flow.

[1339] Step 1:

[1340] The user uploads the document image requiring OCR processing from their device to the server. The user selects the image file (e.g., JPEG, PNG, PDF) via a web browser or dedicated application and clicks the "Upload" button.

[1341] Step 2:

[1342] The server saves the received document image. This ensures that image data is available for use in subsequent processing.

[1343] Step 3:

[1344] The server performs OCR processing on the stored document images. The OCR engine (e.g., Tesseract or Google Cloud Vision API) analyzes the image, recognizes the characters it contains, and extracts the initial text data.

[1345] python

[1346] ocr_result = ocr_engine.process(image_file)

[1347] Step 4:

[1348] The server feeds the extracted text data and the original document images into a generative AI for training. The generative AI learns from this data to understand the patterns and structure of documents.

[1349] python

[1350] gen_ai_model.learn(ocr_result, image_file)

[1351] Step 5:

[1352] The server uses generative AI to supplement the OCR results. Based on the patterns and structures learned by the generative AI model, it supplements unrecognized or misrecognized parts, generating highly accurate text data.

[1353] python

[1354] enhanced_result = gen_ai_model.enhance(ocr_result, image_file)

[1355] Step 6:

[1356] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expressions, voice, or input text to determine whether the user is satisfied, irritated, anxious, etc.

[1357] python

[1358] user_emotion = emotion_engine.analyze(user_input)

[1359] Step 7:

[1360] The server uses a generative AI model to adjust the supplementary text data based on the user's emotions. For example, if the user expresses dissatisfaction, the generative AI adjusts the data to produce more accurate supplementary results.

[1361] python

[1362] adjusted_result = gen_ai_model.adjust_based_on_emotion(enhanced_result, user_emotion)

[1363] Step 8:

[1364] The server provides the user with adjusted, supplementary text data. The final result is returned to the user via an API response or user interface, allowing the user to obtain accurate and emotionally sensitive character recognition results.

[1365] python

[1366] send_response_to_user(adjusted_result)

[1367] Specific example

[1368] This explains the process when a user uploads a JPEG file of a contract to a server. First, the user uploads the contract document using a browser on their computer (Step 1). The server receives it and saves the image (Step 2). Next, the server analyzes the image using an OCR engine and extracts the initial text data (Step 3). Then, the server inputs this text data and the original image data into a generative AI and trains the generative AI model (Step 4). Based on this data, the generative AI fills in any misrecognized parts and generates highly accurate text data (Step 5).

[1369] For this generated text data, the emotion engine analyzes the user's facial expressions, voice, or input text to recognize the user's emotions (Step 6). For example, if the user is expressing dissatisfaction, the generative AI adjusts the results to produce even more accurate results (Step 7). Finally, this adjusted supplementary text data is provided to the user, allowing them to obtain accurate and emotion-sensitive character recognition results (Step 8). This improves the overall user experience by enabling flexible responses that respond to the user's emotions, in addition to improving the accuracy of OCR.

[1370] (Example 2)

[1371] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1372] Optical character recognition (OCR) technology is used in many fields, but it has challenges in recognition accuracy, particularly with misrecognition and missed recognition. Furthermore, recognition results can affect user satisfaction, and there is a need for methods that reflect user emotions regarding the recognition results. Therefore, high-precision OCR processing and the provision of results that take user emotions into consideration are necessary.

[1373] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1374] In this invention, the server includes means for uploading a document image to an information processing device, means for performing optical character recognition processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence model for training, means for supplementing the text data using the generative artificial intelligence model, means for adjusting the supplemented text data based on the user's emotions, and means for providing the adjusted text data to the user. This enables highly accurate OCR processing and the provision of results that take into account the user's emotions.

[1375] A "document image" is image data that is the target of optical character recognition processing, and is generally an image format in which text is embedded (e.g., JPEG, PNG, PDF).

[1376] "Information processing equipment" refers to all devices that perform data input, processing, storage, output, etc., and in this context mainly includes servers and client terminals.

[1377] Optical Character Recognition (OCR) processing is a technology that analyzes characters contained in an image and extracts them as text data.

[1378] "Text data" refers to data representing character information extracted through optical character recognition processing.

[1379] A "generative artificial intelligence model" is an artificial intelligence model that performs data completion or generation based on training data, and in this context, it is used to complete misrecognized or unrecognized parts of OCR results.

[1380] "User" refers to an individual or organization that utilizes optical character recognition (OCR) processing services.

[1381] "Means of adjusting based on emotions" refers to a function that analyzes the user's emotional information and has a generative artificial intelligence model make adjustments in accordance with those emotions.

[1382] "Means of complementarity" refers to a function that complements OCR results based on data learned by a generative artificial intelligence model, thereby generating highly accurate text data.

[1383] This invention relates to a system that combines generative artificial intelligence (AI)-based completion processing with an emotion engine that recognizes user emotions, with the aim of improving the accuracy of optical character recognition (OCR). Specific embodiments of this system are described below.

[1384] System Overview

[1385] The user uploads a document image to an information processing device (server), where OCR processing is performed. This is then supplemented by a generative AI model, and finally, an emotion engine is used to adjust the supplemented result based on the user's emotions before providing it. Specifically, the process proceeds in the following steps:

[1386] Hardware and software

[1387] This system uses the following hardware and software:

[1388] 1. User's device:

[1389] Web browser or dedicated application

[1390] 2. Server:

[1391] Data storage function

[1392] OCR engine (e.g., Tesseract, Google Cloud Vision API)

[1393] Generative AI models

[1394] Emotional Engine

[1395] Data processing and data calculation

[1396] 1. The user uploads a document image from their device to the server.

[1397] 2. The server saves the received document image and extracts the text data using an OCR engine.

[1398] 3. The server inputs the extracted text data and the original document image into a generative AI model for training.

[1399] 4. Generative AI models are used to supplement the OCR results and generate high-accuracy text data.

[1400] 5. The emotion engine analyzes the user's emotional information, and the generative AI model adjusts the completion results based on that information.

[1401] 6. The server provides users with pre-calibrated, high-precision text data.

[1402] Specific example

[1403] Next, we will show a specific example of a user uploading a JPEG file of a contract to the server.

[1404] 1. Upload document images:

[1405] The user accesses the upload page for the JPEG file of the contract using a browser on their computer and clicks the "Upload" button. The image file is then sent to the server.

[1406] 2. Perform OCR processing:

[1407] The server saves the received image file and then performs OCR processing using Tesseract or the Google Cloud Vision API to extract the initial text data.

[1408] 3. Learning using generative AI:

[1409] The server inputs the extracted text data and the original document image into a generative AI program, which then learns to understand the patterns and structure of the document.

[1410] 4. Completion process:

[1411] The generative AI model fills in the gaps in misrecognition and unrecognized parts, generating highly accurate text data.

[1412] 5. Adjustment by the Emotion Engine:

[1413] If a user is dissatisfied, the server will further adjust the completion results based on the user's emotional information (e.g., feedback from input).

[1414] 6. Output of completion results:

[1415] The server provides the user with adjusted, high-precision text data, and the user checks the results in their browser.

[1416] Examples of prompt statements

[1417] The following is a concrete example of a prompt statement for training an AI model to generate this system:

[1418] Image file: Contract_2023.jpeg

[1419] Extracted text: "This agreement is effective as of March 1, 2023…"

[1420] User sentiment: "Satisfied"

[1421] AI-generated output: "This agreement was concluded on March 1, 2023…"

[1422] Adjusted and supplemented result: "This agreement was entered into on March 1, 2023, and is effective as of that date…"

[1423] Image file: receipt_2023.pdf

[1424] Extracted text: "Item Quantity Price"

[1425] User sentiment: "Dissatisfied"

[1426] AI-generated completion result: "Item Quantity Price Total Amount"

[1427] Adjusted and completed result: "Item Quantity Price Total Amount Promotion Applied"

[1428] This enables highly accurate OCR processing and the provision of results that take user emotions into consideration.

[1429] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1430] Step 1:

[1431] The user selects a document image on their device and uploads it to the information processing device (server).

[1432] Input: User-selected document image (e.g., JPEG, PNG, PDF file)

[1433] Output: Document image files saved on the server

[1434] Specific actions:

[1435] The user uses a web browser on their computer to click the file selection button and choose the document image they want to upload. Then, by clicking the "Upload" button, the image file is sent to the server and saved to the server's designated storage.

[1436] Step 2:

[1437] The server performs OCR processing on the saved document images.

[1438] Input: Saved document image file

[1439] Output: Extracted text data

[1440] Specific actions:

[1441] The server uses an OCR engine (e.g., Tesseract or Google Cloud Vision API) to analyze the stored document image files. The OCR engine recognizes the characters in the image and extracts the text data. The extracted text data is stored on the server.

[1442] Step 3:

[1443] The server inputs the extracted text data and the original document image into a generative AI model for training.

[1444] Input: Extracted text data, original document image

[1445] Output: Trained model using a generative AI model

[1446] Specific actions:

[1447] The server inputs the text data obtained through OCR processing and the original document image into a generative AI model. The generative AI model learns and understands the document's patterns and structure based on this data. The trained model is then used for subsequent completion processing.

[1448] Step 4:

[1449] The server uses a generative AI model to supplement the OCR results.

[1450] Input: Pre-trained generative AI model, initial OCR text data

[1451] Output: Interpolated high-precision text data

[1452] Specific actions:

[1453] The server uses a pre-trained generative AI model to perform completion processing on the initial OCR text data. Specifically, misrecognized or unrecognized portions are identified, and these portions are completed by the generative AI model. The completed, high-accuracy text data is stored on the server.

[1454] Step 5:

[1455] The server uses an emotion engine to adjust the completion results based on the user's emotions.

[1456] Input: Completed, high-precision text data, user sentiment information (facial expressions, voice, input characters)

[1457] Output: Text data adjusted based on emotions

[1458] Specific actions:

[1459] Users provide emotional information in real time through their camera and microphone, or input emotional feedback. The server uses an emotion engine to analyze the user's emotional information, has a generative AI model make adjustments based on the emotions, and further refines the completion results.

[1460] Step 6:

[1461] The server then provides the user with the final, adjusted, high-precision text data.

[1462] Input: Text data adjusted based on emotions

[1463] Output: Final high-precision text data provided to the user.

[1464] Specific actions:

[1465] The server sends the adjusted, high-precision text data to the user. The user then downloads or views the results again via a web browser or dedicated application and performs the necessary processing.

[1466] (Application Example 2)

[1467] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1468] Conventional optical character recognition (OCR) systems have suffered from frequent misrecognition in extracting text data from document images, resulting in low accuracy. Furthermore, they often disregarded the user experience, leading to user dissatisfaction. This invention aims not only to improve the accuracy of OCR but also to enhance the user experience by considering user feelings.

[1469] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1470] In this invention, the server includes means for uploading a document image to the server, means for performing optical character recognition (OCR) processing on the document image and extracting text data, means for feeding the extracted text data and the original document image into a generative artificial intelligence (AI) model for training, means for supplementing the text data using the generative AI model, means for adjusting the supplementation results with an emotion engine that analyzes the user's emotional information, and means for providing the supplemented text data to the user. This makes it possible to improve OCR accuracy and provide text data that takes the user's emotions into consideration.

[1471] A "document image" is an image file containing text information provided by the user.

[1472] A "server" is a computer system that provides services to client terminals over a network.

[1473] Optical Character Recognition (OCR) processing is a technology that recognizes characters and numbers in an image and extracts them as text data.

[1474] "Text data" refers to data containing characters and numbers extracted from a document image.

[1475] A "generative artificial intelligence (AI) model" is an artificial intelligence that has the ability to learn from large amounts of data and generate new data.

[1476] An "emotion engine" is software that recognizes a user's emotions and adjusts the processing results based on those emotions.

[1477] "Completion" is the process of filling in omissions and errors in initial text data to make it closer to a complete form.

[1478] "Means of providing to the user" refers to the method of displaying or transmitting the final processing result to the user.

[1479] This invention provides a system that performs high-precision OCR processing on document images while also taking user emotions into consideration. It is basically implemented by the following means.

[1480] Hardware and software

[1481] Hardware: Primarily uses servers, client terminals, and smartphones.

[1482] The server is a central computer system that performs OCR processing, executes generative artificial intelligence (AI) models, and manages data, and is connected to users by applications running on client terminals and smartphones.

[1483] Software: Use the following software.

[1484] OCR engine: Uses tools such as pytesseract and Google Cloud Vision API to perform character recognition from document images.

[1485] Generative artificial intelligence models: Text data is completed using the Transformers library (e.g., GPT-3).

[1486] Emotion Engine: Analyzes user emotions using TextBlob.

[1487] Processing flow

[1488] 1. Uploading Document Images: Users upload document images (e.g., JPEG, PNG, PDF) to the server. Uploads are performed via a dedicated application on a smartphone or client device.

[1489] 2. OCR processing: The server saves the received document image and extracts the text data using an OCR engine (e.g., pytesseract). The extracted text data is saved as an intermediate result.

[1490] 3. Completion by Generative AI: The server feeds the extracted text data and the original document image into a generative AI model (e.g., GPT-3) for training. The generative AI model uses this data to complete the text and generate high-accuracy text data.

[1491] 4. Adjustment by the emotion engine: The server analyzes the user's emotion information using the emotion engine (TextBlob). If the emotion information is "satisfied," the completion result is used as is; if it is "dissatisfied," the text is further modified and completed.

[1492] 5. Provision to the user: The server ultimately provides the user with the completed and adjusted text data. The user obtains highly accurate and emotion-sensitive character recognition results through a dedicated application.

[1493] Specific example

[1494] For example, suppose a user uploads an image of a warranty certificate for a virtual store, and OCR processing is required. In this case, the user takes a picture of the warranty certificate within the virtual store app and sends it to the server. On the server, the OCR engine analyzes the image and extracts initial text data. Then, a generative AI model (e.g., GPT-3) completes this text data, and an emotion engine (e.g., TextBlob) analyzes the user's input text to recognize emotions.

[1495] For example, if a user enters "The contents of this warranty are inaccurate," the emotion engine recognizes the dissatisfaction, and the generative AI model corrects and supplements the text data again. Finally, the completed and adjusted text is returned to the user. An example of the prompt text in this case is as follows:

[1496] "The information on the warranty is inaccurate. Please see below for details:"

[1497] This allows users to obtain highly accurate character recognition results and makes it easier to manage warranty certificates.

[1498] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1499] Step 1:

[1500] The user uploads a document image.

[1501] Input: Document image (JPEG, PNG, PDF, etc.)

[1502] Specific operation: The user launches a dedicated application, takes or selects a document image, and sends it to the server. Clicking the upload button transfers the image to the server.

[1503] Output: Uploaded document image file

[1504] Step 2:

[1505] The server saves the document image and performs OCR processing.

[1506] Input: Uploaded document image file

[1507] Specific operation: The server saves the received image file and runs an OCR engine (e.g., pytesseract) to extract text data from the image. It then performs grayscale conversion, noise reduction, and character recognition on the image.

[1508] Output: Extracted text data

[1509] Step 3:

[1510] The server loads the extracted text data and the original document images into a generative AI model for training.

[1511] Input: Text data, document images

[1512] Specific operation: The server inputs text data and images into a generative artificial intelligence model (e.g., GPT-3) to train it on document patterns and structures. The model then learns to correct misrecognitions based on the input data.

[1513] Output: Generated high-precision text data

[1514] Step 4:

[1515] The server analyzes the user's emotional information and adjusts the completion results accordingly.

[1516] Input: Generated text data, user sentiment information

[1517] Specific operation: The server receives additional input from the user (e.g., "This part is inaccurate") and uses a sentiment engine (e.g., TextBlob) to analyze the user's sentiment. It then further completes or adjusts the text data based on the sentiment.

[1518] Output: Adjusted text data

[1519] Step 5:

[1520] The server provides the user with the final, completed, and adjusted text data.

[1521] Input: Adjusted text data

[1522] Specific operation: The server saves the final adjustment results as a file and sends it to the user through a dedicated application. The user can then view the text data within the application.

[1523] Output: High-precision, emotion-sensitive text data acquired by the user.

[1524] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1525] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1526] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1527] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1528] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1529] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1530] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1531] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1532] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1533] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1534] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1535] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1536] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1537] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1538] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1539] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1540] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1541] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1542] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1543] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1544] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1545] The following is further disclosed regarding the embodiments described above.

[1546] (Claim 1)

[1547] A means of uploading document images to a server,

[1548] A means for performing optical character recognition (OCR) processing on a document image and extracting text data,

[1549] A method for feeding extracted text data and the original document image into a generative artificial intelligence (AI) model for training,

[1550] A method for supplementing text data using generative artificial intelligence models,

[1551] A means of providing the user with completed text data,

[1552] A system that includes this.

[1553] (Claim 2)

[1554] The system according to claim 1, which performs image analysis in optical character recognition processing.

[1555] (Claim 3)

[1556] The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures in text data.

[1557] "Example 1"

[1558] (Claim 1)

[1559] A means of uploading document images to an electronic device,

[1560] A means for performing optical character recognition processing on a document image and extracting character data,

[1561] A method for feeding extracted text data and the original document image into a generative artificial intelligence model for training,

[1562] A method for completing character data using a generative artificial intelligence model,

[1563] A means of providing the user with completed character data,

[1564] A system that includes this.

[1565] (Claim 2)

[1566] The system according to claim 1, which performs image analysis in optical character recognition processing.

[1567] (Claim 3)

[1568] The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures of character data.

[1569] "Application Example 1"

[1570] (Claim 1)

[1571] A means of uploading document images to a server,

[1572] A means for performing optical character recognition (OCR) processing on a document image and extracting text data,

[1573] A method for feeding extracted text data and the original document image into a generative artificial intelligence (AI) model for training,

[1574] A method for supplementing text data using generative artificial intelligence models,

[1575] A means of providing the user with completed text data,

[1576] A means for performing OCR processing and generative AI-based supplementary processing on document images entered via smartphone scanning at a logistics center,

[1577] A system that includes this.

[1578] (Claim 2)

[1579] The system according to claim 1, which performs image analysis in optical character recognition processing.

[1580] (Claim 3)

[1581] The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures in text data.

[1582] "Example 2 of combining an emotion engine"

[1583] (Claim 1)

[1584] A means of uploading document images to an information processing device,

[1585] A means for performing optical character recognition processing on a document image and extracting text data,

[1586] A method for feeding extracted text data and the original document image into a generative artificial intelligence model for training,

[1587] A method for supplementing text data using generative artificial intelligence models,

[1588] A means of adjusting the completed text data based on the user's emotions,

[1589] A means of providing the user with adjusted text data,

[1590] A system that includes this.

[1591] (Claim 2)

[1592] The system according to claim 1, which performs image analysis in optical character recognition processing.

[1593] (Claim 3)

[1594] The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures in text data.

[1595] "Application example 2 when combining with an emotional engine"

[1596] (Claim 1)

[1597] A means of uploading document images to a server,

[1598] A means for performing optical character recognition (OCR) processing on a document image and extracting text data,

[1599] A method for feeding extracted text data and the original document image into a generative artificial intelligence (AI) model for training,

[1600] A method for supplementing text data using generative artificial intelligence models,

[1601] A means of adjusting the completion results using an emotion engine that analyzes the user's emotional information,

[1602] A means of providing the user with completed text data,

[1603] A system that includes this.

[1604] (Claim 2)

[1605] The system according to claim 1, which performs image analysis in optical character recognition processing.

[1606] (Claim 3)

[1607] The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures in text data. [Explanation of Symbols]

[1608] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of uploading document images to a server, A means for performing optical character recognition processing on a document image and extracting text data, A method for feeding extracted text data and the original document image into a generative artificial intelligence model for training, A method for supplementing text data using generative artificial intelligence models, A means of providing the user with completed text data, A system that includes this.

2. The system according to claim 1, which performs image analysis in optical character recognition processing.

3. The system according to claim 1, wherein a generative artificial intelligence model learns patterns and structures in text data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A