System

The system automates character recognition and correction, addressing inefficiencies in conventional OCR systems by enabling efficient and accurate transcription and correction of text data.

JP2026014180APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115177
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional transcription systems using optical character recognition (OCR) technology require manual correction of transcribed text, which is inefficient and time-consuming, especially for large volumes of documents, and lack automated grammar and spelling checks.

Method used

A system that includes an optical character recognition unit, a correction unit, and a storage unit to automatically extract characters, correct grammar and spelling, and display the results on a user terminal, allowing users to review the corrections efficiently.

Benefits of technology

The system automates the transcription and correction process, significantly reducing the time and effort required for users to obtain high-quality text data by providing a comparison of original and corrected text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014180000001_ABST
    Figure 2026014180000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes an optical character recognition means for extracting characters from an image, a correction means for performing text correction of extracted character data, and a means for displaying a correction result on a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional transcription systems using optical character recognition (OCR) technology have left the accuracy and grammar checks of extracted text data to the user, resulting in reduced work efficiency. Furthermore, manually correcting the transcribed text requires a great deal of time and effort, making it particularly inefficient when dealing with large volumes of documents. For this reason, there was a demand for a system that not only transcribes text but also automatically corrects it, making it easy for users to use. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system including an optical character recognition unit that extracts characters from an image, a correction unit that corrects the extracted character data, and a unit that displays the correction results on a user terminal. This system has the function of accurately extracting characters from an image and automatically correcting grammar and spelling errors. It also has a storage unit and a correction unit that temporarily store the extracted character data, check grammar, check spelling, and improve style, and mark up the differences between the corrected character data and the original character data. A user uploads an image from their terminal and sends the uploaded image to a server. The server sends the correction results to the user terminal and displays them to the user. In this way, the user can check the transcription and correction results all at once, improving work efficiency.

[0006] "Optical character recognition" is a technology that extracts character information from an image, and is commonly called OCR (Optical Character Recognition).

[0007] "Character data" refers to text information extracted from an image using optical character recognition means.

[0008] "Text correction" is the process of automatically correcting the grammar, spelling, style, etc. of extracted text data to produce accurate and attractive text.

[0009] A "user terminal" is a device used by a user to access and operate the system, and includes smartphones and personal computers.

[0010] The "storage means" is a part of the system that has the function of temporarily storing extracted character data in storage.

[0011] "Markup" is a technique for visually indicating the differences between corrected text data and the original text data, and generally involves highlighting the differences. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] The present invention relates to a system that extracts characters from an image, automatically corrects the character data, and displays the correction results on a user terminal. This system is composed of several main components, including an optical character recognition (OCR) unit, a text correction unit, and a storage unit.

[0034] First, a user can upload an image from their device. The image is preferably a handwritten note or a printed document. The device then sends the image to the server, which then stores the received image data in storage.

[0035] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data, which is temporarily stored on the server.

[0036] The server then sends this saved text data to the text correction module, which analyzes it for grammar, spelling mistakes, and style improvements and performs corrections. Once the corrections are complete, the text data is returned to the server, which calculates the differences between the original and corrected text data and marks them up.

[0037] Finally, the server sends the results to the user's device, which then displays a comparison of the original text and the corrected text. The user can see the differences on the screen and make further corrections if necessary.

[0038] Specific examples

[0039] User Operation

[0040] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0041] Server Processing

[0042] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0043] Correction process

[0044] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0045] Results display

[0046] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0047] Original text: "Hello. My name is Tanaka."

[0048] Corrected text: "Hello. My name is Tanaka."

[0049] This system automates the entire process from transcription to correction, allowing users to obtain high-quality text data quickly and efficiently.

[0050] The processing flow will be explained below.

[0051] Step 1:

[0052] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0053] Step 2:

[0054] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0055] Step 3:

[0056] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0057] Step 4:

[0058] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0059] Step 5:

[0060] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[0061] Step 6:

[0062] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[0063] Step 7:

[0064] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[0065] Step 8:

[0066] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[0067] Step 9:

[0068] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[0069] Example 1

[0070] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0071] Conventional optical character recognition systems have low accuracy, and corrections to character data extracted from handwritten or printed documents have been performed manually, which requires time and effort. Furthermore, grammar checks and style improvements after character recognition have not been automated, making the system unfriendly to users. The present invention aims to solve these problems by providing a system that performs character recognition and automatic corrections from images with high accuracy and efficiency.

[0072] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0073] In this invention, the server includes a means for users to upload images from their terminals, a means for transmitting the uploaded images to the server, and a means for the server to pass the image data stored therein to an optical character recognition module. This automates the series of processes of character recognition and text correction, enabling users to obtain high-quality text data quickly and efficiently.

[0074] A "terminal" is a device operated by a user, and includes devices such as smartphones, tablets, and PCs.

[0075] A "server" is a computer system that processes, stores, and transmits image and text data.

[0076] An "optical character recognition module" is a piece of software or hardware that extracts text information from an image and uses OCR (optical character recognition) technology.

[0077] A "text correction module" is a piece of software or hardware that checks the grammar, corrects spelling mistakes, and improves style of extracted text data.

[0078] "Storage means" refers to a storage device for temporarily or permanently storing image data or text data, and includes hard disks, SSDs, cloud storage, etc.

[0079] "Upload" refers to the operation in which a user sends data (images or files) from a terminal to another location (such as a server).

[0080] "Markup" is an operation that uses scripts and tags to highlight or decorate specific parts of text data.

[0081] "Comparative display" is an operation that displays two different text data side by side, allowing you to visually confirm the differences between them.

[0082] "Sending the results" means that the server transfers the processed data to the user's terminal, which is usually done over a network.

[0083] This invention is a system that extracts characters from images, automatically corrects the character data, and displays the correction results on the user's terminal. This system is realized mainly by users uploading images from their terminals and the server processing the image data.

[0084] System configuration

[0085] The system includes the following main components:

[0086] 1. User Device

[0087] 2. Server

[0088] 3. Optical Character Recognition Module (OCR Module)

[0089] 4. Text Correction Module

[0090] 5. Preservation means

[0091] Hardware and software used

[0092] User devices: smartphones, tablets, PCs, etc.

[0093] Server: A high-performance computer system

[0094] Optical Character Recognition Module: Google Cloud Vision API, Tesseract, etc.

[0095] Text correction module: Grammarly API, natural language processing model

[0096] Storage method: Cloud storage such as Amazon S3

[0097] Process Flow

[0098] User Operation

[0099] 1. Users take photos of handwritten notes or printed documents using their smartphones or PCs and upload the images through the application.

[0100] 2. The uploaded image is sent from the device to the server.

[0101] Server Processing

[0102] 1. The server temporarily stores the received image data. Cloud storage is recommended as the storage location.

[0103] 2. The stored image data is passed to the OCR module to extract text information. The OCR module uses optical character recognition technology to generate text data from the image.

[0104] 3. The extracted text data is temporarily stored on the server.

[0105] 4. The server sends the saved text data to the writing correction module, which corrects grammar, spelling, and style, and returns the corrected text data to the server.

[0106] 5. The server compares the original text data with the corrected text data and marks up the differences.

[0107] Results display

[0108] 1. The server sends the final correction results to the user's terminal.

[0109] 2. The user's device receives the transmitted data and displays a comparison of the original text and the corrected text, allowing the user to review the results and make further corrections as necessary.

[0110] Specific examples

[0111] Prompt Sentence Examples

[0112] "Upload an image of a handwritten note and automatically extract and correct the text. Example: The image says, 'Hello. My name is Tanaka.' Output the text 'Hello. My name is Tanaka.'"

[0113] This system allows users to obtain high-quality text data quickly and efficiently. In particular, the entire process from transcription to text correction is automated, significantly reducing the time and effort required.

[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0115] Step 1:

[0116] A user can use a device to take a photo of a handwritten note or a printed document and upload the image through the application. The user selects an image and presses the upload button to start the image data upload operation. The input is the image file taken by the user, and the output is the uploaded image file data.

[0117] Step 2:

[0118] The terminal sends the image data uploaded by the user to the server. A secure connection is established using the HTTPS protocol, and the image data is sent to the server as an HTTP request. The input is the image data sent from the user terminal, and the output is the image data received by the server.

[0119] Step 3:

[0120] The server saves the received image data in storage. Cloud storage such as Amazon S3 is used as the storage. A storage path and file name are generated and the image data is saved. The input is the received image data, and the output is the path information for the image file saved in storage.

[0121] Step 4:

[0122] The server obtains the path information of the stored image data and passes it to the OCR module. The connection with the OCR module is made using the API key and authentication information, and the image data is sent to the OCR module's endpoint. The input is the path information of the image data stored in storage, and the output is the image data sent to the OCR module.

[0123] Step 5:

[0124] The OCR module analyzes the received image data and extracts text information from the image. It uses OCR technologies such as Google Cloud Vision API and Tesseract to analyze the text information and generate text data. The input is the image data sent to the OCR module, and the output is the extracted text data.

[0125] Step 6:

[0126] The server receives the text data returned from the OCR module and temporarily stores it. The text data is stored in a database or memory storage within the server. The input is the text data returned from the OCR module, and the output is the text data stored on the server.

[0127] Step 7:

[0128] The server sends the saved text data to the writing correction module. It connects to the writing correction module (e.g., Grammarly API) and sends the text data in JSON format. The input is the text data saved on the server, and the output is the text data sent to the writing correction module.

[0129] Step 8:

[0130] The writing correction module analyzes the received text data and improves grammar, spelling, and style. The corrected text data is returned to the server with markup. The input is the text data sent to the writing correction module, and the output is the corrected text data.

[0131] Step 9:

[0132] The server compares the original text data with the corrected text data and marks up the parts that have been corrected. It calculates the differences and formats them in a user-friendly format. The input is the original text data and the corrected text data, and the output is the marked-up comparison results.

[0133] Step 10:

[0134] The server sends the final marked-up text result to the user's terminal. The result data is formatted in a format such as HTML or PDF and sent as data for display. The input is the marked-up text result, and the output is the result data sent to the user's terminal.

[0135] Step 11:

[0136] The user checks the final result on the terminal and compares the original text with the corrected text. The differences can be seen at a glance on the displayed screen and further corrections can be made if necessary. The input is the result data sent from the server and the output is the displayed data that the user checks.

[0137] (Application example 1)

[0138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0139] In recent years, many product price labels in brick-and-mortar stores are frequently updated, but this process is prone to human error. This can lead to incorrect pricing information, which can not only undermine customer trust but also affect sales. Furthermore, manually checking and correcting price labels is laborious and time-consuming for store staff, so an efficient system is needed. Currently available systems lack the ability to automatically detect and correct price label errors, so a new system is needed to address this issue.

[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0141] In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, means for displaying the correction results on a user terminal, automatic correction means for comparing the extracted character data with the correction results and correcting errors, and means for displaying the results to the user in real time. This allows a store clerk to instantly detect incorrect price information and receive correction suggestions simply by taking a photo of a price label with their smartphone, thereby enabling efficient and accurate price label management.

[0142] The "optical character recognition means" is a means for extracting characters from image data and generating text data.

[0143] The "correction means" is a means for correcting grammar, spelling mistakes, and style of the extracted character data.

[0144] The "means for displaying on the user terminal" refers to a means for displaying the correction results sent from the server on the user's device.

[0145] The "automatic correction means for correcting errors" is an automated means for comparing extracted character data with corrected character data and correcting errors.

[0146] "Means for displaying in real time" refers to means for instantly displaying processed data on a user terminal.

[0147] The "storage means" is a means for temporarily storing image data and extracted character data.

[0148] "Markup means" refers to a means for visually showing the difference between corrected character data and the original character data.

[0149] The "means for detecting errors in price labels and suggesting corrections" is a means for detecting errors in product price information and suggesting correct price information.

[0150] The "means for transmitting to the server" is a means for transferring images uploaded by a user from a terminal to the server.

[0151] The "means for transmitting result data" is a means for transmitting processed data from the server to the user's device.

[0152] The present invention relates to a system for automatically detecting and correcting price label errors in brick-and-mortar stores, which includes optical character recognition, text correction, and automatic correction.

[0153] First, the user takes a picture of the price label using their smartphone, which then sends the image to the server, which then temporarily stores the image data.

[0154] The server then invokes the Optical Character Recognition (OCR) module, passing the image data as input. This OCR module uses OpenCV and Tesseract to extract text data from the image, which is temporarily stored on the server.

[0155] The server then sends the saved text data to a writing correction module, which uses TextBlob to correct grammar, spelling, and style errors, and then returns the corrected text data to the server.

[0156] The server corrects any errors found and marks up the difference between the original and corrected text data. During this process, if the original price information is incorrect, the correct price information is suggested and displayed to the user in real time.

[0157] The user terminal receives and displays the result data sent from the server, allowing the user to immediately check the price label errors in the store and make any necessary corrections.

[0158] For example, the following prompts are used:

[0159] Example prompt sentence:

[0160] "Check the price label on the product. It might say '50,000 yen' instead of '500 yen'."

[0161] This system enables store clerks who manage price labels in physical stores to efficiently and accurately check and correct price information. The hardware used is a smartphone and a camera, and the software uses OpenCV, Tesseract, and TextBlob. This process creates a system that can automatically correct errors in price labels in real time.

[0162] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0163] Step 1:

[0164] A user takes a picture of the price label using a smartphone device.

[0165] Input: Price Label Image

[0166] Output: Image data of the photographed price label

[0167] What happens: The user opens the camera app on their smartphone, focuses on the price label, and takes a photo. The image is then saved to their device.

[0168] Step 2:

[0169] The terminal transmits the captured image to the server.

[0170] Input: Price label image data

[0171] Output: Image data sent to the server

[0172] Specific operation: The device selects the captured image and uploads it to the server via the app. The communication protocol is HTTPS.

[0173] Step 3:

[0174] The server temporarily stores the received image data.

[0175] Input: Received image data

[0176] Output: Saved image data

[0177] Specific operation: The server stores the received image data in a database or temporary file. The storage location is a storage system (e.g., Amazon S3).

[0178] Step 4:

[0179] The server passes the image data to the OCR module to extract the character information.

[0180] Input: Image data

[0181] Output: Extracted text data

[0182] Specific operation: The server uses OpenCV to preprocess the image data, then uses Tesseract to extract text information. The extracted text data is temporarily saved.

[0183] Step 5:

[0184] The server sends the extracted text data to the text correction module.

[0185] Input: Text data

[0186] Output: Corrected text data

[0187] Specific operation: The server analyzes the text data using TextBlob, corrects grammar, spelling mistakes, and style, and saves the corrected text data.

[0188] Step 6:

[0189] The server marks up the differences between the corrected text data and the original text data, and generates correction information for correcting the errors.

[0190] Input: Original text data and corrected text data

[0191] Output: Marked up text data and correction information

[0192] What it does: The server compares each piece of text data and visually displays the differences using a program that highlights the differences (e.g., a diff tool).

[0193] Step 7:

[0194] The server transmits the correction information and the marked-up text data to the user terminal.

[0195] Input: Correction information, marked-up text data

[0196] Output: Information sent to the user's device

[0197] Specific operation: The server prepares the marked-up text data and the correction information for the error correction as a data packet and sends it to the user terminal via the HTTPS protocol.

[0198] Step 8:

[0199] The user terminal displays the received information on the screen, and notifies the user of any price label errors in real time.

[0200] Input: Received correction information, marked-up text data

[0201] Output: Screen showing suggested revisions and original pricing information

[0202] Specific operation: Based on the received information, the user device updates the GUI to display the error and suggested corrections on the screen. The user can check this and manually correct it if necessary.

[0203] This allows store staff to efficiently and accurately detect price label errors and make corrections in real time.

[0204] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0205] This invention relates to a system that extracts characters from images, automatically corrects the character data, recognizes the user's emotions, and reflects them in the correction results. This system is composed of several main components, including an optical character recognition (OCR) means, a text correction means, an emotion engine, and a storage means.

[0206] A user can first upload an image from their device, preferably a handwritten note or a printed document, and the device sends the image to the server, which then stores the received image data in temporary storage.

[0207] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data. This text data is temporarily stored on the server. The server then sends the stored text data to the writing correction module, which analyzes the text for grammar, spelling mistakes, and style improvements and performs the correction process.

[0208] The server then calls an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. For example, it adjusts the correction results based on the user's emotional state, such as whether they are stressed or relaxed.

[0209] The adjusted correction results are sent back to the server, which compares the original text data with the corrected text data, calculates the differences, and marks them up. Finally, the server sends the results to the user's device, which receives them and displays a comparison of the original text and the corrected text. The user can check the differences on the screen and make further corrections if necessary.

[0210] Specific examples

[0211] User Operation

[0212] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0213] Server Processing

[0214] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0215] Correction process

[0216] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0217] Emotion Recognition Processing

[0218] The server calls the emotion engine to analyze the user's emotional state. For example, if the user is nervous, the tone of the corrections will be gentler and the expressions will be gentler. Conversely, if the user is relaxed, more specific comments will be made.

[0219] Results display

[0220] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0221] Original text: "Hello. My name is Tanaka."

[0222] Corrected text: "Hello. My name is Tanaka."

[0223] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[0224] The processing flow will be explained below.

[0225] Step 1:

[0226] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0227] Step 2:

[0228] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0229] Step 3:

[0230] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0231] Step 4:

[0232] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0233] Step 5:

[0234] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[0235] Step 6:

[0236] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[0237] Step 7:

[0238] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[0239] Step 8:

[0240] The server then calls the emotion engine to analyze the user's current emotional state, which uses facial recognition data and previous typing patterns to determine whether the user is tense or relaxed.

[0241] Step 9:

[0242] The server adjusts the tone and style of its corrections based on the data it receives from the emotion engine, for example, increasing kind language and positive feedback if the user is nervous.

[0243] Step 10:

[0244] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[0245] Step 11:

[0246] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[0247] Example 2

[0248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0249] Conventional systems simply extract text from image data and correct the text, but this does not necessarily mean that the corrections are appropriate for the user's feelings. Furthermore, feedback that does not take the user's feelings into account is often ineffective. Furthermore, it is difficult to efficiently process image data uploaded by users and quickly display the results.

[0250] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an optical character recognition means for extracting characters from an image, a correction means for correcting the extracted character data, an emotion recognition means for recognizing the user's emotion and adjusting the correction result based on the emotion, and a means for displaying the correction result on the user terminal. This enables appropriate correction of the text taking the user's emotion into consideration and rapid feedback.

[0251] "Optical character recognition" refers to techniques and devices for extracting character information from image data.

[0252] "Text correction means" refers to the technology and devices that analyze extracted text data from the perspectives of grammar, spelling mistakes, and style, and make appropriate corrections.

[0253] "Emotion recognition means" refers to technology and devices for analyzing a user's emotions and adjusting the results of data processing according to that emotional state.

[0254] "Storage means" refers to devices or technologies for temporarily or permanently storing extracted character data or intermediate processed data.

[0255] "Means for marking up differences" refers to techniques and devices for visually highlighting the differences between the original text data and the modified text data.

[0256] A "user terminal" is a hardware device operated by a user, and includes devices such as smartphones and personal computers.

[0257] "Server" refers to a computer system for processing, storing, and transmitting data.

[0258] This system extracts text from an image, automatically corrects the text, and recognizes the user's emotions and reflects them in the correction results. This system consists of the following main components:

[0259] Key Components

[0260] 1. Optical Character Recognition (OCR):

[0261] This is a technology that extracts text information from image data. Specifically, it uses an OCR engine. For example, Google Cloud Vision API can be used.

[0262] 2. Writing correction methods:

[0263] This technology analyzes the extracted text data from the perspectives of grammar, spelling mistakes, and style, and makes appropriate corrections. For example, you can use a grammar checker API such as Grammarly.

[0264] 3. Emotion recognition means:

[0265] This technology analyzes a user's emotions and adjusts the results of processing data based on their emotional state. It includes algorithms that analyze facial recognition information and text input patterns. This allows it to adjust the tone or expressions to be gentler if the user is nervous, or more specific if the user is relaxed.

[0266] 4. Preservation means:

[0267] This is a technology for temporarily storing extracted text data and intermediate processed data. For example, data can be stored using a cloud storage service.

[0268] 5. Ways to mark up differences:

[0269] This technology visually highlights the differences between the original text and the edited text. It uses a comparison algorithm to highlight the changes.

[0270] Specific examples of operations

[0271] User Operation

[0272] A user can use a smartphone to take a photo of a handwritten note and upload it through the app. For example, consider a handwritten note that reads, "Hello. My name is Tanaka."

[0273] Server Processing and Response

[0274] The server passes the received image to the OCR module. The OCR module extracts character information from the image and generates text data such as "Hello. My name is Tanaka." This text data is saved on the server. The server then sends the saved text data to the writing correction module. The writing correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0275] Performing emotion recognition

[0276] The server calls the emotion engine and analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. If the user is nervous, the corrections will be made in a gentler tone and with gentler expressions. Conversely, if the user is relaxed, more specific corrections will be made.

[0277] Displaying the results

[0278] The server calculates the difference between the original text data and the corrected text data, marks it up, and finally sends the result to the user's terminal, which displays a comparison of the original text and the corrected text on the screen.

[0279] Prompt Sentence Examples

[0280] "Please correct the following sentence: 'Hello. My name is Tanaka.' If the user is relaxed, provide specific feedback."

[0281] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[0282] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0283] Step 1:

[0284] Uploading image data

[0285] User:

[0286] A user uses an application on their smartphone or computer to take an image of a handwritten note or printed document, and then uploads it to a server via the application. For example, consider an image of a handwritten note that reads, "Hello. My name is Tanaka." This image data becomes the input.

[0287] output:

[0288] Uploaded image data.

[0289] Step 2:

[0290] Image data sent to server and saved

[0291] Device:

[0292] The device sends the uploaded image data to the server. Specifically, the device sends the image data to the specified URL of the server via an HTTP POST request.

[0293] server:

[0294] The server stores the received image data in temporary storage. For example, if a cloud storage service is used, the uploaded image is stored in the cloud storage. This image data becomes the input for the next step.

[0295] output:

[0296] Image data stored in temporary storage.

[0297] Step 3:

[0298] Extracting character information using OCR

[0299] server:

[0300] The server calls the OCR module and passes the image data stored in temporary storage as input. The OCR module analyzes the image data and extracts text information. For example, the text data "Hello. My name is Tanaka" is extracted from the image.

[0301] output:

[0302] Extracted text data. The text generated is "Hello. My name is Tanaka."

[0303] Step 4:

[0304] Execution of text correction

[0305] server:

[0306] The server sends the generated text data to a writing correction module. For example, the text data is sent to the Grammarly API, which analyzes it for grammar, spelling mistakes, and style improvements. The writing correction module corrects "Hello" to "Hello" and checks the grammar of "I'm Tanaka."

[0307] output:

[0308] Corrected text data. The corrected text "Hello. My name is Tanaka." is generated.

[0309] Step 5:

[0310] User Emotion Recognition

[0311] server:

[0312] The server calls the emotion engine and analyzes facial recognition information and text input speed information from the user's device to identify the user's emotional state. If the analysis shows that the user is nervous, the correction tone will be softened, and if the user is relaxed, the correction will be more specific.

[0313] output:

[0314] Tailored corrections, such as gentler language in sentences.

[0315] Step 6:

[0316] Diff calculation and markup

[0317] server:

[0318] The server calculates the difference between the original text data and the corrected text data and marks it up, using a comparison algorithm to highlight the part where "Hello" was changed to "Hello."

[0319] output:

[0320] Text data with differences marked up.

[0321] Step 7:

[0322] Display the results on your terminal

[0323] server:

[0324] The server sends the resulting data after the difference calculation to the user terminal. Specifically, it returns JSON format data in an HTTP response.

[0325] Device:

[0326] The user's device receives this and displays the original text and the corrected text on the screen for comparison. For example, the part where "Hello" has been changed to "Hello" is highlighted so that the user can easily confirm it.

[0327] output:

[0328] Text data displayed on the screen before and after correction.

[0329] This series of processes allows the user to quickly and efficiently obtain high-quality text data.

[0330] (Application example 2)

[0331] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0332] Existing optical character recognition technology and text correction systems have difficulty in accurately and efficiently recognizing and correcting text during maintenance work in factories. Furthermore, there is a lack of a means to provide appropriate feedback based on the emotional state of maintenance personnel. As a result, misrecognition and miscorrection occur frequently, leading to a decline in maintenance efficiency and work accuracy.

[0333] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, and emotion recognition means for recognizing the user's emotion and reflecting it in the correction results. This makes it possible to perform accurate character recognition and correction during maintenance work in a factory, and to provide appropriate feedback based on the emotional state of the maintenance worker.

[0334] An "image" is a digital recording of visual information.

[0335] "Characters" are symbols or codes used to represent language.

[0336] "Optical character recognition" is a technology for extracting character information from an image.

[0337] "Character data" is character information extracted by character recognition means expressed in digital form.

[0338] "Text correction means" is a technology for improving grammar, spelling mistakes, and style of text based on extracted character data.

[0339] "Emotion recognition means" refers to technology for analyzing and identifying a user's emotional state.

[0340] A "user terminal" is a device operated by a user, such as a smartphone or tablet.

[0341] "Storage means" refers to a device for temporarily or permanently storing character data.

[0342] A "server" is a computer system that processes and stores data on a network.

[0343] "Temporarily storing" means retaining data for a specific period of time.

[0344] "Grammar" is the rules and structure of a language.

[0345] A "spelling error" is an error in the spelling of a word.

[0346] "Style improvements" are modifications made to improve the readability and appearance of the text.

[0347] "Marking up the differences" means visually indicating the differences between the original character data and the corrected character data.

[0348] "Uploading" means sending data from a user terminal to a server.

[0349] "Result data" refers to data that includes corrected character data and analysis results.

[0350] A "factory robot" is a mechanical device that automatically performs various tasks in a factory.

[0351] An "application" is a software program with a specific function.

[0352] This invention is a system for streamlining maintenance work in factories and providing accurate character recognition and emotion-based feedback. The main components of the system include a server, a terminal (e.g., a factory robot), and a user (a maintenance worker).

[0353] The server is equipped with an optical character recognition unit that extracts characters from images, a text correction unit, and an emotion recognition unit. The terminal is a device operated by a user, which in this case corresponds to a robot deployed in a factory. The user provides the system with the information needed during maintenance work through the terminal.

[0354] Program processing overview

[0355] Hardware and Software

[0356] The server is a computer system equipped with a high-performance processor and sufficient memory, and uses the following software:

[0357] OpenCV: Image processing library

[0358] pytesseract: Optical Character Recognition Library

[0359] TextBlob: A natural language processing library

[0360] EmotionRecognition: A Virtual Library for Emotion Recognition

[0361] System processing flow

[0362] 1. Image capture:

[0363] Factory robots take images of the equipment or parts they are maintaining, for example, by using their cameras to take pictures of wiring diagrams.

[0364] 2. Character extraction:

[0365] The server receives the captured image and extracts character information using pytesseract. For example, it extracts the label "Ryk10 number" from the wiring diagram.

[0366] 3. Text proofreading:

[0367] The server proofreads grammar and spelling mistakes of the extracted character information using TextBlob. For example, it corrects "Ryjk10 number" to "Ryk10 number".

[0368] 4. Emotion recognition:

[0369] The server uses EmotionRecognition to recognize emotions using the facial image of the user captured by the camera mounted on the robot. For example, it detects whether the user is in a stressed state or a relaxed state.

[0370] 5. Feedback adjustment:

[0371] According to the assumed emotional state, the feedback content is adjusted to give kind expressions and specific points. For example, for a user feeling stressed, feedback like "You've had a hard time. It's an excellent job." is provided.

[0372] 6. Result display:

[0373] The server sends the original text, the proofread text, and the further adjusted feedback content to the user terminal and displays them on the screen.

[0374] Addition of specific examples

[0375] For example, if the robot photographs a wiring diagram of a device and recognizes the text as "Ryjk10," it will correct this misrecognition to "Ryk10" using TextBlob. At the same time, if it detects that a maintenance worker is under stress using emotion recognition technology, it will provide feedback such as "Great job, great job."

[0376] Example prompts for generative AI models

[0377] Image: [Link to wiring diagram image]

[0378] Caption: Is there any misidentified text in this wiring diagram? Please correct it, especially the part where it is misidentified as "Ryjk10".

[0379] Emotional state: Stressed

[0380] In this way, the system can streamline maintenance work in factories and provide highly accurate character recognition and emotional feedback.

[0381] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0382] Step 1:

[0383] Image Capture

[0384] The user uses the factory robot's camera to take images of the equipment or parts to be maintained. The input is the image file taken by the user, and the output is the image data. This image data is required for character recognition processing in the subsequent stage.

[0385] Step 2:

[0386] Character extraction

[0387] The server receives the image data sent from the factory robot and extracts character information using the pytesseract library. The input is the image data, and the output is the character data. Specifically, it extracts texts such as the label "Ryk10" written in the wiring diagram within the image.

[0388] Step 3:

[0389] Text proofreading

[0390] The server proofreads grammar and spelling mistakes of the extracted character information using the TextBlob library. The input is the extracted character data, and the output is the proofread character data. As a specific operation, it corrects the misrecognized character "Ryjk10" to "Ryk10".

[0391] Step 4:

[0392] Emotion recognition

[0393] The server uses the EmotionRecognition library to recognize emotions using the face image of the user taken by the camera mounted on the robot. The input is the face image of the user, and the output is the emotion data. Specifically, it analyzes whether the user is in a stressed state or a relaxed state.

[0394] Step 5:

[0395] Feedback adjustment

[0396] The server adjusts the feedback for the proofread character data based on the emotion recognition result. The input is the proofread character data and the emotion data, and the output is the adjusted feedback text. For example, it gives feedback such as "Thank you for your hard work. Great job." to a stressed user.

[0397] Step 6:

[0398] Display of results

[0399] The server sends the original character data, the corrected character data, and the adjusted feedback to the user's terminal and displays them on the screen. The input is the original character data, the corrected character data, and the feedback text, and the output is the result displayed on the user's terminal. Specifically, the original text "Ryjk10" and the corrected text "Ryk10" are compared and displayed, along with the adjusted feedback.

[0400] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0401] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0402] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0403] [Second embodiment]

[0404] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0405] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0406] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0407] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0408] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0410] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0411] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0412] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0413] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0414] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0415] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0416] The present invention relates to a system that extracts characters from an image, automatically corrects the character data, and displays the correction results on a user terminal. This system is composed of several main components, including an optical character recognition (OCR) unit, a text correction unit, and a storage unit.

[0417] First, a user can upload an image from their device. The image is preferably a handwritten note or a printed document. The device then sends the image to the server, which then stores the received image data in storage.

[0418] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data, which is temporarily stored on the server.

[0419] The server then sends this saved text data to the text correction module, which analyzes it for grammar, spelling mistakes, and style improvements and performs corrections. Once the corrections are complete, the text data is returned to the server, which calculates the differences between the original and corrected text data and marks them up.

[0420] Finally, the server sends the results to the user's device, which then displays a comparison of the original text and the corrected text. The user can see the differences on the screen and make further corrections if necessary.

[0421] Specific examples

[0422] User Operation

[0423] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0424] Server Processing

[0425] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0426] Correction process

[0427] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0428] Results display

[0429] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0430] Original text: "Hello. My name is Tanaka."

[0431] Corrected text: "Hello. My name is Tanaka."

[0432] This system automates the entire process from transcription to correction, allowing users to obtain high-quality text data quickly and efficiently.

[0433] The processing flow will be explained below.

[0434] Step 1:

[0435] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0436] Step 2:

[0437] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0438] Step 3:

[0439] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0440] Step 4:

[0441] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0442] Step 5:

[0443] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[0444] Step 6:

[0445] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[0446] Step 7:

[0447] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[0448] Step 8:

[0449] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[0450] Step 9:

[0451] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[0452] Example 1

[0453] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0454] Conventional optical character recognition systems have low accuracy, and corrections to character data extracted from handwritten or printed documents have been performed manually, which requires time and effort. Furthermore, grammar checks and style improvements after character recognition have not been automated, making the system unfriendly to users. The present invention aims to solve these problems by providing a system that performs character recognition and automatic corrections from images with high accuracy and efficiency.

[0455] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0456] In this invention, the server includes a means for users to upload images from their terminals, a means for transmitting the uploaded images to the server, and a means for the server to pass the image data stored therein to an optical character recognition module. This automates the series of processes of character recognition and text correction, enabling users to obtain high-quality text data quickly and efficiently.

[0457] A "terminal" is a device operated by a user, and includes devices such as smartphones, tablets, and PCs.

[0458] A "server" is a computer system that processes, stores, and transmits image and text data.

[0459] An "optical character recognition module" is a piece of software or hardware that extracts text information from an image and uses OCR (optical character recognition) technology.

[0460] A "text correction module" is a piece of software or hardware that checks the grammar, corrects spelling mistakes, and improves style of extracted text data.

[0461] "Storage means" refers to a storage device for temporarily or permanently storing image data or text data, and includes hard disks, SSDs, cloud storage, etc.

[0462] "Upload" refers to the operation in which a user sends data (images or files) from a terminal to another location (such as a server).

[0463] "Markup" is an operation that uses scripts and tags to highlight or decorate specific parts of text data.

[0464] "Comparative display" is an operation that displays two different text data side by side, allowing you to visually confirm the differences between them.

[0465] "Sending the results" means that the server transfers the processed data to the user's terminal, which is usually done over a network.

[0466] This invention is a system that extracts characters from images, automatically corrects the character data, and displays the correction results on the user's terminal. This system is realized mainly by users uploading images from their terminals and the server processing the image data.

[0467] System configuration

[0468] The system includes the following main components:

[0469] 1. User Device

[0470] 2. Server

[0471] 3. Optical Character Recognition Module (OCR Module)

[0472] 4. Text Correction Module

[0473] 5. Preservation means

[0474] Hardware and software used

[0475] User devices: smartphones, tablets, PCs, etc.

[0476] Server: A high-performance computer system

[0477] Optical Character Recognition Module: Google Cloud Vision API, Tesseract, etc.

[0478] Text correction module: Grammarly API, natural language processing model

[0479] Storage method: Cloud storage such as Amazon S3

[0480] Process Flow

[0481] User Operation

[0482] 1. Users take photos of handwritten notes or printed documents using their smartphones or PCs and upload the images through the application.

[0483] 2. The uploaded image is sent from the device to the server.

[0484] Server Processing

[0485] 1. The server temporarily stores the received image data. Cloud storage is recommended as the storage location.

[0486] 2. The stored image data is passed to the OCR module to extract text information. The OCR module uses optical character recognition technology to generate text data from the image.

[0487] 3. The extracted text data is temporarily stored on the server.

[0488] 4. The server sends the saved text data to the writing correction module, which corrects grammar, spelling, and style, and returns the corrected text data to the server.

[0489] 5. The server compares the original text data with the corrected text data and marks up the differences.

[0490] Results display

[0491] 1. The server sends the final correction results to the user's terminal.

[0492] 2. The user's device receives the transmitted data and displays a comparison of the original text and the corrected text, allowing the user to review the results and make further corrections as necessary.

[0493] Specific examples

[0494] Prompt Sentence Examples

[0495] "Upload an image of a handwritten note and automatically extract and correct the text. Example: The image says, 'Hello. My name is Tanaka.' Output the text 'Hello. My name is Tanaka.'"

[0496] This system allows users to obtain high-quality text data quickly and efficiently. In particular, the entire process from transcription to text correction is automated, significantly reducing the time and effort required.

[0497] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0498] Step 1:

[0499] A user can use a device to take a photo of a handwritten note or a printed document and upload the image through the application. The user selects an image and presses the upload button to start the image data upload operation. The input is the image file taken by the user, and the output is the uploaded image file data.

[0500] Step 2:

[0501] The terminal sends the image data uploaded by the user to the server. A secure connection is established using the HTTPS protocol, and the image data is sent to the server as an HTTP request. The input is the image data sent from the user terminal, and the output is the image data received by the server.

[0502] Step 3:

[0503] The server saves the received image data in storage. Cloud storage such as Amazon S3 is used as the storage. A storage path and file name are generated and the image data is saved. The input is the received image data, and the output is the path information for the image file saved in storage.

[0504] Step 4:

[0505] The server obtains the path information of the stored image data and passes it to the OCR module. The connection with the OCR module is made using the API key and authentication information, and the image data is sent to the OCR module's endpoint. The input is the path information of the image data stored in storage, and the output is the image data sent to the OCR module.

[0506] Step 5:

[0507] The OCR module analyzes the received image data and extracts text information from the image. It uses OCR technologies such as Google Cloud Vision API and Tesseract to analyze the text information and generate text data. The input is the image data sent to the OCR module, and the output is the extracted text data.

[0508] Step 6:

[0509] The server receives the text data returned from the OCR module and temporarily stores it. The text data is stored in a database or memory storage within the server. The input is the text data returned from the OCR module, and the output is the text data stored on the server.

[0510] Step 7:

[0511] The server sends the saved text data to the writing correction module. It connects to the writing correction module (e.g., Grammarly API) and sends the text data in JSON format. The input is the text data saved on the server, and the output is the text data sent to the writing correction module.

[0512] Step 8:

[0513] The writing correction module analyzes the received text data and improves grammar, spelling, and style. The corrected text data is returned to the server with markup. The input is the text data sent to the writing correction module, and the output is the corrected text data.

[0514] Step 9:

[0515] The server compares the original text data with the corrected text data and marks up the parts that have been corrected. It calculates the differences and formats them in a user-friendly format. The input is the original text data and the corrected text data, and the output is the marked-up comparison results.

[0516] Step 10:

[0517] The server sends the final marked-up text result to the user's terminal. The result data is formatted in a format such as HTML or PDF and sent as data for display. The input is the marked-up text result, and the output is the result data sent to the user's terminal.

[0518] Step 11:

[0519] The user checks the final result on the terminal and compares the original text with the corrected text. The differences can be seen at a glance on the displayed screen and further corrections can be made if necessary. The input is the result data sent from the server and the output is the displayed data that the user checks.

[0520] (Application example 1)

[0521] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0522] In recent years, many product price labels in brick-and-mortar stores are frequently updated, but this process is prone to human error. This can lead to incorrect pricing information, which can not only undermine customer trust but also affect sales. Furthermore, manually checking and correcting price labels is laborious and time-consuming for store staff, so an efficient system is needed. Currently available systems lack the ability to automatically detect and correct price label errors, so a new system is needed to address this issue.

[0523] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0524] In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, means for displaying the correction results on a user terminal, automatic correction means for comparing the extracted character data with the correction results and correcting errors, and means for displaying the results to the user in real time. This allows a store clerk to instantly detect incorrect price information and receive correction suggestions simply by taking a photo of a price label with their smartphone, thereby enabling efficient and accurate price label management.

[0525] The "optical character recognition means" is a means for extracting characters from image data and generating text data.

[0526] The "correction means" is a means for correcting grammar, spelling mistakes, and style of the extracted character data.

[0527] The "means for displaying on the user terminal" refers to a means for displaying the correction results sent from the server on the user's device.

[0528] The "automatic correction means for correcting errors" is an automated means for comparing extracted character data with corrected character data and correcting errors.

[0529] "Means for displaying in real time" refers to means for instantly displaying processed data on a user terminal.

[0530] The "storage means" is a means for temporarily storing image data and extracted character data.

[0531] "Markup means" refers to a means for visually showing the difference between corrected character data and the original character data.

[0532] The "means for detecting errors in price labels and suggesting corrections" is a means for detecting errors in product price information and suggesting correct price information.

[0533] The "means for transmitting to the server" is a means for transferring images uploaded by a user from a terminal to the server.

[0534] The "means for transmitting result data" is a means for transmitting processed data from the server to the user's device.

[0535] The present invention relates to a system for automatically detecting and correcting price label errors in brick-and-mortar stores, which includes optical character recognition, text correction, and automatic correction.

[0536] First, the user takes a picture of the price label using their smartphone, which then sends the image to the server, which then temporarily stores the image data.

[0537] The server then invokes the Optical Character Recognition (OCR) module, passing the image data as input. This OCR module uses OpenCV and Tesseract to extract text data from the image, which is temporarily stored on the server.

[0538] The server then sends the saved text data to a writing correction module, which uses TextBlob to correct grammar, spelling, and style errors, and then returns the corrected text data to the server.

[0539] The server corrects any errors found and marks up the difference between the original and corrected text data. During this process, if the original price information is incorrect, the correct price information is suggested and displayed to the user in real time.

[0540] The user terminal receives and displays the result data sent from the server, allowing the user to immediately check the price label errors in the store and make any necessary corrections.

[0541] For example, the following prompts are used:

[0542] Example prompt sentence:

[0543] "Check the price label on the product. It might say '50,000 yen' instead of '500 yen'."

[0544] This system enables store clerks who manage price labels in physical stores to efficiently and accurately check and correct price information. The hardware used is a smartphone and a camera, and the software uses OpenCV, Tesseract, and TextBlob. This process creates a system that can automatically correct errors in price labels in real time.

[0545] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0546] Step 1:

[0547] A user takes a picture of the price label using a smartphone device.

[0548] Input: Price Label Image

[0549] Output: Image data of the photographed price label

[0550] What happens: The user opens the camera app on their smartphone, focuses on the price label, and takes a photo. The image is then saved to their device.

[0551] Step 2:

[0552] The terminal transmits the captured image to the server.

[0553] Input: Price label image data

[0554] Output: Image data sent to the server

[0555] Specific operation: The device selects the captured image and uploads it to the server via the app. The communication protocol is HTTPS.

[0556] Step 3:

[0557] The server temporarily stores the received image data.

[0558] Input: Received image data

[0559] Output: Saved image data

[0560] Specific operation: The server stores the received image data in a database or temporary file. The storage location is a storage system (e.g., Amazon S3).

[0561] Step 4:

[0562] The server passes the image data to the OCR module to extract the character information.

[0563] Input: Image data

[0564] Output: Extracted text data

[0565] Specific operation: The server uses OpenCV to preprocess the image data, then uses Tesseract to extract text information. The extracted text data is temporarily saved.

[0566] Step 5:

[0567] The server sends the extracted text data to the text correction module.

[0568] Input: Text data

[0569] Output: Corrected text data

[0570] Specific operation: The server analyzes the text data using TextBlob, corrects grammar, spelling mistakes, and style, and saves the corrected text data.

[0571] Step 6:

[0572] The server marks up the differences between the corrected text data and the original text data, and generates correction information for correcting the errors.

[0573] Input: Original text data and corrected text data

[0574] Output: Marked up text data and correction information

[0575] What it does: The server compares each piece of text data and visually displays the differences using a program that highlights the differences (e.g., a diff tool).

[0576] Step 7:

[0577] The server transmits the correction information and the marked-up text data to the user terminal.

[0578] Input: Correction information, marked-up text data

[0579] Output: Information sent to the user's device

[0580] Specific operation: The server prepares the marked-up text data and the correction information for the error correction as a data packet and sends it to the user terminal via the HTTPS protocol.

[0581] Step 8:

[0582] The user terminal displays the received information on the screen, and notifies the user of any price label errors in real time.

[0583] Input: Received correction information, marked-up text data

[0584] Output: Screen showing suggested revisions and original pricing information

[0585] Specific operation: Based on the received information, the user device updates the GUI to display the error and suggested corrections on the screen. The user can check this and manually correct it if necessary.

[0586] This allows store staff to efficiently and accurately detect price label errors and make corrections in real time.

[0587] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0588] This invention relates to a system that extracts characters from images, automatically corrects the character data, recognizes the user's emotions, and reflects them in the correction results. This system is composed of several main components, including an optical character recognition (OCR) means, a text correction means, an emotion engine, and a storage means.

[0589] A user can first upload an image from their device, preferably a handwritten note or a printed document, and the device sends the image to the server, which then stores the received image data in temporary storage.

[0590] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data. This text data is temporarily stored on the server. The server then sends the stored text data to the writing correction module, which analyzes the text for grammar, spelling mistakes, and style improvements and performs the correction process.

[0591] The server then calls an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. For example, it adjusts the correction results based on the user's emotional state, such as whether they are stressed or relaxed.

[0592] The adjusted correction results are sent back to the server, which compares the original text data with the corrected text data, calculates the differences, and marks them up. Finally, the server sends the results to the user's device, which receives them and displays a comparison of the original text and the corrected text. The user can check the differences on the screen and make further corrections if necessary.

[0593] Specific examples

[0594] User Operation

[0595] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0596] Server Processing

[0597] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0598] Correction process

[0599] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0600] Emotion Recognition Processing

[0601] The server calls the emotion engine to analyze the user's emotional state. For example, if the user is nervous, the tone of the corrections will be gentler and the expressions will be gentler. Conversely, if the user is relaxed, more specific comments will be made.

[0602] Results display

[0603] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0604] Original text: "Hello. My name is Tanaka."

[0605] Corrected text: "Hello. My name is Tanaka."

[0606] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[0607] The processing flow will be explained below.

[0608] Step 1:

[0609] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0610] Step 2:

[0611] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0612] Step 3:

[0613] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0614] Step 4:

[0615] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0616] Step 5:

[0617] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[0618] Step 6:

[0619] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[0620] Step 7:

[0621] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[0622] Step 8:

[0623] The server then calls the emotion engine to analyze the user's current emotional state, which uses facial recognition data and previous typing patterns to determine whether the user is tense or relaxed.

[0624] Step 9:

[0625] The server adjusts the tone and style of its corrections based on the data it receives from the emotion engine, for example, increasing kind language and positive feedback if the user is nervous.

[0626] Step 10:

[0627] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[0628] Step 11:

[0629] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[0630] Example 2

[0631] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0632] Conventional systems simply extract text from image data and correct the text, but this does not necessarily mean that the corrections are appropriate for the user's feelings. Furthermore, feedback that does not take the user's feelings into account is often ineffective. Furthermore, it is difficult to efficiently process image data uploaded by users and quickly display the results.

[0633] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an optical character recognition means for extracting characters from an image, a correction means for correcting the extracted character data, an emotion recognition means for recognizing the user's emotion and adjusting the correction result based on the emotion, and a means for displaying the correction result on the user terminal. This enables appropriate correction of the text taking the user's emotion into consideration and rapid feedback.

[0634] "Optical character recognition" refers to techniques and devices for extracting character information from image data.

[0635] "Text correction means" refers to the technology and devices that analyze extracted text data from the perspectives of grammar, spelling mistakes, and style, and make appropriate corrections.

[0636] "Emotion recognition means" refers to technology and devices for analyzing a user's emotions and adjusting the results of data processing according to that emotional state.

[0637] "Storage means" refers to devices or technologies for temporarily or permanently storing extracted character data or intermediate processed data.

[0638] "Means for marking up differences" refers to techniques and devices for visually highlighting the differences between the original text data and the modified text data.

[0639] A "user terminal" is a hardware device operated by a user, and includes devices such as smartphones and personal computers.

[0640] "Server" refers to a computer system for processing, storing, and transmitting data.

[0641] This system extracts text from an image, automatically corrects the text, and recognizes the user's emotions and reflects them in the correction results. This system consists of the following main components:

[0642] Key Components

[0643] 1. Optical Character Recognition (OCR):

[0644] This is a technology that extracts text information from image data. Specifically, it uses an OCR engine. For example, Google Cloud Vision API can be used.

[0645] 2. Writing correction methods:

[0646] This technology analyzes the extracted text data from the perspectives of grammar, spelling mistakes, and style, and makes appropriate corrections. For example, you can use a grammar checker API such as Grammarly.

[0647] 3. Emotion recognition means:

[0648] This technology analyzes a user's emotions and adjusts the results of processing data based on their emotional state. It includes algorithms that analyze facial recognition information and text input patterns. This allows it to adjust the tone or expressions to be gentler if the user is nervous, or more specific if the user is relaxed.

[0649] 4. Preservation means:

[0650] This is a technology for temporarily storing extracted text data and intermediate processed data. For example, data can be stored using a cloud storage service.

[0651] 5. Ways to mark up differences:

[0652] This technology visually highlights the differences between the original text and the edited text. It uses a comparison algorithm to highlight the changes.

[0653] Specific examples of operations

[0654] User Operation

[0655] A user can use a smartphone to take a photo of a handwritten note and upload it through the app. For example, consider a handwritten note that reads, "Hello. My name is Tanaka."

[0656] Server Processing and Response

[0657] The server passes the received image to the OCR module. The OCR module extracts character information from the image and generates text data such as "Hello. My name is Tanaka." This text data is saved on the server. The server then sends the saved text data to the writing correction module. The writing correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0658] Performing emotion recognition

[0659] The server calls the emotion engine and analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. If the user is nervous, the corrections will be made in a gentler tone and with gentler expressions. Conversely, if the user is relaxed, more specific corrections will be made.

[0660] Displaying the results

[0661] The server calculates the difference between the original text data and the corrected text data, marks it up, and finally sends the result to the user's terminal, which displays a comparison of the original text and the corrected text on the screen.

[0662] Prompt Sentence Examples

[0663] "Please correct the following sentence: 'Hello. My name is Tanaka.' If the user is relaxed, provide specific feedback."

[0664] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[0665] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0666] Step 1:

[0667] Uploading image data

[0668] User:

[0669] A user uses an application on their smartphone or computer to take an image of a handwritten note or printed document, and then uploads it to a server via the application. For example, consider an image of a handwritten note that reads, "Hello. My name is Tanaka." This image data becomes the input.

[0670] output:

[0671] Uploaded image data.

[0672] Step 2:

[0673] Image data sent to server and saved

[0674] Device:

[0675] The device sends the uploaded image data to the server. Specifically, the device sends the image data to the specified URL of the server via an HTTP POST request.

[0676] server:

[0677] The server stores the received image data in temporary storage. For example, if a cloud storage service is used, the uploaded image is stored in the cloud storage. This image data becomes the input for the next step.

[0678] output:

[0679] Image data stored in temporary storage.

[0680] Step 3:

[0681] Extracting character information using OCR

[0682] server:

[0683] The server calls the OCR module and passes the image data stored in temporary storage as input. The OCR module analyzes the image data and extracts text information. For example, the text data "Hello. My name is Tanaka" is extracted from the image.

[0684] output:

[0685] Extracted text data. The text generated is "Hello. My name is Tanaka."

[0686] Step 4:

[0687] Execution of text correction

[0688] server:

[0689] The server sends the generated text data to a writing correction module. For example, the text data is sent to the Grammarly API, which analyzes it for grammar, spelling mistakes, and style improvements. The writing correction module corrects "Hello" to "Hello" and checks the grammar of "I'm Tanaka."

[0690] output:

[0691] Corrected text data. The corrected text "Hello. My name is Tanaka." is generated.

[0692] Step 5:

[0693] User Emotion Recognition

[0694] server:

[0695] The server calls the emotion engine and analyzes facial recognition information and text input speed information from the user's device to identify the user's emotional state. If the analysis shows that the user is nervous, the correction tone will be softened, and if the user is relaxed, the correction will be more specific.

[0696] output:

[0697] Tailored corrections, such as gentler language in sentences.

[0698] Step 6:

[0699] Diff calculation and markup

[0700] server:

[0701] The server calculates the difference between the original text data and the corrected text data and marks it up, using a comparison algorithm to highlight the part where "Hello" was changed to "Hello."

[0702] output:

[0703] Text data with differences marked up.

[0704] Step 7:

[0705] Display the results on your terminal

[0706] server:

[0707] The server sends the resulting data after the difference calculation to the user terminal. Specifically, it returns JSON format data in an HTTP response.

[0708] Device:

[0709] The user's device receives this and displays the original text and the corrected text on the screen for comparison. For example, the part where "Hello" has been changed to "Hello" is highlighted so that the user can easily confirm it.

[0710] output:

[0711] Text data displayed on the screen before and after correction.

[0712] This series of processes allows the user to quickly and efficiently obtain high-quality text data.

[0713] (Application example 2)

[0714] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0715] Existing optical character recognition technology and text correction systems have difficulty in accurately and efficiently recognizing and correcting text during maintenance work in factories. Furthermore, there is a lack of a means to provide appropriate feedback based on the emotional state of maintenance personnel. As a result, misrecognition and miscorrection occur frequently, leading to a decline in maintenance efficiency and work accuracy.

[0716] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, and emotion recognition means for recognizing the user's emotion and reflecting it in the correction results. This makes it possible to perform accurate character recognition and correction during maintenance work in a factory, and to provide appropriate feedback based on the emotional state of the maintenance worker.

[0717] An "image" is a digital recording of visual information.

[0718] "Characters" are symbols or codes used to represent language.

[0719] "Optical character recognition" is a technology for extracting character information from an image.

[0720] "Character data" is character information extracted by character recognition means expressed in digital form.

[0721] "Text correction means" is a technology for improving grammar, spelling mistakes, and style of text based on extracted character data.

[0722] "Emotion recognition means" refers to technology for analyzing and identifying a user's emotional state.

[0723] A "user terminal" is a device operated by a user, such as a smartphone or tablet.

[0724] "Storage means" refers to a device for temporarily or permanently storing character data.

[0725] A "server" is a computer system that processes and stores data on a network.

[0726] "Temporarily storing" means retaining data for a specific period of time.

[0727] "Grammar" is the rules and structure of a language.

[0728] A "spelling error" is an error in the spelling of a word.

[0729] "Style improvements" are modifications made to improve the readability and appearance of the text.

[0730] "Marking up the differences" means visually indicating the differences between the original character data and the corrected character data.

[0731] "Uploading" means sending data from a user terminal to a server.

[0732] "Result data" refers to data that includes corrected character data and analysis results.

[0733] A "factory robot" is a mechanical device that automatically performs various tasks in a factory.

[0734] An "application" is a software program with a specific function.

[0735] This invention is a system for streamlining maintenance work in factories and providing accurate character recognition and emotion-based feedback. The main components of the system include a server, a terminal (e.g., a factory robot), and a user (a maintenance worker).

[0736] The server is equipped with an optical character recognition unit that extracts characters from images, a text correction unit, and an emotion recognition unit. The terminal is a device operated by a user, which in this case corresponds to a robot deployed in a factory. The user provides the system with the information needed during maintenance work through the terminal.

[0737] Overview of Program Processing

[0738] Hardware and Software

[0739] The server is a computer system equipped with a high-performance processor and sufficient memory. Also, the following software is used.

[0740] OpenCV: Image Processing Library

[0741] pytesseract: Optical Character Recognition Library

[0742] TextBlob: Natural Language Processing Library

[0743] EmotionRecognition: Virtual Library for Emotion Recognition

[0744] Processing Flow of the System

[0745] 1. Image Capture:

[0746] The factory robot takes pictures of the equipment and parts to be maintained. For example, it takes a picture of the wiring diagram with the robot's camera.

[0747] 2. Character Extraction:

[0748] The server receives the captured image and extracts character information using pytesseract. For example, it extracts the label "Ryk10 number" from the wiring diagram.

[0749] 3. Text Proofreading:

[0750] The server proofreads grammar and spelling mistakes of the extracted character information using TextBlob. For example, it corrects "Ryjk10 number" to "Ryk10 number".

[0751] 4. Emotion Recognition:

[0752] The server uses EmotionRecognition to recognize emotions from the user's facial images taken by the robot's onboard camera, for example, to detect whether the user is in a stressed or relaxed state.

[0753] 5. Feedback adjustment:

[0754] Depending on the user's expected emotional state, the feedback can be adjusted to provide gentler and more specific feedback. For example, if a user is feeling stressed, the feedback could be, "Thank you for your hard work. You're doing a great job."

[0755] 6. Results display:

[0756] The server sends the original text, the corrected text, and the adjusted feedback to the user's terminal and displays them on the screen.

[0757] Adding specific examples

[0758] For example, if the robot photographs a wiring diagram of a device and recognizes the text as "Ryjk10," it will correct this misrecognition to "Ryk10" using TextBlob. At the same time, if it detects that a maintenance worker is under stress using emotion recognition technology, it will provide feedback such as "Great job, great job."

[0759] Example prompts for generative AI models

[0760] Image: [Link to wiring diagram image]

[0761] Caption: Is there any misidentified text in this wiring diagram? Please correct it, especially the part where it is misidentified as "Ryjk10".

[0762] Emotional state: Stressed

[0763] In this way, the system can streamline maintenance work in the factory and provide highly accurate character recognition and feedback according to emotions.

[0764] The flow of the specific process in Application Example 2 will be described using FIG. 14.

[0765] Step 1:

[0766] Image capture

[0767] The user uses the camera of the factory robot to take pictures of the equipment and parts to be maintained. The input is the image file taken by the user, and the output is this image data. This image data is required for the subsequent character recognition process.

[0768] Step 2:

[0769] Character extraction

[0770] The server receives the image data transmitted from the factory robot and extracts character information using the pytesseract library. The input is the image data, and the output is the character data. Specifically, it extracts texts such as the label "Ryk10" written in the wiring diagram in the image.

[0771] Step 3:

[0772] Text correction

[0773] The server corrects grammar and spelling mistakes in the extracted character information using the TextBlob library. The input is the extracted character data, and the output is the corrected character data. As a specific operation, it corrects the misrecognized character "Ryjk10" to "Ryk10".

[0774] Step 4:

[0775] Emotion recognition

[0776] The server uses the user's facial image captured by the robot's camera to recognize emotions using the EmotionRecognition library. The input is the user's facial image, and the output is emotional data. Specifically, it analyzes whether the user is in a stressed or relaxed state.

[0777] Step 5:

[0778] Feedback Adjustment

[0779] The server adjusts the feedback for the corrected text data based on the emotion recognition results. The input is the corrected text data and emotion data, and the output is the adjusted feedback text. For example, to a user who is under stress, the server may give the feedback, "Thank you for your hard work. You did a great job."

[0780] Step 6:

[0781] Displaying the results

[0782] The server sends the original character data, the corrected character data, and the adjusted feedback to the user's terminal and displays them on the screen. The input is the original character data, the corrected character data, and the feedback text, and the output is the result displayed on the user's terminal. Specifically, the original text "Ryjk10" and the corrected text "Ryk10" are compared and displayed, along with the adjusted feedback.

[0783] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0784] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0785] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0786] [Third embodiment]

[0787] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0788] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0789] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0790] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0791] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0792] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0793] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0794] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0795] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0796] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0797] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0798] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0799] The present invention relates to a system that extracts characters from an image, automatically corrects the character data, and displays the correction results on a user terminal. This system is composed of several main components, including an optical character recognition (OCR) unit, a text correction unit, and a storage unit.

[0800] First, a user can upload an image from their device. The image is preferably a handwritten note or a printed document. The device then sends the image to the server, which then stores the received image data in storage.

[0801] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data, which is temporarily stored on the server.

[0802] The server then sends this saved text data to the text correction module, which analyzes it for grammar, spelling mistakes, and style improvements and performs corrections. Once the corrections are complete, the text data is returned to the server, which calculates the differences between the original and corrected text data and marks them up.

[0803] Finally, the server sends the results to the user's device, which then displays a comparison of the original text and the corrected text. The user can see the differences on the screen and make further corrections if necessary.

[0804] Specific examples

[0805] User Operation

[0806] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0807] Server Processing

[0808] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0809] Correction process

[0810] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0811] Results display

[0812] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0813] Original text: "Hello. My name is Tanaka."

[0814] Corrected text: "Hello. My name is Tanaka."

[0815] This system automates the entire process from transcription to correction, allowing users to obtain high-quality text data quickly and efficiently.

[0816] The processing flow will be explained below.

[0817] Step 1:

[0818] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0819] Step 2:

[0820] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0821] Step 3:

[0822] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0823] Step 4:

[0824] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0825] Step 5:

[0826] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[0827] Step 6:

[0828] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[0829] Step 7:

[0830] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[0831] Step 8:

[0832] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[0833] Step 9:

[0834] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[0835] Example 1

[0836] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0837] Conventional optical character recognition systems have low accuracy, and corrections to character data extracted from handwritten or printed documents have been performed manually, which requires time and effort. Furthermore, grammar checks and style improvements after character recognition have not been automated, making the system unfriendly to users. The present invention aims to solve these problems by providing a system that performs character recognition and automatic corrections from images with high accuracy and efficiency.

[0838] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0839] In this invention, the server includes a means for users to upload images from their terminals, a means for transmitting the uploaded images to the server, and a means for the server to pass the image data stored therein to an optical character recognition module. This automates the series of processes of character recognition and text correction, enabling users to obtain high-quality text data quickly and efficiently.

[0840] A "terminal" is a device operated by a user, and includes devices such as smartphones, tablets, and PCs.

[0841] A "server" is a computer system that processes, stores, and transmits image and text data.

[0842] An "optical character recognition module" is a piece of software or hardware that extracts text information from an image and uses OCR (optical character recognition) technology.

[0843] A "text correction module" is a piece of software or hardware that checks the grammar, corrects spelling mistakes, and improves style of extracted text data.

[0844] "Storage means" refers to a storage device for temporarily or permanently storing image data or text data, and includes hard disks, SSDs, cloud storage, etc.

[0845] "Upload" refers to the operation in which a user sends data (images or files) from a terminal to another location (such as a server).

[0846] "Markup" is an operation that uses scripts and tags to highlight or decorate specific parts of text data.

[0847] "Comparative display" is an operation that displays two different text data side by side, allowing you to visually confirm the differences between them.

[0848] "Sending the results" means that the server transfers the processed data to the user's terminal, which is usually done over a network.

[0849] This invention is a system that extracts characters from images, automatically corrects the character data, and displays the correction results on the user's terminal. This system is realized mainly by users uploading images from their terminals and the server processing the image data.

[0850] System configuration

[0851] The system includes the following main components:

[0852] 1. User Device

[0853] 2. Server

[0854] 3. Optical Character Recognition Module (OCR Module)

[0855] 4. Text Correction Module

[0856] 5. Preservation means

[0857] Hardware and software used

[0858] User devices: smartphones, tablets, PCs, etc.

[0859] Server: A high-performance computer system

[0860] Optical Character Recognition Module: Google Cloud Vision API, Tesseract, etc.

[0861] Text correction module: Grammarly API, natural language processing model

[0862] Storage method: Cloud storage such as Amazon S3

[0863] Process Flow

[0864] User Operation

[0865] 1. Users take photos of handwritten notes or printed documents using their smartphones or PCs and upload the images through the application.

[0866] 2. The uploaded image is sent from the device to the server.

[0867] Server Processing

[0868] 1. The server temporarily stores the received image data. Cloud storage is recommended as the storage location.

[0869] 2. The stored image data is passed to the OCR module to extract text information. The OCR module uses optical character recognition technology to generate text data from the image.

[0870] 3. The extracted text data is temporarily stored on the server.

[0871] 4. The server sends the saved text data to the writing correction module, which corrects grammar, spelling, and style, and returns the corrected text data to the server.

[0872] 5. The server compares the original text data with the corrected text data and marks up the differences.

[0873] Results display

[0874] 1. The server sends the final correction results to the user's terminal.

[0875] 2. The user's device receives the transmitted data and displays a comparison of the original text and the corrected text, allowing the user to review the results and make further corrections as necessary.

[0876] Specific examples

[0877] Prompt Sentence Examples

[0878] "Upload an image of a handwritten note and automatically extract and correct the text. Example: The image says, 'Hello. My name is Tanaka.' Output the text 'Hello. My name is Tanaka.'"

[0879] This system allows users to obtain high-quality text data quickly and efficiently. In particular, the entire process from transcription to text correction is automated, significantly reducing the time and effort required.

[0880] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0881] Step 1:

[0882] A user can use a device to take a photo of a handwritten note or a printed document and upload the image through the application. The user selects an image and presses the upload button to start the image data upload operation. The input is the image file taken by the user, and the output is the uploaded image file data.

[0883] Step 2:

[0884] The terminal sends the image data uploaded by the user to the server. A secure connection is established using the HTTPS protocol, and the image data is sent to the server as an HTTP request. The input is the image data sent from the user terminal, and the output is the image data received by the server.

[0885] Step 3:

[0886] The server saves the received image data in storage. Cloud storage such as Amazon S3 is used as the storage. A storage path and file name are generated and the image data is saved. The input is the received image data, and the output is the path information for the image file saved in storage.

[0887] Step 4:

[0888] The server obtains the path information of the stored image data and passes it to the OCR module. The connection with the OCR module is made using the API key and authentication information, and the image data is sent to the OCR module's endpoint. The input is the path information of the image data stored in storage, and the output is the image data sent to the OCR module.

[0889] Step 5:

[0890] The OCR module analyzes the received image data and extracts text information from the image. It uses OCR technologies such as Google Cloud Vision API and Tesseract to analyze the text information and generate text data. The input is the image data sent to the OCR module, and the output is the extracted text data.

[0891] Step 6:

[0892] The server receives the text data returned from the OCR module and temporarily stores it. The text data is stored in a database or memory storage within the server. The input is the text data returned from the OCR module, and the output is the text data stored on the server.

[0893] Step 7:

[0894] The server sends the saved text data to the writing correction module. It connects to the writing correction module (e.g., Grammarly API) and sends the text data in JSON format. The input is the text data saved on the server, and the output is the text data sent to the writing correction module.

[0895] Step 8:

[0896] The writing correction module analyzes the received text data and improves grammar, spelling, and style. The corrected text data is returned to the server with markup. The input is the text data sent to the writing correction module, and the output is the corrected text data.

[0897] Step 9:

[0898] The server compares the original text data with the corrected text data and marks up the parts that have been corrected. It calculates the differences and formats them in a user-friendly format. The input is the original text data and the corrected text data, and the output is the marked-up comparison results.

[0899] Step 10:

[0900] The server sends the final marked-up text result to the user's terminal. The result data is formatted in a format such as HTML or PDF and sent as data for display. The input is the marked-up text result, and the output is the result data sent to the user's terminal.

[0901] Step 11:

[0902] The user checks the final result on the terminal and compares the original text with the corrected text. The differences can be seen at a glance on the displayed screen and further corrections can be made if necessary. The input is the result data sent from the server and the output is the displayed data that the user checks.

[0903] (Application example 1)

[0904] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0905] In recent years, many product price labels in brick-and-mortar stores are frequently updated, but this process is prone to human error. This can lead to incorrect pricing information, which can not only undermine customer trust but also affect sales. Furthermore, manually checking and correcting price labels is laborious and time-consuming for store staff, so an efficient system is needed. Currently available systems lack the ability to automatically detect and correct price label errors, so a new system is needed to address this issue.

[0906] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0907] In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, means for displaying the correction results on a user terminal, automatic correction means for comparing the extracted character data with the correction results and correcting errors, and means for displaying the results to the user in real time. This allows a store clerk to instantly detect incorrect price information and receive correction suggestions simply by taking a photo of a price label with their smartphone, thereby enabling efficient and accurate price label management.

[0908] The "optical character recognition means" is a means for extracting characters from image data and generating text data.

[0909] The "correction means" is a means for correcting grammar, spelling mistakes, and style of the extracted character data.

[0910] The "means for displaying on the user terminal" refers to a means for displaying the correction results sent from the server on the user's device.

[0911] The "automatic correction means for correcting errors" is an automated means for comparing extracted character data with corrected character data and correcting errors.

[0912] "Means for displaying in real time" refers to means for instantly displaying processed data on a user terminal.

[0913] The "storage means" is a means for temporarily storing image data and extracted character data.

[0914] "Markup means" refers to a means for visually showing the difference between corrected character data and the original character data.

[0915] The "means for detecting errors in price labels and suggesting corrections" is a means for detecting errors in product price information and suggesting correct price information.

[0916] The "means for transmitting to the server" is a means for transferring images uploaded by a user from a terminal to the server.

[0917] The "means for transmitting result data" is a means for transmitting processed data from the server to the user's device.

[0918] The present invention relates to a system for automatically detecting and correcting price label errors in brick-and-mortar stores, which includes optical character recognition, text correction, and automatic correction.

[0919] First, the user takes a picture of the price label using their smartphone, which then sends the image to the server, which then temporarily stores the image data.

[0920] The server then invokes the Optical Character Recognition (OCR) module, passing the image data as input. This OCR module uses OpenCV and Tesseract to extract text data from the image, which is temporarily stored on the server.

[0921] The server then sends the saved text data to a writing correction module, which uses TextBlob to correct grammar, spelling, and style errors, and then returns the corrected text data to the server.

[0922] The server corrects any errors found and marks up the difference between the original and corrected text data. During this process, if the original price information is incorrect, the correct price information is suggested and displayed to the user in real time.

[0923] The user terminal receives and displays the result data sent from the server, allowing the user to immediately check the price label errors in the store and make any necessary corrections.

[0924] For example, the following prompts are used:

[0925] Example prompt sentence:

[0926] "Check the price label on the product. It might say '50,000 yen' instead of '500 yen'."

[0927] This system enables store clerks who manage price labels in physical stores to efficiently and accurately check and correct price information. The hardware used is a smartphone and a camera, and the software uses OpenCV, Tesseract, and TextBlob. This process creates a system that can automatically correct errors in price labels in real time.

[0928] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0929] Step 1:

[0930] A user takes a picture of the price label using a smartphone device.

[0931] Input: Price Label Image

[0932] Output: Image data of the photographed price label

[0933] What happens: The user opens the camera app on their smartphone, focuses on the price label, and takes a photo. The image is then saved to their device.

[0934] Step 2:

[0935] The terminal transmits the captured image to the server.

[0936] Input: Price label image data

[0937] Output: Image data sent to the server

[0938] Specific operation: The device selects the captured image and uploads it to the server via the app. The communication protocol is HTTPS.

[0939] Step 3:

[0940] The server temporarily stores the received image data.

[0941] Input: Received image data

[0942] Output: Saved image data

[0943] Specific operation: The server stores the received image data in a database or temporary file. The storage location is a storage system (e.g., Amazon S3).

[0944] Step 4:

[0945] The server passes the image data to the OCR module to extract the character information.

[0946] Input: Image data

[0947] Output: Extracted text data

[0948] Specific operation: The server uses OpenCV to preprocess the image data, then uses Tesseract to extract text information. The extracted text data is temporarily saved.

[0949] Step 5:

[0950] The server sends the extracted text data to the text correction module.

[0951] Input: Text data

[0952] Output: Corrected text data

[0953] Specific operation: The server analyzes the text data using TextBlob, corrects grammar, spelling mistakes, and style, and saves the corrected text data.

[0954] Step 6:

[0955] The server marks up the differences between the corrected text data and the original text data, and generates correction information for correcting the errors.

[0956] Input: Original text data and corrected text data

[0957] Output: Marked up text data and correction information

[0958] What it does: The server compares each piece of text data and visually displays the differences using a program that highlights the differences (e.g., a diff tool).

[0959] Step 7:

[0960] The server transmits the correction information and the marked-up text data to the user terminal.

[0961] Input: Correction information, marked-up text data

[0962] Output: Information sent to the user's device

[0963] Specific operation: The server prepares the marked-up text data and the correction information for the error correction as a data packet and sends it to the user terminal via the HTTPS protocol.

[0964] Step 8:

[0965] The user terminal displays the received information on the screen, and notifies the user of any price label errors in real time.

[0966] Input: Received correction information, marked-up text data

[0967] Output: Screen showing suggested revisions and original pricing information

[0968] Specific operation: Based on the received information, the user device updates the GUI to display the error and suggested corrections on the screen. The user can check this and manually correct it if necessary.

[0969] This allows store staff to efficiently and accurately detect price label errors and make corrections in real time.

[0970] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0971] This invention relates to a system that extracts characters from images, automatically corrects the character data, recognizes the user's emotions, and reflects them in the correction results. This system is composed of several main components, including an optical character recognition (OCR) means, a text correction means, an emotion engine, and a storage means.

[0972] A user can first upload an image from their device, preferably a handwritten note or a printed document, and the device sends the image to the server, which then stores the received image data in temporary storage.

[0973] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data. This text data is temporarily stored on the server. The server then sends the stored text data to the writing correction module, which analyzes the text for grammar, spelling mistakes, and style improvements and performs the correction process.

[0974] The server then calls an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. For example, it adjusts the correction results based on the user's emotional state, such as whether they are stressed or relaxed.

[0975] The adjusted correction results are sent back to the server, which compares the original text data with the corrected text data, calculates the differences, and marks them up. Finally, the server sends the results to the user's device, which receives them and displays a comparison of the original text and the corrected text. The user can check the differences on the screen and make further corrections if necessary.

[0976] Specific examples

[0977] User Operation

[0978] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[0979] Server Processing

[0980] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[0981] Correction process

[0982] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[0983] Emotion Recognition Processing

[0984] The server calls the emotion engine to analyze the user's emotional state. For example, if the user is nervous, the tone of the corrections will be gentler and the expressions will be gentler. Conversely, if the user is relaxed, more specific comments will be made.

[0985] Results display

[0986] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[0987] Original text: "Hello. My name is Tanaka."

[0988] Corrected text: "Hello. My name is Tanaka."

[0989] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[0990] The processing flow will be explained below.

[0991] Step 1:

[0992] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[0993] Step 2:

[0994] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[0995] Step 3:

[0996] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[0997] Step 4:

[0998] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[0999] Step 5:

[1000] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[1001] Step 6:

[1002] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[1003] Step 7:

[1004] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[1005] Step 8:

[1006] The server then calls the emotion engine to analyze the user's current emotional state, which uses facial recognition data and previous typing patterns to determine whether the user is tense or relaxed.

[1007] Step 9:

[1008] The server adjusts the tone and style of its corrections based on the data it receives from the emotion engine, for example, increasing kind language and positive feedback if the user is nervous.

[1009] Step 10:

[1010] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[1011] Step 11:

[1012] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[1013] Example 2

[1014] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1015] Conventional systems simply extract text from image data and correct the text, but this does not necessarily mean that the corrections are appropriate for the user's feelings. Furthermore, feedback that does not take the user's feelings into account is often ineffective. Furthermore, it is difficult to efficiently process image data uploaded by users and quickly display the results.

[1016] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an optical character recognition means for extracting characters from an image, a correction means for correcting the extracted character data, an emotion recognition means for recognizing the user's emotion and adjusting the correction result based on the emotion, and a means for displaying the correction result on the user terminal. This enables appropriate correction of the text taking the user's emotion into consideration and rapid feedback.

[1017] "Optical character recognition" refers to techniques and devices for extracting character information from image data.

[1018] "Text correction means" refers to the technology and devices that analyze extracted text data from the perspectives of grammar, spelling mistakes, and style, and make appropriate corrections.

[1019] "Emotion recognition means" refers to technology and devices for analyzing a user's emotions and adjusting the results of data processing according to that emotional state.

[1020] "Storage means" refers to devices or technologies for temporarily or permanently storing extracted character data or intermediate processed data.

[1021] "Means for marking up differences" refers to techniques and devices for visually highlighting the differences between the original text data and the modified text data.

[1022] A "user terminal" is a hardware device operated by a user, and includes devices such as smartphones and personal computers.

[1023] "Server" refers to a computer system for processing, storing, and transmitting data.

[1024] This system extracts text from an image, automatically corrects the text, and recognizes the user's emotions and reflects them in the correction results. This system consists of the following main components:

[1025] Key Components

[1026] 1. Optical Character Recognition (OCR):

[1027] This is a technology that extracts text information from image data. Specifically, it uses an OCR engine. For example, Google Cloud Vision API can be used.

[1028] 2. Writing correction methods:

[1029] This technology analyzes the extracted text data from the perspectives of grammar, spelling mistakes, and style, and makes appropriate corrections. For example, you can use a grammar checker API such as Grammarly.

[1030] 3. Emotion recognition means:

[1031] This technology analyzes a user's emotions and adjusts the results of processing data based on their emotional state. It includes algorithms that analyze facial recognition information and text input patterns. This allows it to adjust the tone or expressions to be gentler if the user is nervous, or more specific if the user is relaxed.

[1032] 4. Preservation means:

[1033] This is a technology for temporarily storing extracted text data and intermediate processed data. For example, data can be stored using a cloud storage service.

[1034] 5. Ways to mark up differences:

[1035] This technology visually highlights the differences between the original text and the edited text. It uses a comparison algorithm to highlight the changes.

[1036] Specific examples of operations

[1037] User Operation

[1038] A user can use a smartphone to take a photo of a handwritten note and upload it through the app. For example, consider a handwritten note that reads, "Hello. My name is Tanaka."

[1039] Server Processing and Response

[1040] The server passes the received image to the OCR module. The OCR module extracts character information from the image and generates text data such as "Hello. My name is Tanaka." This text data is saved on the server. The server then sends the saved text data to the writing correction module. The writing correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[1041] Performing emotion recognition

[1042] The server calls the emotion engine and analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. If the user is nervous, the corrections will be made in a gentler tone and with gentler expressions. Conversely, if the user is relaxed, more specific corrections will be made.

[1043] Displaying the results

[1044] The server calculates the difference between the original text data and the corrected text data, marks it up, and finally sends the result to the user's terminal, which displays a comparison of the original text and the corrected text on the screen.

[1045] Prompt Sentence Examples

[1046] "Please correct the following sentence: 'Hello. My name is Tanaka.' If the user is relaxed, provide specific feedback."

[1047] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[1048] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1049] Step 1:

[1050] Uploading image data

[1051] User:

[1052] A user uses an application on their smartphone or computer to take an image of a handwritten note or printed document, and then uploads it to a server via the application. For example, consider an image of a handwritten note that reads, "Hello. My name is Tanaka." This image data becomes the input.

[1053] output:

[1054] Uploaded image data.

[1055] Step 2:

[1056] Image data sent to server and saved

[1057] Device:

[1058] The device sends the uploaded image data to the server. Specifically, the device sends the image data to the specified URL of the server via an HTTP POST request.

[1059] server:

[1060] The server stores the received image data in temporary storage. For example, if a cloud storage service is used, the uploaded image is stored in the cloud storage. This image data becomes the input for the next step.

[1061] output:

[1062] Image data stored in temporary storage.

[1063] Step 3:

[1064] Extracting character information using OCR

[1065] server:

[1066] The server calls the OCR module and passes the image data stored in temporary storage as input. The OCR module analyzes the image data and extracts text information. For example, the text data "Hello. My name is Tanaka" is extracted from the image.

[1067] output:

[1068] Extracted text data. The text generated is "Hello. My name is Tanaka."

[1069] Step 4:

[1070] Execution of text correction

[1071] server:

[1072] The server sends the generated text data to a writing correction module. For example, the text data is sent to the Grammarly API, which analyzes it for grammar, spelling mistakes, and style improvements. The writing correction module corrects "Hello" to "Hello" and checks the grammar of "I'm Tanaka."

[1073] output:

[1074] Corrected text data. The corrected text "Hello. My name is Tanaka." is generated.

[1075] Step 5:

[1076] User Emotion Recognition

[1077] server:

[1078] The server calls the emotion engine and analyzes facial recognition information and text input speed information from the user's device to identify the user's emotional state. If the analysis shows that the user is nervous, the correction tone will be softened, and if the user is relaxed, the correction will be more specific.

[1079] output:

[1080] Tailored corrections, such as gentler language in sentences.

[1081] Step 6:

[1082] Diff calculation and markup

[1083] server:

[1084] The server calculates the difference between the original text data and the corrected text data and marks it up, using a comparison algorithm to highlight the part where "Hello" was changed to "Hello."

[1085] output:

[1086] Text data with differences marked up.

[1087] Step 7:

[1088] Display the results on your terminal

[1089] server:

[1090] The server sends the resulting data after the difference calculation to the user terminal. Specifically, it returns JSON format data in an HTTP response.

[1091] Device:

[1092] The user's device receives this and displays the original text and the corrected text on the screen for comparison. For example, the part where "Hello" has been changed to "Hello" is highlighted so that the user can easily confirm it.

[1093] output:

[1094] Text data displayed on the screen before and after correction.

[1095] This series of processes allows the user to quickly and efficiently obtain high-quality text data.

[1096] (Application example 2)

[1097] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1098] Existing optical character recognition technology and text correction systems have difficulty in accurately and efficiently recognizing and correcting text during maintenance work in factories. Furthermore, there is a lack of a means to provide appropriate feedback based on the emotional state of maintenance personnel. As a result, misrecognition and miscorrection occur frequently, leading to a decline in maintenance efficiency and work accuracy.

[1099] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, and emotion recognition means for recognizing the user's emotion and reflecting it in the correction results. This makes it possible to perform accurate character recognition and correction during maintenance work in a factory, and to provide appropriate feedback based on the emotional state of the maintenance worker.

[1100] An "image" is a digital recording of visual information.

[1101] "Characters" are symbols or codes used to represent language.

[1102] "Optical character recognition" is a technology for extracting character information from an image.

[1103] "Character data" is character information extracted by character recognition means expressed in digital form.

[1104] "Text correction means" is a technology for improving grammar, spelling mistakes, and style of text based on extracted character data.

[1105] "Emotion recognition means" refers to technology for analyzing and identifying a user's emotional state.

[1106] A "user terminal" is a device operated by a user, such as a smartphone or tablet.

[1107] "Storage means" refers to a device for temporarily or permanently storing character data.

[1108] A "server" is a computer system that processes and stores data on a network.

[1109] "Temporarily storing" means retaining data for a specific period of time.

[1110] "Grammar" is the rules and structure of a language.

[1111] A "spelling error" is an error in the spelling of a word.

[1112] "Style improvements" are modifications made to improve the readability and appearance of the text.

[1113] "Marking up the differences" means visually indicating the differences between the original character data and the corrected character data.

[1114] "Uploading" means sending data from a user terminal to a server.

[1115] "Result data" refers to data that includes corrected character data and analysis results.

[1116] A "factory robot" is a mechanical device that automatically performs various tasks in a factory.

[1117] An "application" is a software program with a specific function.

[1118] This invention is a system for streamlining maintenance work in factories and providing accurate character recognition and emotion-based feedback. The main components of the system include a server, a terminal (e.g., a factory robot), and a user (a maintenance worker).

[1119] The server is equipped with an optical character recognition unit that extracts characters from images, a text correction unit, and an emotion recognition unit. The terminal is a device operated by a user, which in this case corresponds to a robot deployed in a factory. The user provides the system with the information needed during maintenance work through the terminal.

[1120] Overview of Program Processing

[1121] Hardware and Software

[1122] The server is a computer system equipped with a high-performance processor and sufficient memory. Additionally, the following software is used.

[1123] OpenCV: Image Processing Library

[1124] pytesseract: Optical Character Recognition Library

[1125] TextBlob: Natural Language Processing Library

[1126] EmotionRecognition: Virtual Library for Emotion Recognition

[1127] Processing Flow of the System

[1128] 1. Image Capture:

[1129] The factory robot takes pictures of the equipment and parts to be maintained. For example, it takes a picture of a wiring diagram with the robot's camera.

[1130] 2. Character Extraction:

[1131] The server receives the captured image and extracts character information using pytesseract. For example, it extracts the label "Ryk10" from the wiring diagram.

[1132] 3. Text Proofreading:

[1133] The server proofreads grammar and spelling mistakes in the extracted character information using TextBlob. For example, it corrects "Ryjk10" to "Ryk10".

[1134] 4. Emotion Recognition:

[1135] The server uses EmotionRecognition to recognize emotions from the user's facial images taken by the robot's onboard camera, for example, to detect whether the user is in a stressed or relaxed state.

[1136] 5. Feedback adjustment:

[1137] Depending on the user's expected emotional state, the feedback can be adjusted to provide gentler and more specific feedback. For example, if a user is feeling stressed, the feedback could be, "Thank you for your hard work. You're doing a great job."

[1138] 6. Results display:

[1139] The server sends the original text, the corrected text, and the adjusted feedback to the user's terminal and displays them on the screen.

[1140] Adding specific examples

[1141] For example, if the robot photographs a wiring diagram of a device and recognizes the text as "Ryjk10," it will correct this misrecognition to "Ryk10" using TextBlob. At the same time, if it detects that a maintenance worker is under stress using emotion recognition technology, it will provide feedback such as "Great job, great job."

[1142] Example prompts for generative AI models

[1143] Image: [Link to wiring diagram image]

[1144] Caption: Is there any misidentified text in this wiring diagram? Please correct it, especially the part where it is misidentified as "Ryjk10".

[1145] Emotional state: Stressed

[1146] In this way, the system can improve the maintenance work in the factory and provide highly accurate character recognition and feedback according to emotions.

[1147] The flow of the specific process in Application Example 2 will be described using FIG. 14.

[1148] Step 1:

[1149] Image capture

[1150] The user uses the camera of the factory robot to take pictures of the equipment and parts to be maintained. The input is the image file captured by the user, and the output is this image data. This image data is required for the subsequent character recognition process.

[1151] Step 2:

[1152] Character extraction

[1153] The server receives the image data sent from the factory robot and extracts character information using the pytesseract library. The input is the image data, and the output is the character data. Specifically, texts such as the label "Ryk10" written in the wiring diagram in the image are extracted.

[1154] Step 3:

[1155] Text proofreading

[1156] The server proofreads grammar and spelling mistakes of the extracted character information using the TextBlob library. The input is the extracted character data, and the output is the proofread character data. As a specific operation, the misrecognized character "Ryjk10" is corrected to "Ryk10".

[1157] Step 4:

[1158] Emotion recognition

[1159] The server uses the user's facial image captured by the robot's camera to recognize emotions using the EmotionRecognition library. The input is the user's facial image, and the output is emotional data. Specifically, it analyzes whether the user is in a stressed or relaxed state.

[1160] Step 5:

[1161] Feedback Adjustment

[1162] The server adjusts the feedback for the corrected text data based on the emotion recognition results. The input is the corrected text data and emotion data, and the output is the adjusted feedback text. For example, to a user who is under stress, the server may give the feedback, "Thank you for your hard work. You did a great job."

[1163] Step 6:

[1164] Displaying the results

[1165] The server sends the original character data, the corrected character data, and the adjusted feedback to the user's terminal and displays them on the screen. The input is the original character data, the corrected character data, and the feedback text, and the output is the result displayed on the user's terminal. Specifically, the original text "Ryjk10" and the corrected text "Ryk10" are compared and displayed, along with the adjusted feedback.

[1166] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1167] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1168] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1169] [Fourth embodiment]

[1170] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1171] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1172] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1173] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1174] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1175] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1176] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1177] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1178] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1179] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1180] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1181] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1182] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1183] The present invention relates to a system that extracts characters from an image, automatically corrects the character data, and displays the correction results on a user terminal. This system is composed of several main components, including an optical character recognition (OCR) unit, a text correction unit, and a storage unit.

[1184] First, a user can upload an image from their device. The image is preferably a handwritten note or a printed document. The device then sends the image to the server, which then stores the received image data in storage.

[1185] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data, which is temporarily stored on the server.

[1186] The server then sends this saved text data to the text correction module, which analyzes it for grammar, spelling mistakes, and style improvements and performs corrections. Once the corrections are complete, the text data is returned to the server, which calculates the differences between the original and corrected text data and marks them up.

[1187] Finally, the server sends the results to the user's device, which then displays a comparison of the original text and the corrected text. The user can see the differences on the screen and make further corrections if necessary.

[1188] Specific examples

[1189] User Operation

[1190] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[1191] Server Processing

[1192] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[1193] Correction process

[1194] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[1195] Results display

[1196] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[1197] Original text: "Hello. My name is Tanaka."

[1198] Corrected text: "Hello. My name is Tanaka."

[1199] This system automates the entire process from transcription to correction, allowing users to obtain high-quality text data quickly and efficiently.

[1200] The processing flow will be explained below.

[1201] Step 1:

[1202] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[1203] Step 2:

[1204] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[1205] Step 3:

[1206] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[1207] Step 4:

[1208] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[1209] Step 5:

[1210] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[1211] Step 6:

[1212] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[1213] Step 7:

[1214] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[1215] Step 8:

[1216] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[1217] Step 9:

[1218] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[1219] Example 1

[1220] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1221] Conventional optical character recognition systems have low accuracy, and corrections to character data extracted from handwritten or printed documents have been performed manually, which requires time and effort. Furthermore, grammar checks and style improvements after character recognition have not been automated, making the system unfriendly to users. The present invention aims to solve these problems by providing a system that performs character recognition and automatic corrections from images with high accuracy and efficiency.

[1222] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1223] In this invention, the server includes a means for users to upload images from their terminals, a means for transmitting the uploaded images to the server, and a means for the server to pass the image data stored therein to an optical character recognition module. This automates the series of processes of character recognition and text correction, enabling users to obtain high-quality text data quickly and efficiently.

[1224] A "terminal" is a device operated by a user, and includes devices such as smartphones, tablets, and PCs.

[1225] A "server" is a computer system that processes, stores, and transmits image and text data.

[1226] An "optical character recognition module" is a piece of software or hardware that extracts text information from an image and uses OCR (optical character recognition) technology.

[1227] A "text correction module" is a piece of software or hardware that checks the grammar, corrects spelling mistakes, and improves style of extracted text data.

[1228] "Storage means" refers to a storage device for temporarily or permanently storing image data or text data, and includes hard disks, SSDs, cloud storage, etc.

[1229] "Upload" refers to the operation in which a user sends data (images or files) from a terminal to another location (such as a server).

[1230] "Markup" is an operation that uses scripts and tags to highlight or decorate specific parts of text data.

[1231] "Comparative display" is an operation that displays two different text data side by side, allowing you to visually confirm the differences between them.

[1232] "Sending the results" means that the server transfers the processed data to the user's terminal, which is usually done over a network.

[1233] This invention is a system that extracts characters from images, automatically corrects the character data, and displays the correction results on the user's terminal. This system is realized mainly by users uploading images from their terminals and the server processing the image data.

[1234] System configuration

[1235] The system includes the following main components:

[1236] 1. User Device

[1237] 2. Server

[1238] 3. Optical Character Recognition Module (OCR Module)

[1239] 4. Text Correction Module

[1240] 5. Preservation means

[1241] Hardware and software used

[1242] User devices: smartphones, tablets, PCs, etc.

[1243] Server: A high-performance computer system

[1244] Optical Character Recognition Module: Google Cloud Vision API, Tesseract, etc.

[1245] Text correction module: Grammarly API, natural language processing model

[1246] Storage method: Cloud storage such as Amazon S3

[1247] Process Flow

[1248] User Operation

[1249] 1. Users take photos of handwritten notes or printed documents using their smartphones or PCs and upload the images through the application.

[1250] 2. The uploaded image is sent from the device to the server.

[1251] Server Processing

[1252] 1. The server temporarily stores the received image data. Cloud storage is recommended as the storage location.

[1253] 2. The stored image data is passed to the OCR module to extract text information. The OCR module uses optical character recognition technology to generate text data from the image.

[1254] 3. The extracted text data is temporarily stored on the server.

[1255] 4. The server sends the saved text data to the writing correction module, which corrects grammar, spelling, and style, and returns the corrected text data to the server.

[1256] 5. The server compares the original text data with the corrected text data and marks up the differences.

[1257] Results display

[1258] 1. The server sends the final correction results to the user's terminal.

[1259] 2. The user's device receives the transmitted data and displays a comparison of the original text and the corrected text, allowing the user to review the results and make further corrections as necessary.

[1260] Specific examples

[1261] Prompt Sentence Examples

[1262] "Upload an image of a handwritten note and automatically extract and correct the text. Example: The image says, 'Hello. My name is Tanaka.' Output the text 'Hello. My name is Tanaka.'"

[1263] This system allows users to obtain high-quality text data quickly and efficiently. In particular, the entire process from transcription to text correction is automated, significantly reducing the time and effort required.

[1264] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1265] Step 1:

[1266] A user can use a device to take a photo of a handwritten note or a printed document and upload the image through the application. The user selects an image and presses the upload button to start the image data upload operation. The input is the image file taken by the user, and the output is the uploaded image file data.

[1267] Step 2:

[1268] The terminal sends the image data uploaded by the user to the server. A secure connection is established using the HTTPS protocol, and the image data is sent to the server as an HTTP request. The input is the image data sent from the user terminal, and the output is the image data received by the server.

[1269] Step 3:

[1270] The server saves the received image data in storage. Cloud storage such as Amazon S3 is used as the storage. A storage path and file name are generated and the image data is saved. The input is the received image data, and the output is the path information for the image file saved in storage.

[1271] Step 4:

[1272] The server obtains the path information of the stored image data and passes it to the OCR module. The connection with the OCR module is made using the API key and authentication information, and the image data is sent to the OCR module's endpoint. The input is the path information of the image data stored in storage, and the output is the image data sent to the OCR module.

[1273] Step 5:

[1274] The OCR module analyzes the received image data and extracts text information from the image. It uses OCR technologies such as Google Cloud Vision API and Tesseract to analyze the text information and generate text data. The input is the image data sent to the OCR module, and the output is the extracted text data.

[1275] Step 6:

[1276] The server receives the text data returned from the OCR module and temporarily stores it. The text data is stored in a database or memory storage within the server. The input is the text data returned from the OCR module, and the output is the text data stored on the server.

[1277] Step 7:

[1278] The server sends the saved text data to the writing correction module. It connects to the writing correction module (e.g., Grammarly API) and sends the text data in JSON format. The input is the text data saved on the server, and the output is the text data sent to the writing correction module.

[1279] Step 8:

[1280] The writing correction module analyzes the received text data and improves grammar, spelling, and style. The corrected text data is returned to the server with markup. The input is the text data sent to the writing correction module, and the output is the corrected text data.

[1281] Step 9:

[1282] The server compares the original text data with the corrected text data and marks up the parts that have been corrected. It calculates the differences and formats them in a user-friendly format. The input is the original text data and the corrected text data, and the output is the marked-up comparison results.

[1283] Step 10:

[1284] The server sends the final marked-up text result to the user's terminal. The result data is formatted in a format such as HTML or PDF and sent as data for display. The input is the marked-up text result, and the output is the result data sent to the user's terminal.

[1285] Step 11:

[1286] The user checks the final result on the terminal and compares the original text with the corrected text. The differences can be seen at a glance on the displayed screen and further corrections can be made if necessary. The input is the result data sent from the server and the output is the displayed data that the user checks.

[1287] (Application example 1)

[1288] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1289] In recent years, many product price labels in brick-and-mortar stores are frequently updated, but this process is prone to human error. This can lead to incorrect pricing information, which can not only undermine customer trust but also affect sales. Furthermore, manually checking and correcting price labels is laborious and time-consuming for store staff, so an efficient system is needed. Currently available systems lack the ability to automatically detect and correct price label errors, so a new system is needed to address this issue.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1291] In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, means for displaying the correction results on a user terminal, automatic correction means for comparing the extracted character data with the correction results and correcting errors, and means for displaying the results to the user in real time. This allows a store clerk to instantly detect incorrect price information and receive correction suggestions simply by taking a photo of a price label with their smartphone, thereby enabling efficient and accurate price label management.

[1292] The "optical character recognition means" is a means for extracting characters from image data and generating text data.

[1293] The "correction means" is a means for correcting grammar, spelling mistakes, and style of the extracted character data.

[1294] The "means for displaying on the user terminal" refers to a means for displaying the correction results sent from the server on the user's device.

[1295] The "automatic correction means for correcting errors" is an automated means for comparing extracted character data with corrected character data and correcting errors.

[1296] "Means for displaying in real time" refers to means for instantly displaying processed data on a user terminal.

[1297] The "storage means" is a means for temporarily storing image data and extracted character data.

[1298] "Markup means" refers to a means for visually showing the difference between corrected character data and the original character data.

[1299] The "means for detecting errors in price labels and suggesting corrections" is a means for detecting errors in product price information and suggesting correct price information.

[1300] The "means for transmitting to the server" is a means for transferring images uploaded by a user from a terminal to the server.

[1301] The "means for transmitting result data" is a means for transmitting processed data from the server to the user's device.

[1302] The present invention relates to a system for automatically detecting and correcting price label errors in brick-and-mortar stores, which includes optical character recognition, text correction, and automatic correction.

[1303] First, the user takes a picture of the price label using their smartphone, which then sends the image to the server, which then temporarily stores the image data.

[1304] The server then invokes the Optical Character Recognition (OCR) module, passing the image data as input. This OCR module uses OpenCV and Tesseract to extract text data from the image, which is temporarily stored on the server.

[1305] The server then sends the saved text data to a writing correction module, which uses TextBlob to correct grammar, spelling, and style errors, and then returns the corrected text data to the server.

[1306] The server corrects any errors found and marks up the difference between the original and corrected text data. During this process, if the original price information is incorrect, the correct price information is suggested and displayed to the user in real time.

[1307] The user terminal receives and displays the result data sent from the server, allowing the user to immediately check the price label errors in the store and make any necessary corrections.

[1308] For example, the following prompts are used:

[1309] Example prompt sentence:

[1310] "Check the price label on the product. It might say '50,000 yen' instead of '500 yen'."

[1311] This system enables store clerks who manage price labels in physical stores to efficiently and accurately check and correct price information. The hardware used is a smartphone and a camera, and the software uses OpenCV, Tesseract, and TextBlob. This process creates a system that can automatically correct errors in price labels in real time.

[1312] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1313] Step 1:

[1314] A user takes a picture of the price label using a smartphone device.

[1315] Input: Price Label Image

[1316] Output: Image data of the photographed price label

[1317] What happens: The user opens the camera app on their smartphone, focuses on the price label, and takes a photo. The image is then saved to their device.

[1318] Step 2:

[1319] The terminal transmits the captured image to the server.

[1320] Input: Price label image data

[1321] Output: Image data sent to the server

[1322] Specific operation: The device selects the captured image and uploads it to the server via the app. The communication protocol is HTTPS.

[1323] Step 3:

[1324] The server temporarily stores the received image data.

[1325] Input: Received image data

[1326] Output: Saved image data

[1327] Specific operation: The server stores the received image data in a database or temporary file. The storage location is a storage system (e.g., Amazon S3).

[1328] Step 4:

[1329] The server passes the image data to the OCR module to extract the character information.

[1330] Input: Image data

[1331] Output: Extracted text data

[1332] Specific operation: The server uses OpenCV to preprocess the image data, then uses Tesseract to extract text information. The extracted text data is temporarily saved.

[1333] Step 5:

[1334] The server sends the extracted text data to the text correction module.

[1335] Input: Text data

[1336] Output: Corrected text data

[1337] Specific operation: The server analyzes the text data using TextBlob, corrects grammar, spelling mistakes, and style, and saves the corrected text data.

[1338] Step 6:

[1339] The server marks up the differences between the corrected text data and the original text data, and generates correction information for correcting the errors.

[1340] Input: Original text data and corrected text data

[1341] Output: Marked up text data and correction information

[1342] What it does: The server compares each piece of text data and visually displays the differences using a program that highlights the differences (e.g., a diff tool).

[1343] Step 7:

[1344] The server transmits the correction information and the marked-up text data to the user terminal.

[1345] Input: Correction information, marked-up text data

[1346] Output: Information sent to the user's device

[1347] Specific operation: The server prepares the marked-up text data and the correction information for the error correction as a data packet and sends it to the user terminal via the HTTPS protocol.

[1348] Step 8:

[1349] The user terminal displays the received information on the screen, and notifies the user of any price label errors in real time.

[1350] Input: Received correction information, marked-up text data

[1351] Output: Screen showing suggested revisions and original pricing information

[1352] Specific operation: Based on the received information, the user device updates the GUI to display the error and suggested corrections on the screen. The user can check this and manually correct it if necessary.

[1353] This allows store staff to efficiently and accurately detect price label errors and make corrections in real time.

[1354] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1355] This invention relates to a system that extracts characters from images, automatically corrects the character data, recognizes the user's emotions, and reflects them in the correction results. This system is composed of several main components, including an optical character recognition (OCR) means, a text correction means, an emotion engine, and a storage means.

[1356] A user can first upload an image from their device, preferably a handwritten note or a printed document, and the device sends the image to the server, which then stores the received image data in temporary storage.

[1357] Next, the server calls the OCR module and passes the image data as input. The OCR module extracts character information from the image and generates text data. This text data is temporarily stored on the server. The server then sends the stored text data to the writing correction module, which analyzes the text for grammar, spelling mistakes, and style improvements and performs the correction process.

[1358] The server then calls an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. For example, it adjusts the correction results based on the user's emotional state, such as whether they are stressed or relaxed.

[1359] The adjusted correction results are sent back to the server, which compares the original text data with the corrected text data, calculates the differences, and marks them up. Finally, the server sends the results to the user's device, which receives them and displays a comparison of the original text and the corrected text. The user can check the differences on the screen and make further corrections if necessary.

[1360] Specific examples

[1361] User Operation

[1362] Users use their smartphones to take photos of their handwritten notes and upload them through the app, which then sends the images to a server.

[1363] Server Processing

[1364] The server passes the received image to the OCR module. The OCR module extracts the text data "Hello. My name is Tanaka." from the image and returns this text data to the server. The server stores this text data.

[1365] Correction process

[1366] The server calls the text correction module and sends the text data. The correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[1367] Emotion Recognition Processing

[1368] The server calls the emotion engine to analyze the user's emotional state. For example, if the user is nervous, the tone of the corrections will be gentler and the expressions will be gentler. Conversely, if the user is relaxed, more specific comments will be made.

[1369] Results display

[1370] The server calculates the difference between the original text and the corrected text and sends the result to the user's terminal, which displays the following on its screen:

[1371] Original text: "Hello. My name is Tanaka."

[1372] Corrected text: "Hello. My name is Tanaka."

[1373] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[1374] The processing flow will be explained below.

[1375] Step 1:

[1376] The user starts the application on the device, selects the image upload function, selects and takes an image using the device's file system or camera, and presses the upload button.

[1377] Step 2:

[1378] The terminal transmits the selected image file to the server, where the image file is temporarily stored.

[1379] Step 3:

[1380] The server generates a request to pass the saved image file to the OCR module, and issues an API request to the OCR module to send the image data.

[1381] Step 4:

[1382] The OCR module analyzes the received image data and converts the characters in the image into text data, which is then returned to the server.

[1383] Step 5:

[1384] The server temporarily stores the text data received from the OCR module, and generates a request to the text correction module based on the stored text data.

[1385] Step 6:

[1386] The server issues an API request to the text correction module and sends the text data. The text correction module analyzes the received text data, extracts grammar, spelling mistakes, and style improvements, and makes corrections.

[1387] Step 7:

[1388] The text data after correction is returned to the server, which compares the uncorrected text data with the corrected text data, calculates the difference, and marks it up.

[1389] Step 8:

[1390] The server then calls the emotion engine to analyze the user's current emotional state, which uses facial recognition data and previous typing patterns to determine whether the user is tense or relaxed.

[1391] Step 9:

[1392] The server adjusts the tone and style of its corrections based on the data it receives from the emotion engine, for example, increasing kind language and positive feedback if the user is nervous.

[1393] Step 10:

[1394] The server compiles the final correction results (the text before correction, the text after correction, and the difference) into a single response and sends it to the user terminal.

[1395] Step 11:

[1396] The user's device analyzes the received result data and displays a comparison between the original text and the corrected text. The user can check the differences between the two on the screen and make further corrections as necessary.

[1397] Example 2

[1398] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1399] Conventional systems simply extract text from image data and correct the text, but this does not necessarily mean that the corrections are appropriate for the user's feelings. Furthermore, feedback that does not take the user's feelings into account is often ineffective. Furthermore, it is difficult to efficiently process image data uploaded by users and quickly display the results.

[1400] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an optical character recognition means for extracting characters from an image, a correction means for correcting the extracted character data, an emotion recognition means for recognizing the user's emotion and adjusting the correction result based on the emotion, and a means for displaying the correction result on the user terminal. This enables appropriate correction of the text taking the user's emotion into consideration and rapid feedback.

[1401] "Optical character recognition" refers to techniques and devices for extracting character information from image data.

[1402] "Text correction means" refers to the technology and devices that analyze extracted text data from the perspectives of grammar, spelling mistakes, and style, and make appropriate corrections.

[1403] "Emotion recognition means" refers to technology and devices for analyzing a user's emotions and adjusting the results of data processing according to that emotional state.

[1404] "Storage means" refers to devices or technologies for temporarily or permanently storing extracted character data or intermediate processed data.

[1405] "Means for marking up differences" refers to techniques and devices for visually highlighting the differences between the original text data and the modified text data.

[1406] A "user terminal" is a hardware device operated by a user, and includes devices such as smartphones and personal computers.

[1407] "Server" refers to a computer system for processing, storing, and transmitting data.

[1408] This system extracts text from an image, automatically corrects the text, and recognizes the user's emotions and reflects them in the correction results. This system consists of the following main components:

[1409] Key Components

[1410] 1. Optical Character Recognition (OCR):

[1411] This is a technology that extracts text information from image data. Specifically, it uses an OCR engine. For example, Google Cloud Vision API can be used.

[1412] 2. Writing correction methods:

[1413] This technology analyzes the extracted text data from the perspectives of grammar, spelling mistakes, and style, and makes appropriate corrections. For example, you can use a grammar checker API such as Grammarly.

[1414] 3. Emotion recognition means:

[1415] This technology analyzes a user's emotions and adjusts the results of processing data based on their emotional state. It includes algorithms that analyze facial recognition information and text input patterns. This allows it to adjust the tone or expressions to be gentler if the user is nervous, or more specific if the user is relaxed.

[1416] 4. Preservation means:

[1417] This is a technology for temporarily storing extracted text data and intermediate processed data. For example, data can be stored using a cloud storage service.

[1418] 5. Ways to mark up differences:

[1419] This technology visually highlights the differences between the original text and the edited text. It uses a comparison algorithm to highlight the changes.

[1420] Specific examples of operations

[1421] User Operation

[1422] A user can use a smartphone to take a photo of a handwritten note and upload it through the app. For example, consider a handwritten note that reads, "Hello. My name is Tanaka."

[1423] Server Processing and Response

[1424] The server passes the received image to the OCR module. The OCR module extracts character information from the image and generates text data such as "Hello. My name is Tanaka." This text data is saved on the server. The server then sends the saved text data to the writing correction module. The writing correction module corrects "Hello." to "Hello." and checks the grammar of "I'm Tanaka." The corrected text data is returned to the server.

[1425] Performing emotion recognition

[1426] The server calls the emotion engine and analyzes the user's facial recognition information and text input patterns to identify the user's emotional state. If the user is nervous, the corrections will be made in a gentler tone and with gentler expressions. Conversely, if the user is relaxed, more specific corrections will be made.

[1427] Displaying the results

[1428] The server calculates the difference between the original text data and the corrected text data, marks it up, and finally sends the result to the user's terminal, which displays a comparison of the original text and the corrected text on the screen.

[1429] Prompt Sentence Examples

[1430] "Please correct the following sentence: 'Hello. My name is Tanaka.' If the user is relaxed, provide specific feedback."

[1431] This system automates the entire process, from transcription to correction and feedback based on the user's emotional state, allowing users to obtain high-quality text data quickly and efficiently.

[1432] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1433] Step 1:

[1434] Uploading image data

[1435] User:

[1436] A user uses an application on their smartphone or computer to take an image of a handwritten note or printed document, and then uploads it to a server via the application. For example, consider an image of a handwritten note that reads, "Hello. My name is Tanaka." This image data becomes the input.

[1437] output:

[1438] Uploaded image data.

[1439] Step 2:

[1440] Image data sent to server and saved

[1441] Device:

[1442] The device sends the uploaded image data to the server. Specifically, the device sends the image data to the specified URL of the server via an HTTP POST request.

[1443] server:

[1444] The server stores the received image data in temporary storage. For example, if a cloud storage service is used, the uploaded image is stored in the cloud storage. This image data becomes the input for the next step.

[1445] output:

[1446] Image data stored in temporary storage.

[1447] Step 3:

[1448] Extracting character information using OCR

[1449] server:

[1450] The server calls the OCR module and passes the image data stored in temporary storage as input. The OCR module analyzes the image data and extracts text information. For example, the text data "Hello. My name is Tanaka" is extracted from the image.

[1451] output:

[1452] Extracted text data. The text generated is "Hello. My name is Tanaka."

[1453] Step 4:

[1454] Execution of text correction

[1455] server:

[1456] The server sends the generated text data to a writing correction module. For example, the text data is sent to the Grammarly API, which analyzes it for grammar, spelling mistakes, and style improvements. The writing correction module corrects "Hello" to "Hello" and checks the grammar of "I'm Tanaka."

[1457] output:

[1458] Corrected text data. The corrected text "Hello. My name is Tanaka." is generated.

[1459] Step 5:

[1460] User Emotion Recognition

[1461] server:

[1462] The server calls the emotion engine and analyzes facial recognition information and text input speed information from the user's device to identify the user's emotional state. If the analysis shows that the user is nervous, the correction tone will be softened, and if the user is relaxed, the correction will be more specific.

[1463] output:

[1464] Tailored corrections, such as gentler language in sentences.

[1465] Step 6:

[1466] Diff calculation and markup

[1467] server:

[1468] The server calculates the difference between the original text data and the corrected text data and marks it up, using a comparison algorithm to highlight the part where "Hello" was changed to "Hello."

[1469] output:

[1470] Text data with differences marked up.

[1471] Step 7:

[1472] Display the results on your terminal

[1473] server:

[1474] The server sends the resulting data after the difference calculation to the user terminal. Specifically, it returns JSON format data in an HTTP response.

[1475] Device:

[1476] The user's device receives this and displays the original text and the corrected text on the screen for comparison. For example, the part where "Hello" has been changed to "Hello" is highlighted so that the user can easily confirm it.

[1477] output:

[1478] Text data displayed on the screen before and after correction.

[1479] This series of processes allows the user to quickly and efficiently obtain high-quality text data.

[1480] (Application example 2)

[1481] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1482] Existing optical character recognition technology and text correction systems have difficulty in accurately and efficiently recognizing and correcting text during maintenance work in factories. Furthermore, there is a lack of a means to provide appropriate feedback based on the emotional state of maintenance personnel. As a result, misrecognition and miscorrection occur frequently, leading to a decline in maintenance efficiency and work accuracy.

[1483] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes optical character recognition means for extracting characters from an image, correction means for correcting the extracted character data, and emotion recognition means for recognizing the user's emotion and reflecting it in the correction results. This makes it possible to perform accurate character recognition and correction during maintenance work in a factory, and to provide appropriate feedback based on the emotional state of the maintenance worker.

[1484] An "image" is a digital recording of visual information.

[1485] "Characters" are symbols or codes used to represent language.

[1486] "Optical character recognition" is a technology for extracting character information from an image.

[1487] "Character data" is character information extracted by character recognition means expressed in digital form.

[1488] "Text correction means" is a technology for improving grammar, spelling mistakes, and style of text based on extracted character data.

[1489] "Emotion recognition means" refers to technology for analyzing and identifying a user's emotional state.

[1490] A "user terminal" is a device operated by a user, such as a smartphone or tablet.

[1491] "Storage means" refers to a device for temporarily or permanently storing character data.

[1492] A "server" is a computer system that processes and stores data on a network.

[1493] "Temporarily storing" means retaining data for a specific period of time.

[1494] "Grammar" is the rules and structure of a language.

[1495] A "spelling error" is an error in the spelling of a word.

[1496] "Style improvements" are modifications made to improve the readability and appearance of the text.

[1497] "Marking up the differences" means visually indicating the differences between the original character data and the corrected character data.

[1498] "Uploading" means sending data from a user terminal to a server.

[1499] "Result data" refers to data that includes corrected character data and analysis results.

[1500] A "factory robot" is a mechanical device that automatically performs various tasks in a factory.

[1501] An "application" is a software program with a specific function.

[1502] This invention is a system for streamlining maintenance work in factories and providing accurate character recognition and emotion-based feedback. The main components of the system include a server, a terminal (e.g., a factory robot), and a user (a maintenance worker).

[1503] The server is equipped with an optical character recognition unit that extracts characters from images, a text correction unit, and an emotion recognition unit. The terminal is a device operated by a user, which in this case corresponds to a robot deployed in a factory. The user provides the system with the information needed during maintenance work through the terminal.

[1504] Program processing overview

[1505] Hardware and Software

[1506] The server is a computer system equipped with a high-performance processor and sufficient memory, and uses the following software:

[1507] OpenCV: Image processing library

[1508] pytesseract: Optical Character Recognition Library

[1509] TextBlob: A natural language processing library

[1510] EmotionRecognition: A Virtual Library for Emotion Recognition

[1511] System processing flow

[1512] 1. Image capture:

[1513] Factory robots take images of the equipment or parts they are maintaining, for example, by using their cameras to take pictures of wiring diagrams.

[1514] 2. Character Extraction:

[1515] The server receives the captured image and extracts character information using pytesseract. For example, it extracts the label "Ryk10" from a wiring diagram.

[1516] 3. Text Proofreading:

[1517] The server proofreads grammar and spelling mistakes in the extracted character information using TextBlob. For example, it corrects "Ryjk10" to "Ryk10".

[1518] 4. Emotion Recognition:

[1519] The server uses EmotionRecognition to recognize emotions using the face image of the user captured by the camera mounted on the robot. For example, it detects whether the user is in a stressed state or a relaxed state.

[1520] 5. Feedback Adjustment:

[1521] According to the assumed emotional state, the feedback content is adjusted to give kind expressions and specific points. For example, for a user feeling stressed, feedback such as "You've had a hard time. It's an excellent job." is provided.

[1522] 6. Result Display:

[1523] The server sends the original text, the proofread text, and the further adjusted feedback content to the user terminal and displays them on the screen.

[1524] Addition of Specific Examples

[1525] For example, if the robot photographs a wiring diagram of a device and recognizes the text as "Ryjk10," it will correct this misrecognition to "Ryk10" using TextBlob. At the same time, if it detects that a maintenance worker is under stress using emotion recognition technology, it will provide feedback such as "Great job, great job."

[1526] Example prompts for generative AI models

[1527] Image: [Link to wiring diagram image]

[1528] Caption: Is there any misidentified text in this wiring diagram? Please correct it, especially the part where it is misidentified as "Ryjk10".

[1529] Emotional state: Stressed

[1530] In this way, the system can streamline maintenance work in factories and provide highly accurate character recognition and emotional feedback.

[1531] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1532] Step 1:

[1533] Image Capture

[1534] The user uses the factory robot's camera to take images of the equipment or parts to be maintained. The input is the image file taken by the user, and the output is the image data. This image data is required for character recognition processing in the subsequent stage.

[1535] Step 2:

[1536] Character extraction

[1537] The server receives the image data transmitted from the factory robot and extracts character information using the pytesseract library. The input is the image data, and the output is the character data. Specifically, it extracts texts such as the label "Ryk10" written on the wiring diagram in the image.

[1538] Step 3:

[1539] Text proofreading

[1540] The server proofreads grammar and spelling mistakes of the extracted character information using the TextBlob library. The input is the extracted character data, and the output is the proofread character data. As a specific operation, it corrects the misrecognized character "Ryjk10" to "Ryk10".

[1541] Step 4:

[1542] Emotion recognition

[1543] The server uses the EmotionRecognition library to recognize emotions using the face image of the user taken by the camera mounted on the robot. The input is the face image of the user, and the output is the emotion data. Specifically, it analyzes whether the user is in a stressed state or a relaxed state.

[1544] Step 5:

[1545] Feedback adjustment

[1546] The server adjusts the feedback for the proofread character data based on the emotion recognition result. The input is the proofread character data and the emotion data, and the output is the adjusted feedback text. For example, it gives feedback such as "Thank you for your hard work. Great job." to a stressed user.

[1547] Step 6:

[1548] Display of results

[1549] The server sends the original character data, the corrected character data, and the adjusted feedback to the user's terminal and displays them on the screen. The input is the original character data, the corrected character data, and the feedback text, and the output is the result displayed on the user's terminal. Specifically, the original text "Ryjk10" and the corrected text "Ryk10" are compared and displayed, along with the adjusted feedback.

[1550] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1551] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1552] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1553] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1554] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1555] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1556] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1557] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1558] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1559] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1560] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1561] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1562] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1563] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1564] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1565] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1566] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1567] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1568] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1569] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1570] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1571] The following is further disclosed regarding the above embodiment.

[1572] (Claim 1)

[1573] Optical character recognition means for extracting characters from an image;

[1574] a correction means for correcting the extracted character data;

[1575] A means for displaying the correction results on a user terminal;

[1576] A system including:

[1577] (Claim 2)

[1578] a storage means for temporarily storing character data obtained by an optical character recognition means for extracting characters from an image;

[1579] A means to improve grammar, spelling, and style when correcting writing;

[1580] A means for marking up the difference between the corrected character data and the original character data;

[1581] The system of claim 1 further comprising:

[1582] (Claim 3)

[1583] a means for a user to upload an image from a terminal;

[1584] means for transmitting the uploaded images to a server;

[1585] means for transmitting result data from the server to the user terminal;

[1586] The system of claim 1 further comprising:

[1587] "Example 1"

[1588] (Claim 1)

[1589] a means for a user to upload an image from a terminal;

[1590] means for transmitting the uploaded images to a server;

[1591] A means for storing the image data received by the server;

[1592] a means for transmitting the stored image data by the server to an optical character recognition module;

[1593] means for an optical character recognition module to extract character information from the image;

[1594] A means for temporarily storing the extracted character information;

[1595] means for transmitting text data to a text correction module;

[1596] A writing correction module provides grammar, spelling, and style improvements;

[1597] A means for returning the corrected text data to the server;

[1598] a means for the server to mark up the differences between the original text and the corrected text;

[1599] means for transmitting results from the server to the user terminal;

[1600] a means for the user terminal to display the original text and the corrected text for comparison;

[1601] A system including:

[1602] (Claim 2)

[1603] means for temporarily storing character data obtained by the optical character recognition module;

[1604] A means for correcting text data for grammar, spelling, and style through a text correction module;

[1605] A means for marking up the differences between the corrected text data and the original text data;

[1606] The system of claim 1 further comprising:

[1607] (Claim 3)

[1608] a means for a user to upload an image from a terminal;

[1609] means for transmitting the uploaded images to a server;

[1610] means for transmitting result data from the server to the user terminal;

[1611] The system of claim 1 further comprising:

[1612] "Application Example 1"

[1613] (Claim 1)

[1614] Optical character recognition means for extracting characters from an image;

[1615] a correction means for correcting the extracted character data;

[1616] A means for displaying the correction results on a user terminal;

[1617] an automatic correction means for comparing the extracted character data with the correction results and correcting errors;

[1618] a means for displaying the results to the user in real time;

[1619] A system including:

[1620] (Claim 2)

[1621] a storage means for temporarily storing character data obtained by an optical character recognition means for extracting characters from an image;

[1622] A means to improve grammar, spelling, and style when correcting writing;

[1623] A means for marking up the difference between the corrected character data and the original character data;

[1624] and means for displaying the corrected price information in real time.

[1625] 10. The system of claim 1.

[1626] (Claim 3)

[1627] a means for a user to upload an image from a terminal;

[1628] means for transmitting the uploaded images to a server;

[1629] means for transmitting result data from the server to the user terminal;

[1630] a means for detecting errors in price labels and suggesting corrections;

[1631] The system of claim 1 further comprising:

[1632] "Example 2: Combining Emotion Engines"

[1633] (Claim 1)

[1634] Optical character recognition means for extracting characters from an image;

[1635] a correction means for correcting the extracted character data;

[1636] emotion recognition means for recognizing the emotion of a user and adjusting the correction result based on the emotion;

[1637] A means for displaying the correction results on a user terminal;

[1638] A system including:

[1639] (Claim 2)

[1640] a storage means for temporarily storing character data obtained by an optical character recognition means for extracting characters from an image;

[1641] A means to improve grammar, spelling, and style when correcting writing;

[1642] A means for marking up the difference between the corrected character data and the original character data;

[1643] The system of claim 1 further comprising:

[1644] (Claim 3)

[1645] a means for a user to upload an image from a terminal;

[1646] means for transmitting the uploaded images to a server;

[1647] means for transmitting result data from the server to the user terminal;

[1648] The system of claim 1 further comprising:

[1649] "Application example 2 when combining emotion engines"

[1650] (Claim 1)

[1651] Optical character recognition means for extracting characters from an image;

[1652] a correction means for correcting the extracted character data;

[1653] emotion recognition means for recognizing the user's emotions and reflecting them in the correction results;

[1654] A means for displaying the correction results on a user terminal;

[1655] A system including:

[1656] (Claim 2)

[1657] a storage means for temporarily storing character data obtained by an optical character recognition means for extracting characters from an image;

[1658] A means to improve grammar, spelling, and style when correcting writing;

[1659] A means for marking up the difference between the corrected character data and the original character data;

[1660] means for adjusting the feedback based on the user's emotional state;

[1661] The system of claim 1 further comprising:

[1662] (Claim 3)

[1663] a means for a user to upload an image from a terminal;

[1664] means for transmitting the uploaded images to a server;

[1665] means for transmitting result data from the server to the user terminal;

[1666] means for applying the application installed on the factory robot;

[1667] The system of claim 1 further comprising: [Explanation of symbols]

[1668] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. Optical character recognition means for extracting characters from an image; a correction means for correcting the extracted character data; A means for displaying the correction results on a user terminal; A system including:

2. a storage means for temporarily storing character data obtained by an optical character recognition means for extracting characters from an image; A means to improve grammar, spelling, and style when correcting writing; A means for marking up the difference between the corrected character data and the original character data; The system of claim 1 further comprising:

3. a means for a user to upload an image from a terminal; means for transmitting the uploaded images to a server; means for transmitting result data from the server to the user terminal; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A