system

A system using character recognition and natural language processing with a self-learning function addresses the challenge of deciphering ancient documents, allowing non-experts to understand them efficiently and improve accuracy through user feedback and emotional analysis.

JP2026123725APending Publication Date: 2026-07-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2025-01-17
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Deciphering ancient documents requires specialized knowledge and is time-consuming, making it difficult for non-experts to access and understand historical materials.

Method used

A system that utilizes character recognition and natural language processing technologies, combined with a self-learning function, to decipher and interpret ancient documents, enabling non-experts to understand them easily and improve interpretation accuracy through user feedback.

Benefits of technology

Enables non-experts to decipher and understand ancient documents efficiently, with continuous improvement in interpretation accuracy through user feedback and emotional state consideration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026123725000001_ABST
    Figure 2026123725000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for extracting handwritten or printed characters from image data using character recognition technology, A means for performing grammatical and syntactic analysis on extracted strings using natural language processing technology, A means of generating interpretation results and providing them to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The deciphering of ancient documents has conventionally been a task that requires highly specialized knowledge and a great deal of time and effort. For this reason, it has been difficult for people other than professional researchers to access, and many historical materials have been left untranslated. In view of such a situation, there is a need to provide a means for more quickly and efficiently deciphering and understanding ancient documents.

Means for Solving the Problems

[0005] This invention solves the above-mentioned problems by providing a system that recognizes characters from image data and analyzes and interprets the context using natural language processing. This system has a self-learning function based on past deciphering history and can further collect user feedback to improve interpretation accuracy. In this way, it realizes an environment in which anyone, even non-experts, can easily decipher and understand ancient documents.

[0006] "Character recognition technology" is a technology that mechanically extracts handwritten or printed characters contained in image data and represents them as digital text.

[0007] "Natural language processing technology" refers to techniques that enable computers to understand and interpret human language and extract its meaning. This includes grammatical analysis and syntactic analysis.

[0008] "Image data" refers to data that stores visual information in a digital format, and is usually represented by pixels.

[0009] "Grammar analysis" is the process of identifying the part of speech and role of words and phrases, which are the constituent elements of a text, and revealing its structure.

[0010] "Syntactic analysis" is the process of analyzing the relationships between words and phrases in order to understand the meaning and structure of a sentence, and breaking down the sentence according to grammatical rules.

[0011] "Interpretation results" refer to the output generated as the final outcome of understanding and extracting meaning from a text.

[0012] "Self-learning function" refers to the algorithms and mechanisms that a system uses to improve its own performance based on experience and feedback.

[0013] "Feedback" refers to user evaluations and opinions on the results generated by a system, and is information used to improve the system. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a method for efficiently deciphering ancient documents in a system consisting of servers and terminals connected via a network. The operation method of this system is described in detail below.

[0036] First, the user takes a picture of the ancient document they wish to decipher with their device and uploads the image data to the system. The device then checks the quality of this image data before sending it to the server.

[0037] The server applies character recognition technology to the received image data to extract character information from the image. This process includes pre-processing such as noise reduction and contrast adjustment to optimize the image and improve recognition accuracy.

[0038] Furthermore, the server utilizes natural language processing technology to perform grammatical and syntactic analysis on the extracted text data, thereby understanding the linguistic structure within the ancient documents. This allows for translation into modern Japanese, taking into account the unique expressions and historical context of the ancient documents.

[0039] The server also performs the process of generating interpretations along with understanding the context. By utilizing self-learning capabilities, drawing on historical background and past interpretation results, it improves the accuracy of its interpretations and derives appropriate results. These interpretation results are then transmitted to the terminal and presented to the user.

[0040] Users can view the decryption results on their device screen and provide feedback on the content. This feedback information is managed on a server and used to improve the accuracy of future decryption processes.

[0041] For example, when deciphering a letter from the Heian period, the user first takes a picture of the letter with their device. The server then extracts the characters, analyzes the grammar and syntax, and provides a translation into modern Japanese and a contextual interpretation. In this way, the system helps people understand historical documents even without specialized knowledge.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user takes a picture of the ancient document they want to decipher using their device. During this process, the device checks the image quality and adjusts the resolution and format as needed.

[0045] Step 2:

[0046] The terminal sends the captured image data to the server. It checks whether the data was successfully transmitted over the network and retransmits it if necessary.

[0047] Step 3:

[0048] The server performs preprocessing on the received image data. Specifically, it performs noise reduction and contrast adjustment to prepare the image for improved character recognition accuracy.

[0049] Step 4:

[0050] The server applies optical character recognition (OCR) technology to pre-processed image data to extract characters. It uses algorithms to handle handwritten characters and special fonts to accurately generate character data.

[0051] Step 5:

[0052] The server performs natural language processing on the extracted text data. Through grammatical and syntactic analysis, it lays the foundation for understanding the context of ancient documents.

[0053] Step 6:

[0054] The server performs a translation into modern Japanese based on the analysis results obtained. This translation takes into account the historical context and background to select appropriate terms.

[0055] Step 7:

[0056] The server generates an interpretation based on the translated text. It utilizes historical data and learning algorithms to derive an accurate interpretation.

[0057] Step 8:

[0058] The server sends the generated interpretation results to the terminal.

[0059] Step 9:

[0060] The terminal displays the interpretation results received from the server to the user. The user reviews the interpretation and provides feedback as needed.

[0061] Step 10:

[0062] User feedback is sent to the server and used to improve decoding accuracy. The server stores this data and incorporates it into subsequent processing.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Traditional methods for deciphering ancient documents heavily relied on the expertise of specialists, making it difficult for ordinary users to easily understand them. Furthermore, improving the accuracy of interpretation and translation quality required considerable effort and time, and collecting feedback and improving analysis accuracy was also challenging. Therefore, there is a need for technology that allows for easy deciphering of ancient documents without specialized knowledge and enables continuous improvement of interpretation accuracy.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for inputting an image of a document using a camera and transmitting it to a data processing device, means for extracting character information from the image data using character recognition technology, and means for grammatical and syntactic analysis of the extracted character information using natural language processing technology. This makes it possible for ordinary users to accurately understand ancient documents without specialized knowledge and to constantly improve the accuracy of their interpretations.

[0068] A "photography device" is a device used to acquire documents and images as digital data and transmit them to a processing device.

[0069] A "data processing device" is a device that analyzes received digital data and extracts or translates textual information.

[0070] "Character recognition technology" is a technology that identifies handwritten or printed characters from image data and converts them into text data.

[0071] "Natural language processing technology" is a technique that analyzes text data based on its grammatical structure and meaning, enabling a machine-like understanding of human language.

[0072] A "generative model" is an artificial intelligence model that generates new data or results based on patterns in existing data.

[0073] "Interpretation results" refer to translations and semantic interpretations generated based on processed data.

[0074] "Methods for collecting opinions" refer to methods for obtaining feedback from users and using it to improve the system.

[0075] "Automatic learning function" is a technology that improves system performance by learning from past data and feedback.

[0076] This invention is an advanced integrated system for deciphering ancient documents without requiring specialized knowledge. The system is centered around terminals, servers, and users, each with clearly defined roles.

[0077] First, the user uses hardware such as a portable camera or tablet camera to acquire image data of the document they wish to decipher. After acquisition, the user uploads this data to the system via the network.

[0078] The device is responsible for evaluating the quality of images provided by the user. Specifically, it uses software to check the image's resolution and brightness and determine if it is suitable for recognition. If the image is unsuitable, it will ask the user to retake it.

[0079] After receiving image data, the server applies our proprietary character recognition technology to extract character information from the image. This process involves pre-processing such as noise reduction and identification of character regions, after which digital text is generated through a recognition algorithm.

[0080] Next, the server uses natural language processing techniques on the extracted text information. It analyzes the grammar and structure unique to ancient documents and translates them into modern Japanese using a generative AI model. This generative model is particularly adept at handling complex linguistic structures, such as those found in documents from the Heian period.

[0081] Users can view the interpretation results sent from the server on their device's display. Based on the decoded results, users can understand the content and provide further feedback. This feedback information is collected by the server and used to improve the accuracy of the analysis.

[0082] A concrete example is a scenario in which a user deciphers a letter from the Heian period. The user takes a picture of the letter and uploads the image to the system. The server recognizes and analyzes the characters and translates them into modern Japanese, taking into account the expressions used in the Heian period. Through this process, the user can understand the content of the letter even without specialized knowledge.

[0083] As an example of a prompt, you can make a request in the format of, "Please translate this Heian period letter, add an interpretation, and present it in an easy-to-understand format."

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] The user takes a picture of the document they wish to decipher using a portable device. The captured image is saved to the device, and then the data format (e.g., JPEG, PNG) and resolution are checked. The output is an image file in the appropriate format.

[0087] Step 2:

[0088] The device evaluates image quality. Specifically, it uses software to analyze the image's brightness, contrast, and resolution to determine if it is suitable for character recognition. If the evaluation results do not meet the criteria, the user is notified to retake the image. The output includes the quality evaluation results or the notification sent to the user.

[0089] Step 3:

[0090] When image data is sent to the server, the server applies OCR technology to extract text information from the image. This process involves noise reduction and identification of text regions, and the recognition algorithm converts the text into digital text. The input is image data, and the output is recognized text information.

[0091] Step 4:

[0092] The server uses natural language processing techniques to perform grammatical and syntactic analysis on the obtained text information. The analysis clarifies the grammatical structure of the document, and then it is translated into modern Japanese. A generative AI model is used to output a modern Japanese translation from the input text information.

[0093] Step 5:

[0094] The server understands the context based on the translation results and uses a generative AI model to generate a detailed interpretation. Historical background information is also incorporated into the contextual interpretation, and data calculations are performed to improve the accuracy of the interpretation. The output is a detailed interpretation result.

[0095] Step 6:

[0096] The interpretation results are sent from the server to the terminal, and the user checks the translation and interpretation on the terminal screen. A user-centered interface clearly presents the decoded content. The displayed interpretation result is obtained as output.

[0097] Step 7:

[0098] Users can provide feedback based on the interpretation results. This feedback is sent to the server, and the system uses this information to improve the accuracy of the decoding process. The input is user feedback, and the output is feedback information.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] Traditionally, paper receipts and invoices have not been digitized, making it difficult to efficiently manage past transactions. Handwritten or outdated formats are particularly difficult to recognize and analyze, highlighting the need for centralized management of data in different formats.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes means for extracting content from image data using character recognition technology, means for analyzing and translating the extracted information using natural language processing technology, and means for structuring the extraction results based on past information. This makes it possible to efficiently digitize information on paper media and manage data in different formats in a unified manner.

[0104] "Character recognition technology" is a technology that converts handwritten or printed character information into digital data.

[0105] "Natural language processing technology" is a technology that aims to analyze and understand human language using computers.

[0106] "Means of analysis and translation" are means that have the ability to analyze extracted information grammatically and semantically and convert it into a specific form or language.

[0107] "Self-learning function" refers to a function that improves processing accuracy by learning on its own based on past data and interpretation results.

[0108] A "data processing system" is an information system designed to efficiently process and manage collected data.

[0109] "Information provision" refers to data and feedback obtained from users and external sources, which are used to improve the system's learning and analysis accuracy.

[0110] To implement this invention, three elements are necessary: ​​a server, a terminal, and a user.

[0111] The server possesses high-performance processing capabilities and is equipped with software that can apply character recognition and natural language processing technologies. Specifically, it extracts character information from captured image data using OCR software, interprets it using an NLP library, and translates it. In addition, the server has the function to structure the data and convert it into a data format for integration with electronic payment systems and other systems as needed. For example, it is used when integrating handwritten invoices as digital data into an accounting system.

[0112] The devices used are smartphones and tablets, which use their cameras to photograph the target items, such as receipts or historical documents. The device checks the quality of the captured image data and sends it to the server. A dedicated application running on the latest mobile OS makes this possible and provides a user-friendly UI / UX.

[0113] Users take and transmit photos through an application on their device and check the interpretation results received from the server. They can also provide feedback on the interpretation to the server, which contributes to improving the overall accuracy of the system.

[0114] As a concrete example, a user takes a picture of a handwritten invoice from the 1940s with their smartphone and uploads the data to a server. This data is then structured through OCR and NLP processing and incorporated into a modern accounting system as a digital transaction history. In this process, an example of a prompt for the generating AI model would be, "I have taken a picture of an old-style handwritten invoice. Please recognize all the text from the photo, extract the date, items, and amounts, and translate them into a modern electronic payment format."

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The user takes a picture of a receipt or invoice using the device. The input is image data captured via the camera sensor. The output is data saved on the device in image file format. This action completes the first step toward digitizing paper-based information.

[0118] Step 2:

[0119] The device automatically checks the quality of the captured image data, verifying that there is no noise or blur. The input is the image file generated in the previous step. As part of the data processing, the image quality is checked, and as output, the image with confirmed quality is sent to the server. This step includes sending feedback to the user prompting them to retake any unchecked images.

[0120] Step 3:

[0121] The server extracts text information from the received image data using OCR software. The input is an image sent from the terminal. As a data calculation, an image processing algorithm scans the text on the paper and generates digital text as output. This process converts handwritten or printed information into readable text.

[0122] Step 4:

[0123] The server uses natural language processing techniques to grammatically and semantically analyze the extracted text. The input is the digital text generated in the previous step, and the output is structured information data. Data analysis is performed, and prompt sentences are applied to the generative AI model to organize the information. This process organizes different formats into a consistent format.

[0124] Step 5:

[0125] The server converts the structured data into a format compatible with the electronic payment system and integrates it with relevant systems as needed. The input is the structured information generated in step 4, and the output is a dataset for system integration. Data conversion and integration processes are performed, and the information is efficiently integrated. In this step, adjustments are made to the new data format.

[0126] Step 6:

[0127] The system presents the user with the interpretation results from the server on their device. The input is the final data processed by the server, and the output is the interpretation displayed on the device's UI. This allows the user to review the results and provide feedback on their satisfaction. This action improves user trust in the system.

[0128] Step 7:

[0129] Users send feedback on the interpretation results to the server via their device. The input consists of user comments and ratings, while the output is the feedback recorded in the server's feedback database. The server uses this data to self-learn and improve future interpretation accuracy. Continuous improvement through user interaction is an integral part of its operation.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention incorporates an emotion engine into a system for deciphering ancient documents, thereby enabling the provision of interpretation results that take into account the user's emotional state. The following describes specific embodiments of this system.

[0132] The user takes a picture of an ancient document with their device and sends the image data to the server. The server applies character recognition technology to the image data and extracts the text information. Then, using natural language processing technology, it performs grammatical and syntactic analysis, translates the content of the ancient document into modern Japanese, and generates an interpretation result.

[0133] In this interpretation process, the emotion engine recognizes the user's emotions. Specifically, it analyzes the user's facial expressions and voice while they are viewing the interpretation results or providing feedback to understand their emotional state. This emotional information is reflected in how the interpretation results are presented and in the provision of additional information. For example, if the user expresses distrust, the system can present alternative interpretations or more detailed background information.

[0134] In this way, interpretations are dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation. Furthermore, the user's emotions are taken into consideration during the feedback process, which can be used to improve the accuracy of future interpretations.

[0135] As a concrete example, when a user deciphers an ancient document from the Sengoku period, the emotion engine detects the user's interest. In this case, the server accordingly provides relevant historical background information and links to similar documents, supporting a deeper understanding. This system enables users to not only decipher the text but also to gain a deeper understanding of its context and background in accordance with their emotions.

[0136] The following describes the processing flow.

[0137] Step 1:

[0138] The user takes a picture of the ancient document they want to decipher with their device. The device checks the image resolution and format and adjusts it appropriately before sending it to the server.

[0139] Step 2:

[0140] The device sends the captured data to the server. The server first performs pre-processing, such as noise reduction, on the received image data.

[0141] Step 3:

[0142] The server applies character recognition technology to the pre-processed image data to extract character information. The algorithm is optimized to handle even unusual fonts in handwritten text.

[0143] Step 4:

[0144] The server applies natural language processing techniques to the character data obtained through character recognition, performing grammatical and syntactic analysis to understand the context.

[0145] Step 5:

[0146] Based on the results of natural language processing, the server translates the content of ancient documents into modern language and generates an interpretation. The accuracy of the translation is improved by considering historical context and linguistic peculiarities.

[0147] Step 6:

[0148] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice while viewing the interpretation results. This information is used to adjust the user interface.

[0149] Step 7:

[0150] On the terminal, the interpretation results generated by the server are displayed to the user. At this time, additional information or alternative interpretations may be presented based on the user's emotional state.

[0151] Step 8:

[0152] The user reviews the interpretation results and provides feedback on them through their device. The emotion engine continuously records the user's feelings during the feedback process.

[0153] Step 9:

[0154] The server accumulates user feedback and sentiment data, using this information to improve the accuracy of future analyses and continuously adjusting the interpretation algorithm.

[0155] (Example 2)

[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0157] When users decipher ancient documents, simple character recognition and translation are insufficient. There is a need for flexible interpretations that reflect the user's emotional state, and for improvements in deciphering accuracy through feedback. However, conventional systems struggle to incorporate emotional information into their interpretations, failing to enhance user understanding and satisfaction.

[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0159] In this invention, the server includes means for receiving image data and extracting character information using character recognition technology; means for applying natural language processing technology to perform grammatical and syntactic analysis on the extracted character information to generate translated data; means for recognizing the user's emotional state using sentiment analysis technology and dynamically adjusting the presentation of interpretation results based on that emotional information; and means for collecting emotional information from the user as feedback data and using it to improve the accuracy of the interpretation results. This enables the provision of appropriate interpretation results according to the user's emotions and continuous improvement of decoding accuracy based on feedback.

[0160] "Image data" refers to data that electronically records the visual information of ancient documents and other types of documents.

[0161] "Character recognition technology" is a technology that reads character information from image data and converts it into digital text.

[0162] "Natural language processing technology" refers to processing techniques that enable computers to understand, analyze, and generate human language.

[0163] "Grammar analysis" is the process of analyzing and understanding the grammatical structure of an input string.

[0164] "Syntax analysis" is the process of identifying the syntactic structure of a string and analyzing its grammatical relationships.

[0165] "Translation data" refers to data that holds information that has been converted from one language to another.

[0166] "Emotional analysis technology" is a method for detecting and classifying an individual's emotional state, and typically uses voice and facial expression data.

[0167] "Interpretation result" refers to the final understanding of the information generated through character recognition and natural language processing.

[0168] "Feedback data" refers to information about user experiences and emotional states collected from users, and is used to improve the system.

[0169] This invention provides a system that offers interpretations of ancient documents while taking into account the user's emotional state. This system operates primarily through the collaboration of a server, a terminal, and the user.

[0170] The user first uses the device's camera function to photograph the ancient document. The image data acquired on the device is then sent to the server using a communication module. To ensure the security of the communication during this process, protocols such as HTTPS are used.

[0171] The server extracts character information from the received image data using character recognition technology. This process utilizes Tesseract OCR, an open-source software. Next, natural language processing techniques are applied to the extracted character information. Specifically, tools such as Python's NLTK library and SpaCy are used to perform grammatical and syntactic analysis and generate translated data.

[0172] Furthermore, the server uses sentiment analysis technology to recognize the user's emotions. The terminal collects the user's facial expressions and voice, which are then analyzed using a cloud-based sentiment analysis service. Based on the user's emotional state derived from the analysis results, the server dynamically adjusts the interpretation and optimizes its presentation. This allows the user to receive a customized learning experience tailored to their individual emotions.

[0173] As a concrete example, when a user deciphers an ancient document from the Sengoku period, if the user's emotional state is analyzed as "interested," the server will additionally present relevant historical background information and links to related documents. In this way, the invention goes beyond mere text deciphering and facilitates understanding by providing context and related knowledge according to the user's emotions.

[0174] An example of a prompt message given to the generating AI model might be: "Take a picture of an ancient document from the Sengoku period and display the interpretation result. If the user's emotion is determined to be 'interested,' also add relevant historical background information." Based on this prompt message, the system will present the user with an appropriate interpretation result.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The device photographs the ancient document. The user takes a picture of the ancient document with the device's camera app and obtains the image data. This image data becomes the input for the next process and is the initial data for the user to view. The device immediately prepares this data for the next step.

[0178] Step 2:

[0179] The device sends image data to the server. The device sends the captured image data to the server using a secure protocol (e.g., HTTPS). This prepares the server to process the data using character recognition technology.

[0180] Step 3:

[0181] The server performs character recognition. The server applies OCR technology (e.g., Tesseract OCR) to extract character information from image data. At this stage, data processing is performed to convert the image into digital text, and text data is output. The server uses this as the basis for translation.

[0182] Step 4:

[0183] The server performs natural language processing. The server applies natural language processing techniques to the extracted text data, performing grammatical and syntactic analysis. Based on these analysis results, the classical Japanese text is translated into modern Japanese, and the interpretation is output. This ensures that the information is presented in a language understandable to the user.

[0184] Step 5:

[0185] The device collects user emotion data. While the user views the interpretation results, the device captures facial expressions with its camera and records audio with its microphone. This data is used as input for emotion analysis.

[0186] Step 6:

[0187] The server analyzes the emotional state. Based on the collected facial and voice data, the server performs emotion analysis technology. This analysis process identifies the user's emotional state, and that state is output. The server then adjusts the interpretation results based on this.

[0188] Step 7:

[0189] The server adjusts the interpretation results. Depending on the user's emotional state, the server dynamically changes how the interpretation results are presented. If the user shows interest, it adds relevant information, for example, to output the most suitable results for the user. It may also provide prompts to the generative AI model to generate new information.

[0190] Step 8:

[0191] Users provide feedback. After receiving the interpretation results, users fill out a feedback form with their experience and suggestions for improvement. The feedback data is used in the next step.

[0192] Step 9:

[0193] The server processes the feedback. It analyzes the collected feedback data and identifies areas for system improvement. This data is used to improve the self-learning algorithm, which helps to improve future interpretation accuracy. This allows the system to continuously evolve and provide interpretations that are more suitable for the user.

[0194] (Application Example 2)

[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0196] In conventional systems, when extracting characters from image data and providing interpretations to users, the user's emotional state was not taken into consideration, making it difficult to provide appropriate interpretation information. In particular, there was a lack of customized information tailored to the user's emotions, resulting in a reduced understanding of the interpretation results. Furthermore, the accuracy of interpretation results based on past decoding history had not been sufficiently improved, and there was a lack of effective means to incorporate user feedback.

[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0198] In this invention, the server includes means for extracting character information from image data using character recognition technology, means for analyzing and interpreting the extracted string using natural language processing technology, means for generating and providing the interpretation result to the user, means for recognizing the user's emotional state using emotion analysis technology, and means for dynamically adjusting the interpretation result based on the user's emotions. This makes it possible to provide interpretation results that correspond to the user's emotional state, thereby improving the understanding of the interpretation. Furthermore, the accuracy of the interpretation result improves through a self-learning function, and the system can be improved by incorporating feedback from the user.

[0199] "Character recognition technology" is a technology that extracts character information from image data and converts it into a digital string of characters.

[0200] "Natural language processing technology" refers to techniques for analyzing extracted strings of text and interpreting them grammatically and syntactically.

[0201] "Interpretation results" refer to the content provided to the user, generated based on the perceived and interpreted information.

[0202] "Emotional analysis technology" is a technology that recognizes a user's emotional state by analyzing their facial expressions and voice.

[0203] "Dynamic adjustment" refers to changing the way interpretation results are presented and their content according to the user's emotional state.

[0204] The "self-learning function" is a feature that improves the interpretation accuracy of the system itself based on past decoding history and feedback from users.

[0205] "Feedback" refers to the reactions and opinions that users give regarding the interpretation results, and is information that can be used to improve the system.

[0206] The system implementing this invention begins with the user using a terminal to photograph an ancient document and sending the image data to a server. The server then extracts character information from the received image data using character recognition technology. This process utilizes image processing libraries such as OpenCV and character recognition models.

[0207] The extracted textual information is subjected to grammatical and syntactic analysis using natural language processing (NLP) techniques. During this process, an NLP library is used to translate the analyzed textual information into modern Japanese and generate an interpretation result.

[0208] During the process of generating interpretation results, emotion analysis technology recognizes the user's emotional state. While the user is reviewing the interpretation results, device data such as camera and microphone data are used to analyze emotions from facial expressions and voice. The emotional state is inferred using a deep learning model based on TENSORFLOW®.

[0209] Based on the user's emotional state, the server dynamically adjusts its interpretations and provides the user with specific examples and additional historical context. This allows the user to understand the information more deeply. For example, if the user expresses interest, historical context based on that interest will be displayed.

[0210] For example, when a user deciphers Edo-period commercial records and expresses surprise at the business practices described, additional information about the underlying economic and social circumstances is provided. An example of a prompt using a generative AI model would be: "Please explain in detail the background of the business practices described in the Edo-period commercial records. Please consider the factors that caused your surprise."

[0211] In this way, the system is designed to function fully through the cooperation of the terminal, server, and user. As a result, interpretations can be dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation.

[0212] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0213] Step 1:

[0214] The user takes a photograph of an ancient document using their device. The captured image data becomes the system's input. The device's camera is used to acquire high-quality images, and the data is sent to a server in a cloud environment.

[0215] Step 2:

[0216] The server initiates a character recognition process on the received image data. Using a character recognition model, it extracts character information from the image data and outputs it as a digital string. Image processing libraries such as OpenCV are utilized in this process.

[0217] Step 3:

[0218] The extracted text information is analyzed on the server using natural language processing (NLP) techniques. Grammatical and syntactic analysis is performed to translate ancient documents into modern Japanese. NLP libraries are utilized to understand the text structure and generate interpretation results. The input is a digital string, and the output is a modern Japanese translation corresponding to the original text.

[0219] Step 4:

[0220] After the interpretation results are generated, the server analyzes the user's emotional state using emotion analysis technology. It uses facial image and audio data obtained from the device as input and infers emotions using deep learning models such as TensorFlow. The output is a category representing the user's emotional state.

[0221] Step 5:

[0222] The server dynamically adjusts the interpretation results based on the recognized emotional state of the user. It adds relevant information and annotations according to the emotional state, preparing a customized interpretation result for the user. For example, if the emotion of surprise is detected, historical background information related to surprise will be added.

[0223] Step 6:

[0224] The terminal provides the user with the final interpretation. The adjusted interpretation sent from the server is displayed, allowing the user to gain a deeper understanding through this information. The output information includes a modern language translation of the text, annotations based on the interpretation, and additional information added according to the sentiment.

[0225] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0226] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0227] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0228] [Second Embodiment]

[0229] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0230] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0231] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0232] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0233] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0234] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0235] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0236] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0237] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0238] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0239] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0240] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0241] This invention provides a method for efficiently deciphering ancient documents in a system consisting of servers and terminals connected via a network. The operation method of this system is described in detail below.

[0242] First, the user takes a picture of the ancient document they wish to decipher with their device and uploads the image data to the system. The device then checks the quality of this image data before sending it to the server.

[0243] The server applies character recognition technology to the received image data to extract character information from the image. This process includes pre-processing such as noise reduction and contrast adjustment to optimize the image and improve recognition accuracy.

[0244] Furthermore, the server utilizes natural language processing technology to perform grammatical and syntactic analysis on the extracted text data, thereby understanding the linguistic structure within the ancient documents. This allows for translation into modern Japanese, taking into account the unique expressions and historical context of the ancient documents.

[0245] The server also performs the process of generating interpretations along with understanding the context. By utilizing self-learning capabilities, drawing on historical background and past interpretation results, it improves the accuracy of its interpretations and derives appropriate results. These interpretation results are then transmitted to the terminal and presented to the user.

[0246] Users can view the decryption results on their device screen and provide feedback on the content. This feedback information is managed on a server and used to improve the accuracy of future decryption processes.

[0247] For example, when deciphering a letter from the Heian period, the user first takes a picture of the letter with their device. The server then extracts the characters, analyzes the grammar and syntax, and provides a translation into modern Japanese and a contextual interpretation. In this way, the system helps people understand historical documents even without specialized knowledge.

[0248] The following describes the processing flow.

[0249] Step 1:

[0250] The user takes a picture of the ancient document they want to decipher using their device. During this process, the device checks the image quality and adjusts the resolution and format as needed.

[0251] Step 2:

[0252] The terminal sends the captured image data to the server. It checks whether the data was successfully transmitted over the network and retransmits it if necessary.

[0253] Step 3:

[0254] The server performs preprocessing on the received image data. Specifically, it performs noise reduction and contrast adjustment to prepare the image for improved character recognition accuracy.

[0255] Step 4:

[0256] The server applies optical character recognition (OCR) technology to pre-processed image data to extract characters. It uses algorithms to handle handwritten characters and special fonts to accurately generate character data.

[0257] Step 5:

[0258] The server performs natural language processing on the extracted text data. Through grammatical and syntactic analysis, it lays the foundation for understanding the context of ancient documents.

[0259] Step 6:

[0260] The server performs a translation into modern Japanese based on the analysis results obtained. This translation takes into account the historical context and background to select appropriate terms.

[0261] Step 7:

[0262] The server generates an interpretation based on the translated text. It utilizes historical data and learning algorithms to derive an accurate interpretation.

[0263] Step 8:

[0264] The server sends the generated interpretation results to the terminal.

[0265] Step 9:

[0266] The terminal displays the interpretation results received from the server to the user. The user reviews the interpretation and provides feedback as needed.

[0267] Step 10:

[0268] User feedback is sent to the server and used to improve decoding accuracy. The server stores this data and incorporates it into subsequent processing.

[0269] (Example 1)

[0270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0271] Traditional methods for deciphering ancient documents heavily relied on the expertise of specialists, making it difficult for ordinary users to easily understand them. Furthermore, improving the accuracy of interpretation and translation quality required considerable effort and time, and collecting feedback and improving analysis accuracy was also challenging. Therefore, there is a need for technology that allows for easy deciphering of ancient documents without specialized knowledge and enables continuous improvement of interpretation accuracy.

[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0273] In this invention, the server includes means for inputting an image of a document using a camera and transmitting it to a data processing device, means for extracting character information from the image data using character recognition technology, and means for grammatical and syntactic analysis of the extracted character information using natural language processing technology. This makes it possible for ordinary users to accurately understand ancient documents without specialized knowledge and to constantly improve the accuracy of their interpretations.

[0274] A "photography device" is a device used to acquire documents and images as digital data and transmit them to a processing device.

[0275] A "data processing device" is a device that analyzes received digital data and extracts or translates textual information.

[0276] "Character recognition technology" is a technology that identifies handwritten or printed characters from image data and converts them into text data.

[0277] "Natural language processing technology" is a technique that analyzes text data based on its grammatical structure and meaning, enabling a machine-like understanding of human language.

[0278] A "generative model" is an artificial intelligence model that generates new data or results based on patterns in existing data.

[0279] "Interpretation results" refer to translations and semantic interpretations generated based on processed data.

[0280] "Methods for collecting opinions" refer to methods for obtaining feedback from users and using it to improve the system.

[0281] "Automatic learning function" is a technology that improves system performance by learning from past data and feedback.

[0282] This invention is an advanced integrated system for deciphering ancient documents without requiring specialized knowledge. The system is centered around terminals, servers, and users, each with clearly defined roles.

[0283] First, the user uses hardware such as a portable camera or tablet camera to acquire image data of the document they wish to decipher. After acquisition, the user uploads this data to the system via the network.

[0284] The terminal is responsible for evaluating the quality of the image provided by the user. Specifically, it uses software to inspect the resolution and brightness of the image and determine whether it is suitable for recognition. If the image is inappropriate, it requests the user to retake the photo.

[0285] After receiving the image data, the server applies the character recognition technology developed by our company to extract character information from the image. In this process, preprocessing such as noise removal and identification of character regions is performed, and then digital text is generated through the recognition algorithm.

[0286] Subsequently, the server uses natural language processing technology for the extracted character information. It analyzes the grammar and structure specific to ancient documents and translates them into modern language using the generated AI model. This generation model is excellent at handling complex language structures, especially those of documents from the Heian period.

[0287] The user can view the interpretation result sent from the server on the display of the terminal. Based on the decoding result, the user can understand the content and can further provide feedback. This feedback information is collected by the server and used to improve the analysis accuracy.

[0288] As a specific example, there is a scenario where a user decodes a letter from the Heian period. The user takes a photo of the letter and uploads the image to the system. The server recognizes and analyzes the characters and translates them into modern language considering the expressions of the Heian period. Through this process, the user can understand the content of the letter without specialized knowledge.

[0289] As an example of the prompt text, it is possible to make a request in the form of "Please translate this letter from the Heian period, add an interpretation, and present it in an easy-to-understand manner."

[0290] The flow of the specific process in Example 1 will be described using FIG. 11.

[0291] Step 1:

[0292] The user takes a picture of the document they wish to decipher using a portable device. The captured image is saved to the device, and then the data format (e.g., JPEG, PNG) and resolution are checked. The output is an image file in the appropriate format.

[0293] Step 2:

[0294] The device evaluates image quality. Specifically, it uses software to analyze the image's brightness, contrast, and resolution to determine if it is suitable for character recognition. If the evaluation results do not meet the criteria, the user is notified to retake the image. The output includes the quality evaluation results or the notification sent to the user.

[0295] Step 3:

[0296] When image data is sent to the server, the server applies OCR technology to extract text information from the image. This process involves noise reduction and identification of text regions, and the recognition algorithm converts the text into digital text. The input is image data, and the output is recognized text information.

[0297] Step 4:

[0298] The server uses natural language processing techniques to perform grammatical and syntactic analysis on the obtained text information. The analysis clarifies the grammatical structure of the document, and then it is translated into modern Japanese. A generative AI model is used to output a modern Japanese translation from the input text information.

[0299] Step 5:

[0300] The server understands the context based on the translation results and uses a generative AI model to generate a detailed interpretation. Historical background information is also incorporated into the contextual interpretation, and data calculations are performed to improve the accuracy of the interpretation. The output is a detailed interpretation result.

[0301] Step 6:

[0302] The interpretation result is sent from the server to the terminal, and the user checks the translation and interpretation on the terminal screen. Through a user-centered interface, the decoded content is presented in an easy-to-understand manner. As output, the displayed interpretation result is obtained.

[0303] Step 7:

[0304] The user can provide feedback based on the interpretation result. This feedback is sent to the server, and the system utilizes this information to improve the accuracy of the decoding process. The input is the user's feedback, and the feedback information is obtained as output.

[0305] (Application Example 1)

[0306] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0307] Conventionally, information on paper receipts and invoices has not been digitized, making it difficult to efficiently manage past transactions. Handwritten or old-format information is particularly difficult to recognize and analyze, and there is a need to manage different forms of data in a unified manner.

[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0309] In this invention, the server includes means for extracting content from image data using character recognition technology, means for analyzing and translating the extracted information using natural language processing technology, and means for structuring the extraction result based on past information. This enables the efficient digitization of information on paper media and the unified management of different forms of data.

[0310] "Character recognition technology" is a technology for converting handwritten or printed character information into digital data.

[0311] "Natural language processing technology" is a technology that aims to analyze and understand human language using computers.

[0312] "Means of analysis and translation" are means that have the ability to analyze extracted information grammatically and semantically and convert it into a specific form or language.

[0313] "Self-learning function" refers to a function that improves processing accuracy by learning on its own based on past data and interpretation results.

[0314] A "data processing system" is an information system designed to efficiently process and manage collected data.

[0315] "Information provision" refers to data and feedback obtained from users and external sources, which are used to improve the system's learning and analysis accuracy.

[0316] To implement this invention, three elements are necessary: ​​a server, a terminal, and a user.

[0317] The server possesses high-performance processing capabilities and is equipped with software that can apply character recognition and natural language processing technologies. Specifically, it extracts character information from captured image data using OCR software, interprets it using an NLP library, and translates it. In addition, the server has the function to structure the data and convert it into a data format for integration with electronic payment systems and other systems as needed. For example, it is used when integrating handwritten invoices as digital data into an accounting system.

[0318] The devices used are smartphones and tablets, which use their cameras to photograph the target items, such as receipts or historical documents. The device checks the quality of the captured image data and sends it to the server. A dedicated application running on the latest mobile OS makes this possible and provides a user-friendly UI / UX.

[0319] Users take and transmit photos through an application on their device and check the interpretation results received from the server. They can also provide feedback on the interpretation to the server, which contributes to improving the overall accuracy of the system.

[0320] As a concrete example, a user takes a picture of a handwritten invoice from the 1940s with their smartphone and uploads the data to a server. This data is then structured through OCR and NLP processing and incorporated into a modern accounting system as a digital transaction history. In this process, an example of a prompt for the generating AI model would be, "I have taken a picture of an old-style handwritten invoice. Please recognize all the text from the photo, extract the date, items, and amounts, and translate them into a modern electronic payment format."

[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0322] Step 1:

[0323] The user takes a picture of a receipt or invoice using the device. The input is image data captured via the camera sensor. The output is data saved on the device in image file format. This action completes the first step toward digitizing paper-based information.

[0324] Step 2:

[0325] The device automatically checks the quality of the captured image data, verifying that there is no noise or blur. The input is the image file generated in the previous step. As part of the data processing, the image quality is checked, and as output, the image with confirmed quality is sent to the server. This step includes sending feedback to the user prompting them to retake any unchecked images.

[0326] Step 3:

[0327] The server extracts text information from the received image data using OCR software. The input is an image sent from the terminal. As a data calculation, an image processing algorithm scans the text on the paper and generates digital text as output. This process converts handwritten or printed information into readable text.

[0328] Step 4:

[0329] The server uses natural language processing techniques to grammatically and semantically analyze the extracted text. The input is the digital text generated in the previous step, and the output is structured information data. Data analysis is performed, and prompt sentences are applied to the generative AI model to organize the information. This process organizes different formats into a consistent format.

[0330] Step 5:

[0331] The server converts the structured data into a format compatible with the electronic payment system and integrates it with relevant systems as needed. The input is the structured information generated in step 4, and the output is a dataset for system integration. Data conversion and integration processes are performed, and the information is efficiently integrated. In this step, adjustments are made to the new data format.

[0332] Step 6:

[0333] The system presents the user with the interpretation results from the server on their device. The input is the final data processed by the server, and the output is the interpretation displayed on the device's UI. This allows the user to review the results and provide feedback on their satisfaction. This action improves user trust in the system.

[0334] Step 7:

[0335] Users send feedback on the interpretation results to the server via their device. The input consists of user comments and ratings, while the output is the feedback recorded in the server's feedback database. The server uses this data to self-learn and improve future interpretation accuracy. Continuous improvement through user interaction is an integral part of its operation.

[0336] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0337] This invention incorporates an emotion engine into a system for deciphering ancient documents, thereby enabling the provision of interpretation results that take into account the user's emotional state. The following describes specific embodiments of this system.

[0338] The user takes a picture of an ancient document with their device and sends the image data to the server. The server applies character recognition technology to the image data and extracts the text information. Then, using natural language processing technology, it performs grammatical and syntactic analysis, translates the content of the ancient document into modern Japanese, and generates an interpretation result.

[0339] In this interpretation process, the emotion engine recognizes the user's emotions. Specifically, it analyzes the user's facial expressions and voice while they are viewing the interpretation results or providing feedback to understand their emotional state. This emotional information is reflected in how the interpretation results are presented and in the provision of additional information. For example, if the user expresses distrust, the system can present alternative interpretations or more detailed background information.

[0340] In this way, interpretations are dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation. Furthermore, the user's emotions are taken into consideration during the feedback process, which can be used to improve the accuracy of future interpretations.

[0341] As a concrete example, when a user deciphers an ancient document from the Sengoku period, the emotion engine detects the user's interest. In this case, the server accordingly provides relevant historical background information and links to similar documents, supporting a deeper understanding. This system enables users to not only decipher the text but also to gain a deeper understanding of its context and background in accordance with their emotions.

[0342] The following describes the processing flow.

[0343] Step 1:

[0344] The user takes a picture of the ancient document they want to decipher with their device. The device checks the image resolution and format and adjusts it appropriately before sending it to the server.

[0345] Step 2:

[0346] The device sends the captured data to the server. The server first performs pre-processing, such as noise reduction, on the received image data.

[0347] Step 3:

[0348] The server applies character recognition technology to the pre-processed image data to extract character information. The algorithm is optimized to handle even unusual fonts in handwritten text.

[0349] Step 4:

[0350] The server applies natural language processing techniques to the character data obtained through character recognition, performing grammatical and syntactic analysis to understand the context.

[0351] Step 5:

[0352] Based on the results of natural language processing, the server translates the content of ancient documents into modern language and generates an interpretation. The accuracy of the translation is improved by considering historical context and linguistic peculiarities.

[0353] Step 6:

[0354] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice while viewing the interpretation results. This information is used to adjust the user interface.

[0355] Step 7:

[0356] On the terminal, the interpretation results generated by the server are displayed to the user. At this time, additional information or alternative interpretations may be presented based on the user's emotional state.

[0357] Step 8:

[0358] The user reviews the interpretation results and provides feedback on them through their device. The emotion engine continuously records the user's feelings during the feedback process.

[0359] Step 9:

[0360] The server accumulates user feedback and sentiment data, using this information to improve the accuracy of future analyses and continuously adjusting the interpretation algorithm.

[0361] (Example 2)

[0362] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0363] When users decipher ancient documents, simple character recognition and translation are insufficient. There is a need for flexible interpretations that reflect the user's emotional state, and for improvements in deciphering accuracy through feedback. However, conventional systems struggle to incorporate emotional information into their interpretations, failing to enhance user understanding and satisfaction.

[0364] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0365] In this invention, the server includes means for receiving image data and extracting character information using character recognition technology; means for applying natural language processing technology to perform grammatical and syntactic analysis on the extracted character information to generate translated data; means for recognizing the user's emotional state using sentiment analysis technology and dynamically adjusting the presentation of interpretation results based on that emotional information; and means for collecting emotional information from the user as feedback data and using it to improve the accuracy of the interpretation results. This enables the provision of appropriate interpretation results according to the user's emotions and continuous improvement of decoding accuracy based on feedback.

[0366] "Image data" refers to data that electronically records the visual information of ancient documents and other types of documents.

[0367] "Character recognition technology" is a technology that reads character information from image data and converts it into digital text.

[0368] "Natural language processing technology" refers to processing techniques that enable computers to understand, analyze, and generate human language.

[0369] "Grammar analysis" is the process of analyzing and understanding the grammatical structure of an input string.

[0370] "Syntax analysis" is the process of identifying the syntactic structure of a string and analyzing its grammatical relationships.

[0371] "Translation data" refers to data that holds information that has been converted from one language to another.

[0372] "Emotional analysis technology" is a method for detecting and classifying an individual's emotional state, and typically uses voice and facial expression data.

[0373] "Interpretation result" refers to the final understanding of the information generated through character recognition and natural language processing.

[0374] "Feedback data" refers to information about user experiences and emotional states collected from users, and is used to improve the system.

[0375] This invention provides a system that offers interpretations of ancient documents while taking into account the user's emotional state. This system operates primarily through the collaboration of a server, a terminal, and the user.

[0376] The user first uses the device's camera function to photograph the ancient document. The image data acquired on the device is then sent to the server using a communication module. To ensure the security of the communication during this process, protocols such as HTTPS are used.

[0377] The server extracts character information from the received image data using character recognition technology. This process utilizes Tesseract OCR, an open-source software. Next, natural language processing techniques are applied to the extracted character information. Specifically, tools such as Python's NLTK library and SpaCy are used to perform grammatical and syntactic analysis and generate translated data.

[0378] Furthermore, the server uses sentiment analysis technology to recognize the user's emotions. The terminal collects the user's facial expressions and voice, which are then analyzed using a cloud-based sentiment analysis service. Based on the user's emotional state derived from the analysis results, the server dynamically adjusts the interpretation and optimizes its presentation. This allows the user to receive a customized learning experience tailored to their individual emotions.

[0379] As a concrete example, when a user deciphers an ancient document from the Sengoku period, if the user's emotional state is analyzed as "interested," the server will additionally present relevant historical background information and links to related documents. In this way, the invention goes beyond mere text deciphering and facilitates understanding by providing context and related knowledge according to the user's emotions.

[0380] An example of a prompt message given to the generating AI model might be: "Take a picture of an ancient document from the Sengoku period and display the interpretation result. If the user's emotion is determined to be 'interested,' also add relevant historical background information." Based on this prompt message, the system will present the user with an appropriate interpretation result.

[0381] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0382] Step 1:

[0383] The device photographs the ancient document. The user takes a picture of the ancient document with the device's camera app and obtains the image data. This image data becomes the input for the next process and is the initial data for the user to view. The device immediately prepares this data for the next step.

[0384] Step 2:

[0385] The device sends image data to the server. The device sends the captured image data to the server using a secure protocol (e.g., HTTPS). This prepares the server to process the data using character recognition technology.

[0386] Step 3:

[0387] The server performs character recognition. The server applies OCR technology (e.g., Tesseract OCR) to extract character information from image data. At this stage, data processing is performed to convert the image into digital text, and text data is output. The server uses this as the basis for translation.

[0388] Step 4:

[0389] The server performs natural language processing. The server applies natural language processing techniques to the extracted text data, performing grammatical and syntactic analysis. Based on these analysis results, the classical Japanese text is translated into modern Japanese, and the interpretation is output. This ensures that the information is presented in a language understandable to the user.

[0390] Step 5:

[0391] The device collects user emotion data. While the user views the interpretation results, the device captures facial expressions with its camera and records audio with its microphone. This data is used as input for emotion analysis.

[0392] Step 6:

[0393] The server analyzes the emotional state. Based on the collected facial and voice data, the server performs emotion analysis technology. This analysis process identifies the user's emotional state, and that state is output. The server then adjusts the interpretation results based on this.

[0394] Step 7:

[0395] The server adjusts the interpretation results. Depending on the user's emotional state, the server dynamically changes how the interpretation results are presented. If the user shows interest, it adds relevant information, for example, to output the most suitable results for the user. It may also provide prompts to the generative AI model to generate new information.

[0396] Step 8:

[0397] Users provide feedback. After receiving the interpretation results, users fill out a feedback form with their experience and suggestions for improvement. The feedback data is used in the next step.

[0398] Step 9:

[0399] The server processes the feedback. It analyzes the collected feedback data and identifies areas for system improvement. This data is used to improve the self-learning algorithm, which helps to improve future interpretation accuracy. This allows the system to continuously evolve and provide interpretations that are more suitable for the user.

[0400] (Application Example 2)

[0401] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0402] In conventional systems, when extracting characters from image data and providing interpretations to users, the user's emotional state was not taken into consideration, making it difficult to provide appropriate interpretation information. In particular, there was a lack of customized information tailored to the user's emotions, resulting in a reduced understanding of the interpretation results. Furthermore, the accuracy of interpretation results based on past decoding history had not been sufficiently improved, and there was a lack of effective means to incorporate user feedback.

[0403] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0404] In this invention, the server includes means for extracting character information from image data using character recognition technology, means for analyzing and interpreting the extracted string using natural language processing technology, means for generating and providing the interpretation result to the user, means for recognizing the user's emotional state using emotion analysis technology, and means for dynamically adjusting the interpretation result based on the user's emotions. This makes it possible to provide interpretation results that correspond to the user's emotional state, thereby improving the understanding of the interpretation. Furthermore, the accuracy of the interpretation result improves through a self-learning function, and the system can be improved by incorporating feedback from the user.

[0405] "Character recognition technology" is a technology that extracts character information from image data and converts it into a digital string of characters.

[0406] "Natural language processing technology" refers to techniques for analyzing extracted strings of text and interpreting them grammatically and syntactically.

[0407] "Interpretation results" refer to the content provided to the user, generated based on the perceived and interpreted information.

[0408] "Emotional analysis technology" is a technology that recognizes a user's emotional state by analyzing their facial expressions and voice.

[0409] "Dynamic adjustment" refers to changing the way interpretation results are presented and their content according to the user's emotional state.

[0410] The "self-learning function" is a feature that improves the interpretation accuracy of the system itself based on past decoding history and feedback from users.

[0411] "Feedback" refers to the reactions and opinions that users give regarding the interpretation results, and is information that can be used to improve the system.

[0412] The system implementing this invention begins with the user using a terminal to photograph an ancient document and sending the image data to a server. The server then extracts character information from the received image data using character recognition technology. This process utilizes image processing libraries such as OpenCV and character recognition models.

[0413] The extracted textual information is subjected to grammatical and syntactic analysis using natural language processing (NLP) techniques. During this process, an NLP library is used to translate the analyzed textual information into modern Japanese and generate an interpretation result.

[0414] During the process of generating interpretation results, emotion analysis technology recognizes the user's emotional state. While the user is reviewing the interpretation results, device data such as the camera and microphone are used to analyze emotions from facial expressions and voice. The emotional state is then inferred using a deep learning model based on TensorFlow.

[0415] Based on the user's emotional state, the server dynamically adjusts its interpretations and provides the user with specific examples and additional historical context. This allows the user to understand the information more deeply. For example, if the user expresses interest, historical context based on that interest will be displayed.

[0416] For example, when a user deciphers Edo-period commercial records and expresses surprise at the business practices described, additional information about the underlying economic and social circumstances is provided. An example of a prompt using a generative AI model would be: "Please explain in detail the background of the business practices described in the Edo-period commercial records. Please consider the factors that caused your surprise."

[0417] In this way, the system is designed to function fully through the cooperation of the terminal, server, and user. As a result, interpretations can be dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation.

[0418] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0419] Step 1:

[0420] The user takes a photograph of an ancient document using their device. The captured image data becomes the system's input. The device's camera is used to acquire high-quality images, and the data is sent to a server in a cloud environment.

[0421] Step 2:

[0422] The server initiates a character recognition process on the received image data. Using a character recognition model, it extracts character information from the image data and outputs it as a digital string. Image processing libraries such as OpenCV are utilized in this process.

[0423] Step 3:

[0424] The extracted text information is analyzed on the server using natural language processing (NLP) techniques. Grammatical and syntactic analysis is performed to translate ancient documents into modern Japanese. NLP libraries are utilized to understand the text structure and generate interpretation results. The input is a digital string, and the output is a modern Japanese translation corresponding to the original text.

[0425] Step 4:

[0426] After the interpretation results are generated, the server analyzes the user's emotional state using emotion analysis technology. It uses facial image and audio data obtained from the device as input and infers emotions using deep learning models such as TensorFlow. The output is a category representing the user's emotional state.

[0427] Step 5:

[0428] The server dynamically adjusts the interpretation results based on the recognized emotional state of the user. It adds relevant information and annotations according to the emotional state, preparing a customized interpretation result for the user. For example, if the emotion of surprise is detected, historical background information related to surprise will be added.

[0429] Step 6:

[0430] The terminal provides the user with the final interpretation. The adjusted interpretation sent from the server is displayed, allowing the user to gain a deeper understanding through this information. The output information includes a modern language translation of the text, annotations based on the interpretation, and additional information added according to the sentiment.

[0431] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0432] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0433] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0434] [Third Embodiment]

[0435] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0436] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0437] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0438] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0439] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0440] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0441] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0442] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0443] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0444] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0445] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0446] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0447] This invention provides a method for efficiently deciphering ancient documents in a system consisting of servers and terminals connected via a network. The operation method of this system is described in detail below.

[0448] First, the user takes a picture of the ancient document they wish to decipher with their device and uploads the image data to the system. The device then checks the quality of this image data before sending it to the server.

[0449] The server applies character recognition technology to the received image data to extract character information from the image. This process includes pre-processing such as noise reduction and contrast adjustment to optimize the image and improve recognition accuracy.

[0450] Furthermore, the server utilizes natural language processing technology to perform grammatical and syntactic analysis on the extracted text data, thereby understanding the linguistic structure within the ancient documents. This allows for translation into modern Japanese, taking into account the unique expressions and historical context of the ancient documents.

[0451] The server also performs the process of generating interpretations along with understanding the context. By utilizing self-learning capabilities, drawing on historical background and past interpretation results, it improves the accuracy of its interpretations and derives appropriate results. These interpretation results are then transmitted to the terminal and presented to the user.

[0452] Users can view the decryption results on their device screen and provide feedback on the content. This feedback information is managed on a server and used to improve the accuracy of future decryption processes.

[0453] For example, when deciphering a letter from the Heian period, the user first takes a picture of the letter with their device. The server then extracts the characters, analyzes the grammar and syntax, and provides a translation into modern Japanese and a contextual interpretation. In this way, the system helps people understand historical documents even without specialized knowledge.

[0454] The following describes the processing flow.

[0455] Step 1:

[0456] The user takes a picture of the ancient document they want to decipher using their device. During this process, the device checks the image quality and adjusts the resolution and format as needed.

[0457] Step 2:

[0458] The terminal sends the captured image data to the server. It checks whether the data was successfully transmitted over the network and retransmits it if necessary.

[0459] Step 3:

[0460] The server performs preprocessing on the received image data. Specifically, it performs noise reduction and contrast adjustment to prepare the image for improved character recognition accuracy.

[0461] Step 4:

[0462] The server applies optical character recognition (OCR) technology to pre-processed image data to extract characters. It uses algorithms to handle handwritten characters and special fonts to accurately generate character data.

[0463] Step 5:

[0464] The server performs natural language processing on the extracted text data. Through grammatical and syntactic analysis, it lays the foundation for understanding the context of ancient documents.

[0465] Step 6:

[0466] The server performs a translation into modern Japanese based on the analysis results obtained. This translation takes into account the historical context and background to select appropriate terms.

[0467] Step 7:

[0468] The server generates an interpretation based on the translated text. It utilizes historical data and learning algorithms to derive an accurate interpretation.

[0469] Step 8:

[0470] The server sends the generated interpretation results to the terminal.

[0471] Step 9:

[0472] The terminal displays the interpretation results received from the server to the user. The user reviews the interpretation and provides feedback as needed.

[0473] Step 10:

[0474] User feedback is sent to the server and used to improve decoding accuracy. The server stores this data and incorporates it into subsequent processing.

[0475] (Example 1)

[0476] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0477] Traditional methods for deciphering ancient documents heavily relied on the expertise of specialists, making it difficult for ordinary users to easily understand them. Furthermore, improving the accuracy of interpretation and translation quality required considerable effort and time, and collecting feedback and improving analysis accuracy was also challenging. Therefore, there is a need for technology that allows for easy deciphering of ancient documents without specialized knowledge and enables continuous improvement of interpretation accuracy.

[0478] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0479] In this invention, the server includes means for inputting an image of a document using a camera and transmitting it to a data processing device, means for extracting character information from the image data using character recognition technology, and means for grammatical and syntactic analysis of the extracted character information using natural language processing technology. This makes it possible for ordinary users to accurately understand ancient documents without specialized knowledge and to constantly improve the accuracy of their interpretations.

[0480] A "photography device" is a device used to acquire documents and images as digital data and transmit them to a processing device.

[0481] A "data processing device" is a device that analyzes received digital data and extracts or translates textual information.

[0482] "Character recognition technology" is a technology that identifies handwritten or printed characters from image data and converts them into text data.

[0483] "Natural language processing technology" is a technique that analyzes text data based on its grammatical structure and meaning, enabling a machine-like understanding of human language.

[0484] A "generative model" is an artificial intelligence model that generates new data or results based on patterns in existing data.

[0485] "Interpretation results" refer to translations and semantic interpretations generated based on processed data.

[0486] "Methods for collecting opinions" refer to methods for obtaining feedback from users and using it to improve the system.

[0487] "Automatic learning function" is a technology that improves system performance by learning from past data and feedback.

[0488] This invention is an advanced integrated system for deciphering ancient documents without requiring specialized knowledge. The system is centered around terminals, servers, and users, each with clearly defined roles.

[0489] First, the user uses hardware such as a portable camera or tablet camera to acquire image data of the document they wish to decipher. After acquisition, the user uploads this data to the system via the network.

[0490] The device is responsible for evaluating the quality of images provided by the user. Specifically, it uses software to check the image's resolution and brightness and determine if it is suitable for recognition. If the image is unsuitable, it will ask the user to retake it.

[0491] After receiving image data, the server applies our proprietary character recognition technology to extract character information from the image. This process involves pre-processing such as noise reduction and identification of character regions, after which digital text is generated through a recognition algorithm.

[0492] Next, the server uses natural language processing techniques on the extracted text information. It analyzes the grammar and structure unique to ancient documents and translates them into modern Japanese using a generative AI model. This generative model is particularly adept at handling complex linguistic structures, such as those found in documents from the Heian period.

[0493] Users can view the interpretation results sent from the server on their device's display. Based on the decoded results, users can understand the content and provide further feedback. This feedback information is collected by the server and used to improve the accuracy of the analysis.

[0494] A concrete example is a scenario in which a user deciphers a letter from the Heian period. The user takes a picture of the letter and uploads the image to the system. The server recognizes and analyzes the characters and translates them into modern Japanese, taking into account the expressions used in the Heian period. Through this process, the user can understand the content of the letter even without specialized knowledge.

[0495] As an example of a prompt, you can make a request in the format of, "Please translate this Heian period letter, add an interpretation, and present it in an easy-to-understand format."

[0496] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0497] Step 1:

[0498] The user takes a picture of the document they wish to decipher using a portable device. The captured image is saved to the device, and then the data format (e.g., JPEG, PNG) and resolution are checked. The output is an image file in the appropriate format.

[0499] Step 2:

[0500] The device evaluates image quality. Specifically, it uses software to analyze the image's brightness, contrast, and resolution to determine if it is suitable for character recognition. If the evaluation results do not meet the criteria, the user is notified to retake the image. The output includes the quality evaluation results or the notification sent to the user.

[0501] Step 3:

[0502] When image data is sent to the server, the server applies OCR technology to extract text information from the image. This process involves noise reduction and identification of text regions, and the recognition algorithm converts the text into digital text. The input is image data, and the output is recognized text information.

[0503] Step 4:

[0504] The server uses natural language processing techniques to perform grammatical and syntactic analysis on the obtained text information. The analysis clarifies the grammatical structure of the document, and then it is translated into modern Japanese. A generative AI model is used to output a modern Japanese translation from the input text information.

[0505] Step 5:

[0506] The server understands the context based on the translation results and uses a generative AI model to generate a detailed interpretation. Historical background information is also incorporated into the contextual interpretation, and data calculations are performed to improve the accuracy of the interpretation. The output is a detailed interpretation result.

[0507] Step 6:

[0508] The interpretation results are sent from the server to the terminal, and the user checks the translation and interpretation on the terminal screen. A user-centered interface clearly presents the decoded content. The displayed interpretation result is obtained as output.

[0509] Step 7:

[0510] Users can provide feedback based on the interpretation results. This feedback is sent to the server, and the system uses this information to improve the accuracy of the decoding process. The input is user feedback, and the output is feedback information.

[0511] (Application Example 1)

[0512] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0513] Traditionally, paper receipts and invoices have not been digitized, making it difficult to efficiently manage past transactions. Handwritten or outdated formats are particularly difficult to recognize and analyze, highlighting the need for centralized management of data in different formats.

[0514] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0515] In this invention, the server includes means for extracting content from image data using character recognition technology, means for analyzing and translating the extracted information using natural language processing technology, and means for structuring the extraction results based on past information. This makes it possible to efficiently digitize information on paper media and manage data in different formats in a unified manner.

[0516] "Character recognition technology" is a technology that converts handwritten or printed character information into digital data.

[0517] "Natural language processing technology" is a technology that aims to analyze and understand human language using computers.

[0518] "Means of analysis and translation" are means that have the ability to analyze extracted information grammatically and semantically and convert it into a specific form or language.

[0519] "Self-learning function" refers to a function that improves processing accuracy by learning on its own based on past data and interpretation results.

[0520] A "data processing system" is an information system designed to efficiently process and manage collected data.

[0521] "Information provision" refers to data and feedback obtained from users and external sources, which are used to improve the system's learning and analysis accuracy.

[0522] To implement this invention, three elements are necessary: ​​a server, a terminal, and a user.

[0523] The server possesses high-performance processing capabilities and is equipped with software that can apply character recognition and natural language processing technologies. Specifically, it extracts character information from captured image data using OCR software, interprets it using an NLP library, and translates it. In addition, the server has the function to structure the data and convert it into a data format for integration with electronic payment systems and other systems as needed. For example, it is used when integrating handwritten invoices as digital data into an accounting system.

[0524] The devices used are smartphones and tablets, which use their cameras to photograph the target items, such as receipts or historical documents. The device checks the quality of the captured image data and sends it to the server. A dedicated application running on the latest mobile OS makes this possible and provides a user-friendly UI / UX.

[0525] Users take and transmit photos through an application on their device and check the interpretation results received from the server. They can also provide feedback on the interpretation to the server, which contributes to improving the overall accuracy of the system.

[0526] As a concrete example, a user takes a picture of a handwritten invoice from the 1940s with their smartphone and uploads the data to a server. This data is then structured through OCR and NLP processing and incorporated into a modern accounting system as a digital transaction history. In this process, an example of a prompt for the generating AI model would be, "I have taken a picture of an old-style handwritten invoice. Please recognize all the text from the photo, extract the date, items, and amounts, and translate them into a modern electronic payment format."

[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0528] Step 1:

[0529] The user takes a picture of a receipt or invoice using the device. The input is image data captured via the camera sensor. The output is data saved on the device in image file format. This action completes the first step toward digitizing paper-based information.

[0530] Step 2:

[0531] The device automatically checks the quality of the captured image data, verifying that there is no noise or blur. The input is the image file generated in the previous step. As part of the data processing, the image quality is checked, and as output, the image with confirmed quality is sent to the server. This step includes sending feedback to the user prompting them to retake any unchecked images.

[0532] Step 3:

[0533] The server extracts text information from the received image data using OCR software. The input is an image sent from the terminal. As a data calculation, an image processing algorithm scans the text on the paper and generates digital text as output. This process converts handwritten or printed information into readable text.

[0534] Step 4:

[0535] The server uses natural language processing techniques to grammatically and semantically analyze the extracted text. The input is the digital text generated in the previous step, and the output is structured information data. Data analysis is performed, and prompt sentences are applied to the generative AI model to organize the information. This process organizes different formats into a consistent format.

[0536] Step 5:

[0537] The server converts the structured data into a format compatible with the electronic payment system and integrates it with relevant systems as needed. The input is the structured information generated in step 4, and the output is a dataset for system integration. Data conversion and integration processes are performed, and the information is efficiently integrated. In this step, adjustments are made to the new data format.

[0538] Step 6:

[0539] The system presents the user with the interpretation results from the server on their device. The input is the final data processed by the server, and the output is the interpretation displayed on the device's UI. This allows the user to review the results and provide feedback on their satisfaction. This action improves user trust in the system.

[0540] Step 7:

[0541] Users send feedback on the interpretation results to the server via their device. The input consists of user comments and ratings, while the output is the feedback recorded in the server's feedback database. The server uses this data to self-learn and improve future interpretation accuracy. Continuous improvement through user interaction is an integral part of its operation.

[0542] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0543] This invention incorporates an emotion engine into a system for deciphering ancient documents, thereby enabling the provision of interpretation results that take into account the user's emotional state. The following describes specific embodiments of this system.

[0544] The user takes a picture of an ancient document with their device and sends the image data to the server. The server applies character recognition technology to the image data and extracts the text information. Then, using natural language processing technology, it performs grammatical and syntactic analysis, translates the content of the ancient document into modern Japanese, and generates an interpretation result.

[0545] In this interpretation process, the emotion engine recognizes the user's emotions. Specifically, it analyzes the user's facial expressions and voice while they are viewing the interpretation results or providing feedback to understand their emotional state. This emotional information is reflected in how the interpretation results are presented and in the provision of additional information. For example, if the user expresses distrust, the system can present alternative interpretations or more detailed background information.

[0546] In this way, interpretations are dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation. Furthermore, the user's emotions are taken into consideration during the feedback process, which can be used to improve the accuracy of future interpretations.

[0547] As a concrete example, when a user deciphers an ancient document from the Sengoku period, the emotion engine detects the user's interest. In this case, the server accordingly provides relevant historical background information and links to similar documents, supporting a deeper understanding. This system enables users to not only decipher the text but also to gain a deeper understanding of its context and background in accordance with their emotions.

[0548] The following describes the processing flow.

[0549] Step 1:

[0550] The user takes a picture of the ancient document they want to decipher with their device. The device checks the image resolution and format and adjusts it appropriately before sending it to the server.

[0551] Step 2:

[0552] The device sends the captured data to the server. The server first performs pre-processing, such as noise reduction, on the received image data.

[0553] Step 3:

[0554] The server applies character recognition technology to the pre-processed image data to extract character information. The algorithm is optimized to handle even unusual fonts in handwritten text.

[0555] Step 4:

[0556] The server applies natural language processing techniques to the character data obtained through character recognition, performing grammatical and syntactic analysis to understand the context.

[0557] Step 5:

[0558] Based on the results of natural language processing, the server translates the content of ancient documents into modern language and generates an interpretation. The accuracy of the translation is improved by considering historical context and linguistic peculiarities.

[0559] Step 6:

[0560] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice while viewing the interpretation results. This information is used to adjust the user interface.

[0561] Step 7:

[0562] On the terminal, the interpretation results generated by the server are displayed to the user. At this time, additional information or alternative interpretations may be presented based on the user's emotional state.

[0563] Step 8:

[0564] The user reviews the interpretation results and provides feedback on them through their device. The emotion engine continuously records the user's feelings during the feedback process.

[0565] Step 9:

[0566] The server accumulates user feedback and sentiment data, using this information to improve the accuracy of future analyses and continuously adjusting the interpretation algorithm.

[0567] (Example 2)

[0568] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0569] When users decipher ancient documents, simple character recognition and translation are insufficient. There is a need for flexible interpretations that reflect the user's emotional state, and for improvements in deciphering accuracy through feedback. However, conventional systems struggle to incorporate emotional information into their interpretations, failing to enhance user understanding and satisfaction.

[0570] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0571] In this invention, the server includes means for receiving image data and extracting character information using character recognition technology; means for applying natural language processing technology to perform grammatical and syntactic analysis on the extracted character information to generate translated data; means for recognizing the user's emotional state using sentiment analysis technology and dynamically adjusting the presentation of interpretation results based on that emotional information; and means for collecting emotional information from the user as feedback data and using it to improve the accuracy of the interpretation results. This enables the provision of appropriate interpretation results according to the user's emotions and continuous improvement of decoding accuracy based on feedback.

[0572] "Image data" refers to data that electronically records the visual information of ancient documents and other types of documents.

[0573] "Character recognition technology" is a technology that reads character information from image data and converts it into digital text.

[0574] "Natural language processing technology" refers to processing techniques that enable computers to understand, analyze, and generate human language.

[0575] "Grammar analysis" is the process of analyzing and understanding the grammatical structure of an input string.

[0576] "Syntax analysis" is the process of identifying the syntactic structure of a string and analyzing its grammatical relationships.

[0577] "Translation data" refers to data that holds information that has been converted from one language to another.

[0578] "Emotional analysis technology" is a method for detecting and classifying an individual's emotional state, and typically uses voice and facial expression data.

[0579] "Interpretation result" refers to the final understanding of the information generated through character recognition and natural language processing.

[0580] "Feedback data" refers to information about user experiences and emotional states collected from users, and is used to improve the system.

[0581] This invention provides a system that offers interpretations of ancient documents while taking into account the user's emotional state. This system operates primarily through the collaboration of a server, a terminal, and the user.

[0582] The user first uses the device's camera function to photograph the ancient document. The image data acquired on the device is then sent to the server using a communication module. To ensure the security of the communication during this process, protocols such as HTTPS are used.

[0583] The server extracts character information from the received image data using character recognition technology. This process utilizes Tesseract OCR, an open-source software. Next, natural language processing techniques are applied to the extracted character information. Specifically, tools such as Python's NLTK library and SpaCy are used to perform grammatical and syntactic analysis and generate translated data.

[0584] Furthermore, the server uses sentiment analysis technology to recognize the user's emotions. The terminal collects the user's facial expressions and voice, which are then analyzed using a cloud-based sentiment analysis service. Based on the user's emotional state derived from the analysis results, the server dynamically adjusts the interpretation and optimizes its presentation. This allows the user to receive a customized learning experience tailored to their individual emotions.

[0585] As a concrete example, when a user deciphers an ancient document from the Sengoku period, if the user's emotional state is analyzed as "interested," the server will additionally present relevant historical background information and links to related documents. In this way, the invention goes beyond mere text deciphering and facilitates understanding by providing context and related knowledge according to the user's emotions.

[0586] An example of a prompt message given to the generating AI model might be: "Take a picture of an ancient document from the Sengoku period and display the interpretation result. If the user's emotion is determined to be 'interested,' also add relevant historical background information." Based on this prompt message, the system will present the user with an appropriate interpretation result.

[0587] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0588] Step 1:

[0589] The device photographs the ancient document. The user takes a picture of the ancient document with the device's camera app and obtains the image data. This image data becomes the input for the next process and is the initial data for the user to view. The device immediately prepares this data for the next step.

[0590] Step 2:

[0591] The device sends image data to the server. The device sends the captured image data to the server using a secure protocol (e.g., HTTPS). This prepares the server to process the data using character recognition technology.

[0592] Step 3:

[0593] The server performs character recognition. The server applies OCR technology (e.g., Tesseract OCR) to extract character information from image data. At this stage, data processing is performed to convert the image into digital text, and text data is output. The server uses this as the basis for translation.

[0594] Step 4:

[0595] The server performs natural language processing. The server applies natural language processing techniques to the extracted text data, performing grammatical and syntactic analysis. Based on these analysis results, the classical Japanese text is translated into modern Japanese, and the interpretation is output. This ensures that the information is presented in a language understandable to the user.

[0596] Step 5:

[0597] The device collects user emotion data. While the user views the interpretation results, the device captures facial expressions with its camera and records audio with its microphone. This data is used as input for emotion analysis.

[0598] Step 6:

[0599] The server analyzes the emotional state. Based on the collected facial and voice data, the server performs emotion analysis technology. This analysis process identifies the user's emotional state, and that state is output. The server then adjusts the interpretation results based on this.

[0600] Step 7:

[0601] The server adjusts the interpretation results. Depending on the user's emotional state, the server dynamically changes how the interpretation results are presented. If the user shows interest, it adds relevant information, for example, to output the most suitable results for the user. It may also provide prompts to the generative AI model to generate new information.

[0602] Step 8:

[0603] Users provide feedback. After receiving the interpretation results, users fill out a feedback form with their experience and suggestions for improvement. The feedback data is used in the next step.

[0604] Step 9:

[0605] The server processes the feedback. It analyzes the collected feedback data and identifies areas for system improvement. This data is used to improve the self-learning algorithm, which helps to improve future interpretation accuracy. This allows the system to continuously evolve and provide interpretations that are more suitable for the user.

[0606] (Application Example 2)

[0607] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0608] In conventional systems, when extracting characters from image data and providing interpretations to users, the user's emotional state was not taken into consideration, making it difficult to provide appropriate interpretation information. In particular, there was a lack of customized information tailored to the user's emotions, resulting in a reduced understanding of the interpretation results. Furthermore, the accuracy of interpretation results based on past decoding history had not been sufficiently improved, and there was a lack of effective means to incorporate user feedback.

[0609] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0610] In this invention, the server includes means for extracting character information from image data using character recognition technology, means for analyzing and interpreting the extracted string using natural language processing technology, means for generating and providing the interpretation result to the user, means for recognizing the user's emotional state using emotion analysis technology, and means for dynamically adjusting the interpretation result based on the user's emotions. This makes it possible to provide interpretation results that correspond to the user's emotional state, thereby improving the understanding of the interpretation. Furthermore, the accuracy of the interpretation result improves through a self-learning function, and the system can be improved by incorporating feedback from the user.

[0611] "Character recognition technology" is a technology that extracts character information from image data and converts it into a digital string of characters.

[0612] "Natural language processing technology" refers to techniques for analyzing extracted strings of text and interpreting them grammatically and syntactically.

[0613] "Interpretation results" refer to the content provided to the user, generated based on the perceived and interpreted information.

[0614] "Emotional analysis technology" is a technology that recognizes a user's emotional state by analyzing their facial expressions and voice.

[0615] "Dynamic adjustment" refers to changing the way interpretation results are presented and their content according to the user's emotional state.

[0616] The "self-learning function" is a feature that improves the interpretation accuracy of the system itself based on past decoding history and feedback from users.

[0617] "Feedback" refers to the reactions and opinions that users give regarding the interpretation results, and is information that can be used to improve the system.

[0618] The system implementing this invention begins with the user using a terminal to photograph an ancient document and sending the image data to a server. The server then extracts character information from the received image data using character recognition technology. This process utilizes image processing libraries such as OpenCV and character recognition models.

[0619] The extracted textual information is subjected to grammatical and syntactic analysis using natural language processing (NLP) techniques. During this process, an NLP library is used to translate the analyzed textual information into modern Japanese and generate an interpretation result.

[0620] During the process of generating interpretation results, emotion analysis technology recognizes the user's emotional state. While the user is reviewing the interpretation results, device data such as the camera and microphone are used to analyze emotions from facial expressions and voice. The emotional state is then inferred using a deep learning model based on TensorFlow.

[0621] Based on the user's emotional state, the server dynamically adjusts its interpretations and provides the user with specific examples and additional historical context. This allows the user to understand the information more deeply. For example, if the user expresses interest, historical context based on that interest will be displayed.

[0622] For example, when a user deciphers Edo-period commercial records and expresses surprise at the business practices described, additional information about the underlying economic and social circumstances is provided. An example of a prompt using a generative AI model would be: "Please explain in detail the background of the business practices described in the Edo-period commercial records. Please consider the factors that caused your surprise."

[0623] In this way, the system is designed to function fully through the cooperation of the terminal, server, and user. As a result, interpretations can be dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation.

[0624] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0625] Step 1:

[0626] The user takes a photograph of an ancient document using their device. The captured image data becomes the system's input. The device's camera is used to acquire high-quality images, and the data is sent to a server in a cloud environment.

[0627] Step 2:

[0628] The server initiates a character recognition process on the received image data. Using a character recognition model, it extracts character information from the image data and outputs it as a digital string. Image processing libraries such as OpenCV are utilized in this process.

[0629] Step 3:

[0630] The extracted text information is analyzed on the server using natural language processing (NLP) techniques. Grammatical and syntactic analysis is performed to translate ancient documents into modern Japanese. NLP libraries are utilized to understand the text structure and generate interpretation results. The input is a digital string, and the output is a modern Japanese translation corresponding to the original text.

[0631] Step 4:

[0632] After the interpretation results are generated, the server analyzes the user's emotional state using emotion analysis technology. It uses facial image and audio data obtained from the device as input and infers emotions using deep learning models such as TensorFlow. The output is a category representing the user's emotional state.

[0633] Step 5:

[0634] The server dynamically adjusts the interpretation results based on the recognized emotional state of the user. It adds relevant information and annotations according to the emotional state, preparing a customized interpretation result for the user. For example, if the emotion of surprise is detected, historical background information related to surprise will be added.

[0635] Step 6:

[0636] The terminal provides the user with the final interpretation. The adjusted interpretation sent from the server is displayed, allowing the user to gain a deeper understanding through this information. The output information includes a modern language translation of the text, annotations based on the interpretation, and additional information added according to the sentiment.

[0637] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0638] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0639] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0640] [Fourth Embodiment]

[0641] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0642] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0643] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0644] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0645] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0646] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0647] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0648] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0649] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0650] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0651] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0652] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0653] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0654] This invention provides a method for efficiently deciphering ancient documents in a system consisting of servers and terminals connected via a network. The operation method of this system is described in detail below.

[0655] First, the user takes a picture of the ancient document they wish to decipher with their device and uploads the image data to the system. The device then checks the quality of this image data before sending it to the server.

[0656] The server applies character recognition technology to the received image data to extract character information from the image. This process includes pre-processing such as noise reduction and contrast adjustment to optimize the image and improve recognition accuracy.

[0657] Furthermore, the server utilizes natural language processing technology to perform grammatical and syntactic analysis on the extracted text data, thereby understanding the linguistic structure within the ancient documents. This allows for translation into modern Japanese, taking into account the unique expressions and historical context of the ancient documents.

[0658] The server also performs the process of generating interpretations along with understanding the context. By utilizing self-learning capabilities, drawing on historical background and past interpretation results, it improves the accuracy of its interpretations and derives appropriate results. These interpretation results are then transmitted to the terminal and presented to the user.

[0659] Users can view the decryption results on their device screen and provide feedback on the content. This feedback information is managed on a server and used to improve the accuracy of future decryption processes.

[0660] For example, when deciphering a letter from the Heian period, the user first takes a picture of the letter with their device. The server then extracts the characters, analyzes the grammar and syntax, and provides a translation into modern Japanese and a contextual interpretation. In this way, the system helps people understand historical documents even without specialized knowledge.

[0661] The following describes the processing flow.

[0662] Step 1:

[0663] The user takes a picture of the ancient document they want to decipher using their device. During this process, the device checks the image quality and adjusts the resolution and format as needed.

[0664] Step 2:

[0665] The terminal sends the captured image data to the server. It checks whether the data was successfully transmitted over the network and retransmits it if necessary.

[0666] Step 3:

[0667] The server performs preprocessing on the received image data. Specifically, it performs noise reduction and contrast adjustment to prepare the image for improved character recognition accuracy.

[0668] Step 4:

[0669] The server applies optical character recognition (OCR) technology to pre-processed image data to extract characters. It uses algorithms to handle handwritten characters and special fonts to accurately generate character data.

[0670] Step 5:

[0671] The server performs natural language processing on the extracted text data. Through grammatical and syntactic analysis, it lays the foundation for understanding the context of ancient documents.

[0672] Step 6:

[0673] The server performs a translation into modern Japanese based on the analysis results obtained. This translation takes into account the historical context and background to select appropriate terms.

[0674] Step 7:

[0675] The server generates an interpretation based on the translated text. It utilizes historical data and learning algorithms to derive an accurate interpretation.

[0676] Step 8:

[0677] The server sends the generated interpretation results to the terminal.

[0678] Step 9:

[0679] The terminal displays the interpretation results received from the server to the user. The user reviews the interpretation and provides feedback as needed.

[0680] Step 10:

[0681] User feedback is sent to the server and used to improve decoding accuracy. The server stores this data and incorporates it into subsequent processing.

[0682] (Example 1)

[0683] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0684] Traditional methods for deciphering ancient documents heavily relied on the expertise of specialists, making it difficult for ordinary users to easily understand them. Furthermore, improving the accuracy of interpretation and translation quality required considerable effort and time, and collecting feedback and improving analysis accuracy was also challenging. Therefore, there is a need for technology that allows for easy deciphering of ancient documents without specialized knowledge and enables continuous improvement of interpretation accuracy.

[0685] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0686] In this invention, the server includes means for inputting an image of a document using a camera and transmitting it to a data processing device, means for extracting character information from the image data using character recognition technology, and means for grammatical and syntactic analysis of the extracted character information using natural language processing technology. This makes it possible for ordinary users to accurately understand ancient documents without specialized knowledge and to constantly improve the accuracy of their interpretations.

[0687] A "photography device" is a device used to acquire documents and images as digital data and transmit them to a processing device.

[0688] A "data processing device" is a device that analyzes received digital data and extracts or translates textual information.

[0689] "Character recognition technology" is a technology that identifies handwritten or printed characters from image data and converts them into text data.

[0690] "Natural language processing technology" is a technique that analyzes text data based on its grammatical structure and meaning, enabling a machine-like understanding of human language.

[0691] A "generative model" is an artificial intelligence model that generates new data or results based on patterns in existing data.

[0692] "Interpretation results" refer to translations and semantic interpretations generated based on processed data.

[0693] "Methods for collecting opinions" refer to methods for obtaining feedback from users and using it to improve the system.

[0694] "Automatic learning function" is a technology that improves system performance by learning from past data and feedback.

[0695] This invention is an advanced integrated system for deciphering ancient documents without requiring specialized knowledge. The system is centered around terminals, servers, and users, each with clearly defined roles.

[0696] First, the user uses hardware such as a portable camera or tablet camera to acquire image data of the document they wish to decipher. After acquisition, the user uploads this data to the system via the network.

[0697] The device is responsible for evaluating the quality of images provided by the user. Specifically, it uses software to check the image's resolution and brightness and determine if it is suitable for recognition. If the image is unsuitable, it will ask the user to retake it.

[0698] After receiving image data, the server applies our proprietary character recognition technology to extract character information from the image. This process involves pre-processing such as noise reduction and identification of character regions, after which digital text is generated through a recognition algorithm.

[0699] Next, the server uses natural language processing techniques on the extracted text information. It analyzes the grammar and structure unique to ancient documents and translates them into modern Japanese using a generative AI model. This generative model is particularly adept at handling complex linguistic structures, such as those found in documents from the Heian period.

[0700] Users can view the interpretation results sent from the server on their device's display. Based on the decoded results, users can understand the content and provide further feedback. This feedback information is collected by the server and used to improve the accuracy of the analysis.

[0701] A concrete example is a scenario in which a user deciphers a letter from the Heian period. The user takes a picture of the letter and uploads the image to the system. The server recognizes and analyzes the characters and translates them into modern Japanese, taking into account the expressions used in the Heian period. Through this process, the user can understand the content of the letter even without specialized knowledge.

[0702] As an example of a prompt, you can make a request in the format of, "Please translate this Heian period letter, add an interpretation, and present it in an easy-to-understand format."

[0703] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0704] Step 1:

[0705] The user takes a picture of the document they wish to decipher using a portable device. The captured image is saved to the device, and then the data format (e.g., JPEG, PNG) and resolution are checked. The output is an image file in the appropriate format.

[0706] Step 2:

[0707] The device evaluates image quality. Specifically, it uses software to analyze the image's brightness, contrast, and resolution to determine if it is suitable for character recognition. If the evaluation results do not meet the criteria, the user is notified to retake the image. The output includes the quality evaluation results or the notification sent to the user.

[0708] Step 3:

[0709] When image data is sent to the server, the server applies OCR technology to extract text information from the image. This process involves noise reduction and identification of text regions, and the recognition algorithm converts the text into digital text. The input is image data, and the output is recognized text information.

[0710] Step 4:

[0711] The server uses natural language processing techniques to perform grammatical and syntactic analysis on the obtained text information. The analysis clarifies the grammatical structure of the document, and then it is translated into modern Japanese. A generative AI model is used to output a modern Japanese translation from the input text information.

[0712] Step 5:

[0713] The server understands the context based on the translation results and uses a generative AI model to generate a detailed interpretation. Historical background information is also incorporated into the contextual interpretation, and data calculations are performed to improve the accuracy of the interpretation. The output is a detailed interpretation result.

[0714] Step 6:

[0715] The interpretation results are sent from the server to the terminal, and the user checks the translation and interpretation on the terminal screen. A user-centered interface clearly presents the decoded content. The displayed interpretation result is obtained as output.

[0716] Step 7:

[0717] Users can provide feedback based on the interpretation results. This feedback is sent to the server, and the system uses this information to improve the accuracy of the decoding process. The input is user feedback, and the output is feedback information.

[0718] (Application Example 1)

[0719] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0720] Traditionally, paper receipts and invoices have not been digitized, making it difficult to efficiently manage past transactions. Handwritten or outdated formats are particularly difficult to recognize and analyze, highlighting the need for centralized management of data in different formats.

[0721] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0722] In this invention, the server includes means for extracting content from image data using character recognition technology, means for analyzing and translating the extracted information using natural language processing technology, and means for structuring the extraction results based on past information. This makes it possible to efficiently digitize information on paper media and manage data in different formats in a unified manner.

[0723] "Character recognition technology" is a technology that converts handwritten or printed character information into digital data.

[0724] "Natural language processing technology" is a technology that aims to analyze and understand human language using computers.

[0725] "Means of analysis and translation" are means that have the ability to analyze extracted information grammatically and semantically and convert it into a specific form or language.

[0726] "Self-learning function" refers to a function that improves processing accuracy by learning on its own based on past data and interpretation results.

[0727] A "data processing system" is an information system designed to efficiently process and manage collected data.

[0728] "Information provision" refers to data and feedback obtained from users and external sources, which are used to improve the system's learning and analysis accuracy.

[0729] To implement this invention, three elements are necessary: ​​a server, a terminal, and a user.

[0730] The server possesses high-performance processing capabilities and is equipped with software that can apply character recognition and natural language processing technologies. Specifically, it extracts character information from captured image data using OCR software, interprets it using an NLP library, and translates it. In addition, the server has the function to structure the data and convert it into a data format for integration with electronic payment systems and other systems as needed. For example, it is used when integrating handwritten invoices as digital data into an accounting system.

[0731] The devices used are smartphones and tablets, which use their cameras to photograph the target items, such as receipts or historical documents. The device checks the quality of the captured image data and sends it to the server. A dedicated application running on the latest mobile OS makes this possible and provides a user-friendly UI / UX.

[0732] Users take and transmit photos through an application on their device and check the interpretation results received from the server. They can also provide feedback on the interpretation to the server, which contributes to improving the overall accuracy of the system.

[0733] As a concrete example, a user takes a picture of a handwritten invoice from the 1940s with their smartphone and uploads the data to a server. This data is then structured through OCR and NLP processing and incorporated into a modern accounting system as a digital transaction history. In this process, an example of a prompt for the generating AI model would be, "I have taken a picture of an old-style handwritten invoice. Please recognize all the text from the photo, extract the date, items, and amounts, and translate them into a modern electronic payment format."

[0734] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0735] Step 1:

[0736] The user takes a picture of a receipt or invoice using the device. The input is image data captured via the camera sensor. The output is data saved on the device in image file format. This action completes the first step toward digitizing paper-based information.

[0737] Step 2:

[0738] The device automatically checks the quality of the captured image data, verifying that there is no noise or blur. The input is the image file generated in the previous step. As part of the data processing, the image quality is checked, and as output, the image with confirmed quality is sent to the server. This step includes sending feedback to the user prompting them to retake any unchecked images.

[0739] Step 3:

[0740] The server extracts text information from the received image data using OCR software. The input is an image sent from the terminal. As a data calculation, an image processing algorithm scans the text on the paper and generates digital text as output. This process converts handwritten or printed information into readable text.

[0741] Step 4:

[0742] The server uses natural language processing techniques to grammatically and semantically analyze the extracted text. The input is the digital text generated in the previous step, and the output is structured information data. Data analysis is performed, and prompt sentences are applied to the generative AI model to organize the information. This process organizes different formats into a consistent format.

[0743] Step 5:

[0744] The server converts the structured data into a format compatible with the electronic payment system and integrates it with relevant systems as needed. The input is the structured information generated in step 4, and the output is a dataset for system integration. Data conversion and integration processes are performed, and the information is efficiently integrated. In this step, adjustments are made to the new data format.

[0745] Step 6:

[0746] The system presents the user with the interpretation results from the server on their device. The input is the final data processed by the server, and the output is the interpretation displayed on the device's UI. This allows the user to review the results and provide feedback on their satisfaction. This action improves user trust in the system.

[0747] Step 7:

[0748] Users send feedback on the interpretation results to the server via their device. The input consists of user comments and ratings, while the output is the feedback recorded in the server's feedback database. The server uses this data to self-learn and improve future interpretation accuracy. Continuous improvement through user interaction is an integral part of its operation.

[0749] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0750] This invention incorporates an emotion engine into a system for deciphering ancient documents, thereby enabling the provision of interpretation results that take into account the user's emotional state. The following describes specific embodiments of this system.

[0751] The user takes a picture of an ancient document with their device and sends the image data to the server. The server applies character recognition technology to the image data and extracts the text information. Then, using natural language processing technology, it performs grammatical and syntactic analysis, translates the content of the ancient document into modern Japanese, and generates an interpretation result.

[0752] In this interpretation process, the emotion engine recognizes the user's emotions. Specifically, it analyzes the user's facial expressions and voice while they are viewing the interpretation results or providing feedback to understand their emotional state. This emotional information is reflected in how the interpretation results are presented and in the provision of additional information. For example, if the user expresses distrust, the system can present alternative interpretations or more detailed background information.

[0753] In this way, interpretations are dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation. Furthermore, the user's emotions are taken into consideration during the feedback process, which can be used to improve the accuracy of future interpretations.

[0754] As a concrete example, when a user deciphers an ancient document from the Sengoku period, the emotion engine detects the user's interest. In this case, the server accordingly provides relevant historical background information and links to similar documents, supporting a deeper understanding. This system enables users to not only decipher the text but also to gain a deeper understanding of its context and background in accordance with their emotions.

[0755] The following describes the processing flow.

[0756] Step 1:

[0757] The user takes a picture of the ancient document they want to decipher with their device. The device checks the image resolution and format and adjusts it appropriately before sending it to the server.

[0758] Step 2:

[0759] The device sends the captured data to the server. The server first performs pre-processing, such as noise reduction, on the received image data.

[0760] Step 3:

[0761] The server applies character recognition technology to the pre-processed image data to extract character information. The algorithm is optimized to handle even unusual fonts in handwritten text.

[0762] Step 4:

[0763] The server applies natural language processing techniques to the character data obtained through character recognition, performing grammatical and syntactic analysis to understand the context.

[0764] Step 5:

[0765] Based on the results of natural language processing, the server translates the content of ancient documents into modern language and generates an interpretation. The accuracy of the translation is improved by considering historical context and linguistic peculiarities.

[0766] Step 6:

[0767] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice while viewing the interpretation results. This information is used to adjust the user interface.

[0768] Step 7:

[0769] On the terminal, the interpretation results generated by the server are displayed to the user. At this time, additional information or alternative interpretations may be presented based on the user's emotional state.

[0770] Step 8:

[0771] The user reviews the interpretation results and provides feedback on them through their device. The emotion engine continuously records the user's feelings during the feedback process.

[0772] Step 9:

[0773] The server accumulates user feedback and sentiment data, using this information to improve the accuracy of future analyses and continuously adjusting the interpretation algorithm.

[0774] (Example 2)

[0775] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0776] When users decipher ancient documents, simple character recognition and translation are insufficient. There is a need for flexible interpretations that reflect the user's emotional state, and for improvements in deciphering accuracy through feedback. However, conventional systems struggle to incorporate emotional information into their interpretations, failing to enhance user understanding and satisfaction.

[0777] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0778] In this invention, the server includes means for receiving image data and extracting character information using character recognition technology; means for applying natural language processing technology to perform grammatical and syntactic analysis on the extracted character information to generate translated data; means for recognizing the user's emotional state using sentiment analysis technology and dynamically adjusting the presentation of interpretation results based on that emotional information; and means for collecting emotional information from the user as feedback data and using it to improve the accuracy of the interpretation results. This enables the provision of appropriate interpretation results according to the user's emotions and continuous improvement of decoding accuracy based on feedback.

[0779] "Image data" refers to data that electronically records the visual information of ancient documents and other types of documents.

[0780] "Character recognition technology" is a technology that reads character information from image data and converts it into digital text.

[0781] "Natural language processing technology" refers to processing techniques that enable computers to understand, analyze, and generate human language.

[0782] "Grammar analysis" is the process of analyzing and understanding the grammatical structure of an input string.

[0783] "Syntax analysis" is the process of identifying the syntactic structure of a string and analyzing its grammatical relationships.

[0784] "Translation data" refers to data that holds information that has been converted from one language to another.

[0785] "Emotional analysis technology" is a method for detecting and classifying an individual's emotional state, and typically uses voice and facial expression data.

[0786] "Interpretation result" refers to the final understanding of the information generated through character recognition and natural language processing.

[0787] "Feedback data" refers to information about user experiences and emotional states collected from users, and is used to improve the system.

[0788] This invention provides a system that offers interpretations of ancient documents while taking into account the user's emotional state. This system operates primarily through the collaboration of a server, a terminal, and the user.

[0789] The user first uses the device's camera function to photograph the ancient document. The image data acquired on the device is then sent to the server using a communication module. To ensure the security of the communication during this process, protocols such as HTTPS are used.

[0790] The server extracts character information from the received image data using character recognition technology. This process utilizes Tesseract OCR, an open-source software. Next, natural language processing techniques are applied to the extracted character information. Specifically, tools such as Python's NLTK library and SpaCy are used to perform grammatical and syntactic analysis and generate translated data.

[0791] Furthermore, the server uses sentiment analysis technology to recognize the user's emotions. The terminal collects the user's facial expressions and voice, which are then analyzed using a cloud-based sentiment analysis service. Based on the user's emotional state derived from the analysis results, the server dynamically adjusts the interpretation and optimizes its presentation. This allows the user to receive a customized learning experience tailored to their individual emotions.

[0792] As a concrete example, when a user deciphers an ancient document from the Sengoku period, if the user's emotional state is analyzed as "interested," the server will additionally present relevant historical background information and links to related documents. In this way, the invention goes beyond mere text deciphering and facilitates understanding by providing context and related knowledge according to the user's emotions.

[0793] An example of a prompt message given to the generating AI model might be: "Take a picture of an ancient document from the Sengoku period and display the interpretation result. If the user's emotion is determined to be 'interested,' also add relevant historical background information." Based on this prompt message, the system will present the user with an appropriate interpretation result.

[0794] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0795] Step 1:

[0796] The device photographs the ancient document. The user takes a picture of the ancient document with the device's camera app and obtains the image data. This image data becomes the input for the next process and is the initial data for the user to view. The device immediately prepares this data for the next step.

[0797] Step 2:

[0798] The device sends image data to the server. The device sends the captured image data to the server using a secure protocol (e.g., HTTPS). This prepares the server to process the data using character recognition technology.

[0799] Step 3:

[0800] The server performs character recognition. The server applies OCR technology (e.g., Tesseract OCR) to extract character information from image data. At this stage, data processing is performed to convert the image into digital text, and text data is output. The server uses this as the basis for translation.

[0801] Step 4:

[0802] The server performs natural language processing. The server applies natural language processing techniques to the extracted text data, performing grammatical and syntactic analysis. Based on these analysis results, the classical Japanese text is translated into modern Japanese, and the interpretation is output. This ensures that the information is presented in a language understandable to the user.

[0803] Step 5:

[0804] The device collects user emotion data. While the user views the interpretation results, the device captures facial expressions with its camera and records audio with its microphone. This data is used as input for emotion analysis.

[0805] Step 6:

[0806] The server analyzes the emotional state. Based on the collected facial and voice data, the server performs emotion analysis technology. This analysis process identifies the user's emotional state, and that state is output. The server then adjusts the interpretation results based on this.

[0807] Step 7:

[0808] The server adjusts the interpretation results. Depending on the user's emotional state, the server dynamically changes how the interpretation results are presented. If the user shows interest, it adds relevant information, for example, to output the most suitable results for the user. It may also provide prompts to the generative AI model to generate new information.

[0809] Step 8:

[0810] Users provide feedback. After receiving the interpretation results, users fill out a feedback form with their experience and suggestions for improvement. The feedback data is used in the next step.

[0811] Step 9:

[0812] The server processes the feedback. It analyzes the collected feedback data and identifies areas for system improvement. This data is used to improve the self-learning algorithm, which helps to improve future interpretation accuracy. This allows the system to continuously evolve and provide interpretations that are more suitable for the user.

[0813] (Application Example 2)

[0814] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0815] In conventional systems, when extracting characters from image data and providing interpretations to users, the user's emotional state was not taken into consideration, making it difficult to provide appropriate interpretation information. In particular, there was a lack of customized information tailored to the user's emotions, resulting in a reduced understanding of the interpretation results. Furthermore, the accuracy of interpretation results based on past decoding history had not been sufficiently improved, and there was a lack of effective means to incorporate user feedback.

[0816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0817] In this invention, the server includes means for extracting character information from image data using character recognition technology, means for analyzing and interpreting the extracted string using natural language processing technology, means for generating and providing the interpretation result to the user, means for recognizing the user's emotional state using emotion analysis technology, and means for dynamically adjusting the interpretation result based on the user's emotions. This makes it possible to provide interpretation results that correspond to the user's emotional state, thereby improving the understanding of the interpretation. Furthermore, the accuracy of the interpretation result improves through a self-learning function, and the system can be improved by incorporating feedback from the user.

[0818] "Character recognition technology" is a technology that extracts character information from image data and converts it into a digital string of characters.

[0819] "Natural language processing technology" refers to techniques for analyzing extracted strings of text and interpreting them grammatically and syntactically.

[0820] "Interpretation results" refer to the content provided to the user, generated based on the perceived and interpreted information.

[0821] "Emotional analysis technology" is a technology that recognizes a user's emotional state by analyzing their facial expressions and voice.

[0822] "Dynamic adjustment" refers to changing the way interpretation results are presented and their content according to the user's emotional state.

[0823] The "self-learning function" is a feature that improves the interpretation accuracy of the system itself based on past decoding history and feedback from users.

[0824] "Feedback" refers to the reactions and opinions that users give regarding the interpretation results, and is information that can be used to improve the system.

[0825] The system implementing this invention begins with the user using a terminal to photograph an ancient document and sending the image data to a server. The server then extracts character information from the received image data using character recognition technology. This process utilizes image processing libraries such as OpenCV and character recognition models.

[0826] The extracted textual information is subjected to grammatical and syntactic analysis using natural language processing (NLP) techniques. During this process, an NLP library is used to translate the analyzed textual information into modern Japanese and generate an interpretation result.

[0827] During the process of generating interpretation results, emotion analysis technology recognizes the user's emotional state. While the user is reviewing the interpretation results, device data such as the camera and microphone are used to analyze emotions from facial expressions and voice. The emotional state is then inferred using a deep learning model based on TensorFlow.

[0828] Based on the user's emotional state, the server dynamically adjusts its interpretations and provides the user with specific examples and additional historical context. This allows the user to understand the information more deeply. For example, if the user expresses interest, historical context based on that interest will be displayed.

[0829] For example, when a user deciphers Edo-period commercial records and expresses surprise at the business practices described, additional information about the underlying economic and social circumstances is provided. An example of a prompt using a generative AI model would be: "Please explain in detail the background of the business practices described in the Edo-period commercial records. Please consider the factors that caused your surprise."

[0830] In this way, the system is designed to function fully through the cooperation of the terminal, server, and user. As a result, interpretations can be dynamically adjusted according to the user's emotional state, improving the understanding of the interpretation.

[0831] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0832] Step 1:

[0833] The user takes a photograph of an ancient document using their device. The captured image data becomes the system's input. The device's camera is used to acquire high-quality images, and the data is sent to a server in a cloud environment.

[0834] Step 2:

[0835] The server initiates a character recognition process on the received image data. Using a character recognition model, it extracts character information from the image data and outputs it as a digital string. Image processing libraries such as OpenCV are utilized in this process.

[0836] Step 3:

[0837] The extracted text information is analyzed on the server using natural language processing (NLP) techniques. Grammatical and syntactic analysis is performed to translate ancient documents into modern Japanese. NLP libraries are utilized to understand the text structure and generate interpretation results. The input is a digital string, and the output is a modern Japanese translation corresponding to the original text.

[0838] Step 4:

[0839] After the interpretation results are generated, the server analyzes the user's emotional state using emotion analysis technology. It uses facial image and audio data obtained from the device as input and infers emotions using deep learning models such as TensorFlow. The output is a category representing the user's emotional state.

[0840] Step 5:

[0841] The server dynamically adjusts the interpretation results based on the recognized emotional state of the user. It adds relevant information and annotations according to the emotional state, preparing a customized interpretation result for the user. For example, if the emotion of surprise is detected, historical background information related to surprise will be added.

[0842] Step 6:

[0843] The terminal provides the user with the final interpretation. The adjusted interpretation sent from the server is displayed, allowing the user to gain a deeper understanding through this information. The output information includes a modern language translation of the text, annotations based on the interpretation, and additional information added according to the sentiment.

[0844] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0845] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0846] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0847] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0848] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0849] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0850] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0851] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0852] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0853] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0854] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0855] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0856] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0857] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0858] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0859] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0860] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0861] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0862] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0863] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0864] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0865] The following is further disclosed regarding the embodiments described above.

[0866] (Claim 1)

[0867] A means for extracting handwritten or printed characters from image data using character recognition technology,

[0868] A means for performing grammatical and syntactic analysis on extracted strings using natural language processing technology,

[0869] A means of generating interpretation results and providing them to the user,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] The system according to claim 1, which has a self-learning function to improve the accuracy of interpretation results based on past decoding history.

[0873] (Claim 3)

[0874] The system according to claim 1, comprising means for collecting user feedback and reflecting that feedback in future decoding accuracy.

[0875] "Example 1"

[0876] (Claim 1)

[0877] A means for inputting an image of a document using a camera and transmitting it to a data processing device,

[0878] A means of extracting character information from image data using character recognition technology,

[0879] A means for performing grammatical and syntactic analysis of extracted character information using natural language processing technology,

[0880] A means of translating the analyzed information into modern language and generating an interpretation result using a generative model,

[0881] A means of presenting the interpretation results and collecting opinions on their content,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, which improves the accuracy of interpretation results by referring to past processing history and utilizing an automatic learning function.

[0885] (Claim 3)

[0886] The system according to claim 1, comprising means for collecting user feedback and improving the accuracy of subsequent analyses based on that information.

[0887] "Application Example 1"

[0888] (Claim 1)

[0889] A means of extracting content from image data using character recognition technology,

[0890] A means of analyzing and translating extracted information using natural language processing technology,

[0891] A means of generating interpretation results and providing them to the user,

[0892] A means of structuring the extraction results based on past information,

[0893] A means of linking it to a data processing system,

[0894] A system that includes this.

[0895] (Claim 2)

[0896] The system according to claim 1, which has a self-learning function to improve the accuracy of interpretation results based on past analysis history.

[0897] (Claim 3)

[0898] The system according to claim 1, comprising means for collecting information provided by users and reflecting that information in future analyses.

[0899] "Example 2 of combining an emotion engine"

[0900] (Claim 1)

[0901] A means for receiving image data and extracting character information using character recognition technology,

[0902] A means for applying natural language processing technology to extract character information, perform grammatical and syntactic analysis, and generate translated data.

[0903] A means for recognizing the user's emotional state using emotion analysis technology and dynamically adjusting the presentation of interpretation results based on that emotional information,

[0904] A means of collecting information about user emotions as feedback data and using it to improve the accuracy of interpretation results,

[0905] A system that includes this.

[0906] (Claim 2)

[0907] The system according to claim 1, having a self-learning function to improve the accuracy of interpretation results based on past decoding and feedback data.

[0908] (Claim 3)

[0909] The system according to claim 1, comprising means for providing additional information to the interpretation result in accordance with the user's emotional information, thereby improving the user's understanding.

[0910] "Application example 2 when combining with an emotional engine"

[0911] (Claim 1)

[0912] A means of extracting character information from image data using character recognition technology,

[0913] A means of analyzing and interpreting extracted strings using natural language processing technology,

[0914] A means of generating interpretation results and providing them to the user,

[0915] A means of recognizing a user's emotional state using emotion analysis technology,

[0916] A means of dynamically adjusting the interpretation results based on the user's emotions,

[0917] A system that includes this.

[0918] (Claim 2)

[0919] The system according to claim 1, having a self-learning function to improve the accuracy of interpretation results based on past decoding history and user emotional feedback.

[0920] (Claim 3)

[0921] The system according to claim 1, comprising means for collecting feedback from users based on their facial expressions and voice, and for reflecting that feedback in future decoding accuracy. [Explanation of Symbols]

[0922] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for extracting handwritten or printed characters from image data using character recognition technology, A means for performing grammatical and syntactic analysis on extracted strings using natural language processing technology, A means of generating interpretation results and providing them to the user, A system that includes this.

2. The system according to claim 1, which has a self-learning function to improve the accuracy of interpretation results based on past decoding history.

3. The system according to claim 1, comprising means for collecting user feedback and reflecting that feedback in future decoding accuracy.