Book-linked AI chat system

The book-linked AI chat system addresses the uncertainty of probabilistic search by using deterministic page specification and multimodal understanding, ensuring accurate and verifiable answers, suitable for high-reliability applications.

JP7791403B1Active Publication Date: 2025-12-24LINGAPORTA CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025110901
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-12-24
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Conventional AI dialogue systems rely on probabilistic search, leading to uncertainty and hallucination in information retrieval, which affects the accuracy and reliability of answers.

Method used

A book-linked AI chat system that uses deterministic lookup based on a specified page in a book, incorporating multimodal understanding to integrate text and visual context, ensuring accurate and reliable answers.

Benefits of technology

The system fundamentally suppresses hallucination and provides inherently verifiable answers by linking responses to a single, unambiguous source, enhancing trust and suitability for regulated industries through auditability and accountability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791403000001_ABST
    Figure 0007791403000001_ABST
Patent Text Reader

Abstract

Conventional search expansion and generation (RAG) techniques rely on probabilistic similarity search. There were two essential problems: information uncertainty (hallucination) and the inability to understand visual elements such as diagrams and layouts of documents. The present invention solves these problems and We provide a new paradigm AI chat system that generates answers interactively with high reliability, verifiability, and auditability. [Solution] The user specifies a book identifier and page number, and using these as the sole key, the system uniquely retrieves a single page image data using a deterministic method that significantly reduces probabilistic searches based on semantic similarity. Multimodal information is then extracted from the retrieved page image data, integrating the complete visual context, including text, figures, and layout structure. Based on this, a large-scale language model generates an answer that reflects the visual structure of the page.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a dialogue system that utilizes artificial intelligence technology, particularly large-scale language models (LLMs) and multimodal document understanding technology. More specifically, the present invention relates to a book-linked AI chat system, its operating method, and program for accurately engaging in dialogue about the content of a specific page of a physical or digital document such as a book, including its visual context. [Background technology]

[0002] In recent years, advances in large-scale language models (LLMs) have led to the widespread adoption of AI chat systems capable of natural conversations with humans. These systems often use Retrieval-Augmented Generation (RAG) technology to generate answers based on specific expertise and up-to-date information. RAG is a technology that searches for relevant information in external knowledge databases and uses that information to generate answers using LLMs. However, conventional RAG systems have inherent problems in their architecture. The core of this problem is that they rely entirely on probabilistic search to identify information sources, which interprets the user's ambiguous question intent and predicts the most appropriate information based on similarity, such as vector search. This probabilistic search does not necessarily guarantee that the most appropriate information will be obtained, and the uncertainty of search results is a major cause of hallucination, where erroneous information is generated. In response to this challenge, research and development in this field has consistently focused on improving the accuracy of probabilistic search. All efforts, such as the development of more powerful embedding models, query expansion, and search result reranking techniques, have focused on how to enable systems to intelligently infer user intent and "discover" more relevant information from a vast information space. This situation suggests that those skilled in the art were trapped by a kind of "technical bias" that focused on solving problems within the framework of probabilistic search. In other words, those skilled in the art were limited to technical thinking about how to improve the performance of probabilistic search, assuming the paradigm of probabilistic search. This prevented them from exploring alternative deterministic architectures that would overturn the paradigm itself. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] US Patent US11301732B2 [Non-patent literature]

[0004] [Non-Patent Document 1] Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020. [Non-patent document 2] Blecher, L. et al., "Nougat: Neural Optical Understanding for Academic Documents", arXiv:2308.13418, 2023. Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention was made in light of the above background technology, and proposes a completely new architectural concept that completely breaks away from the "probabilistic search" paradigm that conventional AI dialogue systems have relied on, and is based on an unshakable information source: a specific page in a book specified by the user. The objective of this invention is to provide a new category of book-linked AI chat system that replaces probabilistic search with deterministic lookup, and furthermore, by having the AI ​​comprehensively understand not only the text information of the referenced page but also the complete visual context, such as charts and layout, it is possible to fundamentally suppress hallucination and generate answers that are highly faithful and reliable, and that are extremely close to human reading comprehension. [Means for solving the problem]

[0006] In order to solve the above problem, the book-linked AI chat system of the present invention comprises a storage means for storing each page of a plurality of books as page image data that is associated one-to-one with a book identifier that identifies the book and a page number that identifies the page, and an input means for accepting questions from a user that include specification of the book identifier and the page number. The most notable feature is that it is equipped with an image acquisition means that uses the specified book identifier and page number as a unique key to deterministically identify and acquire the corresponding single page image data from the storage means without performing a probabilistic search based on semantic similarity. This configuration significantly reduces the generation of misinformation due to uncertainty in the source of information. The system further comprises a document understanding means for extracting multimodal information from the acquired page image data, integrating text information and graphical information indicating the positions and contents of non-text elements, including figures, tables, and mathematical formulas, and an answer generation means for generating an answer that reflects the visual layout and structure of the page, based on the extracted multimodal information and the question, using a large-scale language model. This organic combination of "definitive reference" and "multimodal understanding" is the core of the technical idea of ​​the present invention. [Effects of the Invention]

[0007] According to the present invention, the following remarkable synergistic effects are achieved which cannot be predicted from a simple combination of the constituent elements. (1) Principle-based suppression of hallucination and realization of inherent verifiability: By replacing probabilistic search with "deterministic reference," our invention limits the source of information to a single page image specified by the user. This fundamentally suppresses the occurrence of hallucination due to search uncertainty. Furthermore, when combined with multimodal understanding, this deterministic acquisition produces a qualitatively different effect known as "intrinsic verifiability." That is, because the basis for an AI's answer is always linked to a single, unambiguous authority—"book X, page Y"—users can easily verify the AI's response by referencing the original page. This realizes inherently verifiable AI, something that is impossible to achieve in a probabilistic system where the source of information itself is uncertain, and builds a new relationship of trust between users and AI. (2) Human-like faithful contextual understanding: This invention uses multimodal understanding to analyze a single, reliable page image obtained through deterministic referencing, along with its visual structure. Because the AI ​​retains the complete context, including diagrams, formulas, and layout, which are lost in conventional text fragment (chunk)-based systems, it can generate accurate answers, as if a human were reading, even for advanced questions that assume the visual structure of the page, such as "About the table in the bottom right of this page." This effect is achieved through the synergistic effect of both components, where deep analysis is meaningful only when there is a single, reliable source of information. (3) Suitability for the enterprise sector through auditability and accountability: The inherent verifiability provided by this invention directly supports auditability and accountability, which are crucial for corporate activities. Highly regulated industries, such as financial services, healthcare, and legal services, strongly demand that AI decision-making processes be traceable and their rationale clearly explained. The architecture of this invention consistently links each AI response to auditable evidence—specific pages in specific documents—providing concrete technical solutions to the significant barriers to AI adoption in these industries, namely, the reliability, safety, and compliance challenges. This opens the door to high-value use cases that were previously difficult to apply with traditional, high-risk, "black box" AI. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of a book-linked AI chat system according to an embodiment of the present invention. [Figure 2] Block diagram showing the detailed configuration of the document understanding means [Figure 3] Flowchart showing the operation flow of the present invention [Figure 4] 1 is a flowchart illustrating the information acquisition process of the present invention in comparison with the process of a conventional search expansion generation (RAG) system, which is the background art. [Figure 5] FIG. 1 is a block diagram showing a preferred hardware configuration example of a system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 shows the overall configuration of a book-linked AI chat system 100 according to an embodiment of the present invention. The system 100 includes an input unit 10 , a storage unit 20 , an image acquisition unit 30 , a document understanding unit 40 , an answer generation unit 50 , and an output unit 60 . The storage means 20 stores each page of a book as high-resolution image data, and manages each image data in strict correspondence with a book identifier (eg, ISBN), metadata such as edition information, and page numbers. The input means 10 is an interface that accepts book titles, page numbers, and questions in natural language from the user. The image acquisition means 30 is an important component of the present invention, and uses the input book identifier and page number as keys to uniquely identify and acquire the corresponding single page image data from the storage means 20. The processing here is deterministic, like specifying a file path in a file system, and does not include any probabilistic elements such as similarity search. The document understanding means 40 takes the acquired page image data as input and converts its visual context into structured data. As shown in Figure 2, the text extraction unit 41 extracts all text using OCR processing, the figure / table extraction unit 42 identifies the position and content of figures and tables using image recognition technology, and the layout analysis unit 43 analyzes the structure of the page's paragraphs and headings. These three pieces of information are interrelated and integrated into a single multimodal information. This integrated multimodal information is expressed as a data structure that retains, for example, the content, coordinates, and type (text, figure, table, etc.) of each element, as well as hierarchical relationships between headings and the main text. The answer generation means 50 formats the multimodal information generated by the document understanding means 40 and the user's question as input (prompt) to the large-scale language model (LLM). Based on this rich context information, the LLM generates an answer that takes into account the visual structure of the page. This allows the system to refer to the coordinate information in the structured data and accurately identify the target and generate an answer, even if the user asks a question that includes a relative positional relationship, such as "About the table in the bottom right of the page." The output means 60 presents the generated answer to the user in the form of text, voice, or the like. As shown in the flowchart in Figure 3, the operation of this system consists of a series of clear steps, from accepting user input (S101), to deterministic image data acquisition (S102), multimodal information extraction (S103), answer generation (S104), and answer output (S105). As mentioned above, this embodiment fundamentally solves the problem of search uncertainty that conventional RAG faces by using "deterministic page specification" and further realizes a complete understanding of the page by using "multimodal information extraction." The combination of these two methods produces a synergistic effect that is difficult to predict and is the core of this invention. FIG. 4 is a flowchart showing the information acquisition process of the book-linked AI chat system of the present invention in comparison with the process of a conventional search expansion and generation (RAG) system, which is the background art. A conventional RAG system (Figure 4, top) starts with a query in natural language from the user. The system first interprets the intent of the question and vectorizes its semantic content. Next, this vector is used to calculate the semantic similarity with numerous document chunks in the knowledge database (Similarity Search), and multiple documents that are predicted to be highly relevant are probabilistically retrieved. An answer is generated based on these multiple information sources, including uncertainty. This process, which incorporates "probabilistic search" at its core, is the fundamental cause of hallucinations. In contrast, the system of the present invention (bottom panel of Figure 4) starts the process using a book ID and page number (book ID + page) clearly specified by the user as a single key. Using this key, the system deterministically retrieves the target single page image (single page) directly (Direct Lookup) without going through any probabilistic processes such as calculating semantic similarity. Then, it generates an answer (Generate) based on this single source of information with 100% guaranteed reliability. Thus, Figure 4 clearly shows visually how the present invention transforms the traditional "probabilistic search" paradigm into a "deterministic reference" paradigm, thereby fundamentally resolving the problem of source uncertainty. FIG. 5 is a block diagram showing an example of a suitable hardware configuration for implementing the book-linked AI chat system 100 of the present invention shown in FIG. This system can allocate optimized physical hardware resources to each logical function block in order to realize the logical function block. The index management unit (CPU + RAM) is mainly composed of a CPU and general-purpose memory (RAM). This unit is responsible for quickly resolving the physical storage address of the corresponding image data in the image storage unit from the combination of the book identifier and page number. The deterministic reference processing of the image acquisition means 30 is performed by this index management unit. The image storage unit (SSD / HDD) is composed of a large-capacity storage device such as a solid-state drive (SSD) or a hard disk drive (HDD). As storage means 20, it permanently stores a huge amount of page image data. The document analysis unit (GPU + dedicated memory) is a key feature of this architecture and consists of a graphics processing unit (GPU) and associated high-speed dedicated memory. This unit performs the document understanding process (40), i.e., advanced multimodal analysis (layout analysis, figure extraction, OCR, etc.) of a single acquired page image. The latest document understanding models, such as Nougat, require massive parallel computation for their processing, making the use of a GPU extremely effective. This hardware configuration provides technical evidence for the synergistic effects of the present invention. Deep analysis by the document analysis unit (GPU), which has extremely high computational costs, is possible as an efficient and practical system as a whole precisely because the analysis target is limited to a single reliable page through "deterministic reference" by the index management unit. This demonstrates that the present invention is not simply a combination of concepts, but a concrete and non-obvious system architecture that takes into consideration the efficient use of computing resources. [Industrial Applicability]

[0010] The accuracy, verifiability, and auditability of the answers provided by this invention make it suitable for a wide range of industrial applications requiring extremely high reliability. For example, this technology can be used in educational institutions to provide learning support linked to digital textbooks, libraries to provide interactive reference services for book collections, and publishers to provide reader support. It also provides a level of security and reliability not previously available in high-risk enterprise applications where the "authority" of an answer is as important as the answer itself, such as legal document review, medical record analysis, and financial compliance document queries. This technology will contribute significantly to the development of these industries by providing a level of security and reliability not previously available in the past. [Explanation of symbols]

[0011] 10... Input method 20... Storage means 30... Image acquisition method 40... Document Understanding Methods 41... Text extraction section 42... Chart Extraction Section 43... Layout analysis section 50... Answer generation means 60... Output method 70... Index Management Department 80... Image storage section 90... Document Analysis Section 100... Book-linked AI chat system S101... User input reception S102... Deterministic image data acquisition S103... Multimodal Information Extraction S104... Answer generation S105... Answer output

Claims

1. Claim 1 A book-linked AI chat system that responds to questions about specific pages of a book, a storage means for storing each page of a plurality of books as page image data in one-to-one correspondence with a book identifier for identifying the book and a page number for identifying the page; an input means for receiving a question from a user, the question including the designation of the book identifier and the page number; an image acquisition means for deterministically identifying and acquiring the corresponding single page image data using the specified book identifier and page number as a unique key without performing a probabilistic search based on semantic similarity for all page image data stored in the storage means; a document understanding means for extracting multimodal information integrating text information and diagram information indicating the positions and contents of non-text elements including figures, tables, and mathematical formulas from the acquired page image data; and an answer generation means for generating an answer that reflects the visual layout and structure of the page based on the extracted multimodal information and the question, using a processor that cooperates with a neural network processing unit using a transformer architecture, an image recognition engine, and a text extraction module. A book-linked AI chat system characterized by comprising:

2. The book-linked AI chat system described in claim 1, characterized in that the image acquisition means always acquires the same page image data when the same book identifier and page number are specified, thereby ensuring the reproducibility and verifiability of answer generation.

3. The book-linked AI chat system described in claim 1, characterized in that the document understanding means includes a text extraction unit that extracts text information from the page image data by optical character recognition (OCR) processing, a figure extraction unit that recognizes figures, tables, photographs, and formulas from the page image data along with their position information and extracts figure information, and a layout analysis unit that analyzes paragraph structures, table structures, and heading structures in the page image data and extracts layout information, and integrates the text information, figure information, and layout information to reconstruct the visual context of the page as structured data.

4. The book-linked AI chat system described in claim 1, characterized in that the answer generation means uses the structured data reconstructed by the document understanding means as context information, and generates answers to questions that are based on the diagrams and layout structure of the page by inputting the context information and the question into a large-scale language model.

5. The book-linked AI chat system described in claim 1, characterized in that the storage means stores the image data for multiple books in association with metadata including edition information, publisher information, and ISBN information for each book, and the image acquisition means, in addition to specifying the book identifier and page number, uniquely identifies a specific page of a specific edition by referring to the metadata.

6. The book-linked AI chat system of claim 4, characterized in that the answer generation means responds to questions that are based on the diagrams and layout structure of the page by associating the position information of non-text elements contained in the structured data with linguistic expressions that indicate relative positions, such as "bottom right" or "top," contained in the user's question.

7. The book-linked AI chat system described in claim 3, characterized in that the structured data reconstructed by the document understanding means has a data structure that associates the content of each element (text block, figure, table) within a page, the coordinate information of the element on the page, and the hierarchical structure between the elements.

8. A program for causing a computer to function as the book-linked AI chat system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Book knowledge question-answering system based on generative language model

    CN117421409A

  • Method, program, and information processing device for collecting character information printed on printed material

    JP2025052778A

  • Processing image-bearing electronic documents using a multimodal fusion framework

    US11301732B2