Image processing device

The image processing device effectively generates a table of contents using a large-scale language model to accurately represent document content by acquiring headings from each page, addressing the inadequacy of existing systems.

JP2025185325APending Publication Date: 2025-12-22KYOCERA DOCUMENT SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024093486
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-10
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

Existing electronic devices fail to generate a table of contents that adequately represents the contents of a document.

Method used

An image processing device that includes a document image acquisition unit, character recognition processing unit, and a table of contents generation unit, which uses a large-scale language model to acquire headings from each page and generate a table of contents based on these headings.

Benefits of technology

Generates a table of contents that appropriately represents the document's content, enhancing its representativeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025185325000001_ABST
    Figure 2025185325000001_ABST
Patent Text Reader

Abstract

To obtain an image processing device which generates a table of contents in which the contents of a document are appropriately expressed.SOLUTION: A document image acquisition unit 11 acquires a document image of a plurality of pages about a certain document. A character recognition processing unit 12 executes character recognition processing with respect to the document image of the plurality of pages to acquire text data. A table of contents generation unit 14 (a) acquires a headline of the page from the text data for each page of the plurality of pages, by using a large-scale language model, and (b) generates a table of contents of the document image of the plurality of pages on the basis of the headline for each page of the plurality of pages.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device. [Background technology]

[0002] An electronic device performs character recognition processing on a document image of a document to obtain a group of character strings, detects the first character strings of headings and paragraphs in the group of character strings, and arranges the detected first character strings of headings and paragraphs to generate a table of contents (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-23591 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the electronic devices described above may not generate a table of contents that adequately represents the contents of the document.

[0005] The present invention has been made in view of the above problems, and has as its object to provide an image processing device that generates a table of contents that appropriately represents the contents of a document. [Means for solving the problem]

[0006] The image processing device of the present invention includes a document image acquisition unit that acquires document images of multiple pages of a document, a character recognition processing unit that performs character recognition processing on the document images of the multiple pages to acquire text data, and a table of contents generation unit that (a) acquires a heading for each page of the multiple pages from the text data for that page using a large-scale language model, and (b) generates a table of contents for the document images of the multiple pages based on the heading for each page of the multiple pages. [Effects of the Invention]

[0007] According to the present invention, an image processing device is provided that generates a table of contents that appropriately represents the contents of a document.

[0008] The above and other objects, features and advantages of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing the configuration of an image processing device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating the format of multi-page text data. [Figure 3] FIG. 3 is a diagram illustrating the format of multi-page text data to which a summary and a table of contents have been added. [Figure 4] FIG. 4 shows examples of a table of contents page image and an abstract page image. [Figure 5] FIG. 5 is a flowchart illustrating the operation of the image processing device shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0011] Fig. 1 is a block diagram showing the configuration of an image processing device according to an embodiment of the present invention. The image processing device shown in Fig. 1 is an information processing device such as a personal computer, or an electronic device such as a digital camera or an image forming device (scanner, multifunction peripheral, etc.), and includes an arithmetic processing device 1, a storage device 2, a communication device 3, a display device 4, an input device 5, an internal device 6, etc.

[0012] The arithmetic processing device 1 includes a computer, which executes programs to function as various processing units. Specifically, the computer includes a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and functions as a predetermined processing unit by loading a program stored in the ROM or storage device 2 into the RAM and executing it on the CPU. The arithmetic processing device 1 may also include an ASIC (Application Specific Integrated Circuit) that functions as a specific processing unit.

[0013] The storage device 2 is a non-volatile storage device such as a flash memory, and stores programs and data necessary for the processes described below. The storage device 2 also stores setting data and the like.

[0014] The communication device 3 is a device that performs data communication with external devices, such as a network interface or a peripheral device interface. The display device 4 is a device that displays various information to the user, such as a display panel such as a liquid crystal display. The input device 5 is a device that detects user operations, such as a keyboard or a touch panel.

[0015] The internal device 6 is a device that executes a predetermined function of the image processing device. For example, if the image processing device is an image forming device, the internal device 6 includes an image reading device that optically reads a document image from a document, a printing device that prints an image on a print sheet, etc.

[0016] Here, the processing device 1 operates as the above-mentioned processing units, namely, a document image acquisition unit 11, a character recognition processing unit 12, a text data management unit 13, a table of contents generation unit 14, a summary generation unit 15, and an output processing unit 16.

[0017] The document image acquisition unit 11 acquires document images (image data) of multiple pages of a certain document from the storage device 2, communication device 3, internal device 6, etc., and stores them in RAM, etc. For example, this document image is read from the document by an image reading device. This document may be, for example, a story, a business document, etc.

[0018] The character recognition processing unit 12 executes character recognition processing on the document image of the multiple pages, and acquires text data of the main text of each of the multiple pages.

[0019] The text data management unit 13 associates the text data for each of the above-mentioned multiple pages with the page number of the page.

[0020] Fig. 2 is a diagram illustrating the format of text data for multiple pages. Here, the text data management unit 13 structures the text data of a document image for multiple pages for each page, and generates data in which elements including the page number and the text data XXXi (i = 1, 2,...,N, where N is the number of pages) of the page body are arranged for each page, as shown in Fig. 2, for example.

[0021] The table of contents generation unit 14 (a) acquires the heading of each page from the text data of each of the above-mentioned multiple pages using a large-scale language model such as GPT or PaLM, and (b) generates a table of contents for the above-mentioned multiple-page document image based on the heading of each of the above-mentioned multiple pages.

[0022] Specifically, the table of contents generation unit 14 uses the communication device 3 to access a server on which a large-scale language model is implemented, generates a prompt for each page including the text data of that page and a command for generating a heading, inputs this to the large-scale language model, and obtains the heading (text data) of that page from the large-scale language model.The table of contents generation unit 14 then arranges the page numbers and headings for the multiple pages described above to generate a table of contents for the document image of the multiple pages.

[0023] The summary generation unit 15 (a) acquires a page summary of each of the above-mentioned multiple pages from the text data of each of the multiple pages using a large-scale language model, and (b) generates a summary of the multiple-page document image from the above-mentioned page summaries of the multiple pages.

[0024] Specifically, the summary generation unit 15 uses the communication device 3 to access a server on which a large-scale language model is implemented, generates a prompt for each page that includes the text data of that page and an instruction for generating a summary, inputs this into the large-scale language model, and obtains a summary (text data) of that page from the large-scale language model.

[0025] In this embodiment, the summary generator 15 specifies a writing style for the large-scale language model when obtaining the page summary described above.

[0026] Specifically, the summary generator 15 includes a specification of the writing style in the above-mentioned prompt, such as "writing in a way that even an elementary school student can understand" or "writing in a professional style."

[0027] This writing style may be selected by the user from the list of strings mentioned above, It may be set according to the document type (story, business document, etc.) described above. The document type may be input by the user, or may be automatically determined from a multi-page document image using an existing method.

[0028] Furthermore, a photographing device may be provided that photographs a user of the image processing device (i.e., a user operating the image processing device) and generates a photographed image of the user, and the writing style may be set in accordance with the user's characteristics (such as age) identified from the photographed image. Note that the user's characteristics are identified from the photographed image using existing person recognition processing or the like.

[0029] Figure 3 is a diagram illustrating the format of multi-page text data to which a summary and table of contents have been added. For example, as shown in Figure 3, page headings YYYi and summary ZZZ are inserted into structured multi-page text data. Note that a table of contents may also be inserted into structured multi-page text data in the same way as summary ZZZ.

[0030] The output processing unit 16 generates page images of the generated table of contents and page images of the generated summary, and outputs these page images (prints, transmits data, saves data, etc.).

[0031] FIG. 4 shows examples of a table of contents page image and a summary page image. For example, a table of contents page image and a summary page image as shown in FIG. 4 are generated. The table of contents and summary shown in FIG. 4 are for a document image of a 10-page document, and the document is the fairy tale "Momotaro." In addition, sentences based on text data are drawn in the table of contents page image and the summary page image in a font size selected so that they fit on one page.

[0032] Next, the operation of the image processing device will be described with reference to Fig. 5, which is a flowchart illustrating the operation of the image processing device shown in Fig. 1.

[0033] When the document image acquisition unit 11 acquires document images of multiple pages (step S1), the character recognition processing unit 12 executes character recognition processing on the document images of multiple pages to acquire text data of multiple pages (step S2).

[0034] The text data management unit 13 structures the text data of the document image of multiple pages, for example, as shown in FIG. 2 (step S3).

[0035] Next, the table of contents generator 14 obtains page headings for each page using the large-scale language model, and the summary generator 15 obtains page summaries for each page (step S4). After that, the table of contents generator 14 generates a table of contents from the page headings as described above, and the summary generator 15 generates a summary from the page summaries (step S5).

[0036] Then, the output processing unit 16 generates page images of the generated table of contents and page images of the generated summary (step S6), and outputs these page images (printing, data transmission, data storage, etc.) (step S7).

[0037] As described above, according to the embodiment, the document image acquisition unit 11 acquires document images of multiple pages of a certain document. The character recognition processing unit 12 executes character recognition processing on the document images of the multiple pages to acquire text data. The table of contents generation unit 14 (a) acquires headings for each of the multiple pages from the text data for that page using a large-scale language model, and (b) generates a table of contents for the document images of the multiple pages based on the headings for each of the multiple pages.

[0038] As a result, the table of contents is generated from page headings that reflect the content of each page of the document, so that the table of contents appropriately expresses the content of the document.

[0039] It should be noted that various changes and modifications to the above-described embodiments will be apparent to those skilled in the art. Such changes and modifications may be made without departing from the spirit and scope of the subject matter and without diminishing its intended advantages. In other words, it is intended that such changes and modifications be included within the scope of the claims.

[0040] For example, in the above embodiment, the summary generation unit 15 may obtain a summary of the above-mentioned multi-page document image using the above-mentioned multi-page text data using a large-scale language model. [Industrial Applicability]

[0041] The present invention is applicable, for example, to the automatic generation of a table of contents for a document. [Explanation of symbols]

[0042] 11 Document image acquisition unit 12 Character recognition processing section 13 Text Data Management Department 14 Table of Contents Generation 15 Summary generator

Claims

1. a document image acquisition unit for acquiring document images of multiple pages of a certain document; a character recognition processing unit that performs character recognition processing on the document images of the plurality of pages to acquire text data; (a) a table of contents generation unit that acquires a heading for each page of the plurality of pages from the text data of the plurality of pages using a large-scale language model; and (b) generates a table of contents for the plurality of pages of the document image based on the heading for each page of the plurality of pages. An image processing device comprising:

2. a summary generator; the summary generation unit (a) acquires a page summary of each page from the text data of each page of the plurality of pages using a large-scale language model; and (b) generates a summary of the document images of the plurality of pages from the page summaries of the plurality of pages.

2. The image processing device according to claim 1, wherein:

3. the summary generator specifies a writing style for the large-scale language model when obtaining the page summary; the writing style is set in accordance with the type of the document; 3. The image processing device according to claim 2, wherein:

4. further comprising a photographing device for photographing a user and generating a photographed image of the user; the summary generator specifies a writing style for the large-scale language model when obtaining the page summary; the writing style is set in accordance with the characteristics of the user identified from the captured image; 3. The image processing device according to claim 2, wherein:

Citation Information

Patent Citations

  • Electronic device

    JP2023023591A