Apparatus and method for document type classification through content area extraction

The integration of text and image processing in document analysis addresses the limitations of separate processing, enabling accurate and efficient document type classification and management.

WO2026005099A1PCT designated stage Publication Date: 2026-01-02ALLBIGDAT INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/010841
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2024-07-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing document analysis technologies primarily process images and text separately, failing to adequately reflect interactions between data types, leading to low analysis accuracy and reliability in documents containing diverse formats like academic papers and business reports.

Method used

A document type classification device and method that integrates text and image processing by extracting and matching relevant text and image areas, generating sentences based on these areas, and classifying document types through content area analysis using deep learning algorithms.

Benefits of technology

Enables accurate and efficient document type classification, allowing users to quickly understand and manage diverse documents, enhancing information retrieval, data mining, and automated report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024010841_02012026_PF_FP_ABST
    Figure KR2024010841_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for document type classification through content area extraction according to an embodiment of the present invention, whereby the type of a document is classified by analyzing a text area and an image area constituting the document, comprises: a document data receiving unit for receiving document data from a user terminal used by a user; and an area recognition unit for distinguishing and recognizing the text area and the image area in the received document data.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for classifying document types through content area extraction

[0001] The present invention relates to a document type classification device and method through content area extraction, and more particularly, to a document type classification device and method through content area extraction that extracts text areas and image areas, matches highly relevant text areas and image areas, and classifies the type of the corresponding document based on the matched data.

[0002] This invention was filed with support from the Gyeonggi Province and the Gyeonggi Province Economic and Science Promotion Agency's '2024 Global Startup Commercialization Support Project.'

[0003] The volume of digital documents is constantly increasing, and these documents contain diverse data formats, including text, images, and diagrams. Existing document analysis technologies primarily focus on processing text or images individually, limiting their ability to integrate and analyze diverse data types. For example, academic papers contain text along with visual elements such as tables, images, and graphs, while business reports incorporate text, diagrams, and images. Separating text and images from these documents makes it difficult to fully understand the overall meaning of the document and accurately identify the content of specific sections.

[0004] Current document type classification technologies process images and text separately and then combine the results. However, this approach fails to adequately reflect the interactions between data, resulting in low analysis accuracy and reliability. Therefore, a technology capable of accurately extracting content within a document by integrating image and text processing is needed, and automatically classifying document types based on this information. This technology enables efficient management and analysis of diverse documents, offering significant benefits in areas such as information retrieval, data mining, and automated report generation.

[0005] [Previous literature]

[0006] Registered Patent No. 10-2063036

[0007] The present invention relates to a document type classification device and method through content area extraction, and more particularly, to a document type classification device and method through content area extraction that extracts text areas and image areas, matches highly relevant text areas and image areas, and classifies the type of the corresponding document based on the matched data.

[0008] According to one embodiment of the present invention, a document type classification device through content area extraction that analyzes a text area and an image area constituting a document and classifies the type of the document includes: a document data receiving unit that receives document data from a user terminal used by a user; an area recognition unit that distinguishes and recognizes a text area and an image area in the received document data; a paragraph recognition unit that divides the recognized text area into sentences and non-sentences based on a terminal word, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; a first sentence generation unit that determines a category of a recognized image area based on sample data for a diagram, and generates a first sentence as the content described by the recognized image area based on the determined category; a second sentence generation unit that analyzes paragraphs located within a set range from the recognized image area and generates a second sentence as the content described by each paragraph; a content area generation unit that compares the second sentences with the first sentence to select a second sentence having the most similar meaning, and generates a content area by matching a paragraph including the second sentence with a recognized image area; and a content area generation unit that generates a content area based on the content of the generated content area. It includes a document type determination unit that determines the type of document and transmits the corresponding contents to the user terminal.

[0009] A document type classification device through content area extraction according to one embodiment of the present invention further includes a non-paragraph extraction unit that extracts a non-paragraph when the text area located most adjacent to the front and back of a recognized image area is a non-paragraph, and a summary data determination unit that determines the content of the non-paragraph as summary data of the recognized image when the extracted non-paragraph is located in front of the recognized image, and determines summary data of the recognized image based on the distance between the non-paragraph and the recognized image and the distance between the area below the non-paragraph when the extracted non-paragraph is located in back of the recognized image, and the first sentence generation unit extracts a change value and a corresponding result value from the image area when the recognized image area is a diagram and generates a first sentence based on the change value and the corresponding result value, and when the recognized image area is not a diagram, recognizes an object shown in the image area and generates a first sentence based on the summary data determined for the recognized object and the corresponding image area, and the second sentence generation unit generates paragraphs through a natural language processing technology using a deep learning model algorithm. Create a second sentence that describes what each of them is about.

[0010] The above summary data determination unit determines that the extracted non-paragraph is located behind the recognized image, and if the gap between the non-paragraph and the recognized image is smaller than the gap with the text area or image area below the non-paragraph, the non-paragraph is determined as summary data of the recognized image, and if the gap is larger than or equal to the gap with the text area or image area below the non-paragraph, the summary data of the recognized image is determined to be non-existent.

[0011] According to one embodiment of the present invention, a document type classification device through content area extraction further includes a sentence analysis unit that analyzes first sentences and second sentences corresponding to each of generated content areas to extract a subject, an object, and a complement of the sentences, a word analysis unit that analyzes words corresponding to the extracted subjects, objects, and complements to extract words of a higher concept and words corresponding to lower concepts of words of the higher concept, and a content area classification unit that classifies the generated content areas according to the extracted words into a higher concept level and a lower concept level, wherein the document type determination unit determines the type of a document according to the higher concept level or the lower concept level at which each of the content areas classified by the content area classification unit is classified.

[0012] According to one embodiment of the present invention, a document type classification device through content area extraction further includes a reference data receiving unit that receives reference data diagrammatically indicating whether each content area is a superordinate concept or a subordinate concept for each document type from an administrator terminal, wherein the content area classification unit classifies the generated content areas according to the highest concept level and the subordinate concept level connected from the highest concept level, and the document type determination unit compares the inclusion ratios of content areas classified by the content area classification unit for each subordinate concept level with the inclusion ratios of content areas in the reference data to select the reference data having the smallest difference in ratio, and determines the document type corresponding to the selected reference data as the document type of the received document data.

[0013] According to one embodiment of the present invention, a method for classifying document types through content area extraction using a document type classification device that analyzes text areas and image areas constituting a document and classifies the type of the document, the steps of a document data receiving unit receiving document data from a user terminal used by a user, a step of a region recognition unit recognizing and distinguishing text areas and image areas from the received document data, a step of a paragraph recognition unit dividing the recognized text area into sentences and non-sentences based on terminal words, recognizing an area including at least two consecutive sentences without a line break as a paragraph, and recognizing the remaining area excluding the recognized paragraph as a non-paragraph, a step of a first sentence generation unit determining a category of the recognized image area based on sample data for a diagram, and generating a first sentence as to the content described by the recognized image area based on the determined category, a step of a second sentence generation unit analyzing paragraphs located within a set range from the recognized image area, and generating a second sentence as to the content described by each paragraph, a step of a content area generation unit comparing the second sentences with the first sentence, selecting a second sentence having the most similar meaning, and matching the paragraph including the second sentence with the recognized image area. It includes a step of creating a content area, and a step of determining the type of document based on the contents of the created content area by a document type determination unit and transmitting the contents to a user terminal.

[0014] The present invention generates content areas by matching image areas and paragraphs in document data where it is difficult to intuitively grasp the document type, and divides the content areas into content areas of a higher concept level and content areas of a lower concept level based on the content explained by each of the generated content areas, and compares this with reference data that is diagrammed as upper and lower concepts of the content area according to a previously stored document type, thereby selecting the most similar reference data and determining the document type corresponding to the reference data as the document type of the document data. Accordingly, the user can quickly grasp the document type of the document data through the present invention, thereby understanding the document data more deeply and understanding the information of the document data more accurately and comprehensively.

[0015] FIG. 1 is a block diagram of a document type classification system through content area extraction according to one embodiment of the present invention.

[0016] FIG. 2 is a block diagram of a document type classification device according to one embodiment of the present invention.

[0017] Figure 3 is a flowchart of a document type classification method according to one embodiment of the present invention.

[0018] According to one embodiment of the present invention, a document type classification device through content area extraction that analyzes a text area and an image area constituting a document and classifies the type of the document includes: a document data receiving unit that receives document data from a user terminal used by a user; an area recognition unit that distinguishes and recognizes a text area and an image area in the received document data; a paragraph recognition unit that divides the recognized text area into sentences and non-sentences based on a terminal word, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; a first sentence generation unit that determines a category of a recognized image area based on sample data for a diagram, and generates a first sentence as the content described by the recognized image area based on the determined category; a second sentence generation unit that analyzes paragraphs located within a set range from the recognized image area and generates a second sentence as the content described by each paragraph; a content area generation unit that compares the second sentences with the first sentence to select a second sentence having the most similar meaning, and generates a content area by matching a paragraph including the second sentence with a recognized image area; and a content area generation unit that generates a content area based on the content of the generated content area. It includes a document type determination unit that determines the type of the document.

[0019] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement them. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar reference numerals have been used throughout the specification to indicate similar elements.

[0020] Throughout the specification, when a part is said to be "connected" to another part, this includes not only cases where the parts are "directly connected," but also cases where the parts are "electrically connected" with other elements intervening. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather includes other components, unless otherwise specifically stated. The present invention will now be described in detail with reference to the accompanying drawings.

[0021] FIG. 1 is a block diagram of a document type classification system (1000) through content area extraction according to one embodiment of the present invention.

[0022] Referring to FIG. 1, a document type classification system (1000) through content area extraction according to one embodiment of the present invention may include a document type classification device (200) connected to a user terminal (100) and a network (400).

[0023] The user terminal (100) may be a terminal used by a person seeking to identify the type of document. For example, the user terminal (100) may be a terminal used by a person who references documents to write papers, reports, marketing materials, etc. Through the present invention, users can quickly identify the type of documents and receive assistance in the process of writing their own documents.

[0024] The user terminal (100) may be a smartphone. However, the present invention is not limited thereto, and the user terminal (100) may include electronic devices such as general desktop computers, navigation systems, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc. The electronic device may have one or more general or special purpose processors, memory, storage, and / or networking components (wired or wireless).

[0025] The document type classification device (200) receives document data from a user terminal (100), matches text areas and image areas in the received document data to create a content area, and determines the document type based on the contents of the content area. The document type classification device (200) may be a server or implemented in the form of an application within the user terminal (100). The document type classification device (200) will be described in more detail with reference to FIGS. 2 and 3 .

[0026] The communication method of the network (400) is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, a wired online network, a wireless online network, a broadcasting network) that the network (400) may include, but also short-range wireless communication between devices. For example, the network (400) may include one or more arbitrary networks (400) among networks (400) such as a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), and an online network.

[0027] FIG. 2 is a block diagram of a document type classification device (200) according to one embodiment of the present invention, and FIG. 3 is a flowchart of a document type classification method according to one embodiment of the present invention.

[0028] Referring to FIGS. 2 and 3, a document type classification device (200) according to one embodiment of the present invention may include a document data receiving unit (201), an area recognition unit (202), a paragraph recognition unit (203), a first sentence generation unit (204), a second sentence generation unit (205), a content area generation unit (206), a document type determination unit (207), a non-paragraph extraction unit (208), a summary data determination unit (209), a sentence analysis unit (210), a word analysis unit (211), a content area classification unit (212), and a reference data receiving unit (213).

[0029] The document data receiving unit (201) can receive document data from the user terminal (100). (S11) The present invention can analyze the received document data to determine the type of the corresponding document data.

[0030] The area recognition unit (202) can distinguish and recognize the text area and the image area in the received document data. (S12) The area recognition unit (202) can distinguish and recognize the text area and the image area in the received document data by referring to sample data for the image and sample data for the character.

[0031] The paragraph recognition unit (203) can divide the recognized text area into sentences and non-sentences based on terminal words, recognize an area including at least two consecutive sentences without a line break as a paragraph, and recognize the remaining area excluding the recognized paragraph as a non-paragraph. (S13) Since a sentence has a terminal word such as ~da, the paragraph recognition unit (203) can recognize an area including at least two consecutive sentences as a paragraph by taking this into consideration.

[0032] The first sentence generation unit (204) can determine the category of the recognized image area based on sample data for the diagram, and generate the content described by the recognized image area as a first sentence based on the determined category (S14).

[0033] The first sentence generation unit (204) can extract change values ​​and corresponding result values ​​from the image area when the recognized image area is a diagram and generate a first sentence based on this. For example, when the recognized image area is a line graph, the first sentence generation unit (204) can determine how the result value corresponding to the Y-axis changes according to the change in the value corresponding to the X-axis through line data and generate the first sentence based on this. As an example of the present invention, the first sentence generation unit (204) can generate the first sentence, "As the corresponding value of the X-axis increases, the corresponding value of the Y-axis gradually decreases."

[0034] The first sentence generation unit (204) can recognize an object depicted in the image area if the recognized image area is not a diagram, and generate a first sentence based on the recognized object.

[0035] The first sentence generation unit (204) can recognize people, animals, objects, etc. included in the image area by analyzing the pixel values ​​of an image area other than a diagram using a machine learning algorithm, and can express the content of the image area by generating a first sentence through the recognized objects.

[0036] The first sentence generation unit (204) can utilize summary data, which will be described later, in the process of generating the first sentence for an image area other than a diagram.

[0037] The non-paragraph extraction unit (208) can extract non-paragraphs if the text area located most adjacent to the front and back of the recognized image area is a non-paragraph. Non-paragraphs can generally be subheadings for paragraphs in a document and titles summarizing image areas.

[0038] The summary data determination unit (209) determines the content of the non-paragraph as summary data of the recognized image if the extracted non-paragraph is located in front of the recognized image, and determines the summary data of the recognized image based on the gap between the non-paragraph and the recognized image and the gap with the area below the non-paragraph if the extracted non-paragraph is located in back of the recognized image. Specifically, the summary data determination unit (209) determines the non-paragraph as summary data of the recognized image if the extracted non-paragraph is located in back of the recognized image and the gap between the non-paragraph and the recognized image is smaller than the gap with the text area or image area below the non-paragraph, and determines that there is no summary data of the recognized image if the gap is greater than or equal to the gap with the text area or image area below the non-paragraph. This generally reflects that a title summarizing the image area is displayed very closely to the front or back of the image area.

[0039] If the recognized image area is not a diagram, and the first sentence for the drawing is generated only with the object recognized in the image area, the first sentence for the drawing may be inaccurate. Therefore, the first sentence generation unit (204) can recognize the object depicted in the image area and generate the first sentence based on the recognized object and the summary data determined for the image area.

[0040] The second sentence generation unit (205) can analyze each paragraph located within a set range from the recognized image area and generate the content described by each paragraph as a second sentence (S15). The set range may be paragraphs located between the non-paragraph closest to the upper side from the recognized image area and the corresponding image area, and the second sentence generation unit (205) can analyze the paragraphs in the range and generate the content described by each paragraph as a second sentence. As an example of the present invention, the second sentence generation unit (205) can generate the content described by each paragraph as a second sentence through a known natural language processing technology using a deep learning model algorithm. However, the method of analyzing the paragraphs is not limited to this, and various known methods of summarizing the content of paragraphs can be utilized in the process of generating the second sentence.

[0041] The content area generation unit (206) can compare the second sentences with the first sentence, select the second sentence that has the most similar meaning, and match the paragraph containing the second sentence with the recognized image area to generate the content area. (S16) The content area generation unit (206) can determine whether the two sentences are similar by synthesizing the results of determining whether words corresponding to the same items are similar to each other when the sentence analysis unit (210) to be described later extracts the subject, object, complement, and predicate of the first and second sentences. The content area generation unit (206) can set the highest weights for the subject and predicate in determining whether sentences are similar, and can quantify the similarity between the first and second sentences by setting the next highest weights for the object, and can select the second sentence with the highest similarity value and match the paragraph containing the selected second sentence with the image area to generate the content area.

[0042] The document type determination unit (207) can determine the type of document based on the contents of the generated content area and transmit the contents to the user terminal (100). (S17) The process by which the document type determination unit (207) determines the type of document data will be described in more detail below.

[0043] The sentence analysis unit (210) can analyze the first and second sentences corresponding to each of the generated content areas and extract the subject, object, and complement of the sentences.

[0044] The word analysis unit (211) can analyze words corresponding to the extracted subjects, objects, and complements to extract words corresponding to higher-level concepts and words corresponding to lower-level concepts of words of the higher-level concepts. The words of the lower-level concepts may be words included in units of the higher-level concepts. As an example of the present invention, if the word of the higher-level concept is 'clothing,' the words of the lower-level concepts may be 'T-shirts', 'pants,' etc.

[0045] The content area classification unit (212) can classify the content areas generated based on the extracted words into upper concept levels and lower concept levels. The content area classification unit (212) can classify the generated content areas based on the highest concept level and the lower concept levels connected from the highest concept level.

[0046] The document type determination unit (207) can determine the type of a document based on the upper or lower concept level at which each of the content areas classified by the content area classification unit (212) is classified. In more detail, the document type determination unit (207) can determine the type of document data through the following process.

[0047] The reference data receiving unit (213) can receive reference data diagrammatically indicating whether each document type is a superordinate concept or a subordinate concept of the content area from the administrator terminal of the present invention.

[0048] The document type determination unit (207) compares the inclusion ratios of the content areas classified by the content area classification unit (212) for each lower concept level with the inclusion ratios of the content areas in the reference data, selects the reference data with the smallest difference in ratios, and determines the document type corresponding to the selected reference data as the document type of the received document data. For example, assuming that the ratio of the content area of ​​the highest concept among the content areas of the document data is 20% of the total number of content areas, the ratio of the content area of ​​the lower concept level 1 is 40% of the total, and the ratio of the content area of ​​the lower concept level 2 is 40% of the total, which is expressed as 20:40:40, and in the case of the reference data of an academic paper, it is 19:38:38:5, and in the case of the reference data of marketing materials, it is 50:30:20, in which case the reference data with the smallest difference in ratios is an academic paper, and therefore, in this case, the document type determination unit (207) can determine the type of the document data as an academic paper and transmit it to the user terminal (100).

[0049] In this way, the present invention generates a content area by matching an image area and a paragraph in document data where it is difficult to intuitively grasp the document type, and divides the content areas into a content area of ​​a higher concept level and a content area of ​​a lower concept level based on the content explained by each of the generated content areas, and compares this with reference data that is diagrammed as an upper and lower concept of the content area according to a previously stored document type, thereby selecting the most similar reference data, and determining the document type corresponding to the reference data as the document type of the document data. Accordingly, the user can quickly grasp the document type of the document data through the present invention, thereby understanding the document data more deeply, and understanding the information of the document data more accurately and comprehensively.

[0050] The embodiments described above are provided for illustrative purposes only, and those skilled in the art will readily appreciate that the embodiments described above can be readily modified into other specific forms without altering the technical concepts or essential characteristics of the embodiments described above. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0051] The scope of protection sought through this specification is indicated by the claims described below rather than by the detailed description, and should be interpreted to include all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts.

Claims

1. In a document type classification device through content area extraction that classifies the type of the document by analyzing the text area and image area that constitute the document, A document data receiving unit that receives document data from a user terminal used by a user; An area recognition unit that recognizes and distinguishes between text areas and image areas in received document data; A paragraph recognition unit that divides a recognized text area into sentences and non-sentences based on terminal words, recognizes an area containing at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; A first sentence generation unit that determines a category of a recognized image region based on sample data for a diagram, and generates a first sentence describing the content of the recognized image region based on the determined category; A second sentence generation unit that analyzes each paragraph located within a set range from the recognized image area and generates a second sentence describing the content of each paragraph; A content area generation unit that compares the second sentences with the first sentence, selects the second sentence with the most similar meaning, and creates a content area by matching the paragraph containing the second sentence with the recognized image area; and A document type classification device through content area extraction, including a document type determination unit that determines the type of document based on the contents of the generated content area and transmits the contents to the user terminal.

2. In paragraph 1, A non-paragraph extraction unit that extracts non-paragraphs when the text area located most adjacent to the front and back of the recognized image area is a non-paragraph; and If the extracted non-paragraph is located in front of the recognized image, the content of the non-paragraph is determined as summary data of the recognized image, and if the extracted non-paragraph is located in back of the recognized image, the summary data of the recognized image is determined based on the gap between the non-paragraph and the recognized image and the gap with the area below the non-paragraph, further including a summary data determination unit. The above first sentence generation unit, if the recognized image area is a diagram, extracts a change value and a corresponding result value from the image area and generates a first sentence based on this. If the recognized image area is not a diagram, it recognizes an object depicted in the image area and generates a first sentence based on summary data determined for the recognized object and the image area. The above second sentence generation unit is a document type classification device through content area extraction characterized in that it generates the content described in each paragraph as a second sentence through natural language processing technology using a deep learning model algorithm.

3. In paragraph 2, The above summary data determination unit determines the non-paragraph as summary data of the recognized image when the extracted non-paragraph is located behind the recognized image and the gap between the non-paragraph and the recognized image is smaller than the gap between the text area or image area below the non-paragraph, and determines that there is no summary data of the recognized image when the gap is greater than or equal to the gap between the text area or image area below the non-paragraph.

4. In paragraph 3, A sentence analysis unit that analyzes the first and second sentences corresponding to each of the generated content areas and extracts the subject, object, and complement of the sentences; A word analysis unit that analyzes words corresponding to the extracted subject, object, and complement and extracts words corresponding to upper concept words and lower concept words of the upper concept words; and It further includes a content area classification unit that classifies the content areas generated based on the extracted words into upper concept level and lower concept level, A document type classification device through content area extraction, characterized in that the above document type determination unit determines the type of the document according to the upper concept level or lower concept level at which each of the content areas classified by the content area classification unit is classified.

5. In paragraph 4, Further comprising a reference data receiving unit that receives reference data diagrammatically indicating whether each document type is a superordinate concept or subordinate concept of the content area from the administrator terminal, The above content area classification section classifies the generated content areas according to the highest concept level and the lower concept level connected from the highest concept level. The above document type determination unit compares the inclusion ratios of content areas classified by the above content area classification unit for each lower concept stage with the inclusion ratios of content areas in the above reference data to select the reference data with the smallest difference in ratios, and determines the document type corresponding to the selected reference data as the document type of the received document data through content area extraction.

6. In a document type classification method through content area extraction using a document type classification device that classifies the type of the document by analyzing the text area and image area that constitute the document, A step in which a document data receiving unit receives document data from a user terminal used by a user; A step in which a region recognition unit recognizes and distinguishes between a text region and an image region in received document data; A step in which a paragraph recognition unit divides a recognized text area into sentences and non-sentences based on terminal words, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; A step in which a first sentence generation unit determines a category of a recognized image area based on sample data for a diagram, and generates a first sentence describing the content of the recognized image area based on the determined category; A step in which a second sentence generation unit analyzes each paragraph located within a set range from the recognized image area and generates a second sentence with the content described by each paragraph; A step of generating a content area by comparing the second sentences with the first sentence, selecting the second sentence that has the most similar meaning, and matching the paragraph containing the second sentence with the recognized image area; and A method for classifying document types through content area extraction, including a step of determining the type of a document based on the contents of a content area in which a document type determination unit is created and transmitting the contents to a user terminal.

Citation Information

Patent Citations

  • Multimodal information fusion document content enhancement retrieval system and method

    CN117312601A

  • In-chart text / chart caption / chart legend / chart kind extraction program, computer-readable recording medium for recording extraction program and in-chart text / chart caption / chart legend / chart kind extraction device

    JP2003346161A

  • Document classification device, program and document classification method

    JP2005122550A

  • Method for extracting document structure and document search method

    JP2007286861A

  • Apparatus and method for automatically creating document

    KR101549792B1