Apparatus and method for document analysis through correlation inference between content areas

The document analysis device infers associations between content areas in documents using deep learning, addressing inefficiencies in existing technologies by generating content areas with logical relationships, enhancing data mining and AI performance.

WO2026005100A1PCT designated stage Publication Date: 2026-01-02ALLBIGDAT INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/010842
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2024-07-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing document analysis technologies fail to infer relationships between content areas, leading to inefficient data mining and semantic inference, particularly in AI-based information retrieval and automated report generation, due to a lack of foundational information for correlating text, diagrams, and graphs.

Method used

A document analysis device and method that infers associations between content areas by recognizing text and image areas, generating sentences from paragraphs and diagrams, and grouping related content areas based on logical relationships using deep learning algorithms.

Benefits of technology

Enables accurate and integrated understanding of document information by generating content areas with logical relationships, facilitating efficient data mining and improved AI performance in information retrieval and automated report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024010842_02012026_PF_FP_ABST
    Figure KR2024010842_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for document analysis through correlation inference between content areas according to an embodiment of the present invention analyzes a document by inferring correlations between content areas constituting the document. The apparatus comprises: a document data receiving unit for receiving document data from a user terminal used by a user; and an area recognizing unit for distinguishing and recognizing a text area and an image area from the received document data.
Need to check novelty before this filing date? Find Prior Art

Description

Document analysis device and method through inference of relevance between content areas

[0001] The present invention relates to a document analysis device and method through inference of association between content areas, and more specifically, to a document analysis device and method through inference of association between content areas that accurately analyzes a document by inferring association between contents such as text, diagrams, and graphs based on the structure of the document.

[0002] This invention was filed with support from the Gyeonggi Province and the Gyeonggi Province Economic and Science Promotion Agency's '2024 Global Startup Commercialization Support Project.'

[0003] As businesses and organizations undergo digital transformation and digital asset transformation, demand for information utilization in AI-based information retrieval, data mining, and automated report generation is growing significantly. Existing document analysis technologies focus on extracting content from documents, limiting their ability to infer relationships between content or provide a comprehensive understanding of documents.

[0004] Information bundle organization is crucial for utilizing generative AI, such as large language models (LLMs). However, due to the lack of foundational information for inferring relationships between content, character-based content bundles are most commonly used. This approach, which prioritizes data processing efficiency over actual relationships, often negatively impacts the performance of data mining, semantic inference, and content-generating AI.

[0005] In the present invention, we aim to describe a document analysis device and method through inference of correlation between content areas, which complements these limitations and enables accurate understanding of document information by inferring correlation between contents.

[0006] [Previous literature]

[0007] Registered Patent No. 10-2541414

[0008] The present invention relates to a document analysis device and method through inference of association between content areas, and more specifically, to a document analysis device and method through inference of association between content areas that accurately analyzes a document by inferring association between contents such as text, diagrams, and graphs based on the structure of the document.

[0009] According to one embodiment of the present invention, a document analysis device through inference of association between content areas that analyzes a document through inference of association between content areas constituting the document, a document data receiving unit that receives document data from a user terminal used by a user, an area recognition unit that recognizes and distinguishes a text area and an image area from the received document data, a paragraph recognition unit that distinguishes sentences and non-sentences based on a terminal word in the recognized text area, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph, a content area generation unit that analyzes each of the recognized paragraphs to extract the content described by each paragraph, and creates a content area by matching paragraphs having a logical relationship among paragraphs located within a set range based on the extracted content, a first sentence generation unit that determines a category of a recognized image area based on sample data for a diagram, and creates a first sentence based on the determined category based on the content described by the recognized image area, a second sentence generation unit that analyzes each of the paragraphs located within a set range from the recognized image area to generate a second sentence based on the content described by each of the paragraphs, and generates a second sentence based on the second sentences. The present invention comprises: an image area additional matching unit that selects a second sentence having the most similar meaning to a first sentence and additionally matches a recognized image area to a content area that includes a paragraph that includes the second sentence; a content area analysis unit that derives the contents of each content area by connecting sentences extracted in the process of generating the content area based on the type of logical relationship between paragraphs determined by the content area generation unit; and a document data information transmission unit that infers related content areas between content areas based on the logical relationship between content areas and generates group information that groups them, and transmits the group information together with the derived content contents to a user terminal.

[0010] The above content area generation unit extracts the content described by the recognized paragraph in sentence format, and divides the words included in the subject, object, and complement of the sentences extracted from each paragraph within the set range into words of a higher concept and words of a lower concept that supplement the words of the higher concept, and matches paragraphs that have a logical relationship based on the sentences that include the words of the higher concept and the words of the lower concept to generate a content area, and the types of logical relationships include description and example relationships, cause and effect relationships, and assertion and basis relationships, and the first sentence generation unit, if the recognized image area is a diagram, extracts a change value and a corresponding result value from the image area and generates a first sentence based on this, and if the recognized image area is not a diagram, recognizes an object depicted in the image area and generates a first sentence based on the recognized object, and the second sentence generation unit generates the content described by each paragraph as a second sentence through a natural language processing technology using a deep learning model algorithm, and the words of the lower concept that supplement the words of the higher concept are words that are examples of the words of the higher concept, and the words of the upper concept are words that are examples of the words of the higher concept, and the words of the lower concept are words that are examples of the words of the higher concept. The content area generation unit includes words that further explain the words of the concept, words that serve as causes and grounds for the words of the higher-level concept, and determines the type of logical relationship between the recognized paragraphs based on the characteristics of the words of these lower-level concepts.

[0011] The above content area generation unit extracts the content described by the paragraph, and based on the extracted content, matches paragraphs that have a logical relationship among the non-paragraph closest to the beginning of the paragraph and the paragraph itself, and among the paragraphs that are closest to the end of the paragraph and the paragraph itself, to generate a content area, and the above content area analysis unit analyzes the paragraphs recognized by the content area generation unit and connects the sentences of the content described by each paragraph using a conjunction corresponding to the determined logical relationship to derive the content of each content area.

[0012] According to one embodiment of the present invention, a document analysis device through inference of correlation between content areas further includes a content area counting unit that divides received document data into at least three sections according to a description order and counts and aggregates the number of content areas included in each section, a content area classification unit that counts and aggregates the types of logical relationships of each of the content areas included in each section and classifies them by section, a document data type determination unit that selects a section in which the largest number of content areas are aggregated and derives a logical relationship in which the largest number is aggregated from the content areas of the selected section, and determines the type of received document data based on the derived logical relationship, and the document data information transmission unit transmits information including the type of determined document data and contents aggregated by the content area counting unit and the content area classification unit to a user terminal.

[0013] According to one embodiment of the present invention, a document analysis device through inference of association between content areas further includes a representative content area selection unit that determines a logical relationship with the largest number of aggregated numbers in content areas of each section of document data, and selects a content area with the largest number of included paragraphs among content areas corresponding to the determined logical relationship as a representative content area for each section, and the document data transmission unit generates group information by grouping the representative content areas for each section by inferring related content areas, and transmits content explaining the logical relationship of each of the selected representative content areas together with the group information and information matching the content derived from the selected representative content areas to a user terminal.

[0014] According to one embodiment of the present invention, a method for analyzing a document through inference of association between content areas using a document analysis device that analyzes a document through inference of association between content areas constituting the document comprises the steps of: a document data receiving unit receiving document data from a user terminal used by a user; a step of a region recognition unit recognizing a text area and an image area by distinguishing them from the received document data; a step of a paragraph recognition unit distinguishing sentences and non-sentences based on a terminal word in the recognized text area, recognizing an area including at least two consecutive sentences without a line break as a paragraph, and recognizing the remaining area excluding the recognized paragraph as a non-paragraph; a step of a content area generation unit analyzing each of the recognized paragraphs to extract the content described by each paragraph, and generating a content area by matching paragraphs having a logical relationship among paragraphs located within a set range based on the extracted content; a step of a first sentence generation unit determining a category of the recognized image area based on sample data for a diagram, and generating a first sentence based on the determined category as the content described by the recognized image area; a step of a second sentence generation unit analyzing each of the paragraphs located within a set range from the recognized image area to generate a second sentence based on the content described by each of the paragraphs. A step of generating a sentence, a step of additionally matching the image area by comparing the second sentences with the first sentence and selecting the second sentence having the most similar meaning, and a step of additionally matching the recognized image area to the content area including the paragraph including the second sentence, a step of deriving the content of each content area according to the logical relationship between the text area or image area included in the content area based on the content extracted in sentence form in the process of generating the content area by the content area analysis unit,And the document data information transmission unit includes a step of inferring related content areas between content areas based on the logical relationship between content areas and grouping them to transmit group information together with the derived content to the user terminal.

[0015] The present invention analyzes document data to generate content areas in which paragraphs and image areas are matched, classifies these in various ways according to the sections of the document data, selects a representative content area for each section based on the classification result, groups these, and provides detailed information to the user, thereby enabling the user to efficiently grasp the relationship between content areas considering the structure of the document data, ultimately enabling a more accurate and integrated understanding of information about the document data.

[0016] FIG. 1 is a block diagram of a document analysis system through inference of association between content areas according to one embodiment of the present invention.

[0017] Figure 2 is a block diagram of a document analysis device according to one embodiment of the present invention.

[0018] Figure 3 is a flowchart of a document analysis method according to one embodiment of the present invention.

[0019] According to one embodiment of the present invention, a document analysis device through inference of association between content areas that analyzes a document through inference of association between content areas constituting the document, a document data receiving unit that receives document data from a user terminal used by a user, an area recognition unit that recognizes and distinguishes a text area and an image area from the received document data, a paragraph recognition unit that distinguishes sentences and non-sentences based on a terminal word in the recognized text area, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph, a content area generation unit that analyzes each of the recognized paragraphs to extract the content described by each paragraph, and creates a content area by matching paragraphs having a logical relationship among paragraphs located within a set range based on the extracted content, a first sentence generation unit that determines a category of a recognized image area based on sample data for a diagram, and creates a first sentence based on the determined category based on the content described by the recognized image area, a second sentence generation unit that analyzes each of the paragraphs located within a set range from the recognized image area to generate a second sentence based on the content described by each of the paragraphs, and generates a second sentence based on the second sentences. The present invention comprises: an image area additional matching unit that selects a second sentence having the most similar meaning to a first sentence and additionally matches a recognized image area to a content area that includes a paragraph that includes the second sentence; a content area analysis unit that derives the contents of each content area by connecting sentences extracted in the process of generating the content area based on the type of logical relationship between paragraphs determined by the content area generation unit; and a document data information transmission unit that infers related content areas between content areas based on the logical relationship between content areas and generates group information that groups them, and transmits the group information together with the derived content contents to a user terminal.

[0020] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement them. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar reference numerals have been used throughout the specification to indicate similar elements.

[0021] Throughout the specification, when a part is said to be "connected" to another part, this includes not only cases where the parts are "directly connected," but also cases where the parts are "electrically connected" with other elements intervening. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather includes other components, unless otherwise specifically stated. The present invention will now be described in detail with reference to the accompanying drawings.

[0022] FIG. 1 is a block diagram of a document analysis system (1000) through inference of correlation between content areas according to one embodiment of the present invention.

[0023] Referring to FIG. 1, a document analysis system (1000) through inference of correlation between content areas according to one embodiment of the present invention may include a document analysis device (200) connected to a user terminal (100) and a network (400).

[0024] The user terminal (100) may be a terminal used by someone who wants to quickly understand the type and details of a document. For example, the user terminal (100) may be a terminal used by someone who references documents to write papers, reports, marketing materials, etc. Through the present invention, users can quickly understand the type and details of documents, thereby assisting them in the process of writing their own documents.

[0025] The user terminal (100) may be a smartphone. However, the present invention is not limited thereto, and the user terminal (100) may include electronic devices such as general desktop computers, navigation systems, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc. The electronic device may have one or more general or special purpose processors, memory, storage, and / or networking components (wired or wireless).

[0026] The document analysis device (200) receives document data from a user terminal (100), creates a content area in which paragraphs and image areas are matched in the received document data, and can identify detailed information about the document data based on the contents of the content area. The document analysis device (200) may be a server or implemented as an application within the user terminal (100). The document type classification device will be described in more detail with reference to FIGS. 2 and 3.

[0027] The communication method of the network (400) is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, a wired online network, a wireless online network, a broadcasting network) that the network (400) may include, but also short-range wireless communication between devices. For example, the network (400) may include one or more arbitrary networks (400) among networks (400) such as a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), and an online network.

[0028] FIG. 2 is a block diagram of a document analysis device (200) according to one embodiment of the present invention, and FIG. 3 is a flowchart of a document analysis method according to one embodiment of the present invention.

[0029] Referring to FIGS. 2 and 3, a document type classification device according to one embodiment of the present invention may include a document data receiving unit (201), an area recognition unit (202), a paragraph recognition unit (203), a content area generation unit (204), a first sentence generation unit (205), a second sentence generation unit (206), an image area additional matching unit (207), a content area analysis unit (208), a document data information transmission unit (209), a content area counting unit (210), a content area classification unit (211), a document data type determination unit (212), and a representative content area selection unit (213).

[0030] The document data receiving unit (201) can receive document data from the user terminal (100). (S11) The present invention can analyze the received document data to derive the type and detailed information of the corresponding document data.

[0031] The area recognition unit (202) can distinguish and recognize the text area and the image area in the received document data. (S12) The area recognition unit (202) can distinguish and recognize the text area and the image area in the received document data by referring to sample data for the image and sample data for the character.

[0032] The paragraph recognition unit (203) can divide the recognized text area into sentences and non-sentences based on terminal words, recognize an area including at least two consecutive sentences without a line break as a paragraph, and recognize the remaining area excluding the recognized paragraph as a non-paragraph. (S13) Since a sentence has a terminal word such as ~da, the paragraph recognition unit (203) can recognize an area including at least two consecutive sentences as a paragraph by taking this into consideration.

[0033] The content area generation unit (204) can analyze each recognized paragraph to extract the content described by each paragraph, and based on the extracted content, match paragraphs located within a set range that have a logical relationship to create a content area. (S14) As an example of the present invention, the content area generation unit (204) can extract the content described by each paragraph through a known natural language processing technology using a deep learning model algorithm. However, the method of analyzing paragraphs is not limited to this, and various known methods of summarizing paragraph content can be used to extract the content described by the paragraph.

[0034] The content area generation unit (204) extracts the content described by the recognized paragraph in the form of a sentence, and divides the words included in the subject, object, and complement of the sentences extracted from each paragraph within the set range into words of a higher concept and words of a lower concept that supplement the words of the higher concept, and matches sentences containing words of a higher concept with paragraphs having a logical relationship based on words of a lower concept to generate a content area.

[0035] The words of the lower concept that supplement the words of the upper concept include words that are examples of the words of the upper concept, words that further explain the words of the upper concept, and words that serve as causes and grounds for the words of the upper concept, and the content area generation unit (204) can determine the type of logical relationship between the recognized paragraphs according to the characteristics of the words of the lower concept for the words of the upper concept.

[0036] As an example of the present invention, types of logical relationships may include description and example relationships, cause and effect relationships, and assertion and evidence relationships. However, the types of logical relationships are not limited to these and can vary depending on the administrator's settings.

[0037] Paragraphs within the aforementioned setting range are explained below.

[0038] The content area generation unit (204) can extract the content described in a paragraph, and based on the extracted content, match paragraphs that have a logical relationship among the non-paragraph that is closest to the beginning of the paragraph and the paragraph itself, and the paragraphs that are closest to the end of the paragraph and the paragraph itself, to generate a content area. This is because non-paragraphs generally function as subheadings in a document, and paragraphs that explain the content of the subheading are generally included at the end of the part where the subheading is disclosed, and the content area generation unit (204) can generate a content area by taking this into consideration.

[0039] The first sentence generation unit (205) can determine the category of the recognized image area based on sample data for the diagram, and generate the content described by the recognized image area as a first sentence based on the determined category. (S15)

[0040] The first sentence generation unit (205) can extract change values ​​and corresponding result values ​​from the image area when the recognized image area is a diagram and generate a first sentence based on this. For example, when the recognized image area is a line graph, the first sentence generation unit (205) can determine how the result value corresponding to the Y-axis changes according to the change in the value corresponding to the X-axis through line data and generate a first sentence based on this. As an example of the present invention, the first sentence generation unit (205) can generate the first sentence, "As the corresponding value of the X-axis increases, the corresponding value of the Y-axis gradually decreases."

[0041] The first sentence generation unit (205) can recognize an object depicted in the image area if the recognized image area is not a diagram, and generate a first sentence based on the recognized object.

[0042] The first sentence generation unit (205) can recognize people, animals, objects, etc. included in the image area by analyzing the pixel values ​​of an image area other than a diagram using a machine learning algorithm, and can generate a first sentence to express what the image area represents through the recognized objects.

[0043] The first sentence generation unit (205) can utilize summary data, which will be described later, in the process of generating the first sentence for an image area other than a diagram.

[0044] The non-paragraph extraction unit can extract non-paragraphs if the text area closest to the front and back of the recognized image area is a non-paragraph. Non-paragraphs can typically be subheadings of paragraphs in a document or titles summarizing image areas.

[0045] The summary data determination unit may determine the content of the non-paragraph as summary data of the recognized image if the extracted non-paragraph is located in front of the recognized image, and may determine the summary data of the recognized image based on the gap between the non-paragraph and the recognized image and the gap with the area below the non-paragraph if the extracted non-paragraph is located in back of the recognized image. Specifically, the summary data determination unit may determine the non-paragraph as summary data of the recognized image if the extracted non-paragraph is located in back of the recognized image and the gap between the non-paragraph and the recognized image is smaller than the gap with the text area or image area below the non-paragraph. If the gap is larger than or equal to the gap with the text area or image area below the non-paragraph, the summary data of the recognized image may be determined to not exist. This reflects that a title summarizing the image area is generally displayed very closely to the front or back of the image area.

[0046] If the recognized image area is not a diagram, and the first sentence for the drawing is generated only with the object recognized in the image area, the first sentence for the drawing may be inaccurate. Therefore, the first sentence generation unit (205) can recognize the object depicted in the image area and generate the first sentence based on the recognized object and the summary data determined for the image area.

[0047] The second sentence generation unit (206) can analyze each paragraph located within a set range from the recognized image area and generate the content described by each paragraph as a second sentence (S16). The set range may be paragraphs located between the non-paragraph closest to the upper side from the recognized image area and the corresponding image area, and the second sentence generation unit (206) can analyze the paragraphs in the range and generate the content described by each paragraph as a second sentence. As an example of the present invention, the second sentence generation unit (206) can generate the content described by each paragraph as a second sentence through a known natural language processing technology using a deep learning model algorithm. However, the method of analyzing the paragraphs is not limited to this, and various known methods of summarizing the content of paragraphs can be utilized in the process of generating the second sentence.

[0048] The image area additional matching unit (207) can compare the second sentences with the first sentence, select the second sentence with the most similar meaning, and additionally match the recognized image area to the content area that includes the paragraph containing the second sentence. (S17)

[0049] The image area additional matching unit (207) extracts the subject, object, complement, and predicate of each of the first and second sentences, and determines whether words corresponding to the same item are similar to each other by synthesizing the results to determine whether the two sentences are similar.

[0050] The image area additional matching unit (207) can set the highest weights for the subject and predicate in determining whether there is similarity between sentences, and can quantify the similarity between the first and second sentences by setting the next highest weight for the object, and can select the second sentence with the highest similarity value and additionally match the corresponding image area to the content area that includes the paragraph that includes the selected second sentence, and thus the corresponding image area can be included in the corresponding content area.

[0051] The content area analysis unit (208) can derive the content of each content area by connecting the sentences extracted in the process of creating the content area based on the type of logical relationship between paragraphs determined by the content area generation unit (204). (S18) Specifically, the content area analysis unit (208) can derive the content of each content area by analyzing the sentences extracted in the process of creating the content area, that is, the paragraphs recognized by the content area generation unit (204), and connecting the sentences of the content explained by each paragraph (including the first sentence and the second sentence) using a conjunction corresponding to the determined logical relationship. For example, if the logical relationship of a specific content area is a cause and effect relationship, the content area analyses can derive the content of the corresponding content area by connecting the extracted sentences using a conjunction such as “therefore” or “consequently”.

[0052] The content area counting unit (210) can divide the received document data into at least three sections of different lengths according to the order of description, and count and aggregate the number of content areas included in each section. The sections divided by the content area counting unit (210) may have the same number of lines and thus be sections of the same length. However, this is not limited to this, and the sections may be divided in various ways depending on the administrator's settings.

[0053] The content area classification unit (211) can count and aggregate the types of logical relationships among the content areas included in each section and classify them by section. For example, if document data is divided into a first section, a second section, and a third section, the content area classification unit (211) can count and aggregate the types of logical relationships among the content areas included in each section.

[0054] The document data type determination unit (212) selects a section in which the largest number of content areas are aggregated, derives the logical relationship in which the largest number of content areas are aggregated in the selected section, and determines the type of received document data based on the derived logical relationship. For example, if the logical relationship in which the largest number of content areas are aggregated in the selected section is the description and example relationship, the document data is inferred to be mainly composed of paragraphs of the description and example relationship arranged sequentially, and therefore the document data type determination unit (212) can determine the type of the document data as an explanatory document such as an academic paper.

[0055] The document data information transmission unit (209) can transmit information including the type of determined document data and the content collected by the content area counting unit (210) and the content area classification unit (211) to the user terminal (100). The user can quickly check the type of document data and the overall content of the document data by checking the information on the document data through the user terminal (100).

[0056] The document data information transmission unit (209) can infer related content areas between content areas based on the logical relationship between the content areas and transmit group information that groups these areas together with the derived content to the user terminal (100). The process of inferring related content areas is as follows.

[0057] The representative content area selection unit (213) determines the logical relationship with the largest number of aggregated values ​​in the content areas of each section of document data, and selects the content area with the largest number of paragraphs among the content areas corresponding to the determined logical relationship as the representative content area of ​​each section. The logical relationship with the largest number of aggregated values ​​in the content areas of each section is determined, and the content area with the largest number of paragraphs while having the corresponding logical relationship in each section is likely to be the content area with the largest number of paragraphs in the section. Therefore, the representative content area selection unit (213) selects the representative content area by taking this into consideration.

[0058] The document data information transmission unit (209) can generate group information by grouping related content areas based on the logical relationship between content areas and transmit the group information together with the derived content to the user terminal (100). (S19)

[0059] Specifically, the document data transmission unit can generate group information by grouping representative content areas by section by inferring related content areas, and transmit, together with the group information, content explaining the logical relationship between each of the selected representative content areas and information matching the content derived from the selected representative content areas to the user terminal (100).

[0060] By checking the content explaining the logical relationship between group information and each of the selected representative content areas and the information matching the content derived from the selected representative content areas through the user terminal (100), the user can more accurately and efficiently understand the content of the content areas with high proportions within the document data.

[0061] In this way, the present invention analyzes document data to create content areas in which paragraphs and image areas are matched, classifies these in various ways according to the sections of the document data, selects a representative content area for each section based on the classification result, groups these, and provides detailed information to the user, thereby enabling the user to efficiently grasp the relationship between content areas considering the structure of the document data, ultimately enabling a more accurate and integrated understanding of information about the document data.

[0062] The embodiments described above are provided for illustrative purposes only, and those skilled in the art will readily appreciate that the embodiments described above can be readily modified into other specific forms without altering the technical concepts or essential characteristics of the embodiments described above. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0063] The scope of protection sought through this specification is indicated by the claims described below rather than by the detailed description, and should be interpreted to include all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts.

Claims

1. In a document analysis device that analyzes a document through inference of association between content areas that constitute the document, A document data receiving unit that receives document data from a user terminal used by a user; An area recognition unit that recognizes and distinguishes between text areas and image areas in received document data; A paragraph recognition unit that distinguishes between sentences and non-sentences based on terminal words in a recognized text area, recognizes an area containing at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; A content area generation unit that analyzes each recognized paragraph to extract the content described by each paragraph, and creates a content area by matching paragraphs that have a logical relationship among paragraphs located within a set range based on the extracted content; A first sentence generation unit that determines a category of a recognized image region based on sample data for a diagram, and generates a first sentence describing the content of the recognized image region based on the determined category; A second sentence generation unit that analyzes each paragraph located within a set range from the recognized image area and generates a second sentence describing the content of each paragraph; An image area additional matching unit that compares the second sentences with the first sentence, selects the second sentence with the most similar meaning, and additionally matches the recognized image area with the content area that includes the paragraph containing the second sentence; A content area analysis unit that derives the contents of each content area by connecting the sentences extracted in the process of generating the content area based on the type of logical relationship between paragraphs determined by the content area generation unit; and A document analysis device through inference of association between content areas, including a document data information transmission unit that generates group information by grouping related content areas by inferring related content areas based on the logical relationship between content areas, and transmits the group information together with the derived content to a user terminal.

2. In paragraph 1, The above content area generation unit extracts the content described by the recognized paragraph in the form of sentences, divides the words included in the subject, object, and complement of the sentences extracted from each paragraph within the set range into words of a higher concept and words of a lower concept that supplement the words of the higher concept, and matches sentences containing words of a higher concept with paragraphs that have a logical relationship based on words of a lower concept to generate a content area. Types of logical relationships include explanation and example relationships, cause and effect relationships, and assertion and evidence relationships. The above first sentence generation unit, if the recognized image area is a diagram, extracts the change value and the corresponding result value from the image area and generates the first sentence based on this, and if the recognized image area is not a diagram, recognizes the object depicted in the image area and generates the first sentence based on the recognized object. The second sentence generation unit generates the content described in each paragraph as a second sentence using natural language processing technology using a deep learning model algorithm. A document analysis device through inference of association between content areas, characterized in that the words of the lower concept that supplement the words of the upper concept include words that are examples of the words of the upper concept, words that further explain the words of the upper concept, and words that are causes and grounds for the words of the upper concept, and the content area generation unit determines the type of logical relationship between paragraphs recognized according to the characteristics of the words of the lower concept.

3. In paragraph 2, The above content area generation unit extracts the content described in the paragraph, and based on the extracted content, matches paragraphs that have a logical relationship among the non-paragraph closest to the beginning of the paragraph and the paragraph itself, and among the paragraphs that are closest to the end of the paragraph and the paragraph itself, to generate a content area. The above content area analysis unit is characterized in that the content area generation unit analyzes the recognized paragraphs and connects the sentences of the content described in each paragraph through a conjunction corresponding to the determined logical relationship to derive the content of each content area, thereby providing a document analysis device through inference of the relationship between content areas.

4. In paragraph 3, A content area counting unit that divides the received document data into at least three sections according to the description order and counts and aggregates the number of content areas included in each section; A content area classification unit that counts and aggregates the types of logical relationships of each content area included in each section and classifies them by section; The document data type determination unit further includes a document data type determination unit that selects a section in which the largest number of content areas are aggregated, derives a logical relationship in which the largest number of content areas are aggregated in the selected section, and determines the type of received document data based on the derived logical relationship. A document analysis device through inference of correlation between content areas, characterized in that the document data information transmission unit transmits information including the type of determined document data and the content aggregated by the content area counting unit and the content area classification unit to the user terminal.

5. In paragraph 4, It further includes a representative content area selection unit that determines the logical relationship with the largest number of aggregated content areas in each section of document data, and selects the content area with the largest number of included paragraphs among the content areas corresponding to the determined logical relationship as the representative content area for each section. The above document data transmission unit generates group information by grouping representative content areas by section by inferring related content areas, and transmits to the user terminal, together with the group information, content explaining the logical relationship of each representative content area selected and information matching the content derived from the selected representative content area. A document analysis device through inference of relatedness between content areas.

6. In a document analysis method through inference of association between content areas using a document analysis device that analyzes a document through inference of association between content areas that constitute the document, A step in which a document data receiving unit receives document data from a user terminal used by a user; A step in which a region recognition unit recognizes and distinguishes between a text region and an image region in received document data; A step in which a paragraph recognition unit distinguishes between sentences and non-sentences based on terminal words in a recognized text area, recognizes an area including at least two consecutive sentences without a line break as a paragraph, and recognizes the remaining area excluding the recognized paragraph as a non-paragraph; A step of generating a content area by analyzing each recognized paragraph and extracting the content described by each paragraph, and then matching paragraphs located within a set range that have a logical relationship based on the extracted content to generate a content area; A step of a first sentence generation unit determining a category of a recognized image area based on sample data for a diagram, and generating a first sentence describing the content of the recognized image area based on the determined category; A step in which a second sentence generation unit analyzes each paragraph located within a set range from the recognized image area and generates a second sentence with the content described by each paragraph; A step of additionally matching the image area by comparing the second sentences with the first sentence and selecting the second sentence with the most similar meaning, and additionally matching the recognized image area with the content area that includes the paragraph containing the second sentence; A step in which the content area analysis unit derives the contents of each content area based on the logical relationship between the text areas or image areas included in the content area based on the contents extracted in sentence format during the process of creating the content area; and A method for analyzing a document through inference of association between content areas, including a step in which a document data information transmission unit infers related content areas based on a logical relationship between content areas and groups them and transmits group information, together with the derived content, to a user terminal.

Citation Information

Patent Citations

  • Multimodal information fusion document content enhancement retrieval system and method

    CN117312601A

  • Document classification device, program and document classification method

    JP2005122550A

  • Method for extracting document structure and document search method

    JP2007286861A

  • Apparatus and method for automatically creating document

    KR101549792B1

  • Apparatus for document structure information extraction and document merging using artificial intelligence

    KR102538108B1