AI-powered document structure information extraction and document merging device
The AI-based document merging device leverages character and paragraph attributes to merge documents efficiently, addressing the challenge of integrating diverse documents by determining document types and optimizing content order for coherent output.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2026-03-19
AI Technical Summary
Existing document merging technologies struggle to effectively integrate documents due to differences in document components and lack of understanding the relationship between data structure and text data, necessitating a more sophisticated approach to analyze and merge documents.
An AI-powered document structure information extraction and merging device that utilizes character and paragraph attributes, including size, slant, and hue, to determine document types and merge documents based on similarity and temporal sequence, while optimizing paragraph order and content overlap.
Enables accurate analysis of document attributes and types, allowing for natural document merging with minimal user intervention, ensuring coherent and structured output.
Smart Images

Figure 2026509390000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a document structure information extraction and document merging device that utilizes artificial intelligence. More specifically, it relates to a device that utilizes artificial intelligence to discriminate the attributes of characters, then grasps the document structure information, and based on this, integrates multiple documents into one document to generate a new document
Background Art
[0002] As artificial intelligence has made leaps and bounds, it is trending towards being actively used in various industrial fields.
[0003] Due to the development of such artificial intelligence, various automations have also been carried out in the field of office automation. In particular, a lot of effort has been devoted to recognizing printed content and converting it into data. As a typical example, there is prior research that combines a natural language processing model such as BERT with optical character recognition (OCR) technology to correct recognition results.
[0004] The methods developed so far are in the form of recognizing characters and classifying them according to rules and dictionaries conceptually defined by humans. Therefore, it is difficult to apply a differential analysis method due to differences in the components of the document, and there is a need for technological development that can embody the relationship between the data structure and text data contained in the target document.
[0005] The technology that is the background of this application is disclosed in Registered Patent No. 10-2388781.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The present invention relates to an artificial intelligence-based document structure information extraction and document merging device, and more specifically, to an artificial intelligence-based document structure information extraction and document merging device that uses artificial intelligence to determine character attributes, grasp document structure information from these attributes, and use this as a basis to integrate multiple documents into a single document to generate a new document. [Means for solving the problem]
[0007] An artificial intelligence-powered document structure information extraction and document merging device according to one embodiment of the present invention extracts document structure information through character style information constituting a document, and merges multiple documents from which structure information has been extracted to generate a new document. This device includes a document image receiving unit that receives a document image from a user terminal used by the user, a character recognition unit that applies optical character recognition (OCR) to the received document image to recognize characters contained in the document image, a paragraph recognition unit that recognizes paragraphs based on the vertical spacing of characters recognized in the received document image, and a paragraph attribute determination unit that analyzes the position of recognized paragraphs and the style of characters constituting the paragraph through an artificial intelligence model that utilizes multiple document data as training data to determine the attributes of the paragraph. The character style includes information on size, slant, hue, and typeface, and the paragraph attributes include information on the function of the paragraph.
[0008] An artificial intelligence-based document structure information extraction and document merging device according to one embodiment of the present invention further includes a document type determination unit that determines the type of document composed of a given paragraph based on the attributes of the determined paragraph, and the paragraph attribute determination unit determines the attribute of the recognized paragraph to one of the following: title, subtitle, introduction, main body, or conclusion.
[0009] An artificial intelligence-based document structure information extraction and document merging device according to one embodiment of the present invention further includes: a merge request signal receiving unit that receives a merge request signal from a user terminal, which is a signal requesting the merging of a first document image and a second document image selected from among the document images received from the document image receiving unit; a document type comparison unit that, when the merge request signal is generated, compares the document type determined from the first document image with the document type determined from the second document image, and generates an identical signal if they are the same, and a non-identical signal if they are different; a keyword similarity calculation unit that, when an identical signal is generated, extracts keywords contained in the title paragraph of the first document image and keywords contained in the title paragraph of the second document image, determines whether there is any overlap between the extracted keywords and calculates keyword similarity; and a non-merging signal transmission unit that, when the calculated keyword similarity is less than a set value or a non-identical signal is generated, generates a merge-fail signal and transmits it to the user terminal, wherein the keyword similarity calculation unit extracts keywords contained in the title paragraphs of the first document image and the second document image through an artificial intelligence model that learns from a plurality of paragraph data from which keywords have been extracted.
[0010] An artificial intelligence-based document structure information extraction and document merging device according to one embodiment of the present invention further includes a mergeable signal transmission unit that, when the calculated keyword similarity is equal to or greater than a set value, extracts words contained in the main text paragraphs of the first document image and the second document image, classifies the extracted words in descending order of frequency, and generates a mergeable signal and transmits it to the user terminal if the words of a set rank or higher overlap in both document images at a set ratio or higher, and the non-merging signal transmission unit generates a non-merging signal and transmits it to the user terminal if the words of a set rank or higher overlap in both document images at a set ratio or lower.
[0011] An artificial intelligence-based document structure information extraction and document merging device according to one embodiment of the present invention further includes, when a merging signal is generated, a sequence relationship determination unit that extracts sentences or words representing a point in time from each paragraph of the first document image and the second document image and determines the temporal sequence of the first and second document images from these; a title paragraph determination unit that determines the title paragraph of the document image that is later in time among the first and second document images as the title paragraph of the merged new document based on the sequence relationship determined by the sequence relationship determination unit; and a subtitle paragraph determination unit that determines the title paragraph of the document image that is earlier in time among the first and second document images as the subtitle paragraph of the merged new document based on the sequence relationship determined by the sequence relationship determination unit, wherein the sequence relationship determination unit applies an artificial intelligence model that learns document data containing sentences, conjunctions, and words that can grasp a point in time to each paragraph of the first and second document images to determine the temporal sequence of the first and second document images.
[0012] An artificial intelligence-based document structure information extraction and document merging apparatus according to one embodiment of the present invention includes: an introduction paragraph determination unit that determines the introduction paragraph of a new document by linking the introduction paragraph of the document image that is earlier in time and the introduction paragraph of the document image that is later in time among the first and second document images based on the order relationship determined by the order relationship determination unit; and a main body paragraph of the document image that is earlier in time and the main body paragraph of the document image that is later in time among the first and second document images based on the order relationship determined by the order relationship determination unit. The system further includes a main paragraph determination unit that determines the main paragraphs of the new document formed by linking and merging the two document images, and a conclusion paragraph determination unit that determines the conclusion paragraphs of the new document formed by merging the document image that is later in time from the first document image and the second document image, based on the sequence determined by the sequence determination unit, wherein the main paragraph determination unit determines the linking phrases between the main paragraphs of the first document image and the main paragraphs of the second document image through an artificial intelligence model that learns multiple document data linked by causal relationships, and uses the determined linking phrases to link the two paragraphs.
[0013] An artificial intelligence-based document structure information extraction and document merging device according to one embodiment of the present invention further includes a document deletion unit that identifies sentences contained in the introduction paragraph and the main paragraph determined by the introduction paragraph determination unit and the main paragraph determination unit as subjects, objects, and predicates, and if there are multiple sentences with similar subjects, objects, and predicates, deletes the sentences excluding the sentence with the longest character length among the multiple sentences; a title paragraph determination unit, a subtitle paragraph determination unit, an introduction paragraph determination unit, a main paragraph determination unit, a conclusion paragraph determination unit, and a document deletion unit that merges each paragraph determined by the title paragraph determination unit, the subtitle paragraph determination unit, the introduction paragraph determination unit, the main paragraph determination unit, the conclusion paragraph determination unit, and the document deletion unit to generate new document data in which a first document image and a second document image are merged; and a new document data transmission unit that distinguishes and displays the sentences generated from the first document image and the characters generated from the second document image in the generated new document data and transmits it to the user terminal. [Effects of the Invention]
[0014] This invention can analyze a given document image through an artificial intelligence model to recognize paragraphs, and to understand the individual attributes of the recognized paragraphs and the type of document to which they belong.
[0015] This invention, when there is a request to merge document images, analyzes both document images to determine whether they can be merged, and if it is determined that they can be merged, it determines the order in which the two documents are placed, and generates a new document in which the two documents are merged, taking into account the determined order in which the documents are placed. The user can then obtain a document in which the two documents are naturally merged with only minor modifications to the generated document. [Brief explanation of the drawing]
[0016] [Figure 1] This is a block diagram of a document structure information extraction and document merging system according to one embodiment of the present invention. [Figure 2] This is a block diagram of a document structure information extraction and document merging device according to one embodiment of the present invention. [Figure 3] This is a block diagram of a paragraph determination unit according to one embodiment of the present invention. [Modes for carrying out the invention]
[0017] A document structure information extraction and document merging device utilizing artificial intelligence, which extracts document structure information through character style information constituting a document according to one embodiment of the present invention, and merges multiple documents from which structure information has been extracted to generate a new document, includes a document image receiving unit that receives a document image from a user terminal used by the user, a character recognition unit that applies optical character recognition (OCR) to the received document image to recognize the characters contained in the document image, a paragraph recognition unit that recognizes paragraphs based on the vertical spacing of the characters recognized in the received document image, and a paragraph attribute determination unit that uses multiple document data as training data and analyzes the position of the recognized paragraph and the style of the characters constituting the paragraph through a machine learning algorithm based thereon to determine the attributes of the paragraph, wherein the character style includes information on size, slant, and hue, and the paragraph attributes include information on the function of the paragraph.
[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings so that those with ordinary skill in the art to which the present invention pertains can easily implement it. However, the present invention can be embodied in a variety of different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly illustrate the present invention with the drawings, parts that are not relevant to the description have been omitted, and similar parts throughout the specification have been denoted by similar reference numerals.
[0019] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where it is "directly connected" but also cases where it is "electrically connected" with other elements in between. Furthermore, when a part is described as "containing" some component, this means, unless otherwise stated, that it may contain other components rather than excluding them. The present invention will now be described in detail with reference to the attached drawings.
[0020] Figure 1 is a block diagram of a document structure information extraction and document merging system 1000 according to one embodiment of the present invention.
[0021] Referring to FIG. 1, the document structure information extraction and document merging system 1000 according to an embodiment of the present invention can include a user terminal 100 and a document structure information extraction and document merging apparatus 200 associated with the user terminal 100 via a network 400.
[0022] The user terminal 100 can be a terminal used by a person who intends to extract the structure from a document or merge documents. For example, the user can be a journalist, and in this case, the user can be a person who intends to extract the document structure from another person's article or merge two articles to generate a new single article.
[0023] The user terminal 100 can be a smartphone. However, it is not limited thereto, and the user terminal 100 can include electronic devices such as a general desktop computer, a navigation device, a notebook computer, a digital broadcast terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), and a tablet PC. The electronic device can have one or more general or special-purpose processors, memories, storages, and / or networking components (wired or wireless).
[0024] The document structure information extraction and document merging apparatus 200 receives document image data from the user terminal 100, analyzes it using artificial intelligence to grasp the attributes of paragraphs and the types of documents, and can merge a plurality of documents under specific conditions. The document structure information extraction and document merging apparatus 200 can be a server and can be embodied in the form of an application within the user terminal 100. The detailed content regarding the document structure information extraction and document merging apparatus 200 will be described in more detail in FIGS. 2 and 3.
[0025] Without limitation, as an example of a communication network that the network 400 can include, not only communication methods that utilize a mobile communication network, a wired online network, a wireless online network, and a broadcast network, but also short-range wireless communication between devices can be included. For example, the network 400 can include any one or more of networks such as a PAN (personal area network), a LAN (local area network), a CAN (campus area network), a MAN (metropolitan area network), a WAN (wide area network), a BBN (broadband network), and an online network.
[0026] FIG. 2 is a block diagram of a document structure information extraction and document merging apparatus 200 according to an embodiment of the present invention, and FIG. 3 is a block diagram of a paragraph determination unit 212 according to an embodiment of the present invention.
[0027] Referring to FIGS. 2 and 3, the document structure information extraction and document merging apparatus 200 can include a document image receiving unit 201, a character recognition unit 202, a paragraph recognition unit 203, a paragraph attribute determination unit 204, a document type determination unit 205, a merging request signal receiving unit 206, a document type comparison unit 207, a keyword similarity calculation unit 208, a non-mergable signal transmission unit 209, a mergable signal transmission unit 210, a precedence relationship determination unit 211, a paragraph determination unit 212, a sentence deletion unit 213, a new document data transmission unit 214, and a correction data receiving unit 215.
[0028] The document image receiving unit 201 can receive a document image from the user terminal 100. The document image can be an image file rather than a text document.
[0029] The character recognition unit 202 can apply optical character recognition (OCR) to the received document image to recognize the characters included in the corresponding document image. However, without being limited thereto, the character recognition unit 202 can analyze the corresponding document through not only OCR but also an artificial intelligence deep learning model to recognize the characters included in the document image.
[0030] The paragraph recognition unit 203 can recognize paragraphs based on the vertical spacing between characters recognized in the received document image.
[0031] The spacing between paragraphs is generally wider than the spacing between characters within a single paragraph. The paragraph recognition unit 203 can distinguish paragraphs in the document image where characters are recognized based on this point. However, the paragraph recognition unit 203 does not recognize paragraphs based solely on the vertical spacing between characters, but can distinguish paragraphs in the document image by comprehensively considering the vertical spacing and the fact that the first character of a paragraph begins with an indentation.
[0032] The paragraph attribute determination unit 204 can determine the attributes of a paragraph by analyzing the position of the paragraph recognized through an artificial intelligence model that utilizes multiple document data as training data, and the style of the characters that make up that paragraph.
[0033] The paragraph attribute determination unit 204 learns from multiple document data provided by the administrator through an artificial intelligence model, and can determine the attributes of the relevant paragraph by analyzing the learning results, the position of the recognized paragraph, and the style of the characters that make up the paragraph.
[0034] Character style can include information about character size, slant, hue, and typeface, while paragraph attributes can include information about the paragraph's function. For example, the paragraph attribute determination unit 204 can learn, through an artificial intelligence model, how the characters that make up a title are sized, slant, hue, and typeface, and then search for a recognized paragraph that matches these features to recognize that it is the title paragraph. The ability to recognize paragraph attributes can become more accurate as the amount of learning continuously accumulates.
[0035] The paragraph attribute determination unit 204 can determine the type of document that the paragraph in question is composed of, based on the determined attributes of the paragraph.
[0036] The document type determination unit 205 can determine the type of document composed of a given paragraph based on the attributes of the determined paragraph. The document type determination unit 205 can learn from multiple document data provided by the administrator through an artificial intelligence model and determine the type of document composed of the paragraph whose attributes have been determined based on the learning results. For example, the document type determination unit 205 can learn the arrangement structure of paragraphs that make up a newspaper article through newspaper article document data and can determine whether a given document is a newspaper article based on the learning results. Possible document types include instruction manuals, essays, newspaper articles, and reports. Each of these documents may have a different paragraph arrangement structure.
[0037] Thus, the present invention can analyze a given document image through an artificial intelligence model to recognize paragraphs, and to understand the individual attributes of the recognized paragraphs and the type of document to which they belong.
[0038] The merge request signal receiving unit 206 can receive a merge request signal from the user terminal 100, which is a signal requesting the merging of a first document image and a second document image selected from the document images received from the document image receiving unit 201. The user can input two document images into the user terminal 100 as needed and input a merge request signal to merge them. The input merge request signal can be transmitted to the merge request signal receiving unit 206 through the user terminal 100.
[0039] When a merge request signal is generated, the document type comparison unit 207 compares the document type determined from the first document image with the document type determined from the second document image. If they are the same, it generates an identical signal; otherwise, it generates a non-identical signal. The document type comparison unit 207 can compare the document type of the first document image determined by the document type determination unit 205 described above with the document type determined from the second document image. If the document types match, it generates an identical signal; if the document types do not match, it generates a non-identical signal.
[0040] When identical signals are generated, the keyword similarity calculation unit 208 can extract keywords contained in the title paragraph of the first document image and keywords contained in the title paragraph of the second document image, determine whether there is any overlap between the extracted keywords, and calculate keyword similarity.
[0041] The keyword similarity calculation unit 208 can extract words that are judged to be keywords from the title paragraph of each document image through an artificial intelligence model that extracts keywords from document data, and can calculate keyword similarity based on the overlap ratio by comparing the keywords extracted from both document images.
[0042] The non-merging signal transmission unit 209 can generate a non-merging signal and transmit it to the user terminal 100 if the calculated keyword similarity is less than a set value or if a non-identical signal is generated.
[0043] If there are no overlapping words between the keywords in the title paragraphs of both document images, the keyword similarity can be said to be less than the set value, and in this case, the non-merging signal transmission unit 209 can determine that there are no common points in content between the two documents and generate a non-merging signal. Also, if the document types of the two document images do not match, the paragraph arrangement structure of the two documents will be completely different, and there is a high probability that merging these documents will result in a document with very unnatural content. In this case as well, the non-merging signal transmission unit 209 can generate a non-merging signal. For example, if the two documents are a newspaper article and an essay, the characteristics and paragraph arrangement structure of the two documents are completely different, so the non-merging signal transmission unit 209 can generate a non-merging signal and transmit it to the user terminal 100.
[0044] If the calculated keyword similarity is equal to or greater than a set value, the mergeable signal transmission unit 210 extracts words contained in the main text paragraphs of the first and second document images, classifies the extracted words in descending order of frequency, and if words of a set rank or higher overlap in both document images at a rate greater than or equal to a set ratio, it can generate a mergeable signal and transmit it to the user terminal 100. In other words, if the keyword similarity between the two document images is equal to or greater than a set value, the mergeable signal transmission unit 210 can then compare the main texts of both documents to determine whether the two documents can ultimately be merged. Specifically, the mergeable signal transmission unit 210 extracts words contained in the main text paragraphs of each document image, classifies the extracted words in descending order of frequency, and if words of a set rank or higher overlap in both documents at a rate greater than or equal to a set ratio, it can consider both documents to be of the same type and to contain similar content in their respective main texts, and generate a mergeable signal.
[0045] Conversely, if the two document images have more words than the set rank but less than the set ratio, the two document images are of the same type and their titles are somewhat similar, but the content of the body text is completely different. In this case, the merging of the two documents is deemed impossible, and the merging impossible signal transmission unit 209 can generate a merging impossible signal and transmit it to the user terminal 100.
[0046] The following will explain in detail the process by which a merger signal is generated and the two document images are merged.
[0047] When a mergeable signal is generated, the sequence determination unit 211 extracts sentences or words representing a specific time period from each paragraph of the first document image and the second document image, and from these, it can determine the temporal sequence relationship between the first document image and the second document image.
[0048] The sequence determination unit 211 can determine the temporal sequence of the first and second document images by applying an artificial intelligence model, which learns document data containing sentences, conjunctions, and words that allow for the identification of a specific time period, to each paragraph of the first and second document images.
[0049] For example, if the body of the first document image contains clear information about a specific time point, and the second document image discloses information about the consequences resulting from the content of the first document image, the priority relationship determination unit 211 can determine the priority relationship by recognizing that the time point of the first document image is earlier than the time point of the second document image.
[0050] Once a sequence relationship is determined between the two document images, the present invention can generate a new document by determining the title paragraph, subtitle paragraph, introductory paragraph, main body paragraph, and conclusion paragraph in the following manner.
[0051] The paragraph determination unit 212 may include a title paragraph determination unit 231, a title paragraph determination unit 232, an introduction paragraph determination unit 233, a main body paragraph determination unit 234, and a conclusion paragraph determination unit 235.
[0052] The title paragraph determination unit 231 can determine, based on the sequence of events determined by the sequence determination unit 211, that the title paragraph of the document image that occurs later in time from the first document image and the second document image will be the title paragraph of the merged new document. This can be seen as reflecting the fact that the content occurring later in time is the final result and is more important than the content occurring earlier in time.
[0053] Based on the sequence of events determined by the sequence of events determined by the sequence of events determination unit 211, the title paragraph of the document image that is earlier in time from the first document image and the second document image can be determined as the subtitle paragraph of the merged new document. Since the subtitle of a document must contain content that supports the title, considering that the title paragraph of the document image that is later in time was previously included in the title paragraph, the subtitle paragraph can be seen as reflecting the title paragraph of the document image that is earlier in time that can support the title paragraph.
[0054] The introductory paragraph determination unit 233 can determine the introductory paragraph of the merged new document by linking the introductory paragraph of the document image that is earlier in time and the introductory paragraph of the document image that is later in time, based on the order of events determined by the order of events determination unit 211. In the case of an introductory paragraph, it is a paragraph that briefly introduces the main content before entering the main body of the text, so it can be seen that the content of the introductory paragraphs of both documents is reflected in the introduction of the merged document.
[0055] The paragraph determination unit 234 can determine the paragraphs of the new document by linking the paragraphs of the document image that is earlier in time and the paragraphs of the document image that is later in time, based on the sequence of events determined by first document image and the second document image.
[0056] The paragraph determination unit 234, through an artificial intelligence model that learns from multiple text data linked by causal relationships, determines a connecting phrase that naturally links the main paragraph of the first document image and the main paragraph of the second document image. Using the determined connecting phrase, it can then reflect the results in the main paragraph of the new document created by linking and merging the two paragraphs.
[0057] The conclusion paragraph determination unit 235 can determine the conclusion paragraph of the new document formed by merging the conclusion paragraphs of the first document image and the document image that is later in time, based on the sequence of events determined by the sequence of events determination unit 211. In the case of a conclusion paragraph, the content must reflect the final conclusion, so this can be seen as reflecting the conclusion paragraph of the document that is later in time in the new document formed by merging the two documents.
[0058] The text deletion unit 213 identifies the subjects, objects, and predicates of the sentences contained in the introductory paragraph and the main paragraph determined by the introductory paragraph determination unit 233 and the main paragraph determination unit 234. If there are multiple sentences with similar subjects, objects, and predicates, the text deletion unit 213 can delete all sentences except for the longest sentence among the multiple sentences. This is because duplicate sentences may occur during the document merging process, so the text deletion unit 213 can delete the remaining sentences, except for the longest sentence among such duplicate sentences that is considered to express the content in a more concrete way.
[0059] The new document data transmission unit 214 merges the paragraphs determined by the title paragraph determination unit 231, the title paragraph determination unit 232, the introduction paragraph determination unit 233, the main body paragraph determination unit 234, the conclusion paragraph determination unit 235, and the text deletion unit 213 to generate new document data in which the first document image and the second document image are merged. The generated new document data can then be used to distinguish and display the text generated from the first document image and the characters generated from the second document image, and transmit this data to the user terminal 100.
[0060] The correction data receiving unit 215 can receive correction data from the user terminal 100, which is the text that a user has corrected after viewing the new document data.
[0061] Users can view the new document data, which displays the text generated from the first document image and the text generated from the second document image. This allows them to specifically see how the new document data has been merged, and if there are any parts that need correction, they can directly correct those parts through the user terminal 100 to obtain the final completed new document data.
[0062] Thus, when there is a request to merge document images, the present invention analyzes both document images to determine whether the two documents can be merged. If it is determined that merging is possible, it determines the order in which the two documents are placed, and considering the determined order, generates a new document in which the two documents are merged. The user can then obtain a document in which the two documents are naturally merged with only minor modifications to the generated document.
[0063] The embodiments described above are for illustrative purposes only, and those with ordinary skill in the art to which the embodiments belong will understand that they can be easily modified into other specific forms without altering the technical ideas or essential features of the embodiments. Therefore, the embodiments described above should be understood in all respects as illustrative and not limiting. For example, each component described as a single type may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined form.
[0064] The scope of protection sought through this specification is indicated by the claims set forth below rather than by a detailed description, and should be interpreted as including all modified or altered forms derived from the meaning and scope of the claims and the concept of equivalents thereof.
Claims
1. In an artificial intelligence-powered document structure information extraction and document merging device that extracts document structure information through the style information of characters that make up a document, and merges multiple documents from which structure information has been extracted to generate a new document, A document image receiving unit that receives document images from the user's terminal; A character recognition unit that applies optical character recognition (OCR) to a received document image to recognize the characters contained in that document image; Paragraph recognition unit that recognizes paragraphs based on the vertical spacing of characters recognized in the received document image; and It includes a paragraph attribute determination unit that analyzes the position of a paragraph recognized through an artificial intelligence model that utilizes multiple document data as training data, and the style of the characters that make up that paragraph, in order to determine the attributes of that paragraph. An artificial intelligence-based document structure information extraction and document merging device, characterized in that the character style includes information on size, slant, hue, and typeface, and the paragraph attributes include information on the function of the paragraph.
2. It further includes a document type determination unit that determines the type of document composed of the paragraph based on the attributes of the determined paragraph, The paragraph attribute determination unit determines the attribute of the recognized paragraph to one of the following: title, subtitle, introduction, body, or conclusion, as described in claim 1, a document structure information extraction and document merging device utilizing artificial intelligence.
3. A merge request signal receiving unit receives a merge request signal from a user terminal, which is a signal requesting the merger of a first document image and a second document image selected from the document images received from the document image receiving unit; When the merge request signal is generated, a document type comparison unit compares the document type determined from the first document image with the document type determined from the second document image, and generates an identical signal if they are the same, and a non-identical signal if they are different; A keyword similarity calculation unit that, when identical signals are generated, extracts keywords contained in the title paragraph of the first document image and keywords contained in the title paragraph of the second document image, determines whether there is any overlap between the extracted keywords, and calculates keyword similarity; and The system further includes a non-merging signal transmission unit that generates a non-merging signal and transmits it to the user terminal if the calculated keyword similarity is less than a set value or if non-identical signals are generated. The document structure information extraction and document merging apparatus utilizing artificial intelligence according to claim 2, characterized in that the keyword similarity calculation unit extracts keywords contained in the title paragraphs of the first document image and the second document image through an artificial intelligence model that learns from a plurality of paragraph data from which keywords have been extracted.
4. The system further includes a mergeable signal transmission unit that, if the calculated keyword similarity is equal to or greater than a set value, extracts words contained in the main text paragraphs of the first and second document images, classifies the extracted words in descending order of frequency, and generates a mergeable signal and transmits it to the user terminal if there are words of a set rank or higher that overlap in both document images at a ratio greater than or equal to a set value. The document structure information extraction and document merging device utilizing artificial intelligence according to claim 3, characterized in that the non-merging signal transmission unit generates a non-merging signal and transmits it to the user terminal when the number of words in both document images that are of a set rank or higher overlaps in a ratio less than that of a set rank.
5. When the mergeable signal is generated, a sequence determination unit extracts sentences or words representing time points from paragraphs in the first and second document images, and determines the temporal sequence between the first and second document images from these extractions; A title paragraph determination unit determines, based on the sequence determined by the sequence determination unit, the title paragraph of the document image that is later in time among the first and second document images as the title paragraph of the merged new document; and The system further includes a subtitle paragraph determination unit that, based on the prior-order relationship determined by the aforementioned priority relationship determination unit, determines the title paragraph of the document image that is earlier in time among the first and second document images as the subtitle paragraph of the merged new document. The document structure information extraction and document merging device utilizing artificial intelligence according to claim 4, characterized in that the sequence relationship determination unit applies an artificial intelligence model, which learns document data containing sentences, conjunctions, and words that allow for the determination of a time point, to each paragraph of the first document image and the second document image to determine the temporal sequence relationship between the first document image and the second document image.
6. An introductory paragraph determination unit determines the introductory paragraph of a new document by linking the introductory paragraph of the document image that is earlier in time and the introductory paragraph of the document image that is later in time among the first and second document images, based on the order relationship determined by the aforementioned order relationship determination unit; A paragraph determination unit determines the paragraphs of the new document formed by linking the paragraphs of the first document image that is earlier in time and the paragraphs of the second document image that is later in time, based on the prior-sequence relationship determined by the aforementioned priority relationship determination unit; and The system further includes a conclusion paragraph determination unit that determines, based on the sequence of events determined by the sequence determination unit, the conclusion paragraph of the new document formed by merging the conclusion paragraph of the document image that is later in time from the first document image and the second document image. The paragraph determination unit determines the linking phrase between the paragraphs of the first document image and the paragraphs of the second document image through an artificial intelligence model that learns multiple text data linked by causal relationships, and uses the determined linking phrase to merge the two paragraphs, as described in claim 5, for the extraction of document structure information and document merging device utilizing artificial intelligence.
7. A sentence deletion unit that identifies the sentences contained in the introduction paragraph and the main paragraph determined by the introduction paragraph determination unit and the main paragraph determination unit as subjects, objects, and predicates, and deletes the sentences that have the longest character length among the multiple sentences in which the specified subjects, objects, and predicates exist, and The artificial intelligence-based document structure information extraction and document merging apparatus according to claim 6, further comprising a new document data transmission unit that merges the paragraphs determined by the title paragraph determination unit, the subtitle paragraph determination unit, the introduction paragraph determination unit, the main body paragraph determination unit, the conclusion paragraph determination unit, and the text deletion unit to generate new document data in which the first document image and the second document image are merged, and transmits the generated new document data to the user terminal by distinguishing and displaying the text generated from the first document image and the characters generated from the second document image.