Document processor based on artificial intelligence (AI)

Through the document processor based on artificial intelligence AI, traditional document tools are solved in terms of insufficient intelligence, multi-modal processing capabilities, personalized adaptation and collaboration efficiency, and fast and accurate document information extraction and processing are achieved, improving work efficiency and document quality.

CN120449831APending Publication Date: 2025-08-08DIGITAL (SHANGHAI) ENTERPRISE DEV CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510335232.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional document processing tools and existing AI document tools have shortcomings in terms of intelligence, multimodal processing capabilities, personalized adaptation, collaboration efficiency and real-time processing, and cannot meet the growing needs of users.

Method used

The document processor based on artificial intelligence AI is adopted, including input module, preprocessing module, document analysis module, document classification module, multi-modal analysis module, semantic graph construction module, information extraction module, information editing module, dynamic optimization module, document temporary storage module and output module. Natural language processing and computer vision technology are used for document analysis and processing, realizing joint analysis and semantic association of text, images, and tables, and providing personalized processing solutions and a friendly user interface.

Benefits of technology

It realizes the rapid and accurate extraction of key information from massive documents, improves work efficiency and scientific decision-making, enhances the ability to integrate multimodal information, supports real-time collaboration and personalized processing, reduces manual operations, and improves the speed and quality of document processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449831A_ABST
    Figure CN120449831A_ABST
Patent Text Reader

Abstract

The invention discloses a document processor based on artificial intelligence (AI), which belongs to the technical field of artificial intelligence and natural language processing, and comprises an input module used for receiving documents from various sources; the preprocessing module is used for carrying out format unification and preprocessing operation on the input document; the document analysis module is used for analyzing documents by using a natural language processing technology and a computer vision technology; a document classification module; a multi-modal analysis module; a semantic map construction module; the information extraction module is used for extracting key information in the document; an information editing module; a document temporary storage module; an output module; and a user interface module. According to the method, key information such as texts, images and table contents can be rapidly and accurately extracted from massive documents, data can be rapidly identified and sorted, complexity and errors caused by manual extraction are avoided, accurate and timely data support is provided for enterprise decision making, and the working efficiency and decision making scientificity are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and natural language processing, and in particular to a document processor based on artificial intelligence (AI). Background Art

[0002] With the rapid development of information technology, document processing has become an essential component of modern office work. Traditional document processing tools offer certain capabilities for editing and formatting, but they still fall short in terms of intelligence, user experience, and efficiency. This is particularly true in areas such as document content generation, semantic understanding, contextual association, and intelligent recommendations.

[0003] Traditional document processing tools rely mainly on manual operations and have the following limitations:

[0004] 1) Low intelligence: Unable to automatically identify the document's logical structure, semantic associations, and potential errors (such as format conflicts and missing citations);

[0005] 2) Weak multimodal processing capabilities: Insufficient ability to jointly analyze mixed text and graphics, tabular data, and formula symbols;

[0006] 3) Insufficient personalized adaptation: It is difficult to dynamically adjust output according to different industry standards (such as APA paper format and legal document templates);

[0007] 4) Inefficient collaboration: The version conflict detection and automatic merging functions are not perfect when multiple people collaborate;

[0008] Although existing AI document tools partially solve the problems of grammar checking and content generation, they still have the following shortcomings:

[0009] 1) Insufficient semantic depth: insufficient understanding of the contextual relevance of professional terms and legal clauses;

[0010] 2) Poor cross-document consistency: Unable to ensure terminology uniformity and logical coherence across multiple documents;

[0011] 3) Real-time processing delay: The response time for complex document processing exceeds 5 seconds, making it difficult to meet real-time interaction requirements;

[0012] In summary, traditional document processing tools and existing AI document tools are difficult to meet the growing needs of users; therefore, we propose an artificial intelligence (AI)-based document processor to solve this problem. Summary of the Invention

[0013] The purpose of the present invention is to provide a document processor based on artificial intelligence (AI) to solve the problems raised in the above background technology.

[0014] In order to achieve the above object, the present invention adopts the following technical solutions:

[0015] Artificial intelligence AI-based document processor, including:

[0016] Input module, used to receive documents from various sources;

[0017] The preprocessing module unifies the format and performs preprocessing operations on the input documents;

[0018] Document analysis module, which uses natural language processing technology and computer vision technology to analyze documents;

[0019] Document classification module, which classifies documents according to preset classification rules and machine learning models;

[0020] Multimodal parsing module, which performs joint parsing and semantic association on heterogeneous data such as text, images, tables, and formulas in documents;

[0021] The semantic graph construction module builds semantic associations, causal relationships, contrast relationships, reference links, and format dependencies based on entities, paragraphs, and diagrams in the document;

[0022] Information extraction module, used to extract key information in the document;

[0023] Information editing module, which edits content according to user needs, such as extracting outlines and continuing writing;

[0024] Dynamic optimization module generates personalized processing solutions based on user historical operation data;

[0025] The document storage module is used to temporarily save and manage documents that users are editing or processing in various application scenarios to ensure data security and convenience;

[0026] Output module, outputs the processed documents in the manner required by the user;

[0027] The user interface module provides a friendly user interaction interface and supports text input, voice input and image input.

[0028] Preferably, the multimodal analysis module includes:

[0029] Text parsing unit: uses the BERT-DeepStruct model to extract paragraph topics, logical relationships, and entity labels;

[0030] Image parsing unit: detects chart areas based on YOLOv7, and generates image-text descriptions using CLIP-ViT;

[0031] Table parsing unit: reconstructs table data structure through TabNet and connects to external database to verify data consistency;

[0032] Preferably, the pre-processing module includes:

[0033] Format conversion unit, converting all documents into editable text format;

[0034] The noise reduction unit performs noise reduction processing on the document to remove irrelevant elements in the document, such as advertisements, headers and footers;

[0035] The image enhancement unit performs image enhancement processing on the images in the document to improve the clarity and recognizability of the images.

[0036] Preferably, the input module is connected to the preprocessing module, the preprocessing module is connected to the document analysis module, the document classification module, the multimodal parsing module and the document temporary storage module, the document analysis module, the document classification module and the multimodal parsing module are connected to the semantic graph construction module, the semantic graph construction module is connected to the information extraction module, the information editing module and the dynamic optimization module, and the information extraction module, the information editing module, the document temporary storage module and the dynamic optimization module are connected to the output module.

[0037] Preferably, the semantic graph construction module includes:

[0038] The entity recognition and extraction unit identifies and extracts various entities from documents, including names of people, places, organizations, time, date, currency, percentages, etc. The entity recognition and extraction unit uses named entity recognition technology to label and identify entities in documents by building an entity dictionary and applying machine learning algorithms.

[0039] The relationship identification and extraction unit identifies relationships between entities in a document, such as relationships between people, relationships between organizations, and associations between things. The relationship identification unit uses relationship extraction technology to extract and identify relationships between entities by analyzing the grammatical structure and semantic information of sentences, combining machine learning algorithms and manual rules.

[0040] The attribute recognition and extraction unit identifies and extracts attribute information of entities in documents, such as a person's age, gender, and occupation, and an organization's founding date, registered capital, and business scope. The attribute recognition unit uses attribute extraction technology to extract and identify entity attributes by analyzing the entity's description and context, combined with machine learning algorithms and knowledge bases.

[0041] The knowledge fusion and reasoning unit fuses and integrates the identified entities, relationships, and attribute information to build a knowledge base of semantic graphs;

[0042] Semantic Query and Application Unit: Provides semantic query and application services based on the constructed semantic graph. By inputting natural language query statements, it searches for relevant entities, relationships, and attribute information in the semantic graph.

[0043] Preferably, the information extraction module includes:

[0044] Key information extraction unit, which extracts key information from documents, such as the article's metadata information such as author, publication date, abstract, keywords, etc.

[0045] Entity information extraction unit: extracts other entity information from the document, such as product name, product specifications, product price, and supplier information;

[0046] Relationship information extraction unit:

[0047] Extract relationship information between entities in documents;

[0048] Event information extraction unit: extracts event information from documents;

[0049] Sentiment information extraction: Analyze the text content in the document and extract the author's emotional tendencies or opinions and attitudes.

[0050] Preferably, the document analysis module performs word segmentation on the text in the document, decomposes the text into words, phrases or vocabulary units, selects appropriate word segmentation algorithms and dictionaries according to different language characteristics and application scenarios, and improves the accuracy and efficiency of word segmentation; analyzes the grammatical structure of sentences in the document, and determines the relationship between each component in the sentence, such as subject, predicate, object, attributive, adverbial, complement, etc.; performs semantic understanding on the text in the document on the basis of lexical analysis and syntactic analysis, and identifies the meaning of entities, concepts, events, relationships, etc. in the text; judges the emotional tendency in the document, that is, whether the text is positive, negative or neutral; evaluates the emotional tendency of the document by analyzing factors such as vocabulary, sentence structure, rhetoric, etc. in the text, combined with emotional dictionaries and machine learning algorithms; extracts the theme or keyword of the document, that is, the core content or topic discussed in the document; Determine the topic of a document using methods such as word frequency, keyword co-occurrence, and text summarization; generate a summary of the document, which is a brief summary of the document's content; identify image content in the document, including objects, scenes, people, text, and other information in the image; extract and classify images using the convolutional neural network (CNN) model in deep learning to identify the categories and attributes of objects in the image; segment images in the document into different regions or objects so that these regions or objects can be processed and analyzed separately; extract feature vectors of images in the document for image recognition, classification, and similarity comparison; convert image text in the document into editable text; understand the content and meaning of images in the document, including scenes, events, behaviors, and other information in the image; and analyze and interpret image content through a combination of computer vision technology and natural language processing technology.

[0051] The beneficial effects of the present invention are:

[0052] 1. The document processor based on artificial intelligence (AI) in the present invention can quickly and accurately extract key information from massive documents, whether it is text, images or table content. It can quickly identify and organize data, avoiding the tediousness and errors of manual extraction, providing accurate and timely data support for corporate decision-making, and greatly improving work efficiency and the scientific nature of decision-making;

[0053] 2. In the present invention, the AI-based document processor effectively integrates multimodal information such as text, images, and tables in documents. For example, in project document management, it combines text descriptions with relevant charts and pictures to make the information presentation more complete and intuitive. This integration helps to fully understand the document content, explore the potential connections between different modal information, and provide a richer perspective for business analysis and innovation.

[0054] 3. In the present invention, the document processor based on artificial intelligence (AI) intelligently classifies documents by using machine learning algorithms, and can also quickly retrieve relevant documents based on user query requirements. In scientific research institutions, it can help researchers quickly find the required research papers from a large amount of literature, saving time and energy, improving the efficiency of research and learning, and promoting the acquisition and dissemination of knowledge.

[0055] 4. In the present invention, the document processor based on artificial intelligence (AI) can achieve a high degree of automation in all aspects, from format conversion and content editing to review and proofreading. For example, in the field of news editing, it can automatically perform typeset, grammar check and content optimization on news articles, reducing the workload of manual operations and lowering labor costs, while improving the speed and quality of document processing, making the document production process more efficient and smooth.

[0056] 5. In the present invention, the document processor based on artificial intelligence (AI) facilitates communication and collaboration among team members by recording the operations and modifications of multiple users on the same document in real time. In corporate offices, employees from different departments can edit and discuss a document at the same time. The system will display everyone's work progress in real time, promote the efficiency of team collaboration, break down information silos, and realize knowledge sharing and innovation. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a system block diagram of the document processor based on artificial intelligence (AI) proposed in the present invention.

[0058] In the figure: 1. Input module; 2. Preprocessing module; 3. Document analysis module; 4. Document classification module; 5. Multimodal parsing module; 6. Semantic graph construction module; 7. Information extraction module; 8. Information editing module; 9. Dynamic optimization module; 10. Document temporary storage module; 11. Output module; 12. User interface module. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0060] Reference Figure 1 , an AI-based document processor, including:

[0061] Input module 1, for receiving documents from various sources;

[0062] Preprocessing module 2, performs format unification and preprocessing operations on the input documents;

[0063] Document analysis module 3 uses natural language processing technology and computer vision technology to analyze documents;

[0064] Document classification module 4, classifies documents according to preset classification rules and machine learning models;

[0065] Multimodal parsing module 5, which performs joint parsing and semantic association on heterogeneous data such as text, images, tables, and formulas in documents;

[0066] Semantic graph construction module 6: constructs semantic associations, causal relationships, contrast relationships, reference links, and format dependencies based on entities, paragraphs, and charts in the document;

[0067] Information extraction module 7, used to extract key information in the document;

[0068] Information editing module 8, which edits content according to user needs, such as extracting outlines and continuing writing;

[0069] Dynamic optimization module 9 generates personalized processing solutions based on user historical operation data;

[0070] The document temporary storage module 10 is used to temporarily save and manage documents being edited or processed by users in various application scenarios to ensure data security and convenience;

[0071] Output module 11, outputs the processed document in the manner required by the user;

[0072] The user interface module 12 provides a user-friendly interactive interface and supports text input, voice input, and image input.

[0073] In this embodiment, the multimodal analysis module 5 includes:

[0074] Text parsing unit: uses the BERT-DeepStruct model to extract paragraph topics, logical relationships, and entity labels;

[0075] Image parsing unit: detects chart areas based on YOLOv7, and generates image-text descriptions using CLIP-ViT;

[0076] Table parsing unit: reconstructs the table data structure through TabNet and connects to the external database to verify data consistency.

[0077] In this embodiment, the preprocessing module 2 includes:

[0078] Format conversion unit, converting all documents into editable text format;

[0079] The noise reduction unit performs noise reduction processing on the document to remove irrelevant elements in the document, such as advertisements, headers and footers;

[0080] The image enhancement unit performs image enhancement processing on the images in the document to improve the clarity and recognizability of the images.

[0081] In this embodiment, the input module 1 is connected to the preprocessing module 2, the preprocessing module 2 is connected to the document analysis module 3, the document classification module 4, the multimodal parsing module 5 and the document temporary storage module 10, the document analysis module 3, the document classification module 4 and the multimodal parsing module 5 are connected to the semantic graph construction module 6, the semantic graph construction module 6 is connected to the information extraction module 7, the information editing module 8 and the dynamic optimization module 9, and the information extraction module 7, the information editing module 8, the document temporary storage module 10 and the dynamic optimization module 9 are connected to the output module 11.

[0082] In this embodiment, the semantic graph construction module 6 includes:

[0083] The entity recognition and extraction unit identifies and extracts various entities from documents, including names of people, places, organizations, time, date, currency, percentages, etc. The entity recognition and extraction unit uses named entity recognition technology to label and identify entities in documents by building an entity dictionary and applying machine learning algorithms.

[0084] The relationship identification and extraction unit identifies relationships between entities in a document, such as relationships between people, relationships between organizations, and associations between things. The relationship identification unit uses relationship extraction technology to extract and identify relationships between entities by analyzing the grammatical structure and semantic information of sentences, combining machine learning algorithms and manual rules.

[0085] The attribute recognition and extraction unit identifies and extracts attribute information of entities in documents, such as a person's age, gender, and occupation, and an organization's founding date, registered capital, and business scope. The attribute recognition unit uses attribute extraction technology to extract and identify entity attributes by analyzing the entity's description and context, combined with machine learning algorithms and knowledge bases.

[0086] The knowledge fusion and reasoning unit fuses and integrates the identified entities, relationships, and attribute information to build a knowledge base of semantic graphs;

[0087] Semantic Query and Application Unit: Provides semantic query and application services based on the constructed semantic graph. By inputting natural language query statements, it searches for relevant entities, relationships, and attribute information in the semantic graph.

[0088] In this embodiment, the information extraction module 7 includes:

[0089] The key information extraction unit extracts key information from the document, such as the article's AI-based document processor, author, publication date, abstract, keywords and other metadata information. This key information can help users quickly understand the basic situation and main content of the document, making it easier for users to filter and sort the documents. For different types of documents, the key information extraction method and focus may vary. For example, for news report documents, the focus is on key information such as AI-based document processor, introduction, and event subject; for academic paper documents, the focus is on key information such as AI-based document processor, author, journal name, research question, research method, and main conclusion; for business contract documents, the focus is on key information such as the two parties to the contract, signing date, contract subject, and amount;

[0090] Entity information extraction unit: Extracts other entity information from documents, such as product name, product specifications, product price, and supplier information. This entity information is of great reference value for the company's market research, product development, supply chain management, and other aspects. For example, in market research reports, it can extract competitor product information and compare the company's own product advantages; in the product development process, it can extract raw material supplier information and product performance indicators, etc. The entity information extraction unit is implemented through natural language processing technology and machine learning algorithms. For example, it uses named entity recognition (NER) technology combined with specific industry dictionaries and training models to accurately extract and classify entities in documents. At the same time, text mining technology can also be used to analyze and mine the large amount of documents accumulated by the company itself to extract valuable entity information and knowledge.

[0091] Relationship information extraction unit:

[0092] Extract the relationship information between entities in the document, such as the cooperative relationship between a company and its customers, the composition relationship between products and parts, etc. Relationship information extraction can help enterprises build and improve their own knowledge graphs and enterprise resource planning (ERP) systems. For example, in the enterprise's ERP system, record the order relationship between customers and enterprises, the procurement relationship between suppliers and enterprises, etc.; in the product development process, record the assembly relationship between products and parts, etc. Relationship information extraction can be achieved through relationship extraction (RE) technology. By analyzing the grammatical structure and semantic information of the sentence, combined with machine learning algorithms and manual rule templates, the entity relationships in the document are extracted and identified. Relationship information extraction can provide strong support for the company's decision-making, business process optimization, etc. For example, by analyzing the changing trends of the cooperative relationship between customers and enterprises in sales data, enterprises can adjust their sales strategies and customer relationship management plans; by analyzing the changes in the composition relationship between parts in product development documents, enterprises can optimize product design and production processes, etc.

[0093] Event information extraction unit: extracts event information from documents, such as the time, location, and list of participants of the meeting; the start time, expected completion time, and progress of the project. Event information extraction is of great significance to the project management, schedule scheduling, and other aspects of the enterprise. For example, in project management, timely understanding of the progress of the project and key node events; in the daily operation of the enterprise, reasonable arrangement of meeting time and participants, etc. Event information extraction can be achieved through event extraction (EE) technology. Based on pre-defined event templates and extraction rules, combined with natural language processing technology and machine learning algorithms, event information in documents is identified and extracted. Event information extraction can help enterprises keep abreast of important events in a timely manner and improve the operational efficiency and management level of the enterprise;

[0094] Sentiment extraction: Analyzes the text content of a document to extract the author's sentiment or opinions, such as positive, negative, or neutral sentiment. Sentiment extraction is valuable for businesses in market research, customer satisfaction surveys, and brand building. For example, by analyzing the sentiment of customer feedback, companies can understand customer satisfaction with their products and services. By analyzing the changing sentiment trends of market opinion, companies can adjust their marketing and public relations strategies. Sentiment extraction can be achieved through sentiment analysis. Using dictionary-based methods or machine learning algorithms, sentiment classification models are constructed to determine and classify the sentiment of text within a document. Semantic analysis techniques can also be combined to further understand the emotional connotation and underlying meaning of the text. For example, the author's sentiment can be determined by analyzing the modifiers and sentiment words in a sentence, as well as the overall semantic and logical relationships within the sentence. The overall sentiment can be determined by analyzing the sentiment correlations between multiple related sentences.

[0095] In this embodiment, the document analysis module 3 performs word segmentation on the text in the document, decomposes the text into words, phrases or vocabulary units, selects appropriate word segmentation algorithms and dictionaries according to different language characteristics and application scenarios, and improves the accuracy and efficiency of word segmentation; analyzes the grammatical structure of sentences in the document, and determines the relationship between each component in the sentence, such as subject, predicate, object, attributive, adverbial, complement, etc.; based on lexical analysis and syntactic analysis, performs semantic understanding on the text in the document, and identifies the meaning of entities, concepts, events, relationships, etc. in the text; and judges the emotional tendency in the document, that is, whether the text is positive, negative or neutral. By analyzing factors such as vocabulary, sentence structure, and rhetoric in the text, combined with sentiment dictionaries and machine learning algorithms, the sentiment of the document is assessed. The document's theme or key words, i.e., the core content or topic discussed in the document, are extracted. The document's theme is determined by analyzing word frequency, keyword co-occurrence, and text summaries. A summary of the document, a brief summary of the document's content, is generated. Image content within the document is identified, including information such as objects, scenes, people, and text within the image. Image feature extraction and classification using a deep learning convolutional neural network (CNN) model identify the categories and attributes of objects within the image. Images within the document are segmented into distinct regions or objects for separate processing and analysis. Feature vectors are extracted from the images within the document for use in image recognition, classification, and similarity comparison. Image text within the document is converted into editable text. The content and meaning of images within the document are understood, including information such as scenes, events, and behaviors within the image. Image content is analyzed and interpreted through a combination of computer vision and natural language processing techniques.

[0096] In this embodiment, when in use, documents from various sources are received through the input module 1, and the format conversion unit converts all documents into an editable text format; the noise reduction unit performs noise reduction processing on the document to remove irrelevant elements in the document, such as noise information in advertisements, headers and footers; the image enhancement unit performs image enhancement processing on the image in the document to improve the clarity and recognizability of the image; the document analysis module 3 uses natural language processing technology and computer vision technology to analyze the document, and the document classification module 4 classifies the document according to preset classification rules and machine learning models; the text parsing unit uses the BERT-DeepStruct model to extract paragraph topics, logical relationships and entity labels; the image parsing unit detects chart areas based on YOLOv7, and CLIP-ViT generates image-text association descriptions; the table parsing unit reconstructs the table data structure through TabNet and associates it with an external database to verify data consistency;

[0097] The entity recognition and extraction unit recognizes and extracts various entities from the document, and labels and identifies the entities in the document by constructing an entity dictionary and using machine learning algorithms; the relationship recognition and extraction unit recognizes the relationship between entities in the document; the attribute recognition and extraction unit recognizes and extracts the attribute information of entities in the document; the knowledge fusion and reasoning unit fuses and integrates the recognized entities, relationships and attribute information to construct a knowledge base of semantic graphs; the semantic query and application unit provides semantic query and application services based on the constructed semantic graph; the information editing module 8 edits content according to user needs, such as extracting outlines and continuing writing; the dynamic optimization module 9 generates personalized processing solutions based on the user's historical operation data; the output module 11 outputs the processed documents in the manner required by the user.

[0098] The above is a detailed introduction to the document processor based on artificial intelligence (AI) provided by the present invention. Specific embodiments are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be noted that, for those skilled in the art, various improvements and modifications may be made to the present invention without departing from the principles of the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A document processor based on artificial intelligence (AI), characterized by: include: An input module (1) for receiving documents from various sources; Preprocessing module (2), performs format unification and preprocessing operations on the input documents; Document analysis module (3), which uses natural language processing technology and computer vision technology to analyze documents; Document classification module (4), classifies documents according to preset classification rules and machine learning models; Multimodal parsing module (5) performs joint parsing and semantic association on heterogeneous data such as text, images, tables, and formulas in documents; Semantic graph construction module (6), which builds semantic associations, causal relationships, contrast relationships, reference links and format dependencies based on entities, paragraphs and diagrams in the document; Information extraction module (7), used to extract key information in the document; Information editing module (8), which edits content according to user needs, such as extracting outlines and continuing writing; Dynamic optimization module (9), generates personalized processing solutions based on user historical operation data; A document temporary storage module (10) is used to temporarily save and manage documents being edited or processed by users in various application scenarios to ensure data security and convenience; Output module (11), outputs the processed document in the manner required by the user; The user interface module (12) provides a friendly user interaction interface and supports text input, voice input and image input.

2. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The multimodal analysis module (5) comprises: Text parsing unit: uses the BERT-DeepStruct model to extract paragraph topics, logical relationships, and entity labels; Image parsing unit: detects chart areas based on YOLOv7, and generates image-text descriptions using CLIP-ViT; Table parsing unit: reconstructs the table data structure through TabNet and connects to the external database to verify data consistency.

3. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The pre-processing module (2) comprises: Format conversion unit, converting all documents into editable text format; The noise reduction unit performs noise reduction processing on the document to remove irrelevant elements in the document, such as advertisements, headers and footers; The image enhancement unit performs image enhancement processing on the images in the document to improve the clarity and recognizability of the images.

4. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The input module (1) is connected to the preprocessing module (2), the preprocessing module (2) is connected to the document analysis module (3), the document classification module (4), the multimodal parsing module (5) and the document temporary storage module (10), the document analysis module (3), the document classification module (4) and the multimodal parsing module (5) are connected to the semantic graph construction module (6), the semantic graph construction module (6) is connected to the information extraction module (7), the information editing module (8) and the dynamic optimization module (9), and the information extraction module (7), the information editing module (8), the document temporary storage module (10) and the dynamic optimization module (9) are connected to the output module (11).

5. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The semantic graph construction module (6) includes: The entity recognition and extraction unit identifies and extracts various entities from documents, including names of people, places, organizations, time, date, currency, percentages, etc. The entity recognition and extraction unit uses named entity recognition technology to label and identify entities in documents by building an entity dictionary and applying machine learning algorithms. The relationship identification and extraction unit identifies relationships between entities in a document, such as relationships between people, relationships between organizations, and associations between things. The relationship identification unit uses relationship extraction technology to extract and identify relationships between entities by analyzing the grammatical structure and semantic information of sentences, combining machine learning algorithms and manual rules. The attribute recognition and extraction unit identifies and extracts attribute information of entities in documents, such as a person's age, gender, and occupation, and an organization's founding date, registered capital, and business scope. The attribute recognition unit uses attribute extraction technology to extract and identify entity attributes by analyzing the entity's description and context, combined with machine learning algorithms and knowledge bases. The knowledge fusion and reasoning unit fuses and integrates the identified entities, relationships, and attribute information to build a knowledge base of semantic graphs; Semantic query and application unit: Based on the constructed semantic graph, it provides semantic query and application services. By inputting natural language query statements, it searches for relevant entities, relationships and attribute information in the semantic graph.

6. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The information extraction module (7) comprises: Key information extraction unit, which extracts key information from documents, such as the article's metadata information such as author, publication date, abstract, keywords, etc. Entity information extraction unit: extracts other entity information from the document, such as product name, product specifications, product price, and supplier information; Relationship information extraction unit: Extract relationship information between entities in documents; Event information extraction unit: extracts event information from documents; Sentiment information extraction: Analyze the text content in the document and extract the author's emotional tendencies or opinions and attitudes.

7. The document processor based on artificial intelligence (AI) according to claim 1, characterized in that: The document analysis module (3) performs word segmentation processing on the text in the document, decomposing the text into words, phrases or vocabulary units, and selects appropriate word segmentation algorithms and dictionaries according to different language characteristics and application scenarios to improve the accuracy and efficiency of word segmentation; analyzes the grammatical structure of sentences in the document, and determines the relationship between various components in the sentence, such as subject, predicate, object, attributive, adverbial, complement, etc.; Based on lexical analysis and syntactic analysis, the text in the document is semantically understood to identify the meaning of entities, concepts, events, relationships, etc. in the text; the emotional tendency in the document is judged, that is, whether the text is positive, negative or neutral; by analyzing the vocabulary, sentence structure, rhetoric and other factors in the text, combined with the sentiment dictionary and machine learning algorithm, the emotional tendency of the document is evaluated; the theme or subject words of the document are extracted, that is, the core content or topic discussed in the document; the theme of the document is determined by analyzing the word frequency, keyword co-occurrence relationship, text summary and other methods in the text; the summary of the document is generated, that is, a brief summary of the document content; the image content in the document is identified, including images Objects, scenes, people, text and other information in the film; through the convolutional neural network (CNN) model in deep learning, feature extraction and classification of images are carried out to identify the categories and attributes of objects in the image; images in documents are divided into different areas or objects so that these areas or objects can be processed and analyzed separately; feature vectors of images in documents are extracted for image recognition, classification and similarity comparison; image text in documents is converted into editable text; the content and meaning of images in documents, including scenes, events, behaviors and other information in the images, are understood; image content is analyzed and interpreted through the combination of computer vision technology and natural language processing technology.

Citation Information

Cited By

  • Industrial standard document-oriented deep semantic entity and relation automatic extraction method

    CN120832418A