Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

423 results about "Structured document" patented technology

A structured document is an electronic document where some method such as markup or embedded coding, is used to identify the whole and parts of the document as having various meanings beyond their formatting. For example, a structured document might identify a certain portion as a "chapter title" (or "code sample" or "quatrain") rather than as "Helvetica bold 24" or "indented Courier". Such portions in general are commonly called "components" or "elements" of a document.

System and method for adaptive semantic parsing and structured data transformation of digitized documents

A computing system is disclosed for transforming document data into schema-conformant structured outputs. The system obtains document data comprising multi-format structured documents and classifies each document by type and class using vector-based modeling and structural feature analysis. An extraction configuration is selected for each document, the configuration comprising machine-executable instructions for parsing based on semantic and layout characteristics. The system extracts semantic data using structured inference, transforms the semantic data into schema-conformant outputs, and validates the outputs using temporal and domain-specific constraints. Validated structured data may be used for downstream processing, visualizations, or optimization based on performance metrics.
Owner:ALTHQ INC

Defect detection method and system based on joint distribution optimization and structural knowledge guidance

The invention discloses a defect detection method and system based on joint distribution optimization and structural knowledge guidance. The method comprises the following steps: analyzing an equipment structured document to construct an equipment physical topological graph containing a physical connection relationship; processing the topological graph by adopting a graph convolutional network to generate a structured feature vector, and extracting a sensing feature vector of the multi-modal sensing data; constructing a knowledge-enhanced attention model, and guiding a perception feature vector to perform cross-modal fusion by using a structured feature vector as a key and a value; and training the model through a combined optimization target containing joint distribution divergence loss and Lipschitz stability constraint so as to ensure the consistency and physical authenticity of the multi-modal features in the shared semantic space. According to the method, the defect detection accuracy can be remarkably improved, and the root cause diagnosis report with physical interpretability is generated.
Owner:NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI

Multi-model collaborative knowledge graph construction method, system and equipment and storage medium

The invention provides a multi-model collaborative knowledge graph construction method, system and device and a storage medium, and the method comprises the steps that a routing engine receives a knowledge graph construction request, matches a preset rule base according to target domain parameters, and generates a node assembly sequence; the scheduling engine constructs a task execution directed acyclic graph according to the task execution directed acyclic graph; the execution engine preprocesses the original document according to the graph to generate a structured document block set with metadata; calling a pre-training model in the small model resource pool to perform entity and relation extraction to form a preliminary entity set and an association relation set; attribute completion and implicit relation reasoning are carried out on the preliminary entity set, and a completion entity attribute set and a newly-added relation set are generated; performing entity alignment on the preliminary entity set and the complemented entity attribute set by the large model to obtain a fused entity set; and performing conflict resolution on the incidence relation set and the newly-added relation set to obtain a fusion relation set. And storing the fusion entity set and the fusion relationship set to a knowledge graph database.
Owner:XIAMEN YUANTING INFORMATION TECH CO LTD

Project research and development data key information processing method and device

The embodiment of the invention provides a project research and development data key information processing method and device. A research and development content recognition system and a self-adaptive research and development knowledge graph are constructed. Unified processing and time sequence alignment of text, voice and image contents are realized through multi-modal information decomposition and fusion. Semantic completion and error correction are carried out based on research and development of a semantic analysis model, a knowledge graph structure is dynamically constructed and optimized, and a structured document with a traceability relation is generated. The system adopts a deep neural network model to extract research and development key information, constructs a multi-level document framework, and realizes intelligent conversion from research and development data to a project application document. According to the method, the defects of the traditional technology in the aspects of multi-modal information processing and knowledge structure optimization are effectively overcome, and the research and development data management and project declaration efficiency is remarkably improved.
Owner:ZHEJIANG WANCHUANG HUILI TECHNOLOGY SERVICE CO LTD

Structured document generation using document-scale embeddings

A method and related system for generating document embeddings within an embedding space based on a set of structured documents by determining (i) a first vector based on a first segment of a first document and (ii) a second vector based on a second segment of the first document and updating association vectors indicating the second segment based on a distance between the first and second vectors. The method also includes generating a document embedding based on the association vectors, generating a candidate vector based on a candidate document, and determining a result indicating that a second distance between the candidate vector and a first document embedding satisfies a document embedding distance threshold. The method may also include generating a new document by providing, to a text generation model, a portion of the candidate document and a portion of the second segment of the first document.
Owner:CAPITAL ONE SERVICES LLC

Intelligent correction method and device based on semantic analysis, equipment and medium

The invention relates to the technical field of semantic analysis, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses an intelligent correction method and device based on semantic analysis, equipment and a medium. The method comprises the following steps: extracting a structured document content from an explanation data source, extracting an auxiliary explanation text corresponding to each unit in the structured document content from the explanation data source, carrying out semantic understanding processing on the structured document content and the auxiliary explanation texts by utilizing a semantic analysis module to generate a semantic analysis result, and executing a correction operation through an intelligent correction module based on the semantic analysis result. And after the correction operation is completed, the structured document content and the auxiliary explanation text are input into a format reconstruction module, and a specification-combined document is generated and output. According to the method, semantic understanding is carried out by combining structured document content and the auxiliary explanation text, and an automatic and standardized document compliance processing flow is realized based on intelligent correction and format reconstruction driven by a compliance semantic analysis result.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Document processing method and system based on text content extraction

The invention relates to a document processing method and system based on text content extraction. The method comprises the steps that an original document containing text, image and format information is received, the encoding format of the document is automatically detected, character set conversion is executed, and hierarchical indexes including page numbers, paragraphs and tables are established for an unstructured document; the method comprises the following steps: synchronously processing text content and visual layout through a pre-trained visual-language model, extracting word-level and sentence-level semantic features by a text stream embedding layer, analyzing spatial distribution features of document elements by a visual encoder, and fusing text and visual features through a cross-modal attention mechanism; and loading the domain knowledge graph matched with the document type, and executing entity linking to associate the text mentions to the knowledge nodes. According to the document processing method and system based on text content extraction, through the synergistic effect of vision-text joint coding and knowledge enhancement, the accuracy of financial contract key clause recognition tasks is improved, the error rate is lower than that of industry benchmark products, and the semantic understanding precision is remarkably improved.
Owner:WIN THE BID HUIKANG TECH CO LTD

Knowledge base intelligent deep auditing method and system based on AI large model

The invention provides a knowledge base intelligent deep auditing method and system based on an AI large model, and the method comprises the steps: obtaining an unstructured file uploaded by a user, analyzing the unstructured file, extracting a semantic auditing unit, and constructing a semantic auditing unit set; constructing a graph structure based on context dependence according to the semantic auditing unit set, marking continuous, neighbor, reference and reverse logic edge types, and fusing a compliance attention mechanism to update node representation to obtain a node semantic representation set; identifying candidate problem items through a double-layer risk discrimination function by combining the graph structure and the node semantic representation set design according to the introduced law and regulation knowledge embedding library; and executing causal chain backtracking on the candidate problem items to generate a structured auditing conclusion. The auditing accuracy and efficiency are remarkably improved, and the manual rechecking burden is reduced.
Owner:GUANGZHOU TAIXIN INFORMATION TECH CO LTD

Chatbot System For Structured And Unstructured Data

Techniques for operating a chatbot system for enterprise-level conversational agents are disclosed. These techniques are performed by an application or cloud service executing on one or more computing devices. An enterprise system can deploy conversational agents onto user devices to run as chat interfaces for logging analytics question-answering. One example application or cloud service may be a multi-model chat mechanism configured to support these chat interfaces with backend functionality. In response to an incoming question, the chat mechanism first consolidates the question with any conversation history and then, classifies the user's question as either a question regarding unstructured document data, a question regarding structured log data, or a hybrid question. Based on the classification, the chat mechanism can generate a proper large language model (LLM) response.
Owner:ORACLE INT CORP

Text information structured recovery method and system based on large language model and application

The invention provides a text information structured recovery method and system based on a large language model and application. The text information structured recovery method comprises the following steps: S1, extracting original text content from a webpage or an unstructured document; s2, designing a cue word template according to different scenes and target structures, and generating cue words; s3, guiding a large language model to analyze the original text content and the cue word in the step S2, and generating a text result with a hierarchical structure; s4, analyzing a text result in the step S3, constructing a semantic structure tree, and forming a multi-layer nested structure; and S5, applying the multi-layer nested structure in the step S4 to database modeling and content indexing, compared with the defects of non-uniform structure loss, poor universality and semantic understanding intelligence deficiency in the prior art, the manual processing cost can be remarkably reduced, the data structuring efficiency and accuracy are improved, and the semantic understanding intellectuality is improved. And a stable and high-quality structured text support is provided for a large model ecological system.
Owner:SHENZHEN NAT HEALTH CULTURE COMM CO LTD

File conversion method, device, equipment and program product

The invention provides a file conversion method and device, equipment and a program product, and the method comprises the steps: obtaining a to-be-converted file which is an image file or a portable document format file, i.e., a PDF file; identifying elements in the to-be-converted file, determining attributes of the identified elements based on an identification result, and performing logic layering on the identified elements; the elements comprise texts; generating a structured document based on the logical layering result and the attributes of the identification elements; and converting the structured document or the edited structured document into an editable file in a portable document format, so that the PDF can be edited. According to the method and the device, the automatic conversion from the image or the PDF to the editable PDF is realized, and in the conversion process, through the structured document generated in the middle, while the logic level of the document is improved, a user is supported to carry out instant editing, and the editing efficiency and convenience are improved.
Owner:UCWEB

Information extraction system for unstructured documents using retrieval augmentation providing source traceability and error control

A system for extracting a number of data elements from one or more unstructured data sources. The system may separate the text from the tables in a document, such that only the table data may be sent to the large language model (LLM), when the LLM only needs to review the table data. The system generates chunks from the document. The system associates unique identifiers with each chunk to provide traceability. The system identifies relevant chunks from the documents and includes the relevant chunks with a request to extract the data elements in a prompt to the LLM. The system also includes a request for the LLM to report the chunks used during extraction of the data elements. The reported chunks are stored with the extracted data for verification, auditing, and error control.
Owner:AMERICAN INTERNATIONAL GROUP INC

Multi-modal index knowledge base, construction method thereof and question and answer processing method

The invention discloses a multi-modal index knowledge base and a construction method thereof. The construction method comprises the following steps: processing a heterogeneous document to obtain a semi-structured document; identifying a title hierarchical relationship of the document to construct a document logic structure; carrying out minimum chapter blocking on the text of the semi-structured document to obtain logic blocks; performing semantic segmentation on each logic block to obtain text blocks; the method comprises the following steps: constructing text block nodes by meta-information of text blocks, constructing non-text block nodes by meta-information of non-text elements, extracting document nodes, chapter nodes and chapter-chapter inclusion relationships according to a document logic structure, respectively extracting semantic information from the text blocks and the non-text elements, and storing the semantic information in a database; recording the corresponding relationship between the text block nodes and the semantic information and between the non-text block nodes and the semantic information; constructing a knowledge graph based on each node and relationship and storing the knowledge graph into a graph database; and constructing semantic knowledge based on the semantic information and storing the semantic knowledge into a vector database. According to the scheme, lossless retention of multi-modal information and structured organization of document logic are realized, and efficient indexing and accurate recall are facilitated.
Owner:浙江泰隆商业银行股份有限公司

Document analysis method and device, equipment and storage medium

The invention provides a document analysis method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of deep learning, data processing and document analysis. According to the specific implementation scheme, at least one layout area divided by an article to which the document image belongs is determined according to the layout of the document image; identifying element contents of a plurality of layout elements in the document image; sorting the reading sequence of the layout elements in the same layout area; and obtaining structured document information according to the sorting result and the corresponding element content. According to the technical scheme, different article areas on the same page can be accurately identified and separated, the respective reading sequence is correctly reconstructed on the basis, and the content in the document image is converted into structured information.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Intelligent analysis method and device for unstructured PDF document, equipment and medium

The invention discloses an intelligent analysis method and device for an unstructured PDF document, equipment and a medium, and relates to the field of document analysis, the method comprises the steps of obtaining a to-be-analyzed PDF document, analyzing page elements in the PDF document, and generating a document metadata dictionary; if the PDF document does not contain the extractable text, converting the PDF document into an image and performing optical character recognition to generate first structured data; if the PDF document contains the extractable text, judging whether the PDF document contains a table or not; if the PDF document does not contain the table, a PDFMiner is adopted to extract the text, and second structured data is generated; if the PDF document contains the table, performing multi-modal feature extraction and feature fusion on the PDF document according to the document metadata dictionary to obtain multi-modal fusion features, and generating third structured data according to the multi-modal fusion features; according to the method and the device, the analysis precision and efficiency of the PDF document are improved.
Owner:LU ZE TECH CO LTD

Generating structured documents with traceable source lineage

Systems and methods disclosed herein are enabled to dynamically generate structured documents using one or more artificial intelligence models. A computing device receives an output generation request and uses a first AI model to retrieve data chunks from source documents and applicable templates. A second AI model ranks the retrieved chunks based on one or more metrics, such as vector similarity, keyword density, and temporal relevance. A third AI model subsequently generates a response using the ranked chunks, templates, and predefined operational boundaries for each chunk. The generated response is tagged with source identifiers to enable the traceability of the response back to corresponding chunks. The system transmits, via the computing device, the response, the retrieved chunks, and / or the source identifiers.
Owner:CITIBANK N A

Aero-engine system demand servitization collaborative management method

The invention discloses an aero-engine system demand servitization collaborative management method, which comprises the steps of constructing a semantic mapping rule of an aero-engine demand document based on an OSLC specification, and converting an unstructured document into a standardized data model; standardized access of elements in the demand document is achieved through resource identification and a servitization interface, and a Web-based demand document interoperation mechanism is constructed; multidisciplinary collaborative design is carried out based on a demand document interoperation mechanism, and a demand document is updated in real time to form a complete design demand document; constructing an intelligent change analysis mechanism based on project tracking to realize multi-version comparison and change propagation path visualization of the design requirement document; constructing a fine-grained version evolution and conflict management mechanism based on the demand entry influence network, and executing demand version merging and conflict automatic detection in a parallel development scene; according to the method, aero-engine system demand document entry level collaborative editing and standardized service interface can be realized, and the development efficiency is remarkably improved.
Owner:BEIJING INST OF TECH

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Dynamic rule generation and self-adaptive auditing system and method for material management

The invention relates to the technical field of material management, and discloses a dynamic rule generation and self-adaptive auditing system and method for material management, and the system comprises a rule intelligent extraction module, a rule management knowledge base, an enhanced auditing engine, a man-machine cooperation calibration module and a self-adaptive execution module. The method comprises the steps of automatic rule extraction, rule storage and management, enhanced auditing and reasoning, man-machine collaborative calibration, knowledge base real-time optimization and adaptive routing execution. According to the method, the rule is automatically extracted from the unstructured document, the problem that a traditional system rule depends on manpower and is lagged in updating is solved, dynamic optimization of the rule and confidence is achieved by introducing a man-machine collaborative feedback closed loop, the accuracy and transparency of an audit decision are improved by enhancing reasoning and explainable decision technologies, and the audit efficiency is improved. And the optimal balance between auditing efficiency and risk control is realized through self-adaptive routing execution based on credibility.
Owner:PANGU CLOUD CHAIN (TIANJIN) DIGITAL TECH CO LTD

Multi-agent power grid project intelligent monitoring, control and evaluation system and method

The invention relates to the technical field of project intelligent management and control, and discloses a multi-agent power grid project intelligent monitoring, management and control and evaluation system and method, and the system comprises a mixed data resource library, a multi-agent system, an index calculation module and a process control module. The mixed data resource library stores structured business data and unstructured documents; the multi-agent system comprises an information extraction agent, an index design agent, a code generation and treatment agent and the like, works cooperatively, and converts a high-level natural language management and control rule into an executable code; the index calculation module is responsible for executing codes according to a scheduling strategy and calculating quantitative indexes; and the process control module automatically identifies risks and executes management and control operations such as process locking according to the index result and a preset threshold value. According to the invention, the complex business logic can be automated and coded, the accuracy and timeliness of risk identification are improved, and refined and prospective intelligent management and control of the power grid project are realized.
Owner:BEIJING JINGHANG TIANLI TECH CO LTD

Hierarchical Tree-Based Attention for Computationally Efficient Language Processing

This invention introduces a Hierarchical Tree-Based Attention (HTA) mechanism to optimize transformer-based large language models (LLMs) for processing hierarchical documents. HTA leverages a lineage-based approach to model parent-child and sibling relationships, preserving document hierarchy while reducing memory and computational demands. A novel data processing pipeline segments content into blocks, establishes hierarchical relationships, and produces annotated input for LLMs. During attention calculation, embeddings for lineage-related blocks compress information outside the immediate hierarchy, ensuring scalability without sacrificing accuracy. HTA enables efficient applications in structured document processing, such as legal, healthcare, and education, while improving generative tasks like summarization and question answering. This approach advances hierarchical NLP with superior fidelity and reduced latency.
Owner:PIERIS HIMAKARA NAYANAJITH

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Enabling an efficient understanding of contents of a large document without structuring or consuming the large document

The system obtains a record in a database and a property associated with the record in the database, where the record includes a large document, and where the large document is unstructured or semi-structured. The system receives an input indicating a type of analysis to perform associated with the record and performs, using an artificial intelligence, the analysis associated with the record to obtain an output. The type of analysis to obtain the output includes generating a document describing contents of the record, where the document describing the contents of the record is smaller than the record. The system stores the output as the property in the database and enables access to the database based on the property, thereby enabling an efficient understanding of contents of the document without consuming the document.
Owner:NOTION LABS INC

Document generation method based on agent collaboration

The invention discloses a document generation method based on agent collaboration, which comprises the following steps: acquiring initial document data through a content collection agent, and standardizing the initial document data to obtain input content; the input content is distributed to different intelligent agents; performing text generation on the standardized document data through a content generation agent and a multi-task parallel processing mechanism to obtain preliminary document content; performing structure planning coordination on the preliminary document content according to framework requirements through a structure planning agent to obtain a structured document draft; performing language retouching optimization on the structured document draft through a language retouching agent according to a retouching requirement to obtain retouched document content; and optimizing the different intelligent agents, integrating output results of the optimized intelligent agents to obtain final document content, and storing the final document content to obtain final document output.
Owner:SHANGHAI SOURCE CODE BANG DIGITAL TECHNOLOGY CO LTD +2

Information extraction system for unstructured documents using independent tabular and textual retrieval augmentation

A system for extracting a number of data elements from one or more data sources. The system may separate the text from the tables in a document, such that only the table data may be sent to the large language model (LLM), when the LLM only needs to review the table data. The system may include converting a PDF to text, and separating the tables form the document text using markdown language from converting the PDF. The system may form table chunks and text chunks, index the chunks using a vector embedding and store a chunk identifier, document identifier, and or a page identifier with the chunk to provide result traceability. The system may, in response to a prompt, retrieve and send the targeted table chunks or text chunks to the LLM to extract the data elements. The system populates an ontological data store based on the LLM response.
Owner:AMERICAN INTERNATIONAL GROUP INC

Multi-modal PDF document analysis method and device, equipment and medium

The invention provides a multi-mode PDF (Portable Document Format) document analysis method, device and equipment and a medium, and the method comprises the following steps: loading a PDF document and carrying out preprocessing, including page splitting and content cleaning, to generate standardized document data; dynamically extracting contents in the standardized document data, wherein the contents comprise paragraphs, pictures and table elements; processing the dynamically extracted pictures, including shielding meaningless pictures based on a preset rule, and analyzing picture contents by using a multi-modal model to generate readable picture information; processing the dynamically extracted table, including optimizing and merging the table structure into a single element format to generate structured table data; combining paragraphs, readable picture information and structured table data, converting the paragraphs, the readable picture information and the structured table data into a complete structured document format, and inserting in key positions to enhance coherence context description; and outputting the complete structured document format as a final analysis result. According to the method, the analysis precision and the knowledge extraction efficiency of the complex PDF document can be remarkably improved.
Owner:深圳市和讯华谷信息技术有限公司

Document information input method and device and storage medium

The invention discloses a receipt information input method and device and a storage medium, and relates to the technical field of electric digital data processing, the receipt information input method comprises the steps that structured receipt data of a receipt image is extracted through a prior model, and the prior model is pre-trained through receipt layout data; windowing the structured document data according to fields to generate a detection task; verifying the target document data in the detection task through a rule engine to obtain a grammar result, and verifying the target document data in the detection task through a knowledge graph to obtain a semantic result; and if the grammar result and the semantic result are correct, storing the target document data to a document database. The technical problem of low error detection rate of content contradiction before and after receipt input due to the fact that the related technology focuses on text extraction and cannot understand context logic is solved, and the technical effect of improving the error detection rate is achieved.
Owner:GUANGZHOU PINGYUN CRAFTSMAN TECH CO LTD

Structured document generation using different embedding space regions

A method and related system for generating a document using different portions of an embedding space includes obtaining a related document based on a first text, generating first vectors in an embedding space based on the first text and second vectors in the embedding space based on the related document, and determining a first region in the embedding space based on the first vectors and a second region in the embedding space based on the second vectors. The method further includes generating a first portion of a structured document based on the first vectors and third vectors in a third region within the first region but not within the second region. The method further includes generating a second portion of the structured document based on the first and second vectors and the first portion of the structured document.
Owner:CAPITAL ONE SERVICES LLC

Question and answer method based on structured document retrieval enhancement

The invention discloses a question answering method based on structured document retrieval enhancement, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing structured information analysis and blocking processing on a structured document to construct a tree structure reflecting a hierarchical relationship of the document; based on a cross-document association relationship, combining the tree structures corresponding to the plurality of structured documents to form a multi-document tree structure network; and when a question request is received, generating corresponding retrieval enhancement information through the tree structure network, and inputting the retrieval enhancement information into the question and answer model to obtain answer information matched with the question request. According to the method, the multi-document tree structure network which takes the tree hierarchical relationship of the structured document as a core and fuses cross-document topic clustering and reference mapping is constructed, so that efficient retrieval enhanced questions and answers oriented to the structured document are realized, and the accuracy, context consistency and traceability of answers are remarkably improved.
Owner:GRG BANKING EQUIPMENT CO LTD

Character-based representation learning for information extraction using artificial intelligence techniques

Methods, apparatus, and processor-readable storage media for character-based representation learning for information extraction using artificial intelligence techniques are provided herein. An example computer-implemented method includes identifying, from unstructured documents, words and corresponding document position information using artificial intelligence-based text extraction techniques; generating an intermediate output by implementing at least one character embedding with respect to the unstructured documents using at least one artificial intelligence-based encoder; determining structure-related information for at least a portion of the unstructured documents using one or more artificial intelligence-based graph-related techniques; generating a character-based representation of at least a portion of the unstructured documents using at least one artificial intelligence-based decoder; classifying one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and performing one or more automated actions based on the classifying.
Owner:DELL PROD LP