Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

142 results about "Document level" patented technology

Document Level. Also sometimes referred to as Code Behind. Consists of a single assembly associated with a single workbook, document or template. The code is inside an assembly that is then linked to the particular office file.

Multilevel Data Analysis

A graphical, hierarchical document stream browser and environment for semantic (e.g. framing) and performance data analysis and interactive visualization integrates three scales: entities (competitive), entity (diachronic), and document (linguistic). The document level includes annotation and computational linguistics facilities; the entity level has calendrical and time-series focus. All levels emphasize deep linkage and network (i.e. connective / relational space) view of objects, with user-configurable connectivity. Large language model (LLM) integrations provide synthetic advisories, public opinions, reports, plot insights, comparisons; traditional natural language processing techniques and neural models are also employed. A smart plot system includes a “plot cart” and interpreter with an analysis snippet library. Graph structure may arise via adjustable blending or perceptual optimization of canned attribute-related distance functions or via link-induction query language with deep “semantic stored procedure” subexpressions, or feed into graph neural network-style inference for predictions. Most non-LLM ongoing computational load is client-side, using precomputed hierarchical summary files.
Owner:PONTIMYX CORP

System and method for automatically generating SysML model based on mixed AI and domain knowledge

The invention discloses a SysML model automatic generation system based on mixed AI and domain knowledge, and the system comprises a preprocessing module which is used for carrying out the text preprocessing and structural enhancement of an engineering document of a PDF or Word version; the NLP extraction module is used for identifying six types of core entities by adopting aviation corpus fine tuning BERT, constructing a document-level relational graph by utilizing GNN, modeling a cross-paragraph dependency relationship, calling LLM for semantic fuzzy sentences to generate a thinking chain, extracting a reasoning path and solving ambiguity; the rule conversion engine module is used for mapping the entity relation graph into a SysML memory object tree; and the controllable generation module is used for carrying out limited decoding on the LLM by utilizing a Guidance framework. The invention further discloses an automatic SysML model generation method based on the mixed AI and domain knowledge. According to the method, the problems of low manual modeling efficiency and poor semantic consistency in traditional MBSE implementation are solved.
Owner:SHANGHAI LINGSHU INTELLIGENT TECH CO LTD +2

Automated question-answer generation system for documents

A system and method for generating question-answer pairs is disclosed. The system and method can receive a document. A sentence and / or a further sentence in the document may be identified. A syntactic map for the sentence and / or the further sentence may be generated. Noun phrases and prepositional phrases may be identified based on the syntactic map. Sentence level questions may be generated based on phrases identified using natural language processing (NLP) techniques. Document level questions can also be generated based on syntactic maps generated and NLP techniques.
Owner:AMERICAN EXPRESS (INDIA) PTE LTD

Training of an electronic document extraction model

Systems and methods are disclosed for training an electronic document extraction model, including the generation of the training data to train the model based on sampling a pool of electronic documents based on a rareness metric of the documents. Each electronic document has a document-level rareness metric generated, with the document-level rareness metric being based on one or more of a structural rareness metric or a content rareness metric of the document. The structural rareness metric measures the rareness of the document structure, which may be irrespective of the text content of the document. The content rareness metric measures the rareness of the document content, which may be irrespective of the document structure. The electronic documents are sampled based on the document-level rareness metrics to increase the number of rare documents in the training data without unduly biasing the sampling to optimize the training data for training the extraction model.
Owner:INTUIT INC

Entity linking method and system based on large language model

The invention discloses an entity linking method and system based on a large language model, and the method comprises the steps: carrying out the enhancement of the context of a given entity reference item in an entity document through a large language model, generating a candidate entity list for the entity reference item, and generating description information for each candidate entity in the candidate entity list; constructing a question and answer task of a single choice question for each entity reference item; constructing a reference graph by utilizing the entity reference items and the candidate entity list to obtain association degrees among the entity reference items; all the entity reference items are sorted, question and answer pairs of the preset number of entity reference items with the highest association degree of the current entity reference items are selected as dialogue contexts according to the sorting result, single choice questions are sequentially input into a large language model for answering, and the entity linking result of each entity reference item is obtained. According to the method, the potential relation between the entity reference items can be effectively captured, and the method is good in performance on the entity link task of the document level.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Document retrieval method and system based on electric power semantic enhancement and electronic equipment

The invention relates to a document retrieval method and system based on electric power semantic enhancement and electronic equipment, belongs to the technical field of natural language processing, and solves the problem of low retrieval accuracy caused by low complex knowledge utilization rate and insufficient electric power professional semantic understanding in the prior art. Comprising the following steps: receiving user query content, and obtaining a query embedding vector by utilizing a modal joint embedding model; based on the electric power knowledge graph, utilizing a large language model and a text embedding model to obtain a structured query vector of user query content; according to the query embedded vector and the structured query vector, obtaining a plurality of candidate documents and document-level similarity scores and page-level similarity scores thereof, and further obtaining a comprehensive similarity score of each candidate document by using a double-path prediction model; and obtaining a total score according to the document-level similarity score, the page-level similarity score and the comprehensive similarity score of each candidate document, and selecting a plurality of candidate documents with the highest total score as a retrieval result. And the retrieval precision is improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Question and answer pair generation method and system

The invention provides a question and answer pair generation method and system, and the method comprises the steps: constructing a knowledge base according to obtained knowledge data, the knowledge base comprises a text data set of a document level and a text block data set of a unit level, and constructing knowledge maps corresponding to the text data set and the text block data set, the knowledge graph is used for representing relevance among different documents in a data set, the data set comprises the text data set and the text block data set, the documents in the text data set and the text block data set are sampled respectively to obtain sampled documents, and the sampled documents are stored in the text data set and the text block data set; and according to the sampling document and the knowledge graph corresponding to the sampling document, generating a question and answer pair. The richness, comprehensiveness and diversity of automatically generated question and answer pairs are improved.
Owner:ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Document-level event extraction method and device based on heterogeneous graph interactive learning and public context sensing fusion

The invention provides a document level event extraction method and device based on heterogeneous graph interactive learning and public context sensing fusion. The method comprises the following steps: acquiring a to-be-detected document; inputting a to-be-detected document into the trained event extraction model to generate an extraction result; comprising the steps that sentences in an input document and entity mentions extracted from the sentences are coded, and initial representation vectors of the sentences and the entity mentions are generated; constructing a heterogeneous graph of the input document, performing information interaction on nodes in the heterogeneous graph according to the initial representation vectors of the nodes, and generating global representation vectors mentioned by sentences and entities; performing binary classification on each predefined event type according to the global representation vector of the sentence and predicting the occurrence probability of each predefined event type; extracting a public context between any two entities according to the global representation vectors mentioned by all the entities so as to predict an adjacent matrix between any two entities, and extracting an entity combination; and combining and pairing the predicted event type and the extracted entity to generate an event record.
Owner:HENAN UNIVERSITY

Small sample unified granularity relation extraction method based on large language model

The invention discloses a small sample unified granularity relation extraction method based on a large language model, which comprises the following steps of: firstly, giving a specific task description as a part of an input context of the large language model; thirdly, giving an analogous example to the large language model as context demonstration; in order to better prompt the position information of the entity of the large language model in the context, performing entity enhancement on the context input into the large language model; and finally, a mode for serializing the relation triad is defined, and thinking chain reasoning information is fused in the mode, so that a large language model can be helped to perform relation extraction by utilizing thinking chain prompts. According to the method, context learning, thinking chain and entity enhancement technologies are introduced for unified granularity relation extraction tasks including a sentence level, a document level and a cross-document level, the powerful reasoning ability of a large language model is fully played, and the effectiveness of the model in the unified granularity relation extraction task, especially in a small sample scene, is improved.
Owner:NANJING UNIV +1

End-side adaptive document structure understanding method and system

The invention provides an end-side adaptive document structure understanding method, which comprises the following steps of: uniformly rendering and normalizing a to-be-analyzed document, and outputting a page-level pixel grid and basic metadata; executing lightweight layout analysis and region classification to obtain a bounding box, a reading sequence and a region type label of each region in the page; each document area is routed to a corresponding special analysis channel for parallel analysis, and each analysis channel outputs a structured intermediate result and confidence; performing consistency verification and completion reasoning on intermediate results output by each channel, and generating a traceable verification evidence chain for low-confidence fragments; all channel results after verification are fused into a unified document-level structured output; and for a new document type or a continuous low-confidence mode, starting an adaptive process of a parameter efficient fine tuning technology to generate a channel-level increment weight packet, and updating model parameters of an analysis channel. The method has the beneficial effect that parallel accurate analysis of different elements such as tables, formulas, texts and the like can be realized.
Owner:SHENZHEN XINGSHENG DIGITAL TECH CO LTD

Systems and methods for double level ranking

This invention relates to systems and methods for performing double-level ranking of documents. The system implements methods for retrieving documents based on a pre-processed user query to generate a document level ranking of one or more documents that are determined to be relevant. The system implements methods for aggregating one or more sub-topic snippets from the document level ranked documents. The system further implements methods for generating a topic level ranking of the one or more sub-topic snippets. Once topic level ranking of the one or more sub-topic snippets has been performed, the system implements methods for transmitting the topic level ranked one or more sub-topic snippets to a user associated with the user query.
Owner:INTUIT INC

Document-level relation extraction method and system based on information gain and prototype comparative learning

The invention belongs to the field of natural language processing in computer intelligent information processing, and discloses a document level relation extraction method and system based on information gain and prototype comparative learning. The invention provides a document-level relation extraction model based on a graph structure, which considers two aspects of extracting more accurate node features and relieving data imbalance. The problem that an existing document-level relation extraction model generally adopts a graph-based model and faces inherent data imbalance is solved. At present, the problems that noise interference is caused by irrelevant nodes and edges in the node feature updating process, the learning ability of a model to a real relation is insufficient due to too many negative samples in a document, and all different relation types cannot be accurately predicted through multi-label classification exist in research.
Owner:YANBIAN UNIV

Layout-aware multimodal pretraining for multimodal document understanding

Systems and methods for document processing that can process and understand the layout, text size, text style, and multimedia of a document can generate more accurate and informed document representations. The layout of a document paired with text size and style can indicate what portions of a document are possibly more important, and the understanding of that importance can help with understanding of the document. Systems and methods utilizing a hierarchical framework that processes the block-level and the document-level of a document can capitalize on these indicators to generate a better document representation.
Owner:GOOGLE LLC

Document-level event extraction method based on interleaving argument association matching algorithm

The invention discloses a document-level event extraction method based on an interleaving argument association matching algorithm, which relates to the field of natural language processing and comprises the following steps of: performing data enhancement on original data of a document-level event to generate an enhanced training data set; defining event types and argument roles corresponding to the event types, and constructing a structured event template; constructing a prompt template, and performing multi-task optimization on the UIE model in combination with the enhanced training data set; extracting an entity, a trigger word and an argument in the target document, and generating a structured recognition result; and carrying out structure recombination by adopting an interleaving argument association matching algorithm to generate a document-level event extraction result. According to the method, through data enhancement, prompt template construction and an interleaving argument association matching algorithm, the adaptability of the model under the condition of data deficiency is improved, error accumulation is reduced, cross-paragraph argument association is accurately recognized, and the accuracy and integrity of document-level event extraction are remarkably improved.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

Iterative graph neural network-based event causal identification method, apparatus and device, and medium

The invention discloses an event causal identification method and device based on an iterative graph neural network, equipment and a medium, and relates to the technical field of artificial intelligence and machine learning, and the method comprises the steps: carrying out the sentence coding and event extraction of an input text, and obtaining a sentence embedding and event mention result; then, sentence-level embedding and document-level embedding of the event are generated using a multi-granularity context awareness mechanism. Then, constructing an initial event causal graph structure, and encoding the initial event causal graph structure to obtain graph embedding; and finally, dynamically updating an event causal graph structure by combining sentence-level embedding, document-level embedding and graph embedding through an iterative graph optimization mechanism, and realizing accurate identification of the event causal relationship. Through a multi-granularity context perception mechanism and an iterative graph optimization mechanism, local and global context information is effectively integrated, the accuracy and robustness of document-level event causal relationship recognition are improved, and the method is particularly excellent in performance when processing long texts and cross-sentence causal relationships and can better adapt to complex document structures.
Owner:NAT UNIV OF DEFENSE TECH

Interactive neural-symbolic orchestration of subjective stored procedures

A graphical, hierarchical document stream browser and environment for semantic (e.g. framing) and performance data analysis and interactive visualization integrates three scales: entities (competitive), entity (diachronic), and document (linguistic). The document level includes annotation and computational linguistics facilities; the entity level has calendrical and time-series focus. All levels emphasize deep linkage and network (i.e. connective / relational space) view of objects, with user-configurable connectivity. Large language model (LLM) integrations provide synthetic advisories, public opinions, reports, plot insights, comparisons; traditional natural language processing techniques and neural models are also employed. A smart plot system includes a “plot cart” and interpreter with an analysis snippet library. Graph structure may arise via adjustable blending or perceptual optimization of canned attribute-related distance functions or via link-induction query language with deep “semantic stored procedure” subexpressions, or feed into graph neural network-style inference for predictions. Most non-LLM ongoing computational load is client-side, using precomputed hierarchical summary files.
Owner:PONTIMYX CORP

Digital management method and system for supply chain receipts

The invention relates to a supply chain document digital management method and system, and the method comprises the steps: obtaining a supply chain document image, and carrying out the element recognition, and obtaining a document element set; judging the ownership state of the receipt element set to generate ownership information; performing receipt identification and storage analysis on the receipt element set based on the ownership information to obtain receipt level information; and storing and recording the receipt element set according to the receipt level information to obtain a secure storage log. According to the invention, the traceability and auditing integrity of the document data can be enhanced.
Owner:SHENZHEN QIANHAIZEJIN IND & FINANCE TECH CO LTD

Entity pair guided scientific and technical literature document level relation extraction method and system

The invention provides an entity pair guided scientific and technical literature document level relation extraction method and system, and the method comprises the steps: carrying out the entity recognition of an input scientific and technical literature document, and obtaining an entity set in the scientific and technical literature document; based on an entity pair pre-screening mechanism of multiple sampling and similarity verification, screening out a candidate entity pair set from all possible entity pairs of the entity set; then generating enhanced relation description fusing corresponding entity type information and relation semantics between the entity pairs; based on a pre-constructed relation semantic knowledge base, a double-layer filtering mechanism is adopted, and a corresponding fine screening candidate relation set is retrieved for each enhanced relation description; and guiding the large language model to perform triple fact judgment by using detailed semantic description of the candidate relationship to obtain an output result. According to the method, high-precision relation extraction is ensured, and meanwhile, the calculation overhead of long text processing is remarkably reduced, so that scientific and technical literature document-level relation extraction is more accurate and efficient.
Owner:CHENGDU DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

A SysML Model Automatic Generation System and Method Based on Hybrid AI and Domain Knowledge

This invention discloses an automatic SysML model generation system based on hybrid AI and domain knowledge, comprising: a preprocessing module for text preprocessing and structure enhancement of PDF or Word versions of engineering documents; an NLP extraction module for identifying six core entities by fine-tuning BERT using aviation corpus, constructing a document-level relationship graph using GNN, modeling cross-paragraph dependencies, and generating "thought chains" by calling LLM for semantically ambiguous sentences to extract inference paths and resolve ambiguities; a rule transformation engine module for mapping entity relationship graphs to SysML in-memory object trees; and a controllable generation module for performing restricted decoding of LLM using the Guidance framework. An automatic SysML model generation method based on hybrid AI and domain knowledge is also disclosed. This invention solves the problems of low efficiency and poor semantic consistency in traditional MBSE implementations involving manual modeling.
Owner:SHANGHAI LINGSHU INTELLIGENT TECH CO LTD +2

Graphics-Informed Data-Contextual Report Compilation via Automated Point-of-Interest Detection in Visualization Inventories

A graphical, hierarchical document stream browser and environment for semantic (e.g. framing) and performance data analysis and interactive visualization integrates three scales: entities (competitive), entity (diachronic), and document (linguistic). The document level includes annotation and computational linguistics facilities; the entity level has calendrical and time-series focus. All levels emphasize deep linkage and network (i.e. connective / relational space) view of objects, with user-configurable connectivity. Large language model (LLM) integrations provide synthetic advisories, public opinions, reports, plot insights, comparisons; traditional natural language processing techniques and neural models are also employed. A smart plot system includes a “plot cart” and interpreter with an analysis snippet library. Graph structure may arise via adjustable blending or perceptual optimization of canned attribute-related distance functions or via link-induction query language with deep “semantic stored procedure” subexpressions, or feed into graph neural network-style inference for predictions. Most non-LLM ongoing computational load is client-side, using precomputed hierarchical summary files.
Owner:PONTIMYX CORP

Document-level entity relationship extraction method and device, electronic equipment and storage medium

The invention provides a document-level entity relationship extraction method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The document-level entity relationship extraction method comprises the following steps: acquiring text data, encoding by utilizing a pre-training language model, labeling entity mention positions, fusing characteristics, determining relevance weights and relevance scores, generating a relationship characteristic graph, constructing a relationship matrix and processing by using a multi-label classifier. According to the method and the device, various information in the document can be comprehensively considered, accurate extraction of the document-level entity relationship is realized, and the problem that global information of the document is neglected while only sentence-level relationship extraction is emphasized is solved. According to the method, multiple mentions of the entity words, the relevance between the entity words and the text, the relevance strength between the entity pairs and other factors are comprehensively considered, the semantic structure of the text can be more accurately understood, the context relevance between the entities is captured, and the accuracy of entity relation extraction in the to-be-processed document is improved.
Owner:CHINA SHIPBUILDING ZHIHAI INNOVATION RES INST CO LTD +1

Duplicate checking method and system for products in various stages of software research and development

The invention discloses a duplicate checking method and system for products in all stages of software research and development, and belongs to the technical field of software engineering. In order to solve the problems of high efficiency and accuracy of code, document and function duplicate checking, the technical means of abstract syntax tree analysis, hash index construction, deep learning semantic analysis, image feature extraction, table structure comparison, semantic matching calculation and the like are mainly adopted. According to the method, code-level structured duplicate checking, document-level multi-modal comparison and functional-level semantic analysis can be realized, and the accuracy and efficiency of product duplicate checking in the software development process are improved.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Document-level event argument extraction method based on enhanced AMR graph

The invention discloses a document-level event argument extraction method based on an enhanced AMR graph. The method comprises the following steps: 1, splicing a document text, an event type and a role type into an input text according to a predetermined label; 2, the text feature tensor is converted into a text feature tensor through an encoder; 3, selecting candidate arguments, and inputting the candidate arguments into a scoring device to calculate scores; 4, selecting the first N candidate arguments with the highest score to enhance the AMR graph; 5, extracting word node features in the graph by using the graph convolutional network; and 6, obtaining an event argument and a role set thereof through a role classifier in combination with the text features and the word node features. According to the method, the event type and role information key information required by an event argument extraction task lacked in the AMR graph are enhanced; besides, the sensitivity of the graph structure to noise interference contained in candidate arguments is reduced by limiting the information flow direction in the graph, so that the enhanced AMR graph can keep the integrity of the semantic relationship between entities in the structure dynamic adjustment process, and the accuracy of current event argument extraction can be improved.
Owner:HEFEI UNIV OF TECH

Hybrid expert routing method and system based on trusted RAG

PendingCN121980043AMaintain computational efficiencyStay scalableMetadata multimedia retrievalMachine learningRouting decisionEngineering
The invention relates to the technical field of artificial intelligence, and discloses a credible RAG-based hybrid expert routing method and system, and the method comprises the steps: obtaining metadata of a document associated with an input text to generate a document-level credible RAG vector; obtaining a hidden representation of each token according to the input text and the text content of the associated document, obtaining a credible vector corresponding to each token based on the document-level credible RAG vector, and splicing the hidden representation of each token and the corresponding credible vector to generate an enhanced input vector of each token; taking the enhanced input vector as an input of a gating network, calculating an expert score and selecting at least one expert network for weighted fusion to generate a final output; through the method, the whole routing decision process of organically fusing the credible hierarchical information of the external knowledge into the hybrid expert model is displayed, the spanning from semantic driving to semantic-credible cooperative driving is realized, and the routing accuracy is remarkably improved.
Owner:HANGZHOU HAINAJIN FUSHUI INTELLIGENT TECHNOLOGY CO LTD

Document-level relation extraction method based on multi-level feature collaborative modeling

The invention relates to the field of natural language processing (NLP), in particular to a document level relation extraction technology. A traditional relation extraction method has the problems of insufficient local semantics, insufficient global semantic modeling, difficulty in reasoning complex relations and the like when processing long documents, cross-sentence relations and long-distance dependence. In order to solve the technical problem, the invention provides a document-level relation extraction method based on multi-level feature collaborative modeling. The method comprises the following steps: acquiring context semantic representation of a document by utilizing a pre-training language model, and constructing entity representation through dynamic aggregation of multiple mentions of an entity; neighborhood fine-grained interaction features between entity pairs are captured in combination with a local interaction convolution module, and key information contexts related to entity relation inference are extracted through a global attention mechanism. Furthermore, a multi-level stacked feature fusion structure is designed, and progressive collaborative modeling of local semantics and global semantics is realized, so that the expression ability of the model to a cross-sentence relationship, long-distance reasoning and a complex relationship is enhanced. In addition, the invention provides a hybrid adaptive loss function to improve the robustness of the model to difficult-to-classify samples and low-frequency relationships; and a teacher-student type knowledge distillation mechanism is introduced, and the student model learning is guided by using pseudo labels and evidence distribution, so that the overall relationship inference performance is improved. The method has high relation modeling ability, reasoning ability and generalization ability, and can be widely applied to tasks such as knowledge graph construction and information extraction.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Document-level relation extraction method and system based on knowledge enhancement and evidence guidance

The invention discloses a document-level relation extraction method and system based on knowledge enhancement and evidence guidance, and belongs to the technical field of natural language processing and information extraction. According to the invention, three core modules are mainly used for cooperative work: a document graph enhancement module is used for constructing a hierarchical heterogeneous graph and fusing co-reference analysis to enhance semantic representation; the knowledge enhancement module introduces an external knowledge base and adopts a confidence coefficient filtering mechanism to reduce knowledge noise; the evidence guidance reasoning module realizes multi-hop reasoning through axial attention and evidence supervision, and solves the technical problems of decentralized modeling of reasoning capability, large knowledge integration noise, insufficient evidence guidance and limited long-range dependence capture capability in the existing method. Experiments show that the method can effectively capture inter-sentence dependence, suppress knowledge noise and improve multi-hop reasoning stability, and can be widely applied to scenes such as knowledge graph construction, intelligent question and answer and information retrieval.
Owner:DALIAN MARITIME UNIVERSITY

Large language model event extraction method fusing label reconstruction and multi-dimensional instruction set

The invention relates to the technical field of natural language processing, and provides a large language model event extraction method fusing label reconstruction and a multi-dimensional instruction set, which comprises the following steps of: 1, reconstructing a label; 2, constructing a multi-dimensional instruction set; constructing two different multi-dimensional instruction sets for the multi-dimensional instruction library by adopting a multi-dimensional layered alternate combination strategy through a second-level instruction architecture and a third-level instruction architecture; 3, fine adjustment of the model; taking the reconstructed document-level event text and the event record sequences in two different formats as input and output of fine tuning respectively, and performing fine tuning on an LLaMA-3. 2-1B model by using a multi-dimensional instruction set and adopting a LoRA technology; 4, event extraction; and extracting events by using a large model for fusing label reconstruction and multi-dimensional instruction set fine tuning. According to the method, the data labeling cost is reduced through label reconstruction, the semantic analysis capability of a large language model on financial field events is enhanced by utilizing a multi-dimensional hierarchical instruction set, and the financial field event extraction performance is improved.
Owner:QINGHAI NORMAL UNIV

Document-level relation extraction method for fusing subgraph and displaying and constructing reasoning path

The invention provides a document-level relation extraction method for fusing subgraphs and displaying a constructed reasoning path, and belongs to the technical field of natural language processing. The method comprises the steps that an input text sequence is converted into a word vector sequence through an encoder; constructing a heterogeneous graph comprising entity nodes, mention nodes and sentence nodes through a document graph construction layer; explicitly defining three reasoning paths in sentences, between sentences and integration; extracting a sub-graph from the periphery of the target entity based on the document graph and the reasoning path, and reasoning by applying an R-GCN network; global encoder feature information and local subgraph feature information are sent to a fusion relation classification layer for relation distribution probability prediction; and optimizing a weighted adaptive loss function, and dynamically adjusting loss contribution degrees of different types of samples. According to the method, the reasoning ability for complex relations, the long-distance dependence capture effect and the accuracy of long-tail relation recognition are improved, and meanwhile, the accuracy, efficiency and generalization ability of document-level relation extraction are remarkably improved.
Owner:HUBEI UNIV

Long text abstract generation method based on hierarchical graph comparison theme

PendingCN122021560ASemantic analysisText processingDocument representationInformation coverage
The invention discloses a long text abstract generation method based on hierarchical graph comparison themes, which comprises the following steps of: 1, preprocessing an original document, dividing sentence sequences, and obtaining global context-aware sentence and document representation through a hierarchical encoder network; 2, deducing document-level and sentence-level topic distribution by using a neural topic model; and 3, constructing a supervision graph based on a standard abstract to perform graph comparison learning so as to close the topic representation of a document and a key sentence and push redundant information. According to the method, the deep semantic structure of the long document can be effectively captured, so that the theme consistency and the information coverage degree of the abstract can be improved, and the redundancy is reduced.
Owner:ANHUI AGRICULTURAL UNIVERSITY