Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Document summarization" patented technology

A key event based multi-document summarization generation method

ActiveCN117874220Bimprove consistencyImprove information coverageSemantic analysisText database indexingDocument summarizationGeneration process
The application discloses a kind of multi-document abstract generation methods based on key event, first by internet collection under the same theme multiple media articles, and in document set basis, according to artificial rule, standard abstract is generated to construct sample dataset;Then the dataset is preprocessed, and the input data of training model is generated;Then construct the sequence-to-sequence multi-document abstract generation model based on key event fusion;Finally, based on the model after training, the output model is constructed, and the document set to be summarized is automatically summarized using the output model.The application uses event extraction technology to extract key events containing dynamic and static information from multiple documents to mine the relationship of multiple documents, which can guide the abstract generation process at multiple levels, thereby improving the information coverage and factual consistency of the abstract result, highlighting the event information in the original text, and enhancing the logicality of the abstract.
Owner:SOUTHEAST UNIV

Explainable and efficient text summarization

ActiveUS12632659B2Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Model globalization for long document summarization

ActiveUS12675517B2Document summarizationAlgorithm
A summarization system includes: K embedding modules configured to: receive K blocks of text, respectively, of a document to be summarized; and generate K first representations based on the K blocks of text, respectively, where K is an integer greater than 2; a first propagation module configured to generate second representations based on the K first representations; a second propagation module configured to generate third representations based on the second representations; an output module configured to select ones of the K blocks based on the third representations; and a summary module configured to generate a summary of the document from text of the selected ones of the K blocks.
Owner:NAVER CORP

Cross-language long document summarization method based on dynamic latent key information constraints

ActiveCN121542420BNatural language translationSemantic analysisDocument summarizationData set
The present application relates to a cross-language long document summarization method based on dynamic latent key information constraint, and belongs to the technical field of natural language processing. The method comprises: cross-language long document summarization dataset construction; cross-language long document summarization model construction based on dynamic latent key information constraint; cross-language long document summarization model training based on dynamic latent key information constraint; and cross-language summarization generation for long document content. According to the four-part method process, a cross-language long document summarization device based on dynamic latent key information constraint is modularly manufactured; the present application can effectively mine deep semantic association and structural features in long documents, and solve the semantic drift and structural defocus problems existing in the traditional method in the cross-language scene. In long document processing, the present application exhibits a significant performance advantage, and the performance of the present application in the ROUGE-L index is improved compared with the baseline model, thereby providing technical support for multilingual information integration, international knowledge sharing and the like.
Owner:KUNMING UNIV OF SCI & TECH

A generative multi-document summarization method using entity explicit graph

The present application provides a generative multi-document summarization method using Entity explicit graph to solve the problems of the prior art. The Entity graph based on SimBert is used instead of the traditional document similarity graph to express the connection between documents. Entity extraction is performed on each paragraph, and the Entity of each paragraph is spliced into a sentence. Then, the cosine similarity between all paragraphs is calculated using SimBert. After threshold filtering, an explicit graph representing the connection between paragraphs is obtained. The graphing method based on neural networks and aimed at the entities within the document is more effective than the inflexible rules defined by humans. In addition, the present application improves the attention mechanism and hierarchical graph attention mechanism of the graph perception, introduces a gating mechanism and a residual connection, so that the explicit graph and implicit graph information can be better integrated, and the implicit relationship learned by the attention mechanism is reserved with a residual path to ensure that the information learned by the network occupies a more important position, thereby guiding the generative multi-document summarization.
Owner:WUHAN UNIV

Explainable and efficient text summarization

PendingUS20260127373A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Explainable and efficient text summarization

PendingUS20260127376A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Explainable and efficient text summarization

PendingUS20260127375A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

A multi-document summarization generation method, system and device

This invention provides a method, system, and device for generating multi-document summaries, relating to the field of computer technology. The method includes: employing an improved Siamese neural network to evaluate the relevance scores between document paragraphs and headings in a multi-document context; combining a hierarchical clustering deduplication method based on relevance scores to extract a subset of paragraphs from the document paragraph set whose importance is greater than a set importance level and whose redundancy is less than a set redundancy level; and processing the paragraph subset using a multi-level Transformer-based summarization device to generate multi-document summary text. The multi-level Transformer-based summarization device has cross-document learning capabilities. This invention can generate concise, high-quality multi-document summary text.
Owner:MINZU UNIVERSITY OF CHINA

Document summarization comparison

PendingUS20260099669A1Natural language data processingDocument summarizationTheoretical computer science
Systems and techniques that facilitate comparisons of machine learning model generated summaries are provided. For example, one or more embodiments described herein can comprise a system, which can comprise a memory that can store computer executable components. The system can also comprise a processor, operably coupled to the memory that can execute the computer executable components stored in memory. The computer executable components can comprise an answer component that generates a first answer to a question of a set of questions based on a document, wherein the document is associated with the set of questions and generates a second answer to the question based on a summary document; and a similarity component that updates a similarity score of the document and the summary document, based on a comparison of the first answer to the second answer and a similarity threshold.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Document processing method and related device

PCT designated stageWO2026012150A1Semantic analysisText processingDocument summarizationLinguistic model
Disclosed are a document processing method and a related device, relating to the field of artificial intelligence. The method comprises: acquiring a document processing request; and acquiring a first fragment set of a first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. The first fragment set is from a part of a fragment graph network constructed from a plurality of fragments of the first document, and the fragment graph network is configured for indicating an association relationship between the plurality of fragments of the first document. Therefore, performing document summarization by selecting from the fragment graph network a small number of representative fragments capable of expressing content of the first document ensures document summary quality, reduces the fragments input into a large language model, shortens the duration for processing long documents, and improves processing efficiency.
Owner:HUAWEI TECH CO LTD

Cross-language long document abstracting method based on dynamic potential key information constraint

ActiveCN121542420ANatural language translationSemantic analysisDocument summarizationData set
The invention relates to a cross-language long document abstracting method based on dynamic potential key information constraint, and belongs to the technical field of natural language processing. The method comprises the following steps: constructing a cross-language long document abstract data set; constructing a cross-language long document abstract model based on dynamic potential key information constraint; training a cross-language long document abstract model based on dynamic potential key information constraint; and performing cross-language abstract generation on the long document content. A cross-language long document abstract device based on dynamic potential key information constraint is modularly manufactured according to the four parts of method processes; according to the method, deep semantic association and structural features in a long document can be effectively mined, and the problems of semantic drift and structural defocus existing in a cross-language scene in a traditional method are solved. Compared with a base line model, the method has the advantages that obvious performance advantages are shown in long document processing, compared with the base line model, the ROUGE-L index performance is improved, and technical support is provided for scenes such as multilingual information integration and international knowledge sharing.
Owner:KUNMING UNIV OF SCI & TECH

Long document summarization generation method based on RNN and sparse self-attention mechanism

The application relates to the field of natural language processing, deep learning and text abstract generation, in particular to a long document abstract generation method based on RNN and a sparse self-attention mechanism, which comprises the following steps: text data is segmented or padded in a data padding mode to fill the data into a fixed length L, word embedding conversion is carried out on the L-length segments to obtain word vector representation; the word vector segments are taken as the input of an encoder, and the encoder encodes to obtain the context representation corresponding to the word vector segments; in the decoding stage, the word vector representation and the corresponding context representation are input into a decoder to obtain the final hidden feature; the hidden feature passes through a layer of Softmax layer to obtain the predicted output; the application effectively utilizes a hierarchical framework model to model text abstract related factors, so that the readability, relevance and accuracy of the abstract generated by the model are improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A multi-document summarization method based on knowledge graph and BART semantics

ActiveCN116860960BSemantic analysisText processingDocument summarizationData set
The application belongs to the technical field of natural language processing, and particularly relates to a multi-document summarization method based on a knowledge graph and BART semantics. The method comprises the following steps: constructing a multi-document summarization training data set; constructing a knowledge graph for multi-document summarization; constructing a multi-document summarization model fusing knowledge and graph attention; training the multi-document summarization model and generating a summary. The semantic knowledge graph fusing external knowledge strengthens the connection of distant entities, the method of fusing knowledge graph and BART semantic information makes the model better combine the attention of the knowledge graph and the text sequence, makes up for the shortcomings of the deep learning model, reduces the dependence of the model on large-scale labeled samples, and generates higher-quality summary content.
Owner:SHANXI UNIV

Document processing method and related equipment

PendingCN121301288ASemantic analysisText processingDocument summarizationLinguistic model
The invention discloses a document processing method and related equipment, and relates to the field of artificial intelligence. The method comprises the steps of obtaining a document processing request; obtaining a first fragment set of the first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. Wherein the first fragment set is from a part of a fragment graph network constructed by a plurality of fragments of the first document, and the fragment graph network is used for indicating the association relationship among the plurality of fragments of the first document. Therefore, a small number of representative fragments capable of expressing the content of the first document are selected from the fragment graph network for document summarization, the document summarization quality is ensured, fragments input to a large language model are reduced, the long document processing time is shortened, and the processing efficiency is improved.
Owner:HUAWEI TECH CO LTD

Multi-document summarization extraction method and system based on capsule-biGRU network and event automatic classification

The application discloses a multi-document summarization extraction method and system based on a Capsule-BiGRU network and event automatic classification, and belongs to the technical field of natural language processing. Firstly, text keywords are extracted, and preliminary clustering is completed according to the extraction result; a local feature matrix of the text is extracted by using a capsule network, a global feature matrix of the text is extracted by using a bidirectional gated recurrent unit network, text similarity fusion analysis is carried out according to the extracted local feature matrix and global feature matrix, and a multi-level similarity vector of the text is obtained; then, text similarity is determined, and accurate clustering is carried out according to the text similarity determination result; finally, the calculation of a minimal dominating set is carried out on each type of document in the text clustering result, the theme and the semantics are fused, and a multi-document summarization extraction result is obtained. The method fuses multi-layer clustering, global features and local features to improve the accuracy of multi-document summarization extraction, and can solve the problem that accurate clustering is difficult in the current multi-document summarization extraction process.
Owner:XI AN JIAOTONG UNIV

Enhanced search result generation using multi-document summarization

ActiveUS12561375B2Special data processing applicationsDocument management systemsDocument summarizationMulti-document summarization
Enhanced search results are generated using multi-document summarization. A multi-document summarization system receives a search query from a user and retrieves a plurality of search result documents based on the search query. The summarization system generates a summary of each of the plurality of search result documents using distinct per-document summarization machine learning models, where the distinct per-document summarization machine learning models are trained on a training dataset. The summarization system synthesizes the summary of each of the plurality of search result documents into a single-consolidated answer responsive to the received search query. The multi-document summarization system formats the single-consolidated answer to include citations to the plurality of search result documents.
Owner:SNOWFLAKE INC

A language model construction method for long text reasoning tasks

PendingCN122311425ADocument summarizationCode generation
This invention relates to the field of computer data processing and discloses a method for constructing a language model for long text inference tasks. The method first constructs a Transformer decoder infrastructure and introduces a calibration requirement predictor at each layer to dynamically determine whether global calibration needs to be triggered. Based on the determination result, the model adaptively selects sliding window attention for local computation to improve efficiency, or uses full attention for global calibration to integrate long-range dependencies, thus forming a dynamically hybrid attention computation path. After training on long-chain inference data, the model can be configured with parameters and deployed according to task requirements to perform tasks such as document summarization and code generation. This invention effectively balances the computational efficiency and modeling accuracy of long text processing, significantly improving inference speed and reducing memory consumption.
Owner:SHENZHEN LUXI TECHNOLOGY CO LTD

Document summarization apparatus, method, and non-transitory computer readable medium

PendingUS20260073124A1Semantic analysisDocument summarizationTheoretical computer science
According to one embodiment, a document summarization apparatus includes a processor. The processor performs natural language processing on text included in a document to extract a plurality of linguistic representations and a first semantic relationship between the linguistic representations from the text. The processor classifies the linguistic representations into a plurality of clusters by semantic similarity. The processor determines a second semantic relationship between the clusters based on the first semantic relationship. The processor generates a graph representing the clusters and the second semantic relationship.
Owner:KK TOSHIBA

Document summarization device, method, and program

PendingJP2026049958ASemantic analysisDocument summarizationVerbal expression
To summarize a document more concisely. [Solution] The document summarization device according to the embodiment comprises an extraction unit, a classification unit, a determination unit, and a generation unit. The extraction unit performs natural language processing on the text contained in the document to extract a plurality of linguistic expressions and a first semantic relationship between the plurality of linguistic expressions from the text. The classification unit classifies the plurality of linguistic expressions into a plurality of clusters based on semantic similarity. The determination unit determines a second semantic relationship between the plurality of clusters based on the first semantic relationship. The generation unit generates a graph representing the plurality of clusters and the second semantic relationship.
Owner:KK TOSHIBA

General opinion summaries for medical decision-making

PendingJP2026500911ANatural language translationTherapiesDocument summarizationEncoder decoder
A method and system for document summarization includes dividing a document into sentences (602) and sorting the sentences by a metric that promotes the spread of review opinions from the document to generate a ranked list of sentences. Groups of sentences with similar embeddings are formed (614), and a trained generalization encoder-decoder model is applied (616) to output common generalizations for the sentences in each group. Sentences are added to the summary in rank order (608) from generalizations corresponding to sentences in the ranked list until a target summary length is reached. Actions are performed (612) in response to the summaries.
Owner:NEC LABORATORIES AMERICA INC

A document abstract optimization generation method and system

PendingCN122262325ASemantic analysisBiological modelsDocument summarizationEngineering
This invention proposes a document summarization optimization generation method and system, belonging to the field of data processing technology. The method includes: based on a historical document dataset, segmenting sentence sequences and constructing a document graph structure; combining a graph neural network to obtain target node features; then combining a contrastive learning model to construct positive and negative sample pairs and calculate the contrastive loss; inputting the target node features into a preset large model to generate a candidate summary sequence; constructing a joint objective optimization function and optimizing it to obtain an optimized graph neural network, an optimized contrastive learning model, and an initial optimized large model; performing knowledge distillation on the initial optimized large model to obtain a target optimized large model; and generating an optimized document summary from the document to be processed using the optimized model. This invention achieves structured processing of documents and fully utilizes structured information through joint optimization and knowledge distillation, fully considering the logical structure of document sentences and the deep semantic relationships between sentences to generate accurate document summaries.
Owner:GUANGDONG POWER GRID CO LTD

Explainable and efficient text summarization

PendingUS20260127374A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP