Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

35 results about "Document summarization" patented technology

A key event based multi-document summarization generation method

ActiveCN117874220Bimprove consistencyImprove information coverageSemantic analysisText database indexingDocument summarizationGeneration process
The application discloses a kind of multi-document abstract generation methods based on key event, first by internet collection under the same theme multiple media articles, and in document set basis, according to artificial rule, standard abstract is generated to construct sample dataset;Then the dataset is preprocessed, and the input data of training model is generated;Then construct the sequence-to-sequence multi-document abstract generation model based on key event fusion;Finally, based on the model after training, the output model is constructed, and the document set to be summarized is automatically summarized using the output model.The application uses event extraction technology to extract key events containing dynamic and static information from multiple documents to mine the relationship of multiple documents, which can guide the abstract generation process at multiple levels, thereby improving the information coverage and factual consistency of the abstract result, highlighting the event information in the original text, and enhancing the logicality of the abstract.
Owner:SOUTHEAST UNIV

Explainable and efficient text summarization

ActiveUS12632659B2Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Model globalization for long document summarization

ActiveUS12675517B2Document summarizationAlgorithm
A summarization system includes: K embedding modules configured to: receive K blocks of text, respectively, of a document to be summarized; and generate K first representations based on the K blocks of text, respectively, where K is an integer greater than 2; a first propagation module configured to generate second representations based on the K first representations; a second propagation module configured to generate third representations based on the second representations; an output module configured to select ones of the K blocks based on the third representations; and a summary module configured to generate a summary of the document from text of the selected ones of the K blocks.
Owner:NAVER CORP

Cross-language long document summarization method based on dynamic latent key information constraints

ActiveCN121542420BNatural language translationSemantic analysisDocument summarizationData set
The present application relates to a cross-language long document summarization method based on dynamic latent key information constraint, and belongs to the technical field of natural language processing. The method comprises: cross-language long document summarization dataset construction; cross-language long document summarization model construction based on dynamic latent key information constraint; cross-language long document summarization model training based on dynamic latent key information constraint; and cross-language summarization generation for long document content. According to the four-part method process, a cross-language long document summarization device based on dynamic latent key information constraint is modularly manufactured; the present application can effectively mine deep semantic association and structural features in long documents, and solve the semantic drift and structural defocus problems existing in the traditional method in the cross-language scene. In long document processing, the present application exhibits a significant performance advantage, and the performance of the present application in the ROUGE-L index is improved compared with the baseline model, thereby providing technical support for multilingual information integration, international knowledge sharing and the like.
Owner:KUNMING UNIV OF SCI & TECH

A generative multi-document summarization method using entity explicit graph

The present application provides a generative multi-document summarization method using Entity explicit graph to solve the problems of the prior art. The Entity graph based on SimBert is used instead of the traditional document similarity graph to express the connection between documents. Entity extraction is performed on each paragraph, and the Entity of each paragraph is spliced into a sentence. Then, the cosine similarity between all paragraphs is calculated using SimBert. After threshold filtering, an explicit graph representing the connection between paragraphs is obtained. The graphing method based on neural networks and aimed at the entities within the document is more effective than the inflexible rules defined by humans. In addition, the present application improves the attention mechanism and hierarchical graph attention mechanism of the graph perception, introduces a gating mechanism and a residual connection, so that the explicit graph and implicit graph information can be better integrated, and the implicit relationship learned by the attention mechanism is reserved with a residual path to ensure that the information learned by the network occupies a more important position, thereby guiding the generative multi-document summarization.
Owner:WUHAN UNIV

Explainable and efficient text summarization

PendingUS20260127373A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Systems and methods for parameter ensembling for reducing hallucination in abstractive summarization

PendingUS20250307532A1Natural language analysisBiological modelsDocument summarizationData set
Embodiments described herein provide a document summarization framework that employs an ensemble of summarization models, each of which is a modified version of a base summarization model to control hallucination. For example, a base summarization model may first be trained on a full training data set. The trained base summarization model is then fine-tuned using a first filtered subset of the training data which contains noisy data, resulting in an “anti-expert” model. The parameters of the anti-expert model are subtracted from the parameters of the trained base model to produce a final summarization model which yields robust factual performance.
Owner:SALESFORCE INC

Explainable and efficient text summarization

PendingUS20260127376A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

Systems and methods for risk factor predictive modeling with document summarization

A system and method for document summarization generates summarized articles and risk factor categorizations for display at a graphical user interface (GUI) dashboard. A transaction monitoring system includes an adverse media dashboard pipeline for processing risk factor alerts and generating document summarizations for display at a user device. Document summarization extracts several sentences from a source text and stacks the sentences to create a summary. The method creates a vector representation of each sentence using a machine learning word embedding model and generates a sentence similarity matrix by computing cosine similarity values. A sentence graph creation algorithm creates a graph corresponding to the sentence similarity matrix and calculates importance scores used in selecting sentences for the document summary. The GUI dashboard includes first, second, third and fourth dashboard regions for displaying alert report records, media records, and graphical user interface layouts of document summaries and risk factor visualizations.
Owner:BANK OF MONTREAL

Explainable and efficient text summarization

PendingUS20260127375A1Natural language translationSemantic analysisDocument summarizationEngineering
A computer-implemented, machine learning method for generating explainable text summaries includes extracting a subset of sentences from an input document as an extractive summary and adding context to the extracted sentences to generate a prompt. A fluent summary is generated by using the prompt as input to a generative language model. Source information for a sentence from the fluent summary is determined by mapping the sentence from the fluent summary to a sentence in the extractive summary and the sentence from the extractive summary to a sentence from the input document. A transparent summary view is generated showing the sentence from the fluent summary along with the source information from the extractive summary and the input document for display on a user interface. The method has applications including, but not limited to medical AI, public safety and other machine learning applications for reliable and explainable document summarization.
Owner:NEC CORP

A multi-document summarization generation method, system and device

This invention provides a method, system, and device for generating multi-document summaries, relating to the field of computer technology. The method includes: employing an improved Siamese neural network to evaluate the relevance scores between document paragraphs and headings in a multi-document context; combining a hierarchical clustering deduplication method based on relevance scores to extract a subset of paragraphs from the document paragraph set whose importance is greater than a set importance level and whose redundancy is less than a set redundancy level; and processing the paragraph subset using a multi-level Transformer-based summarization device to generate multi-document summary text. The multi-level Transformer-based summarization device has cross-document learning capabilities. This invention can generate concise, high-quality multi-document summary text.
Owner:MINZU UNIVERSITY OF CHINA

Document summarization comparison

PendingUS20260099669A1Natural language data processingDocument summarizationTheoretical computer science
Systems and techniques that facilitate comparisons of machine learning model generated summaries are provided. For example, one or more embodiments described herein can comprise a system, which can comprise a memory that can store computer executable components. The system can also comprise a processor, operably coupled to the memory that can execute the computer executable components stored in memory. The computer executable components can comprise an answer component that generates a first answer to a question of a set of questions based on a document, wherein the document is associated with the set of questions and generates a second answer to the question based on a summary document; and a similarity component that updates a similarity score of the document and the summary document, based on a comparison of the first answer to the second answer and a similarity threshold.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Document processing method and related device

PCT designated stageWO2026012150A1Semantic analysisText processingDocument summarizationLinguistic model
Disclosed are a document processing method and a related device, relating to the field of artificial intelligence. The method comprises: acquiring a document processing request; and acquiring a first fragment set of a first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. The first fragment set is from a part of a fragment graph network constructed from a plurality of fragments of the first document, and the fragment graph network is configured for indicating an association relationship between the plurality of fragments of the first document. Therefore, performing document summarization by selecting from the fragment graph network a small number of representative fragments capable of expressing content of the first document ensures document summary quality, reduces the fragments input into a large language model, shortens the duration for processing long documents, and improves processing efficiency.
Owner:HUAWEI TECH CO LTD

Cross-language long document abstracting method based on dynamic potential key information constraint

ActiveCN121542420ANatural language translationSemantic analysisDocument summarizationData set
The invention relates to a cross-language long document abstracting method based on dynamic potential key information constraint, and belongs to the technical field of natural language processing. The method comprises the following steps: constructing a cross-language long document abstract data set; constructing a cross-language long document abstract model based on dynamic potential key information constraint; training a cross-language long document abstract model based on dynamic potential key information constraint; and performing cross-language abstract generation on the long document content. A cross-language long document abstract device based on dynamic potential key information constraint is modularly manufactured according to the four parts of method processes; according to the method, deep semantic association and structural features in a long document can be effectively mined, and the problems of semantic drift and structural defocus existing in a cross-language scene in a traditional method are solved. Compared with a base line model, the method has the advantages that obvious performance advantages are shown in long document processing, compared with the base line model, the ROUGE-L index performance is improved, and technical support is provided for scenes such as multilingual information integration and international knowledge sharing.
Owner:KUNMING UNIV OF SCI & TECH

Long document summarization generation method based on RNN and sparse self-attention mechanism

The application relates to the field of natural language processing, deep learning and text abstract generation, in particular to a long document abstract generation method based on RNN and a sparse self-attention mechanism, which comprises the following steps: text data is segmented or padded in a data padding mode to fill the data into a fixed length L, word embedding conversion is carried out on the L-length segments to obtain word vector representation; the word vector segments are taken as the input of an encoder, and the encoder encodes to obtain the context representation corresponding to the word vector segments; in the decoding stage, the word vector representation and the corresponding context representation are input into a decoder to obtain the final hidden feature; the hidden feature passes through a layer of Softmax layer to obtain the predicted output; the application effectively utilizes a hierarchical framework model to model text abstract related factors, so that the readability, relevance and accuracy of the abstract generated by the model are improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A multi-document summarization method based on knowledge graph and BART semantics

ActiveCN116860960BSemantic analysisText processingDocument summarizationData set
The application belongs to the technical field of natural language processing, and particularly relates to a multi-document summarization method based on a knowledge graph and BART semantics. The method comprises the following steps: constructing a multi-document summarization training data set; constructing a knowledge graph for multi-document summarization; constructing a multi-document summarization model fusing knowledge and graph attention; training the multi-document summarization model and generating a summary. The semantic knowledge graph fusing external knowledge strengthens the connection of distant entities, the method of fusing knowledge graph and BART semantic information makes the model better combine the attention of the knowledge graph and the text sequence, makes up for the shortcomings of the deep learning model, reduces the dependence of the model on large-scale labeled samples, and generates higher-quality summary content.
Owner:SHANXI UNIV

Document processing method and related equipment

PendingCN121301288ASemantic analysisText processingDocument summarizationLinguistic model
The invention discloses a document processing method and related equipment, and relates to the field of artificial intelligence. The method comprises the steps of obtaining a document processing request; obtaining a first fragment set of the first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. Wherein the first fragment set is from a part of a fragment graph network constructed by a plurality of fragments of the first document, and the fragment graph network is used for indicating the association relationship among the plurality of fragments of the first document. Therefore, a small number of representative fragments capable of expressing the content of the first document are selected from the fragment graph network for document summarization, the document summarization quality is ensured, fragments input to a large language model are reduced, the long document processing time is shortened, and the processing efficiency is improved.
Owner:HUAWEI TECH CO LTD

Multi-document summarization extraction method and system based on capsule-biGRU network and event automatic classification

The application discloses a multi-document summarization extraction method and system based on a Capsule-BiGRU network and event automatic classification, and belongs to the technical field of natural language processing. Firstly, text keywords are extracted, and preliminary clustering is completed according to the extraction result; a local feature matrix of the text is extracted by using a capsule network, a global feature matrix of the text is extracted by using a bidirectional gated recurrent unit network, text similarity fusion analysis is carried out according to the extracted local feature matrix and global feature matrix, and a multi-level similarity vector of the text is obtained; then, text similarity is determined, and accurate clustering is carried out according to the text similarity determination result; finally, the calculation of a minimal dominating set is carried out on each type of document in the text clustering result, the theme and the semantics are fused, and a multi-document summarization extraction result is obtained. The method fuses multi-layer clustering, global features and local features to improve the accuracy of multi-document summarization extraction, and can solve the problem that accurate clustering is difficult in the current multi-document summarization extraction process.
Owner:XI AN JIAOTONG UNIV

Document summarization method and device, and storage medium

PCT designated stageWO2025251960A1Biological modelsNatural language data processingDocument summarizationDocumentation
Provided are a document summarization method and device, and a storage medium. The method comprises: in response to an access operation for a document, displaying a summary control in a page of the document (S201); and in response to the document meeting a preset condition or receiving a trigger operation on the summary control, displaying summary content in a document text area associated with the summary control, wherein the summary content is generated on the basis of the content of the document (S202).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Multi-document summarization method and system incorporating explicit and implicit variational augmentation

ActiveCN116304003BSemantic analysisEnergy efficient computingDocument summarizationAlgorithm
The application discloses a kind of multi-document literature abstract methods and systems combined with explicit and implicit variation enhancement, the method of the application includes: input document is captured using neural topic model explicit language topic representation, with initial abstract sentence or output last sentence abstract sentence fusion to obtain explicit fusion features, then using implicit variable model to obtain implicit upper feature, implicit upper feature is used Sentence decoder predicts the sentence planning of next sentence;Initial abstract sentence or output last sentence abstract sentence is input word decoder, combined with the sentence planning of next sentence to obtain next word until the sentence planning of next sentence is the terminator.The application can realize sentence planning control abstract generation, overcome the problems of internal decoding error and incoherent sentence context in decoding process, realize accurate and coherent abstract, improve the quality of abstract generation.
Owner:NAT UNIV OF DEFENSE TECH

SUMMARY OF PREVIOUS OPINIONS ON MEDICAL DECISION-MAKING

PendingDE112024000839T5Natural language translationTherapiesDocument summarizationEncoder decoder
Document summarization procedures and systems involve splitting (602) documents into sentences and sorting the sentences according to a metric that promotes the prevalence of evaluative opinions from the documents to create a ranked list of sentences. Groups of sentences with similar embeddings are formed (614), and a trained generalization-encoder-decoder model is applied (616) to output a common generalization of the sentences in each group. Sentences are added to a summary of the generalizations (608) that correspond to the sentences in the ranked list, in order of their rank, until a target summary length is reached. An action is then taken in response to the summary (612).
Owner:NEC LABORATORIES AMERICA INC

Enhanced search result generation using multi-document summarization

Enhanced search results are generated using multi-document summarization. A multi-document summarization system receives a search query from a user and retrieves a plurality of search result documents based on the search query. The summarization system generates a summary of each of the plurality of search result documents using distinct per-document summarization machine learning models, where the distinct per-document summarization machine learning models are trained on a training dataset. The summarization system synthesizes the summary of each of the plurality of search result documents into a single-consolidated answer responsive to the received search query. The multi-document summarization system formats the single-consolidated answer to include citations to the plurality of search result documents.
Owner:SNOWFLAKE INC

System and method for annotation-guided document summarization in the generation of multiple summaries through generative artificial intelligence

PCT designated stageWO2025254980A1Natural language translationNatural language analysisDocument summarizationEngineering
A computing device operating with generative artificial intelligence (Al) logic for condensing content of a document to produce a summary is described. The computing device comprises at least a processor and a non-transitory storage medium coupled to the processor. The non-transitory storage medium includes an artificial intelligence (Al) summarization workflow software tool configured to identify and extract annotations associated with a document, generate one or more prompts including the annotations and content associated with the document, and output the one or more prompts to generative Al logic. The computing device is further configured to receive from generative Al logic a plurality of summaries in response to the one or more prompts, where a first summary of the plurality of summaries is formed with at least a different writing style than a second summary of the plurality of summaries.
Owner:MH SUB I LLC

A language model construction method for long text reasoning tasks

PendingCN122311425ADocument summarizationCode generation
This invention relates to the field of computer data processing and discloses a method for constructing a language model for long text inference tasks. The method first constructs a Transformer decoder infrastructure and introduces a calibration requirement predictor at each layer to dynamically determine whether global calibration needs to be triggered. Based on the determination result, the model adaptively selects sliding window attention for local computation to improve efficiency, or uses full attention for global calibration to integrate long-range dependencies, thus forming a dynamically hybrid attention computation path. After training on long-chain inference data, the model can be configured with parameters and deployed according to task requirements to perform tasks such as document summarization and code generation. This invention effectively balances the computational efficiency and modeling accuracy of long text processing, significantly improving inference speed and reducing memory consumption.
Owner:SHENZHEN LUXI TECHNOLOGY CO LTD

Document summarization apparatus, method, and non-transitory computer readable medium

PendingUS20260073124A1Semantic analysisDocument summarizationTheoretical computer science
According to one embodiment, a document summarization apparatus includes a processor. The processor performs natural language processing on text included in a document to extract a plurality of linguistic representations and a first semantic relationship between the linguistic representations from the text. The processor classifies the linguistic representations into a plurality of clusters by semantic similarity. The processor determines a second semantic relationship between the clusters based on the first semantic relationship. The processor generates a graph representing the clusters and the second semantic relationship.
Owner:KK TOSHIBA

Document summarization device, method, and program

PendingJP2026049958ASemantic analysisDocument summarizationVerbal expression
To summarize a document more concisely. [Solution] The document summarization device according to the embodiment comprises an extraction unit, a classification unit, a determination unit, and a generation unit. The extraction unit performs natural language processing on the text contained in the document to extract a plurality of linguistic expressions and a first semantic relationship between the plurality of linguistic expressions from the text. The classification unit classifies the plurality of linguistic expressions into a plurality of clusters based on semantic similarity. The determination unit determines a second semantic relationship between the plurality of clusters based on the first semantic relationship. The generation unit generates a graph representing the plurality of clusters and the second semantic relationship.
Owner:KK TOSHIBA

General opinion summaries for medical decision-making

PendingJP2026500911ANatural language translationTherapiesDocument summarizationEncoder decoder
A method and system for document summarization includes dividing a document into sentences (602) and sorting the sentences by a metric that promotes the spread of review opinions from the document to generate a ranked list of sentences. Groups of sentences with similar embeddings are formed (614), and a trained generalization encoder-decoder model is applied (616) to output common generalizations for the sentences in each group. Sentences are added to the summary in rank order (608) from generalizations corresponding to sentences in the ranked list until a target summary length is reached. Actions are performed (612) in response to the summaries.
Owner:NEC LABORATORIES AMERICA INC

System and method for annotation-guided document summarization through generative artificial intelligence

PCT designated stageWO2025254979A1Natural language translationNatural language analysisDocument summarizationEngineering
A computing device operating with generative artificial intelligence (AI) logic for condensing content of a document to produce a summary is described. The computing device features at least a processor and a non-transitory storage medium coupled to the processor. The non-transitory storage medium includes an AI summarization workflow software tool that, when executed, is configured to identify and extract annotations associated with a document, generate a prompt including the annotations and content associated with the document, and output the prompt to generative AI logic to enable generation of at least a summary of the document based on the annotations.
Owner:MH SUB I LLC

System and method for annotation-guided document summarization through generative artificial intelligence

PendingUS20250371265A1Natural language translationDocument summarizationEngineering
A computing device operating with generative artificial intelligence (AI) logic for condensing content of a document to produce a summary is described. The computing device features at least a processor and a non-transitory storage medium coupled to the processor. The non-transitory storage medium includes an AI summarization workflow software tool that, when executed, is configured to identify and extract annotations associated with a document, generate a prompt including the annotations and content associated with the document, and output the prompt to generative AI logic to enable generation of at least a summary of the document based on the annotations.
Owner:MH SUB I LLC