Large language model-driven theme-oriented multi-document abstract generation method

Through the three-stage large language model summary generation method, the problems of input length limitation and insufficient topic orientation in the prior art during multi-document summary generation are solved, and high-quality and focused multi-document summary generation are achieved.

CN119988604APending Publication Date: 2025-05-13THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510066406.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing large language models face problems such as input length limitation, loss of long context information and lack of topic orientation when dealing with multi-document summary generation tasks, resulting in low summary quality.

Method used

A three-stage big model abstract generation method is adopted, including data preprocessing, topic-oriented abstract generation and clustering analysis. The specific steps include receiving multiple documents, generating a summary and topic of each document, grouping documents through a clustering algorithm, and generating a comprehensive summary within each category, and finally generating a final summary through fusion.

Benefits of technology

It breaks through the input length limit of the large model, improves the adaptability to massive documents, enhances the consistency and focus of the abstract, improves the readability and topic tracking of the abstract, and meets the personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005244388230000051
    Figure BDA0005244388230000051
  • Figure BDA0005244388230000061
    Figure BDA0005244388230000061
  • Figure BDA0005244388230000071
    Figure BDA0005244388230000071
Patent Text Reader

Abstract

The invention relates to a theme-oriented multi-document abstract generation method driven by a large language model, and belongs to the technical field of natural language processing. The method comprises the following steps: executing multi-document data preprocessing; a large language model is applied to generate a concise abstract and a recognition theme for each preprocessed single document; digests and topics of the documents are converted into vector representation, and a plurality of categories based on content similarity are formed through a clustering algorithm; for the documents in each category, generating a comprehensive abstract by using a large language model; and fusing the comprehensive abstracts of all categories, and generating a final abstract by using the large language model again. Through staged processing and topic vector injection, the problems of input length limitation and lack of topic coherence when massive documents are processed by a traditional method are solved, and the accuracy and the focusing degree of the abstract are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a topic-oriented multi-document summary generation method driven by a large language model. Background Art

[0002] In recent years, with the rapid development of deep learning and natural language processing technologies, large language models (such as ChatGPT, Claude, Qwen, etc.) have made significant progress in text summarization tasks. Through large-scale pre-training and fine-tuning, these models can generate high-quality single-document summaries and show excellent performance in news summaries, scientific literature research, business analysis reports and other fields. However, when faced with a large number of multi-document summary generation tasks, existing methods still face many challenges.

[0003] Large language models are usually limited to fixed input lengths. For example, GPT-4 supports a context length of 8K tokens, but it is still difficult to handle very long documents or multi-document combinations. To adapt to input requirements, existing methods often use truncation or segmentation, which may cause important information to be lost and affect the quality of the summary. In addition, most multi-document summarization methods focus on extracting key information from a single document and ignore the topic relevance between documents, resulting in the generated summary being not focused enough and unable to highlight the core content.

[0004] In summary, existing large language models face problems such as input length limitation, loss of long context information and lack of topic orientation when dealing with large-scale multi-document summary generation tasks, which limits their widespread application in this field. Summary of the invention

[0005] In view of this, the present invention provides a topic-oriented multi-document summary generation method driven by a large language model. The present invention effectively solves the problems of input length limitation, lack of topic orientation and insufficient customization of summary generation in the prior art when processing massive documents through a three-stage large model summary generation method.

[0006] The technical solution adopted by the present invention is:

[0007] A topic-oriented multi-document summary generation method driven by a large language model comprises the following steps:

[0008] S1, receiving multiple documents as input and performing data preprocessing on the multiple documents;

[0009] S2, for each preprocessed document, a large language model is applied to generate a summary and topic of the document;

[0010] S3, converts the summary and topic of each document into vector representation, and then uses clustering algorithms based on these vectors to group all documents into multiple categories;

[0011] S4, for each category, the large language model is called again to generate a comprehensive summary based on the content of all documents in the category;

[0012] S5, fuses the comprehensive summaries of each category and processes them again through the large language model to generate the final summary.

[0013] Furthermore, the specific method of step S1 is:

[0014] S101, Acquiring Documents from Various Sources;

[0015] S102, performing document cleaning to remove non-text elements in the document and remove irrelevant and redundant text content;

[0016] S103, converting documents of different formats into the same representation form to ensure that all documents have the same structure and encoding method.

[0017] Furthermore, the specific method of step S3 is:

[0018] S301, converting the summary and topic of the document into high-dimensional feature vectors using TF-IDF, Word2Vec or BERT embedding to capture the semantic information of the document, and then combining the summary vector and the topic vector by direct concatenation, element-level operation, average pooling, weight weighting, neural network layer processing and attention mechanism;

[0019] S302, using a clustering algorithm to group the document vectors. The goal of clustering is to classify similar documents into the same category and to classify different documents into different categories.

[0020] Furthermore, the specific method of step S4 is:

[0021] S401, concatenating summary texts and subject texts of all documents in the same category to form a unified input document;

[0022] S402, input the input document into the large language model to generate a comprehensive summary, thereby merging the same topics to form a unified topic and summary.

[0023] Furthermore, the specific method of step S5 is:

[0024] The comprehensive summaries of each category are merged to form a complete text, and the merged text is input into the large language model again to generate the final summary.

[0025] The beneficial effects achieved by the present invention are:

[0026] 1. The large model summary and topic generation of a single document of the present invention breaks through the input length limitation of the large model through segmentation processing and topic generation, and improves the adaptability to massive documents.

[0027] 2. The large model summary generation of single-category documents of the present invention enhances the relevance of documents with similar themes and improves the coherence and focus of the summary through topic vector injection and cluster analysis.

[0028] 3. The fusion generation of multi-category document summaries of the present invention improves the readability and topic tracking of summaries through customized summary generation, thus meeting the personalized needs of users.

[0029] 4. The present invention not only achieves innovation in the large language model summary generation method, but also can provide users and enterprises with high-quality service experience and significant economic benefits. DETAILED DESCRIPTION

[0030] The cases described below are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operating procedures, but do not limit the protection scope of the patent of the present invention. Any technical solution obtained in the form of equivalent replacement or equivalent transformation should fall within the protection scope of the present invention.

[0031] A topic-oriented multi-document summary generation method driven by a large language model comprises the following steps:

[0032] S1, multi-document data preprocessing: The system receives multiple documents as input and performs a series of data preprocessing operations on these documents. The data preprocessing operation process includes document acquisition, document cleaning and document coding standardization. The details are as follows:

[0033] Document acquisition: The system can obtain documents from a variety of sources, including but not limited to the Internet, local file systems, databases, or crawling from third-party platforms (such as news websites, academic journals, etc.) through API interfaces.

[0034] This embodiment calls a third-party web crawling API to obtain multiple documents D = {d1, d2, ..., d n}, where n is the number of multiple documents.

[0035] Document cleaning: removes non-text elements in documents, such as HTML tags, script code, style information, etc., while clearing any irrelevant or redundant text content to improve the efficiency and accuracy of subsequent processing.

[0036] This embodiment parses HTML tags to remove document data content containing non-text content such as tables and pictures as much as possible.

[0037] Document encoding standardization: Convert documents in different formats into a standard internal representation to ensure that all documents have the same structure and encoding method.

[0038] This embodiment converts all documents into a plain text format encoded in UTF-8 to facilitate subsequent natural language processing tasks.

[0039] S2, single-document large-model summary and topic generation: For each preprocessed document, the system calls a large language model to generate a summary of the document and the topic of the document.

[0040] This step mainly includes large model summary generation and large model topic generation. The details are as follows:

[0041] Large model summary generation: The preprocessed document is input into the large language model, and the model will generate a concise summary based on the document content. The summary should cover the main points, conclusions or important information of the document, and the length can be adjusted according to actual needs.

[0042] Large model topic generation: In addition to generating summaries, large language models can also identify themes or keywords of documents. These themes can be core concepts, events, characters, etc. that appear repeatedly in user-defined documents, or they can be abstract themes automatically generated by the model. The identification of themes helps with subsequent clustering analysis.

[0043] This example uses the Qwen2.5-7B-Instruct large model, and the temperature parameter is set to 0.1. The prompt word template example is:

[0044]

[0045]

[0046] S3, vector representation and clustering based on single document summaries and topics: The system converts the summary of each document and its topic obtained from step S2 into vector representations, and then uses a clustering algorithm based on these vectors to group the documents into multiple categories.

[0047] This step includes single document vector generation and multi-document clustering classification. The specific process is as follows:

[0048] Single document vector generation: In single document vector generation, the summary and topic of the document are first converted into high-dimensional feature vectors using technologies such as TF-IDF, Word2Vec, or BERT embedding to capture the semantic information of the text. Subsequently, a variety of strategies are used to combine the summary vector and the topic vector, including direct concatenation, element-level operations (addition / multiplication), average pooling, weighted weighting, neural network layer processing, and attention mechanisms. The combined document feature vector not only retains the uniqueness of the summary, but also incorporates the category information of the topic, so that documents with similar topics can be more closely clustered in the vector space. This is conducive to clustering analysis and comprehensive summary generation.

[0049] This embodiment adopts the BAAI / bge-m3 text vector model, and the vector dimension is 1024. The document vector combination method adopted is direct concatenation, and the single document d i The summary vector is s i , the topic vector is t i , then the combined document vector is c i =[s i ,t i ].

[0050] Category division for multi-document clustering: Use clustering algorithms (such as K-means, DBSCAN, hierarchical clustering, etc.) to group document vectors. The goal of clustering is to classify similar documents into the same category and different documents into different categories. The result of clustering can be a fixed number of categories or a dynamically determined number of categories, depending on the needs of the application scenario. Based on the clustering results, the system divides the documents into multiple categories. Based on the above summary vector and topic vector combination method, the documents within each category have high similarity in content and topic, while there are large differences between documents in different categories. The results of category division provide a basis for the subsequent generation of comprehensive summaries.

[0051] This embodiment uses the Kmeans++ algorithm for clustering and the Silhouette Score clustering effect evaluation index to find the optimal number of clusters, and the hyperparameter of the maximum number of clusters is set to 10.

[0052] S4, large model summary generation for single-category documents: For each category formed in step S3, the system calls the large language model again to generate a comprehensive summary based on the content of all documents in the category.

[0053] This step includes single-category document integration and single-category large model comprehensive summary generation. The specific process is as follows:

[0054] Document integration within a single category: All document contents within the same category, including document summary text and subject text, are concatenated to form a unified input document.

[0055] Comprehensive summary generation: The integrated document text is input into the large language model to generate a comprehensive summary. This summary combines the same topics to form a unified topic and summary, integrating the similarities and differences of the content of multiple topic-related documents. The length of the comprehensive summary can be adjusted according to actual needs, usually longer than the summary of a single document, but still concise and clear.

[0056] This embodiment uses the same large language model as step S2. The comprehensive summary prompt word template example is:

[0057]

[0058]

[0059] S5, fusion generation of multi-category document summaries: The system fuses the comprehensive summaries generated by each category in step S4, and generates the final summary through further large language model processing. The specific process of this step is as follows:

[0060] Combine the comprehensive summaries of each category to form a complete text. Input the fused text into the large language model again to generate the final summary. This final summary should comprehensively summarize the content of all categories, highlighting the key points of each category while maintaining overall coherence and completeness. The selection and order of the final summary category topics can be customized according to the user's input. The length of the final summary can also be adjusted according to the user's needs.

[0061] This embodiment uses the same large language model as step S2. The comprehensive summary prompt word template example is:

[0062]

[0063]

[0064] In summary, the present invention solves the problems of input length limitation and lack of topic coherence encountered by traditional methods when processing massive documents through phased processing and topic vector injection, and improves the accuracy and focus of the summary.

Claims

1. A topic-oriented multi-document summarization method driven by a large language model, characterized in that: The following steps are involved: S1, receiving multiple documents as input and performing data preprocessing on the multiple documents; S2, for each preprocessed document, a large language model is applied to generate a summary and topic of the document; S3, converts the summary and topic of each document into vector representation, and then uses clustering algorithms based on these vectors to group all documents into multiple categories; S4, for each category, the large language model is called again to generate a comprehensive summary based on the content of all documents in the category; S5, fuses the comprehensive summaries of each category and processes them again through the large language model to generate the final summary.

2. According to claim 1, a topic-oriented multi-document summarization method driven by a large language model is characterized in that: The specific method of step S1 is: S101, Acquiring Documents from Various Sources; S102, performing document cleaning to remove non-text elements in the document and remove irrelevant and redundant text content; S103, converting documents of different formats into the same representation form to ensure that all documents have the same structure and encoding method.

3. The topic-oriented multi-document summarization method driven by a large language model according to claim 1, characterized in that: The specific method of step S3 is: S301, converting the summary and topic of the document into high-dimensional feature vectors using TF-IDF, Word2Vec or BERT embedding to capture the semantic information of the document, and then combining the summary vector and the topic vector by direct concatenation, element-level operation, average pooling, weight weighting, neural network layer processing and attention mechanism; S302, using a clustering algorithm to group the document vectors. The goal of clustering is to classify similar documents into the same category and to classify different documents into different categories.

4. The method for generating topic-oriented multi-document summarization driven by a large language model according to claim 1, characterized in that: The specific method of step S4 is: S401, concatenating summary texts and subject texts of all documents in the same category to form a unified input document; S402, input the input document into the large language model to generate a comprehensive summary, thereby merging the same topics to form a unified topic and summary.

5. The method for generating topic-oriented multi-document summarization driven by a large language model according to claim 1, characterized in that: The specific method of step S5 is: The comprehensive summaries of each category are merged to form a complete text, and the merged text is input into the large language model again to generate the final summary.