Multi-document semantic similarity comparison system and method based on knowledge vector library
Through a multi-document semantic similarity comparison system based on the knowledge vector library, combined with large language model and multi-level abstract technology, the problems of information overload and insufficient deep semantic understanding in multi-document similarity comparison are solved, and efficient and accurate multi-document processing and innovative evaluation are achieved.
Patent Information
- Application Number
- CN202510069763.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
In the comparison of multiple documents similarity, the prior art has problems such as information overload, insufficient deep semantic understanding, poor adaptability in professional fields, and difficult to achieve balance of interpretability and efficiency.
Design a multi-document semantic similarity comparison system based on knowledge vector library, and generate detailed similarity analysis reports through document preprocessing, key element recognition, multi-level summary generation and knowledge vector library construction, combined with large language models for vectorization comparison and in-depth analysis.
It realizes a deep semantic understanding of documents, accurately identify key information, evaluates innovation, and provides users with high-quality and interpretable similarity analysis results. It is suitable for structured professional documents, improving the efficiency and accuracy of multi-document processing.
Smart Images

Figure CN119990093A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data comparison, and in particular relates to a multi-document semantic similarity comparison system and method based on a knowledge vector library. Background Art
[0002] In the current era of rapid development of information technology and popular application of large language models, research on document similarity comparison technology has made significant progress. However, with the explosive growth of digital content, similarity comparison between multiple documents has become an increasingly important and challenging problem.
[0003] Early research mainly relied on rule-based methods and statistical methods, such as edit distance algorithm, TF-IDF and cosine similarity. Although these methods are computationally efficient, they mainly focus on surface text matching and are difficult to capture deep semantics, especially when dealing with complex semantic relationships of multiple documents. With the development of machine learning, technologies such as latent semantic analysis, topic models and support vector machines have been introduced, which have improved semantic understanding capabilities to a certain extent. These methods have begun to be able to handle the semantic relationships of multiple documents, but still face challenges in efficiency and accuracy, especially in large-scale document collections. In recent years, breakthroughs in deep learning in the field of natural language processing have brought greater progress. Word embedding technology provides a new paradigm for text representation, and recurrent neural networks and convolutional neural networks have shown strong capabilities in text classification and similarity judgment. These technologies provide a better semantic understanding foundation for multi-document similarity comparison. The latest research focuses on document similarity comparison using large-scale pre-trained language models. BERT and its variants are widely used in text encoding and similarity calculation. Sentence-BERT proposes a twin network structure specifically for sentence embedding, while the GPT series of models demonstrate powerful text generation and understanding capabilities. These models have demonstrated unprecedented performance in dealing with multi-document similarity, but have also brought new challenges in computing resources and efficiency.
[0004] To address the challenges of document processing in professional fields, researchers have begun to explore domain adaptation and transfer learning techniques, such as domain-specific pre-training and fine-tuning strategies, and text representation methods combined with domain knowledge graphs. These methods have achieved remarkable results in multi-document similarity comparison in specific fields. However, although previous scholars have made certain contributions in document similarity recognition, there are still some shortcomings in multi-document similarity comparison:
[0005] 1. Information overload and innovation identification: The generation of massive documents makes it increasingly difficult to filter out valuable information from multiple documents. Users not only need to quickly locate relevant content in the information flood, but also need to be able to identify truly innovative ideas and methods. This puts forward dual requirements for multi-document comparison technology: it must be able to efficiently process large amounts of data, and it must also have the ability to identify and evaluate innovations.
[0006] 2. Deep semantic understanding and adaptation to professional fields: With the advancement of natural language processing technology and the increasing degree of specialization in various industries, simple keyword matching can no longer meet the needs. Multi-document comparison technology needs to achieve deep semantic understanding and context analysis, while also being able to adapt to specific terms, formats and key elements in different professional fields to accurately process structured professional documents such as policy documents, legal documents, and scientific papers.
[0007] 3. Balance between explainability and efficiency: In scenarios such as decision support and academic research, users not only need to know the similarities between multiple documents, but also need to understand the specific reasons for the similarities or differences. This requires multi-document comparison tools to provide detailed and explainable analysis results. However, improving the depth and explainability of the analysis often increases computational complexity. How to maintain high efficiency while ensuring the quality of the analysis has become a key challenge.
[0008] 4. Construction and application of knowledge vector library: Although the concept of knowledge vector library has been proposed, how to efficiently construct the knowledge vector library based on a limited number of documents to be compared and use this library to achieve similarity comparison between multiple documents is still an unsolved problem. The knowledge vector library needs to be able to accurately capture the semantic information of the document and support fast similarity calculation and update.
[0009] In this context, developing an intelligent document similarity comparison solution that can deeply understand text semantics, identify innovations, adapt to professional fields, provide interpretable results, and realize efficient multi-document processing capabilities based on the knowledge vector library has important theoretical significance and practical application value. This can not only help users manage and utilize information resources more effectively, but also provide powerful tool support for innovative research, knowledge management, decision support and other fields. Summary of the invention
[0010] The purpose of the present invention is to solve the above-mentioned shortcomings in the prior art and propose a multi-document semantic similarity comparison system and method based on a knowledge vector library. It aims to provide an intelligent, efficient, accurate and interpretable document similarity comparison solution that can deeply understand document semantics, accurately identify key information, evaluate innovation, and provide users with detailed analysis reports, effectively responding to document processing needs in the era of information explosion.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] Design a multi-document semantic similarity comparison system based on a knowledge vector library, including the construction of a knowledge vector library and document similarity comparison. The construction of the knowledge vector library includes:
[0013] A document preprocessing module, which is used to receive input documents, perform format unification processing, and pass the processed documents to subsequent modules;
[0014] A key element identification module, wherein the key element identification module identifies key elements through two submodules, namely, a large language model analysis submodule and an expert knowledge base submodule, and fuses the output results of the two submodules to generate a final key element list;
[0015] A multi-level summary generation module, wherein the multi-level summary generation module generates a parent summary and a child summary of a document through multiple generators, and forms a multi-level semantic representation of the document;
[0016] A knowledge vector library generation module, wherein the knowledge vector library generation module is used to convert the generated parent abstract and child abstract into vectors and construct a generated knowledge vector library;
[0017] The document similarity comparison includes a document preprocessing module, a key element identification module, a multi-level summary generation module and a knowledge vector library generation module, and also includes:
[0018] A vectorized comparison module, which calculates the similarity between the summary generated by the document similarity comparison unit and the pre-built knowledge vector library, and selects N knowledge fragments with the highest similarity for each summary;
[0019] A large language model analysis module, which performs similarity analysis on the parent summary, the child summary and the corresponding TopN knowledge fragments through multiple parallel analysis processes, and generates a similarity analysis result;
[0020] A result comprehensive reporting module receives all results from the large language model analysis module and generates a comprehensive report.
[0021] Furthermore, the document preprocessing module performs format unification processing including text extraction, encoding conversion, and special character removal operations.
[0022] Furthermore, the multi-level summary generation module includes a parent summary generator and a child summary generator. The parent summary generator generates an overall summary containing the main content and key elements of the document based on the entire document content using a large language model. The child summary generator is used to generate corresponding local summaries for each paragraph or chapter of the document. The two-level summaries of the parent summary generator and the child summary generator together constitute a multi-level semantic representation of the document.
[0023] Furthermore, the knowledge vector library generation module includes a parent summary vectorization unit and a child summary vectorization unit, the parent summary vectorization unit is used to convert the parent summary into a high-dimensional vector, and the child summary vectorization unit is used to convert each child summary into a high-dimensional vector.
[0024] Furthermore, the vectorized comparison module includes a similarity calculation unit and a TopN selection unit. The similarity calculation unit uses a cosine similarity algorithm to calculate the similarity between the summary vector and the knowledge base vector. The TopN selection unit selects the N knowledge fragments with the highest similarity for each summary.
[0025] Furthermore, the large language model analysis module includes two parallel analysis processes:
[0026] a. Parent summary analysis process: Input the parent summary and its corresponding TopN knowledge fragments into the large language model for similarity analysis to generate the overall similarity analysis results;
[0027] b. Sub-summary analysis process: Each sub-summary and its corresponding TopN knowledge fragment are input into the large language model for similarity analysis to generate local similarity analysis results;
[0028] The analysis results of both the parent summary analysis process and the child summary analysis process include similarity scores and detailed comparisons of key elements.
[0029] Furthermore, the result comprehensive report module includes a data integration unit, a comprehensive analysis unit and a report generation unit;
[0030] The data integration unit is used to summarize the analysis results of the parent abstract and all sub-abstracts;
[0031] The comprehensive analysis unit is used to input the integrated data into the large language model again to generate a final comprehensive analysis report;
[0032] The report generation unit is used to generate a detailed report including overall similarity evaluation, key element comparison, and innovation analysis based on the comprehensive analysis results.
[0033] In order to solve the above technical problems, the present invention also proposes a multi-document semantic similarity comparison method based on a knowledge vector library, which adopts the multi-document semantic similarity comparison system based on a knowledge vector library, and specifically includes the following steps:
[0034] Step 1: The user inputs the document to be compared into the system;
[0035] Step 2: After the document is preprocessed by the document preprocessing module, the key element identification module automatically identifies the document type and key elements and generates a key element list;
[0036] Step 3: The multi-level summary generation module generates a multi-level summary, and the knowledge vector library generation module constructs a knowledge vector library to complete the construction of the knowledge vector library;
[0037] Step 4: The vectorized comparison module performs vectorized comparison on the summary generated by the document similarity comparison unit and the pre-built knowledge vector library, and performs in-depth analysis on the comparison results through the large language model analysis module to complete the document similarity comparison;
[0038] Step 5: The result comprehensive report module generates a system comprehensive report and presents it to the user.
[0039] Compared with the prior art, the multi-document semantic similarity comparison system and method based on the knowledge vector library proposed in the present invention has the following beneficial effects: the present invention realizes deep semantic understanding and accurate similarity analysis of documents by combining large language models, vectorized comparison and multi-level summarization technology. The system can not only process general texts, but is also particularly suitable for structured professional documents, such as policy documents, legal documents, etc. Through automatic identification and directional comparison of key elements, the system can capture the essential differences between documents and provide users with high-quality and explainable similarity analysis results. Specifically:
[0040] (1) The present invention solves the technical problems of deep semantic understanding and structured document processing by proposing a multi-level summary generation and analysis mechanism.
[0041] Specific measures:
[0042] a) Generate overall document summary (parent summary) and paragraph-level summary (sub-summary) using a large language model;
[0043] b) Perform multi-level similarity comparison based on the summary to capture the overall semantics and local details of the document;
[0044] This improvement helps to more comprehensively understand the content of documents, especially significantly improving the processing capabilities of long texts and complex professional documents.
[0045] (2) The present invention proposes the identification and comparison of domain-specific key elements to solve the technical problem of accurate analysis of documents in professional fields.
[0046] Specific measures:
[0047] a) Combine large language models and domain expert knowledge to automatically identify key elements of documents in a specific domain;
[0048] b) Conduct targeted comparison and analysis of these key elements;
[0049] This improvement enables the system to better adapt to the needs of different professional fields and improves the processing accuracy of professional documents such as policy documents and legal documents.
[0050] (3) The present invention proposes to solve the technical problem of balancing innovative content identification, result interpretability and processing efficiency by achieving comprehensive optimization of innovative evaluation, explainability analysis and efficient processing.
[0051] Specific measures:
[0052] a) Leverage the knowledge base of a large language model to develop an innovative scoring mechanism to identify and quantify novel ideas or methods in documents;
[0053] b) Generate a detailed similarity analysis report, including overall assessment and key element comparison, and provide human-readable explanations;
[0054] c) Adopt hierarchical processing strategies and parallel technology to improve the efficiency of large-scale document processing while ensuring the quality of analysis;
[0055] This comprehensive improvement can help users filter out innovative content from massive amounts of information, understand the basis for similarity judgments, and effectively respond to the efficiency requirements of large-scale document comparison tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0057] Figure 1 It is a schematic block diagram of the overall system architecture of the present invention;
[0058] Figure 2 It is a schematic diagram of the system implementation framework of the present invention;
[0059] Figure 3 is a schematic block diagram of an element identification module in the present invention;
[0060] Figure 4 is a schematic block diagram of a multi-level summary generation module in the present invention;
[0061] Figure 5 It is a schematic block diagram of a vectorized comparison module in the present invention;
[0062] Figure 6 is a schematic block diagram of a large language model analysis module in the present invention;
[0063] Figure 7 It is a schematic block diagram of a result comprehensive reporting module in the present invention;
[0064] The markings in the figure are: 101, document preprocessing module; 102, key element identification module, 201, large language model analysis submodule, 202, expert knowledge base submodule; 103, multi-level summary generation module, 301, parent summary generator, 302, child summary generator; 4, knowledge vector library generation module, 104, vectorization comparison module, 401, knowledge vector library, 402, parent summary vectorization unit, 403, child summary vectorization unit, 404, similarity calculation unit, 405, TopN selection unit; 105, large language model analysis module, 501, parent summary analysis process, 502, child summary analysis process; 106, result comprehensive report module, 601, data integration unit, 602, comprehensive analysis unit, 603, report generation unit. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention; it is obvious that the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments, and all other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making creative work are within the scope of protection of the present invention.
[0066] The structural features of the present invention are now described in detail with reference to the accompanying drawings.
[0067] See also Figure 1-Figure 2 , a multi-document semantic similarity comparison system based on a knowledge vector library, including the construction of a knowledge vector library: a document preprocessing module 101, a key element identification module 102, a multi-level summary generation module 103, and a knowledge vector library generation module 4; document similarity comparison: a vectorized comparison module 104, a large language model analysis module 105, and a result comprehensive report module 106.
[0068] The construction of the knowledge vector library includes:
[0069] The document preprocessing module 101 is responsible for receiving input documents and performing format unification processing, including but not limited to text extraction, encoding conversion, removal of special characters, etc. The processed documents are passed to subsequent modules for analysis.
[0070] See also Figure 3 , a key element identification module 102, which implements the identification of key elements through two submodules:
[0071] a. Large language model analysis submodule 201: Based on the professional field background of the document, the pre-trained large language model is used to automatically identify the key elements that need to be paid attention to in the document similarity comparison task in this field.
[0072] b. Expert knowledge base submodule 202: integrates the key elements list given by experts in the field based on past experience.
[0073] The output results of the two sub-modules are fused to generate the final list of key elements.
[0074] See also Figure 4 , a multi-level summary generation module 103, which includes two main parts:
[0075] a. Parent summary generator 301: Based on the entire document content, a large language model is used to generate an overall summary containing the main content and key elements of the document.
[0076] b. Sub-summary generator 302: Generates a corresponding local summary for each paragraph or chapter of the document.
[0077] These two levels of summaries together constitute a multi-level semantic representation of the document.
[0078] See also Figure 5 , a knowledge vector library generation module 4, which converts the generated summary into a vector and generates a knowledge vector library 401:
[0079] a. Parent summary vectorization unit 402: converts the parent summary into a high-dimensional vector.
[0080] b. Sub-summary vectorization unit 403: converts each sub-summary into a high-dimensional vector.
[0081] Document similarity comparison includes:
[0082] The document preprocessing module 101 is responsible for receiving input documents and performing format unification processing, including but not limited to text extraction, encoding conversion, removal of special characters, etc. The processed documents are passed to subsequent modules for analysis.
[0083] See also Figure 3 , a key element identification module 102, which implements the identification of key elements through two submodules:
[0084] a. Large language model analysis submodule 201: Based on the professional field background of the document, the pre-trained large language model is used to automatically identify the key elements that need to be paid attention to in the document similarity comparison task in this field.
[0085] b. Expert knowledge base submodule 202: integrates the key elements list given by experts in the field based on past experience.
[0086] The output results of the two sub-modules are fused to generate the final list of key elements.
[0087] See also Figure 4 , a multi-level summary generation module 103, which includes two main parts:
[0088] a. Parent summary generator 301: Based on the entire document content, a large language model is used to generate an overall summary containing the main content and key elements of the document.
[0089] b. Sub-summary generator 302: Generates a corresponding local summary for each paragraph or chapter of the document.
[0090] These two levels of summaries together constitute a multi-level semantic representation of the document.
[0091] See also Figure 5 , a vectorized comparison module 104, which calculates the similarity between the generated summary and the pre-built knowledge vector library 401:
[0092] a. Parent summary vectorization unit 402: converts the parent summary into a high-dimensional vector.
[0093] b. Sub-summary vectorization unit 403: converts each sub-summary into a high-dimensional vector.
[0094] c. Similarity calculation unit 404: uses algorithms such as cosine similarity to calculate the similarity between the summary vector and the knowledge base vector.
[0095] d. TopN selection unit 405: selects N knowledge fragments with the highest similarity for each summary.
[0096] See also Figure 6 , a large language model analysis module 105, which includes two parallel analysis processes:
[0097] a. Parent summary analysis process 501: Input the parent summary and its corresponding TopN knowledge fragments into the large language model to generate an overall similarity analysis result.
[0098] b. Sub-abstract analysis process 502: Perform similarity analysis on each sub-abstract and its corresponding TopN knowledge fragment.
[0099] The analysis results of both processes include similarity scores and detailed comparisons of key elements.
[0100] See also Figure 7 , result comprehensive reporting module 106, which receives all the results from the large language model analysis module:
[0101] a. Data integration unit 601: Summarizes the analysis results of the parent summary and all child summaries.
[0102] B. Comprehensive analysis unit 602: inputs the integrated data into the large language model again to generate a final comprehensive analysis report.
[0103] c. Report generation unit 603: Based on the comprehensive analysis results, a detailed report including overall similarity evaluation, key element comparison, innovation analysis, etc. is generated.
[0104] The present invention also provides a multi-document semantic similarity comparison method based on a knowledge vector library, which specifically includes the following steps:
[0105] S1. The user inputs the document to be compared into the system.
[0106] S2. After the document is preprocessed, the system automatically identifies the document type and key elements.
[0107] Specifically, after the document is preprocessed by the document preprocessing module 101, the key element identification module 102 automatically identifies the document type and key elements, and generates a key element list.
[0108] S3. The system generates a multi-level summary and performs a vectorized comparison with the knowledge base.
[0109] Specifically, the multi-level summary generation module 103 generates a multi-level summary, and the knowledge vector library generation module 4 constructs a knowledge vector library 401.
[0110] S4. Large language model comparison and in-depth analysis of the results.
[0111] Specifically, the vectorized comparison module 104 performs vectorized comparison on the summary generated by the document similarity comparison unit and the pre-built knowledge vector library 401 , and performs in-depth analysis on the comparison result through the large language model analysis module 105 .
[0112] S5. The system generates a comprehensive report and presents it to the user.
[0113] Specifically, the result summary report module 106 generates a system summary report.
[0114] For further explanation, this embodiment prepares two text cases for practice.
[0115] Document 1: Impacts and responses to climate change;
[0116] Document 2: Global warming: causes, effects and solutions.
[0117] Our ultimate goal is to achieve a similarity comparison between two texts and provide a detailed comparison analysis report.
[0118] Document 1: Impacts and responses to climate change.
[0119] The continued rise in global temperatures is having a profound impact on our planet. Scientists have observed that the Earth's average temperature has increased significantly over the past century, a phenomenon that is largely attributed to greenhouse gas emissions from human activities. These gases, especially carbon dioxide, form an insulating layer in the atmosphere, causing more heat to be trapped at the Earth's surface. The consequences of climate change are multifaceted. Rising sea levels threaten coastal areas and island nations, and extreme weather events such as hurricanes and droughts are becoming more frequent and severe. Ecosystems are changing, and many species are at risk of extinction. Agricultural production is affected, which could lead to food security issues.
[0120] To address this global challenge, governments and organizations are taking action. Reducing the use of fossil fuels, developing renewable energy, and improving energy efficiency are key strategies. At the same time, adaptive measures such as improving infrastructure and adjusting agricultural practices are also important. On a personal level, everyone can contribute by changing their lifestyles, such as choosing more environmentally friendly modes of transportation and reducing waste. Only by working together globally can we mitigate the effects of climate change and create a sustainable future for future generations.
[0121] Climate change has also had a profound impact on the global economy. Some traditional industries are under pressure to transform, while new green industries are emerging. The insurance industry must cope with the increasing risk of natural disasters, and the tourism industry needs to replan destinations affected by climate change. At the same time, climate change has also exacerbated global inequality, as developing countries often lack the resources and technology to cope with the impacts of climate change.
[0122] International cooperation plays a key role in addressing climate change. The Paris Agreement is an important milestone, marking the world's determination to jointly address climate change. However, there are still many challenges in achieving the goals of the agreement, and countries need to reach more consensus on emission reduction commitments, technology transfer and financial support. Climate diplomacy is becoming an increasingly important issue in international relations.
[0123] Document 2: Global warming: causes, effects and solutions.
[0124] Our planet is facing a major challenge: global warming. Over the past century, the significant rise in the temperature of the Earth's surface has attracted widespread attention from the scientific community. Studies have shown that this phenomenon is closely related to human activities, especially the large amount of greenhouse gases emitted during the industrialization process. These gases accumulate in the atmosphere, forming a greenhouse-like effect that prevents heat from escaping into space, thus causing the global average temperature to rise.
[0125] This temperature rise brings a series of serious consequences. Coastal areas are facing the threat of rising sea levels, and some low-lying island nations may even be submerged in the future. Weather patterns have become more erratic, leading to more extreme climate events such as severe storms and persistent droughts. Biodiversity is also threatened, with many plant and animal species having difficulty adapting to the rapidly changing environment. In addition, the increased instability of agricultural production could trigger a food crisis on a global scale.
[0126] Faced with this global problem, countries around the world are actively seeking solutions. Key strategies include turning to clean energy sources, such as solar and wind power, while improving energy efficiency. Governments are formulating policies to encourage the development and application of green technologies. Adaptive measures are also being implemented, including transforming urban infrastructure to cope with extreme weather and adjusting agricultural production methods. At the individual level, the public is also encouraged to take environmentally friendly actions, such as using public transportation and reducing the use of disposable products. Only through global collaboration and everyone's efforts can we effectively respond to climate change and protect our common home.
[0127] The impact of global warming on the Arctic is particularly significant. The rapid melting of the Arctic ice sheet not only threatens the local ecosystem, but is also changing global climate patterns. The opening of the Arctic shipping route has brought new economic opportunities, but also triggered geopolitical disputes. In addition, the melting of permafrost may release a large amount of methane, further exacerbating the greenhouse effect and forming a vicious cycle.
[0128] Technological innovation plays an increasingly important role in addressing global warming. Carbon capture and storage technologies are expected to reduce the amount of carbon dioxide in the atmosphere, while artificial intelligence and big data analysis are being used to optimize energy use and predict climate change trends. The development of new materials is also helping to improve the energy efficiency of buildings and vehicles. However, the large-scale application of these technologies still faces challenges such as cost and technical maturity.
[0129] Experimental process:
[0130] Experimental setup:
[0131] (1) Experimental environment:
[0132] hardware:
[0133] IntelXeonE5-2680v4@2.40GHz,128GBRAM,NVIDIA TeslaV100GPU.
[0134] Software: Python 3.8, PyTorch 1.9, Transformers 4.10.0.
[0135] (2) Dataset:
[0136] Document 1: “Impacts and Responses to Climate Change” (1,024 words).
[0137] Document 2: “Global Warming: Causes, Effects and Solutions” (956 words).
[0138] Additional test set: 50 pairs of scientific literature abstracts on related topics (500-1000 words each).
[0139] (3) Baseline model:
[0140] TF-IDF+cosine similarity.
[0141] BERT (bert-base-uncased) + average pooling.
[0142] Sentence-BERT(stsb-roberta-large).
[0143] Proposed model: This embodiment proposes a multi-document semantic similarity comparison system based on a knowledge vector library.
[0144] Experimental steps:
[0145] (1) Knowledge vector library construction:
[0146] The GPT-4o model is used to generate multi-level summaries for two main documents and 50 pairs of test documents. The BERT-large model is used to convert the summaries into 512-dimensional vectors, and a knowledge vector library is constructed, which contains the parent and child summary vectors of all documents.
[0147] (2) Identification of key elements:
[0148] The GPT-4o model is used to identify key elements in the document and combined with the predefined expert knowledge base in the field of climate change to generate the final list of key elements.
[0149] (3) Multi-level summary generation:
[0150] Generate a parent summary and multiple sub-summaries for each document:
[0151] Parent Abstract: 150-200 words.
[0152] Sub-abstracts: 50-100 words per paragraph.
[0153] (4) Vectorized comparison:
[0154] Calculate the cosine similarity between the summary vector and the knowledge base vector, and select the top-5 most similar knowledge fragments.
[0155] (5) Large language model analysis:
[0156] The summary and its corresponding Top-5 knowledge fragments are input into the GPT-4o model to generate detailed similarity analysis results.
[0157] (6) Result synthesis and report generation:
[0158] Integrate all analysis results to generate a final comprehensive analysis report.
[0159] The specific experimental process of the two core documents:
[0160] In order to more intuitively demonstrate the workflow and effects of this system, this example conducts detailed experiments on "The Impact and Response to Climate Change" (Document 1) and "Global Warming: Causes, Impacts and Solutions" (Document 2). The following are the specific steps:
[0161] (1) Knowledge vector library construction:
[0162] Use the GPT-4o model to generate multi-level summaries for two documents:
[0163] Document 1 Parent Summary:
[0164] “This article discusses the global impact of climate change and response strategies. Greenhouse gas emissions caused by human activities have caused global temperatures to rise, resulting in rising sea levels and frequent extreme weather. The article emphasizes the importance of reducing the use of fossil fuels and developing renewable energy, and points out the key role of individual behavior changes and international cooperation in addressing climate change. At the same time, the article also explores the far-reaching impact of climate change on the global economy.”
[0165] Document 2 Parent Summary:
[0166] “The article elaborates on the causes, impacts and possible solutions of global warming. Greenhouse gas emissions during industrialization are considered the main cause. Impacts include rising sea levels, increased extreme weather, and threats to biodiversity. The article proposes solutions such as switching to clean energy and improving energy efficiency, and emphasizes the importance of changing personal behavior. In particular, it points out the significant impact of global warming on the Arctic region and the key role of scientific and technological innovation in addressing climate change.”
[0167] The summary is converted into a 512-dimensional vector using the BERT-large model.
[0168] (2) Identification of key elements:
[0169] Using the GPT-4o model to identify key elements in a document:
[0170] Greenhouse gas emissions, sea level rise, extreme weather, ecosystem impacts, energy policy, international cooperation, individual behavior changes, economic impacts, Arctic impacts, and scientific and technological innovation.
[0171] Combined with the predefined expert knowledge base in the field of climate change, key elements are identified and supplemented.
[0172] (3) Multi-level summary generation:
[0173] Generate a sub-summary for each document. The following is a partial sub-summary:
[0174] Document 1, first paragraph summary:
[0175] "Global temperatures continue to rise, mainly due to greenhouse gas emissions from human activities. Over the past century, the average temperature of the Earth has increased significantly. Gases such as carbon dioxide form an insulating layer in the atmosphere, causing more heat to be trapped on the Earth's surface."
[0176] Summary of the first paragraph of Document 2:
[0177] "Global warming is a major challenge we are facing today. Over the past century, the temperature of the earth's surface has risen significantly, which is closely related to human activities. The large amount of greenhouse gases emitted during the industrialization process have accumulated in the atmosphere, forming a greenhouse effect, preventing heat from dissipating, and causing the global average temperature to rise."
[0178] (4) Vectorized comparison:
[0179] Calculate the cosine similarity between the parent summary vectors of two documents and all vectors in the knowledge base.
[0180] Select the Top-5 most similar knowledge fragments.
[0181] (5) Large language model analysis:
[0182] Input the summaries of the two documents and their corresponding Top-5 knowledge fragments into the GPT-4o model to generate detailed similarity analysis results:
[0183] Similarity analysis report:
[0184] 1. Overall similarity: 0.875.
[0185] The two documents are highly related in topic and content, both focusing on the causes, impacts and countermeasures of climate change / global warming.
[0186] 2. Comparison of key elements:
[0187] a. Greenhouse gas emissions: Both documents emphasize that this is the main reason, with a high similarity (0.91).
[0188] b. Sea level rise: Both documents mention this impact, but document 2 describes it in more detail and has a higher similarity (0.86).
[0189] c. Extreme weather: Both documents emphasize this point, and the similarity is very high (0.94).
[0190] d. Ecosystem impact: Both are mentioned, but document 1 is more detailed and has a higher similarity (0.84).
[0191] e. International cooperation: Document 1 emphasizes this more, while Document 2 involves it less, with a medium similarity (0.82).
[0192] 3. Unique viewpoint:
[0193] Document 1: Discusses in detail the impact of climate change on the global economy.
[0194] Document 2: Special emphasis is placed on the impact of global warming on the Arctic region.
[0195] 4. Innovation analysis:
[0196] Document 1 demonstrated innovation in economic impact analysis (score: 0.78).
[0197] Document 2 showed innovation in terms of Arctic impact and application of scientific and technological innovation (score: 0.82).
[0198] 5. Writing style:
[0199] Both documents use an objective, information-dense style, but Document 2 goes into more detail in the solution section.
[0200] Summary: The two documents are highly similar in theme and core content, but each has its own focus and unique perspective, and together they constitute a comprehensive analysis of the climate change / global warming issue.
[0201] Evaluation Metrics:
[0202] Similarity accuracy: Pearson correlation coefficient with the human-annotated similarity scores.
[0203] Key feature identification F1 score: F1 score of key features annotated manually.
[0204] Processing time: The time required to complete the entire analysis process.
[0205] Explainability score: Domain experts rate the reports generated by the system (on a scale of 1-10).
[0206] Experimental results:
[0207] Similarity evaluation results
[0208]
[0209]
[0210] Case analysis: Similarity comparison between document 1 and document 2:
[0211] (1) Overall similarity score:
[0212] This paper’s method: 0.875.
[0213] Sentence-BERT: 0.831.
[0214] BERT+Average Pooling: 0.802.
[0215] TF-IDF+cosine similarity: 0.763.
[0216] (2) Comparison of key elements:
[0217]
[0218] (3) Innovation analysis:
[0219] Document 1 Unique perspective: The impact of climate change on the global economy (Innovation score: 0.78).
[0220] Document 2: Unique perspective: Special impacts in the Arctic region (Innovation score: 0.82).
[0221] (4) Processing time analysis:
[0222]
[0223] (5) Explainability evaluation: 10 domain experts rated the reports generated by the system, with an average score of 9.1 / 10. Main evaluation:
[0224] Depth of analysis: 9.3 / 10.
[0225] Logical clarity: 9.2 / 10.
[0226] Key information capture: 9.4 / 10.
[0227] Detailed comparison results of the two core documents.
[0228] (6) Overall similarity score:
[0229] This paper’s method: 0.875.
[0230] Sentence-BERT: 0.831.
[0231] BERT+Average Pooling: 0.802.
[0232] TF-IDF+cosine similarity: 0.763.
[0233] (7) Comparison details of key elements:
[0234]
[0235]
[0236] (8) Details of innovation analysis:
[0237] Document 1: Unique perspective: The impact of climate change on the global economy; Innovation score: 0.78.
[0238] Analysis: This paper explores in detail how climate change will affect the global economic structure, the insurance and tourism industries, and the exacerbation of global inequality. This comprehensive analysis of economic impacts is relatively rare in the climate change discussion.
[0239] Document 2 Unique perspective: Special impacts in the Arctic region; Innovation score: 0.82.
[0240] Analysis: The special impacts of global warming on the Arctic region are discussed in depth, including ice sheet melting, the opening of shipping routes, and geopolitical impacts. This analysis of regional special impacts provides a new perspective for the discussion of global warming.
[0241] 6. Results and discussion:
[0242] 1. Comparative analysis of two core documents: A detailed comparison of the two documents "Impacts and Responses to Climate Change" and "Global Warming: Causes, Impacts and Solutions" fully demonstrates the superiority of this system in practical applications:
[0243] a. High-precision similarity evaluation: The similarity score of 0.875 given by the system is not only higher than other benchmark models, but also provides sufficient supporting evidence through detailed key element comparison. This proves the advantage of this system in capturing deep semantic relationships.
[0244] b. Fine-grained key element analysis: The system not only identifies 10 key elements, but also quantitatively compares the relevance of each element in the two documents. This detailed analysis provides users with in-depth insights and helps them understand the subtle differences and commonalities between documents.
[0245] c. Innovation recognition capability: The system successfully captured the unique perspectives of the two documents (economic impact and Arctic impact) and gave a reasonable innovation score. This capability is extremely important for academic research, patent analysis and other fields.
[0246] d. Efficient processing: Despite the in-depth analysis, the system only takes about 5.6 seconds to complete the entire process, demonstrating its efficiency in practical applications.
[0247] e. Comprehensive interpretability: The detailed report generated by the system not only gives a quantitative similarity score, but also provides qualitative analysis in multiple dimensions, including a comparison of writing styles. This comprehensive interpretability greatly enhances the credibility and practicality of the analysis results.
[0248] 2. Similarity accuracy: The method proposed in this example significantly outperforms the baseline model in similarity evaluation, with an improvement of 4.5%-19.1%. This shows that the combination of multi-level summarization and knowledge vector library can more accurately capture the semantic relationship between documents.
[0249] 3. Key element identification: The F1 score is significantly improved, 15.6 percentage points higher than the second-best model. This confirms the advantage of this system in identifying and aligning key elements of multiple documents.
[0250] 4. Processing time: Although the processing time of this method is longer, this is an acceptable trade-off considering the depth of analysis and the quality of results. For scenarios that require high-quality analysis, such as academic research or policy making, this extra time is worth it.
[0251] 5. Interpretability: This method performs well in terms of interpretability, with a score much higher than other models. This is extremely important for application scenarios that require a detailed understanding of the basis for similarity judgment.
[0252] 6. Case Study: A detailed comparison of Document 1 and Document 2 demonstrates the system’s ability to identify subtle differences and unique perspectives. The quantitative comparison of key elements provides users with in-depth insights.
[0253] 7. Innovation Identification: The system successfully captures the unique viewpoints in the two documents, which is of great value for academic research and innovation management.
[0254] 8. Processing time analysis: The time distribution of each stage is reasonable. The construction of the knowledge vector library and the generation of multi-level summaries take up more time, which is in line with expectations because these steps lay the foundation for subsequent analysis.
[0255] Overall, the experimental results strongly support the effectiveness and superiority of the present invention. The system has shown significant advantages in accuracy, comprehensiveness and interpretability, and is particularly suitable for application scenarios that require deep text understanding and accurate similarity assessment.
[0256] The multi-document semantic similarity comparison system based on the knowledge vector library of the present invention realizes deep semantic understanding and accurate similarity analysis of documents by combining large language models, vectorized comparison and multi-level summarization technology. The system can not only process general texts, but is also particularly suitable for structured professional documents, such as policy documents, legal documents, etc. Through automatic identification and directional comparison of key elements, the system can capture the essential differences between documents and provide users with high-quality and interpretable similarity analysis results. Specifically:
[0257] 1. Efficient multi-document processing and deep semantic understanding:
[0258] By innovatively applying knowledge vector library technology, this system achieves efficient multi-document semantic representation and fast similarity calculation. Combined with large language models and multi-level summarization technology, the system can process multiple documents at the same time and achieve deep semantic understanding across documents. This not only greatly improves the processing speed of multi-document similarity comparison, but also significantly improves the accuracy of the comparison. Especially for large-scale document sets, the system can quickly identify content with the same meaning using different expressions, effectively responding to the needs of multi-document processing in the era of information explosion.
[0259] 2. Intelligent professional document analysis and directional comparison:
[0260] This system integrates an intelligent key element recognition mechanism based on a knowledge vector library and a large language model, which can automatically identify and align common key elements in multiple professional documents. This enables the system to achieve more accurate multi-document directional comparison when processing structured professional documents (such as policy documents, legal documents, and scientific papers). By enhancing the semantic representation of key elements through the knowledge vector library, the system has greatly improved the processing capabilities and comparison accuracy of multiple documents in professional fields, providing powerful tool support for document management and analysis in various industries.
[0261] 3. Innovation evaluation and multi-dimensional analysis:
[0262] This system innovatively combines the knowledge vector library with a large language model to develop a multi-document innovation scoring mechanism. This enables the system to effectively identify and quantify the innovative content in the document set while performing multi-document similarity comparisons. Combined with the rapid retrieval capability and hierarchical processing strategy of the knowledge vector library, the system significantly improves the efficiency of multi-document innovation evaluation while ensuring the quality of analysis. This function has important application value for scientific research institutions, patent offices, and innovative enterprises, and can greatly improve the efficiency of innovation management and evaluation.
[0263] 4. High interpretability and wide adaptability:
[0264] The detailed multi-document analysis report generated by the system not only provides an overall similarity assessment, but also includes a specific comparison of key elements between multiple documents and a human-understandable explanation. This multi-dimensional, visual analysis result greatly improves the credibility and practicality of the multi-document comparison results. At the same time, the modular design and adaptive capabilities based on the knowledge vector library enable the system to be flexibly applied to different professional fields, and can adapt to various types of multi-document comparison requirements without a lot of customization. The system's knowledge vector library can also be continuously updated and expanded, enabling it to continuously adapt to new document types and domain knowledge, with extremely strong scalability and broad application prospects.
[0265] 5. Resource optimization and sustainable development:
[0266] Through the innovative application of the knowledge vector library, this system achieves efficient storage and retrieval of document semantic information. This method not only greatly reduces repeated calculations, but also significantly reduces the system's computing resource requirements. Compared with traditional methods, this system can achieve higher performance with lower energy consumption when processing large-scale multi-document sets, which is in line with the concept of green computing. At the same time, the sustainable update mechanism of the knowledge vector library ensures that the system can continuously adapt to new language patterns and knowledge structures, providing a solid foundation for the long-term development and continuous optimization of the system.
[0267] These beneficial effects fully reflect the innovative value of the present invention in the field of multi-document similarity comparison and knowledge management. By combining the knowledge vector library and large language model technology, this system not only solves the efficiency and accuracy problems faced by traditional methods in processing multiple documents, but also provides a new technical path for the intelligent management, analysis and innovative evaluation of large-scale documents, which has important theoretical significance and wide practical application value.
[0268] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-document semantic similarity comparison system based on a knowledge vector library, including the construction of a knowledge vector library and document similarity comparison, characterized in that: The construction of the knowledge vector library includes: A document preprocessing module (101), the document preprocessing module (101) is used to receive input documents, perform format unification processing, and pass the processed documents to subsequent modules; A key element identification module (102), wherein the key element identification module (102) identifies key elements through two submodules, namely, a large language model analysis submodule (201) and an expert knowledge base submodule (202), and fuses the output results of the two submodules to generate a final key element list; A multi-level summary generation module (103), wherein the multi-level summary generation module (103) generates a parent summary and a child summary of a document through multiple generators, and forms a multi-level semantic representation of the document; A knowledge vector library generation module (4), wherein the knowledge vector library generation module (4) is used to convert the generated parent abstract and child abstract into vectors, and construct a generated knowledge vector library (401); The document similarity comparison includes a document preprocessing module (101), a key element identification module (102), a multi-level summary generation module (103) and a knowledge vector library generation module (4), and also includes: A vectorized comparison module (104), wherein the vectorized comparison module (104) performs similarity calculation between the abstract generated by the document similarity comparison unit and the pre-built knowledge vector library (401), and selects N knowledge fragments with the highest similarity for each abstract; A large language model analysis module (105), wherein the large language model analysis module (105) performs similarity analysis on the parent abstract and the child abstract as well as the corresponding TopN knowledge fragments through multiple parallel analysis processes, and generates a similarity analysis result; A result comprehensive reporting module (106), wherein the result comprehensive reporting module (106) receives all the results from the large language model analysis module (105) and generates a comprehensive report.
2. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 1 is characterized in that: The document preprocessing module (101) performs format unification processing including text extraction, encoding conversion, and special character removal operations.
3. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 1 is characterized in that: The multi-level summary generation module (103) includes a parent summary generator (301) and a child summary generator (302). The parent summary generator (301) generates an overall summary including the main content and key elements of the document based on the entire document content using a large language model. The child summary generator (302) is used to generate a corresponding local summary for each paragraph or chapter of the document. The two-level summaries of the parent summary generator (301) and the child summary generator (302) together constitute a multi-level semantic representation of the document.
4. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 3 is characterized in that: The knowledge vector library generation module (4) comprises a parent summary vectorization unit (402) and a child summary vectorization unit (403), wherein the parent summary vectorization unit (402) is used to convert a parent summary into a high-dimensional vector, and the child summary vectorization unit (403) is used to convert each child summary into a high-dimensional vector.
5. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 4 is characterized in that: The vectorized comparison module (104) comprises a similarity calculation unit (404) and a TopN selection unit (405). The similarity calculation unit (404) uses a cosine similarity algorithm to calculate the similarity between the abstract vector and the knowledge base vector. The TopN selection unit (405) selects N knowledge fragments with the highest similarity for each abstract.
6. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 5 is characterized in that: The large language model analysis module (105) includes two parallel analysis processes: a. Parent summary analysis process (501): input the parent summary and its corresponding TopN knowledge fragments into the large language model for similarity analysis to generate an overall similarity analysis result; b. Sub-abstract analysis process (502): input each sub-abstract and its corresponding TopN knowledge fragment into the large language model for similarity analysis to generate local similarity analysis results; The analysis results of both the parent summary analysis process (501) and the child summary analysis process (502) include similarity scores and detailed comparisons of key elements.
7. The multi-document semantic similarity comparison system based on the knowledge vector library according to claim 6 is characterized in that: The result comprehensive report module (106) includes a data integration unit (601), a comprehensive analysis unit (602) and a report generation unit (603); The data integration unit (601) is used to summarize the analysis results of the parent abstract and all sub-abstracts; The comprehensive analysis unit (602) is used to input the integrated data into the large language model again to generate a final comprehensive analysis report; The report generation unit (603) is used to generate a detailed report including overall similarity evaluation, key element comparison, and innovation analysis based on the comprehensive analysis results.
8. A method for comparing the semantic similarity of multiple documents based on a knowledge vector library, using the system for comparing the semantic similarity of multiple documents based on a knowledge vector library as claimed in any one of claims 1 to 7, characterized in that: The steps include: Step 1: The user inputs the document to be compared into the system; Step 2: After the document is preprocessed by the document preprocessing module (101), the key element identification module (102) automatically identifies the document type and key elements, and generates a key element list; Step 3, the multi-level summary generation module (103) generates a multi-level summary, and the knowledge vector library generation module (4) constructs a knowledge vector library (401); Step 4: The vectorized comparison module (104) performs vectorized comparison on the summary generated by the document similarity comparison unit and the pre-built knowledge vector library (401), and performs in-depth analysis on the comparison result through the large language model analysis module (105); Step 5: The result comprehensive report module (106) generates a system comprehensive report and presents it to the user.
Citation Information
Cited By
Literature review generation method, electronic equipment, storage medium and program product
CN120781846A
Semantic focusing test question duplicate checking method based on large language model
CN121031565A