Innovative and automatic literature subject term extraction method based on large language model

Through the automated document subject word extraction method based on large language model, the problem of insufficient efficiency and accuracy in the analysis of large-scale document data is solved, and the rapid and accurate processing of massive document data is achieved, and the efficiency and quality of information retrieval is improved.

CN120068864APending Publication Date: 2025-05-30天津仁爱学院

Patent Information

Application Number
CN202510088121.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has insufficient efficiency and accuracy in the rapid and accurate analysis of large-scale literature data, especially in the fields of academic research, knowledge management and information retrieval, which affects the availability of information and decision-making efficiency.

Method used

An automated document subject word extraction method based on large language model (LLM) is adopted to achieve rapid and precise processing of massive document data through steps such as data collection, preprocessing, few sample learning, subject word list sorting, iterative optimization and subject word screening.

Benefits of technology

This method can efficiently process large-scale literature data, improve the accuracy and relevance of subject word extraction, lower technical thresholds, enhance the efficiency and quality of information retrieval, and promote cross-domain research and knowledge management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068864A_ABST
    Figure CN120068864A_ABST
Patent Text Reader

Abstract

The invention provides an innovative and automatic literature subject term extraction method based on a large language model, and relates to the technical field of natural language processing and information retrieval, and the method comprises the steps of S1, data collection, S2, data preprocessing, S3, few-sample learning and subject term list extraction, S4, subject term list sorting, S5, iterative optimization, and S6, subject term screening. The subject term list is processed and generated by using the large language model, compared with the use of BERTopic, large-scale literature sets can be processed and analyzed more efficiently at a time, huge data sets can be effectively managed and analyzed, subject terms generated by each literature can be stored through continuous iteration and optimization processes, and the method is suitable for large-scale literature processing and analysis. Compared with a traditional LDA method, the subject terms are more accurate. Therefore, the continuity and depth of research can be ensured, a reliable basis is provided for subsequent research, the pre-trained large language model is directly utilized to perform semantic analysis and subject term generation of literatures, and a tedious model training process is omitted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of natural language processing and information retrieval. More specifically, it particularly relates to an innovative automated method for extracting subject terms from documents based on large language models. Background Art

[0002] In the current field of informatics and related fields, the significant deficiencies of traditional subject term extraction methods in terms of efficiency and accuracy when faced with a large collection of documents; For example, Chinese Patent Publication No. CN118193730A classifies patent texts by using the International Patent Classification (IPC) number and creates a tag set for each field. Then, it analyzes the abstract part of each patent specification, extracts the theme through the LDA model for classification. Based on the classified patent texts and theme results, it further identifies the key common technical features in each patent. Then, it constructs a deep learning network and trains it with these technical features and high-quality patent texts to effectively identify and evaluate new patent texts. Although this method combines the advantages of deep learning and the theme model and can identify technical trends and innovations from complex data, it faces two major limitations in practical applications. First, the effectiveness of this method highly depends on the quality of the input data. If there are errors in the classification or theme extraction of the patent text, or if the text information is incomplete, it may seriously affect the recognition accuracy and the practicality of the model. Second, the limitation of the model's generalization ability is also an important issue. Due to the limitation of the training data, the model may not be able to effectively process new data or rare technology types that are significantly different from the training set, resulting in a decline in the recognition effect; Chinese Patent also discloses CN117725212A, which is a method for identifying technical themes of scientific and technological projects using BERTopic. Among them, BERTopic is a tool for topic modeling based on the BERT model. It converts documents into dense vectors, uses UMAP for dimensionality reduction, and then uses HDBSCAN for clustering and c-TF-ICF to identify the themes in the document set. This method relies on deep learning to understand the deep semantics of the text and thus attempts to capture more detailed topic details during the topic modeling process. However, BERTopic has several limitations when dealing with large-scale data sets. First, due to the themes after clustering, each time the BERT model is optimized, different topic structures may be obtained, and the entire process needs to be redone. When dealing with a vast amount of literature, even after clustering, c-TF-ICF needs to process a huge vocabulary. For each word, its frequency of occurrence in each category and its inverse frequency in all categories need to be calculated. This scale of data requires a large amount of memory and computing power. At the same time, many words may only appear in a few documents or categories, resulting in very sparse c-TF-ICF values calculated, which affects the quality and usability of the final subject terms; The objective of the present invention is to address the significant deficiencies in the efficiency and accuracy of traditional subject term extraction methods when faced with a vast collection of documents in the current field of informatics and related areas. In view of the limitations of existing technologies in rapidly and accurately analyzing large-scale document data, this challenge is particularly prominent in multiple fields such as academic research, knowledge management, and information retrieval, seriously affecting the usability of information and the decision-making efficiency.

[0003] To this end, the present invention proposes an innovative automated document subject term extraction method based on a large language model (LLM). This method leverages the deep semantic understanding ability of the large language model and combines the few-shot learning mechanism to achieve rapid and precise processing of massive document data. Through continuous self-iteration, this method can effectively enhance the accuracy and relevance of subject term extraction.

[0004] The present invention provides an academic and industrial community with a more efficient and accurate system for extracting document subject terms to address the challenge of extracting subject terms from massive document data. It can promote the rapid dissemination and utilization of knowledge and accelerate the process of scientific research and technological innovation. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides an innovative automated document subject term extraction method based on a large language model to solve the above problems.

[0006] An innovative automated document subject term extraction method based on a large language model includes the following steps: S1: Data collection, obtaining a large amount of document data from multiple sources, and batch obtaining relevant document data using the officially open interfaces; S2: Data preprocessing, including applying an efficient text data deduplication technique to deduplicate the document data obtained from each platform, uniformly converting the document formats from different sources into a processable text format, and removing the noise information in the documents; S3: Few-shot learning and subject term list extraction, randomly selecting several documents from the preprocessed document data for precise manual annotation, and extracting a key subject term list as few-shot learning examples for the large language model; S4: Subject term list sorting, calculating the weights of subject terms using methods such as mutual information, and applying a sorting algorithm to sort the subject terms; S5: Iterative optimization, including initial annotation, automatic update, and performance evaluation, regularly using the subject term list generated by the model and the corresponding documents as new few-shot examples to replace the old examples, and regularly evaluating the extraction effect of the model; S6: Subject term screening. Collect and summarize the subject terms of all documents, then perform vectorization, index construction, and clustering. Evaluate and extract the most important topics using indicators such as term frequency and information entropy.

[0007] Preferably, the sources of data collection include academic databases, internal literature databases of universities and research institutions, and online literature storage platforms.

[0008] Preferably, the deduplication technique in data preprocessing is to use the SHA-256 hash algorithm.

[0009] Preferably, the large language models in few-shot learning and subject term list extraction include GLM or DeepSeek.

[0010] Preferably, the sorting algorithm in subject term list sorting includes quicksort.

[0011] Preferably, the performance evaluation method in iterative optimization is to test the subject terms of manually annotated documents by automatically updating and extracting few-shot examples at fixed intervals. If the score is lower than 100 points, the few-shot examples are rectified again.

[0012] Preferably, the vectorization in subject term screening uses pre-trained vectorization functions, including Word2Vec or GloVe.

[0013] Preferably, the clustering in subject term screening uses the HNSW algorithm in the FAISS library.

[0014] Preferably, the clustering objective function in subject term screening is to minimize the distance between each vector and its nearest neighbor.

[0015] Preferably, the results in subject term screening are presented by visualization methods such as bar charts and heatmaps.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. Ability to process a large number of documents: This solution can process and generate subject term lists using large language models. Compared with using BERTopic, it can process and analyze large-scale document collections more efficiently at one time, and can effectively manage and analyze large data sets.

[0017] 2. Preservation and accuracy of intermediate results: Through continuous iteration and optimization in the solution, the subject terms generated for each document can be preserved. These subject terms are more accurate than traditional LDA methods. This not only helps to ensure the coherence and depth of research, but also provides a reliable basis for subsequent research.

[0018] 3. Rely on pre-trained semantic understanding models without additional training: This solution directly uses pre-trained large language models for semantic parsing and subject term generation of literature, eliminating the cumbersome model training process. This method not only lowers the technical threshold but also ensures the use of off-the-shelf high-quality models, avoiding potential data quality issues when training models independently.

[0019] 4. Improve the efficiency and quality of information retrieval: This method can accurately extract the subject terms of literature, enabling users to find the required literature more accurately and quickly during information retrieval. Through accurate subject term annotation, the retrieval system can better understand the user's needs and provide more relevant and targeted retrieval results. This not only saves the user's time but also increases the probability of obtaining effective information, helping users quickly locate valuable content in a vast amount of literature and promoting the effective dissemination and utilization of knowledge.

[0020] 5. Facilitate cross-disciplinary research and knowledge integration: By extracting and analyzing subject terms from a large number of literature from different fields and sources, potential associations and intersections between different fields can be discovered. This helps break down disciplinary barriers and promotes cross-disciplinary research cooperation and knowledge integration. Researchers can more clearly see the similarities and differences in research topics between different fields, thus inspiring innovative thinking and driving the generation of comprehensive and innovative research results.

[0021] 6. Enhance knowledge management and decision support capabilities: For academic institutions, enterprises, government departments, etc., accurately extracting the subject terms of literature helps establish a more efficient knowledge management system. It can systematically organize and classify internal and external literature resources, facilitating the storage, retrieval, and sharing of knowledge. At the same time, based on the analysis of the subject terms of a large number of literature, it can provide strong support for decision-making, helping decision-makers more comprehensively understand the research status and trends in related fields, and thus making more informed and forward-looking decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the flowchart of the data collection part in the present invention; Figure 2 is the flowchart of the LLM extracting the subject term list from the literature in the present invention; Figure 3 is the flowchart of extracting subject terms from the set of subject terms extracted from all the literature in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and examples. The following examples are used to illustrate the present invention but cannot be used to limit the scope of the present invention.

[0024] Please refer to Figures 1 - 3, the present invention provides an innovative automated method for extracting subject terms from literature based on large language models (LLMs). The purpose of the present invention is to address the significant deficiencies in efficiency and accuracy of traditional subject term extraction methods when faced with a large corpus of literature in current informatics and related fields. In view of the limitations of existing technologies in rapidly and accurately analyzing large-scale literature data, this challenge is particularly prominent in multiple fields such as academic research, knowledge management, and information retrieval, seriously affecting the usability of information and decision-making efficiency.

[0025] S1 Data collection: Obtain a large amount of literature data from multiple sources, including but not limited to academic databases (such as Google Scholar, etc.), internal literature databases of universities and research institutions, and online literature storage platforms (such as arXiv, etc.). Use the officially open interfaces to batch obtain relevant literature data to ensure the diversity and comprehensiveness of the data.

[0026] S2 Data preprocessing: 1. Duplicate removal: Apply an efficient text data duplicate removal technique (such as using the SHA-256 hash algorithm) to remove duplicate literature records from the literature data obtained from each platform to prevent duplicate data from affecting the effect of subsequent processing steps; 2. Literature format unification: Unify the literature formats from different sources (such as PDF, HTML, TXT, etc.) into a processable text format; 3. Literature cleaning: Remove the noise information in the literature, including advertisements, copyright statements, and references, etc., and use custom regular expressions and text cleaning algorithms to ensure obtaining clean text data.

[0027] S3 Few-shot learning and subject term list extraction: Randomly select several pieces of literature from the preprocessed literature data for precise manual annotation, and extract a key subject term list as few-shot learning examples for large language models (such as GLM or DeepSeek, etc.), enabling the LLM to quickly learn based on a small number of samples and possess the ability to extract subject term lists.

[0028] S4 Subject term list sorting: In order to improve the relevance and accuracy of subject terms, it is also necessary to sort the initially generated subject term list. And use methods such as Mutual Information that can be used to measure the correlation between words and document topics to calculate the weights of subject terms; The mutual information calculation formula is: ; where p(x,y) represents the probability that word x and document topic y appear simultaneously, and p(x) and p(y) represent the marginal probabilities of word x and document topic y respectively; Sort the subject terms using the calculated weights and apply a sorting algorithm (such as but not limited to quicksort, etc.) to select the most relevant subject term.

[0029] S5 Iterative Optimization: Continuously improve the quality and relevance of subject term extraction through model self-optimization and regular performance evaluation. The specific steps include: First step: Initial annotation: Use manually annotated documents and subject terms as few-shot examples; Second step: Automatic update: Regularly use the subject term list generated by the model and the corresponding documents as new few-shot examples to replace the old examples, continuously improving the model's extraction ability; Third step: Performance evaluation: Regularly evaluate the extraction effect of the model to ensure that the introduction of new few-shot examples can improve the accuracy and relevance of subject term extraction.

[0030] Among them, the performance evaluation part can be carried out in the following way: Test the subject terms of the manually annotated documents with the few-shot examples extracted by automatic update at fixed intervals. If the score is lower than 100 points, rectify the few-shot examples again. The score formula is as follows: ; Where is the number of correct subject terms output by the large model, is the number of the total test set.

[0031] S6 Subject Term Screening: The subject terms of all documents will be collected and summarized, and further screening and clustering will be carried out. To achieve efficient processing, approximate algorithms such as FAISS (an efficient similarity search library) are mainly used for fast approximate nearest neighbor clustering. This method not only improves the processing speed but also reduces resource consumption. The main steps include: Vectorization: Each subject term w is converted into a vector vw = V(w), where V is a pre-trained vectorization function such as Word2Vec or GloVe; Index construction and clustering: Add all subject term vectors to the index established by FAISS. FAISS accelerates the similarity search between vectors by establishing an efficient index structure, making the clustering process faster. The goal of clustering is to group subject terms with high similarity together to form several clear subject categories; Clustering objective function: During the clustering process, we can determine the nearest neighbor of each subject term vector through nearest neighbor search, and then cluster based on this information. The clustering objective function can be described as minimizing the distance between each vector and its nearest neighbor: ; where \(C_i\) is the set of vectors in the \(i\)-th cluster, and \(\) is the central vector of \(C_i\); Result sorting and output: Metrics such as term frequency and information entropy will be used to evaluate and extract the most important topics. The term frequency \(f\) is defined as: ; where, is the number of occurrences of the topic word in the cluster.

[0032] In addition, the information entropy \(H(C)\) of each cluster will be calculated to evaluate the purity of the cluster: ; where, is the probability of occurrence of any topic word in the \(i\)-th cluster; The final results can be presented in various ways, including visualization methods such as bar charts and heatmaps.

[0033] Working principle: Step 1: Download relevant literature from a literature website that can perform API access as the main data source. Use the hashlib library in Python to implement the SHA-256 hashing algorithm to generate a unique hash value for each piece of literature and remove duplicate literature records. And remove non-core content such as advertisements and copyright statements through regular expressions, and at the same time use a custom text cleaning function to remove noise in the literature; Step 2: Randomly select 5 pieces of the collected literature for detailed manual annotation, including all possible topic words of the literature, and make them into few-shot samples to provide knowledge for the large language model to be able to understand semantics and generate a list of topic words; Use these annotated literatures as examples and input them together with the unannotated literatures into the large language model to automatically generate a list of topic words for the remaining literatures. And sort the list of topic words using algorithms such as mutual information to obtain the 3 most relevant topic words. And through setting the number of iterations, such as replacing one example in the few-shot samples with the data generated by the model every 20 iterations, to achieve continuous optimization of the model; Step 3: Convert the subject terms extracted from the literature into vector form using a pre-trained vectorization model, such as Word2Vec, and import all the subject term vectors into the FAISS library. Use the HNSW (Hierarchical Navigable Small World) algorithm in the FAISS library for fast similarity search. For example, set k = 100, that is, retrieve 100 most similar subject term vectors each time. Perform mean processing on these retrieved vectors to obtain a new central vector, representing the clustering center of these 100 vectors. This method effectively reduces the large number of subject terms by 100 times, reducing the computational amount for the next clustering process; Step 4: After processing by the HNSW algorithm, only perform clustering analysis on the new vectors obtained after mean processing, such as using the fusion of word frequency and information entropy. This step will obtain an integrated list of subject terms.

[0034] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the present invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described to better explain the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and design various embodiments with various modifications suitable for a particular purpose.

Claims

1. An innovative automatic document subject word extraction method based on a large language model, characterized by: The following steps are involved: S1: Data collection; S2: data preprocessing; S3: Few-shot learning and keyword list extraction; S4: Ranking of subject heading list; S5: Iterative optimization; S6: Subject word screening.

2. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The sources of data collection include academic databases, internal literature databases of universities and research institutions, and online literature storage platforms.

3. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The deduplication technology in the data preprocessing is to use the SHA-256 hash algorithm.

4. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: Large language models for few-shot learning and word list extraction include GLM or DeepSeek.

5. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The sorting algorithms used in sorting keyword lists include quick sort.

6. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The performance evaluation method in iterative optimization is to test the keywords of manually annotated documents by automatically updating the extracted few-shot examples at fixed intervals. If the score is lower than 100 points, the few-shot examples are revised.

7. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The vectorization in topic word screening uses pre-trained vectorization functions, including Word2Vec or GloVe.

8. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The clustering in keyword screening uses the HNSW algorithm in the FAISS library.

9. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The clustering objective function in keyword screening is to minimize the distance between each vector and its nearest neighbor.

10. The innovative automatic document subject word extraction method based on a large language model as claimed in claim 1, characterized in that: The results of the subject term screening are presented through visualization methods such as bar charts and heat maps.

Citation Information

Patent Citations

  • Science and technology project technology theme identification method based on BERTopic theme identification model

    CN117725212A

  • Technical patent identification method based on deep learning and topic model

    CN118193730A

Cited By

  • Theme recognition method and system for large-scale text data and readable medium

    CN120745647A

  • A method, system and readable medium for topic identification of large scale text data

    CN120745647B