Method and system for analyzing and predicting theme trend of scientific and technical literature
Through pre-trained semantic analysis models and topic modeling technology, deep semantic analysis and topic prediction of scientific and technological literature are solved, and the problem of difficult to understand deep semantics and predict future trends in the existing technology is solved, and a more accurate and forward-looking analysis of scientific and technological literature is achieved.
Patent Information
- Application Number
- CN202510550556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing technology is difficult to accurately understand the deep semantics of scientific and technological literature texts, identify key areas of attention in literature, and lack of effective prediction models for future literature evolution trends, resulting in limited support for government scientific and technological literature research.
The pre-trained semantic analysis model is used to encode the text data of scientific and technological literature, generate document semantic vectors, and construct the theme hierarchical structure through topic clustering and hierarchical clustering. At the same time, through time window division and dynamic topic modeling, the changing characteristics of the topic are analyzed, and the intensity change trend of the topic is predicted based on these characteristics.
It has achieved an accurate understanding of the deep semantics of scientific and technological literature texts, identified key areas of concern to the literature, and provided effective predictions of future evolution trends of literature, enhancing the support ability of scientific and technological literature research to government scientific and technological decision-making.
Smart Images

Figure CN120068882A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text mining and document analysis, and in particular to a method and system for analyzing and predicting the subject trend of scientific and technological documents. Background Art
[0002] In recent years, scientific and technological literature research has ushered in new development opportunities, but at the same time it faces the challenges of explosive growth in scientific and technological literature information and the iteration of research methods.
[0003] Traditional scientific and technological literature analysis mainly relies on manual reading, summarization and qualitative analysis methods, such as content analysis and comparative research. Although these methods can provide a deep understanding of literature texts, they are inefficient when faced with large-scale literature documents, and it is difficult to objectively quantify the evolution of literature. In recent years, text mining methods based on topic models such as LDA (latent Dirichlet allocation) have been introduced into the field of literature analysis, which can automatically identify topics and cluster them from a large number of literature texts.
[0004] The most relevant technology currently is the literature topic analysis method based on text mining. This method uses a specific algorithm to preprocess, extract features and model topics of literature texts, thereby revealing the distribution and evolution of literature topics. Its technical principle is to use computer linguistics and statistical methods to transform unstructured literature texts into structured data that can be quantified and analyzed, and then mine the semantic information and topic patterns contained in the text.
[0005] However, existing technologies have obvious limitations. First, the content of documents is heterogeneous and fragmented, and traditional topic models are difficult to accurately understand the deep semantics of document texts; second, existing methods do not clearly define the priority of cutting-edge scientific and technological development, making it difficult to accurately identify the key areas of focus of documents; in addition, there is a lack of effective prediction models for the future evolution trend of documents, and it is impossible to provide forward-looking guidance for scientific and technological decision-making. These problems seriously restrict the ability of scientific and technological literature research to support government scientific and technological decision-making. Summary of the invention
[0006] In view of this, the present application provides a method and system for analyzing and predicting document topic trends, which solves the problems in the prior art that traditional topic models are difficult to accurately understand the deep semantics of document texts, difficult to accurately identify key areas of document focus, and lack an effective model for predicting future document evolution trends.
[0007] The present application embodiment provides a method for analyzing and predicting the trend of scientific and technological literature topics, including: Collecting scientific and technological literature text data, and preprocessing the scientific and technological literature text data to obtain target text data; Encoding the target text data using a pre-trained semantic analysis model to obtain a document semantic vector; Perform topic clustering on the target text data according to the document semantic vector, and extract multiple topics and the topic representations of each topic; Perform hierarchical clustering on the topics according to the topic representations to form a topic hierarchical structure; Divide the target text data into multiple time windows, analyze the change characteristics of the topics in each time window, and construct a topic evolution time series according to the change characteristics, where the change characteristics include occurrence frequency and intensity change; Predict the intensity change trend of the topic according to the topic evolution time series.
[0008] Encoding the target text data using a pre-trained semantic analysis model to obtain a document semantic vector, including: Perform segmentation processing on the target text data to obtain text segments of a specified length; Use the semantic analysis model to obtain the semantic representations of the text segments; Integrate the semantic representations of the text segments into a document semantic vector, and the document semantic vector is used to indicate the semantic information and context relationship of the target text data.
[0009] The performing topic clustering on the target text data according to the document semantic vector, and extracting multiple topics and the topic representations of each topic, includes: Perform dimensionality reduction processing on the document semantic vector using a preset dimensionality reduction model to obtain a low-dimensional semantic vector; Based on the low-dimensional semantic vector, use a preset clustering model to perform density clustering on the target text data to form multiple topics; Calculate the word feature values of the words in each topic, and extract keywords according to the word feature values, where the word feature values include word frequency and inverse document frequency; Optimize the keywords using a preset maximum similarity matching algorithm to obtain the topic representation of the topic.
[0010] The performing hierarchical clustering on the topics according to the topic representations to form a topic hierarchical structure, includes: Calculate semantic similarity based on the topic representations, and construct a similarity matrix between the topics according to the topic representations and the semantic similarity; Use a preset hierarchical clustering model to perform hierarchical clustering on the topics to obtain a hierarchical system represented as a tree structure, and the hierarchical system includes multiple first-level topics and multiple second-level topics corresponding to the first-level topics; Optimize and adjust the hierarchical system to form a topic hierarchical structure.
[0011] Dividing the target text data into multiple time windows, analyzing the change characteristics of the theme in each time window, and constructing a theme evolution time series based on the change characteristics, includes: Dividing the target text data in terms of time dimension according to the publication time of the target text data to obtain a plurality of consecutive time windows; Using a preset theme mining model to mine the theme of the target text data within the time window; According to the theme mining result, calculating the theme characteristics of each theme in different time windows, where the theme characteristics include appearance frequency and intensity; Using a preset dynamic theme model to analyze the theme characteristics to obtain the change characteristics of the theme, where the change characteristics are used to indicate the semantic change of the theme over time; Generating the theme evolution time series according to the change characteristics.
[0012] Preprocessing the scientific and technological literature text data to obtain target text data, includes: Using a preset large language model to evaluate the scientific and technological literature text data to obtain a complexity score, where the complexity score is used to indicate the professionalism, intersection, and structural complexity of the scientific and technological literature text data; Classifying the scientific and technological literature text data according to the complexity score to obtain a first type of text and a second type of text, where the first type of text is used to indicate that the complexity score is lower than a preset score threshold, and the second type of text is used to indicate that the complexity score is higher than the score threshold; Performing lightweight analysis on the first type of text and performing in-depth analysis on the second type of text to obtain a preliminary processing result; Obtaining the target text data according to the preliminary processing result.
[0013] Using a preset large language model to evaluate the scientific and technological literature text data to obtain a complexity score, includes: Collecting training texts with complexity annotations, fine-tuning and training a preset initial large language model to obtain a large language model; Using the large language model to calculate the analysis characteristics of the scientific and technological literature text data, where the analysis characteristics include professional vocabulary density, domain coverage, and logical hierarchy depth; Calculating a text complexity score according to the analysis characteristics and generating an evaluation basis description according to the text complexity score; Obtaining the complexity score of the scientific and technological literature text data according to the evaluation basis description.
[0014] Performing topic clustering on the target text data according to the document semantic vector, extracting multiple topics and the topic representations of each topic, further comprising: Based on the document semantic vector, establishing a scoring index system including literature level classification, publication time, and citation frequency; According to the document semantic vector and the scoring index system, constructing a cost-sensitive decision tree, which is used to indicate increasing the weight assigned to samples whose importance reaches a preset importance threshold; Extracting classification rules from the cost-sensitive decision tree, where the classification rules include feature thresholds and classification paths; Optimizing the initial center point distribution of topic clustering according to the classification rules, and performing topic clustering on the target text data according to the optimization result.
[0015] After predicting the intensity change trend of the topic according to the topic evolution time series, further comprising: Based on the topic intensity change trend, constructing a knowledge graph including literature entities, influence paths, and effect quantification indicators; Using the knowledge graph to establish a multi-level analysis framework; According to the multi-level analysis framework, establishing a retrieval and reasoning system, where the retrieval and reasoning system includes an inference chain from literature to influence; Adjusting the effect quantification indicators, and performing multi-scenario prediction by combining the adjusted effect quantification indicators and the retrieval and reasoning system to obtain development trends under different conditions; Among them, the inference chain is established through the following steps, including: Based on the multi-level analysis framework, decomposing the influence analysis task into multiple subtasks with logical dependencies; Using the large language model to process the subtasks to obtain a progressive inference chain; Retrieving associated data related to each node of the progressive inference chain in the knowledge graph; Combining the associated data and the inference process indicated by the large language model to verify and correct the inference result, and generating an inference chain.
[0016] An embodiment of the present application further provides a device for analyzing and predicting the topic trend of scientific and technological literature, including: A data preprocessing module, configured to collect scientific and technological literature text data, and preprocess the scientific and technological literature text data to obtain target text data; A semantic encoding module, configured to encode the target text data using a pre-trained semantic analysis model to obtain a document semantic vector; A topic clustering module, configured to perform topic clustering on the target text data according to the document semantic vectors, and extract multiple topics and the topic representations of each topic; A hierarchical analysis module, configured to perform hierarchical clustering on the topics according to the topic representations to form a topic hierarchical structure; A time evolution analysis module, configured to divide the target text data into multiple time windows, analyze the change characteristics of the topics in each time window, and construct a topic evolution time series according to the change characteristics, where the change characteristics include occurrence frequency and intensity change; A trend prediction module, configured to predict the intensity change trend of the topics according to the topic evolution time series.
[0017] An embodiment of the present application further provides a computer device, where the computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for analyzing and predicting the trends of scientific and technological literature topics described above.
[0018] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method for analyzing and predicting the trends of scientific and technological literature topics described above.
[0019] An embodiment of the present application further provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of the method for analyzing and predicting the trends of scientific and technological literature topics described above are implemented.
[0020] The present application has the following technical effects: Innovatively combines the BERT semantic model with the topic model to construct a BERTopic scientific and technological literature analysis framework, which can capture the deep semantic relationships of literature texts more accurately than traditional topic models such as LDA; designs a multi-level topic classification system, organizes the topics through a hierarchical clustering algorithm, and realizes an all-round analysis from the macroscopic literature orientation to the microscopic literature measures; proposes a method for dynamically analyzing the evolution of literature topics, and realizes a quantitative analysis of the evolution law of literature topics through time window division and dynamic topic modeling; develops a literature trend prediction model based on the topic evolution pattern, combines time series analysis and deep learning methods, and provides a forward-looking reference for future literature orientation. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required for the embodiments. The drawings herein are incorporated into the specification and form a part of this specification. These drawings show the embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic flowchart of the method for analyzing and predicting the theme trend of scientific and technological literature provided by the embodiments of the present application; Figure 2 It is a schematic flowchart of the implementation of the text representation module based on BERT semantic embedding provided by the embodiments of the present application; Figure 3 It is a schematic flowchart of the implementation of the scientific and technological literature theme mining module based on BERTopic provided by the embodiments of the present application; Figure 4 It is a schematic flowchart of the implementation of the literature theme hierarchical analysis module provided by the embodiments of the present application; Figure 5 It is a schematic flowchart of the implementation of the literature theme time evolution analysis module provided by the embodiments of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all of them. Usually, the components of the embodiments of the present disclosure described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure to be protected, but only represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0024] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0025] In this text, the term "and / or" only describes an association relationship, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the term "at least one" in this text means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0026] As Figure 1 shown, an embodiment of this application provides a method for analyzing and predicting the theme trend of scientific and technological literature, including: S1: Collect scientific and technological literature text data, preprocess the scientific and technological literature text data, and obtain target text data; The system first collects scientific and technological literature text data from various channels such as national and local government websites and official websites of science and technology departments through web crawler technology. The collected data includes core information such as literature titles, publishing institutions, release times, and full text of the literature, ensuring the authority and comprehensiveness of the data source. After collection, the system constructs a structured literature text database to provide basic data support for subsequent analysis. Next, the system conducts a comprehensive cleaning process on the collected literature text, including removing special characters, punctuation marks, and stop words in the text, unifying the text format, and performing Chinese word segmentation using professional word segmentation tools. At the same time, the system also extracts metadata information of the literature text, such as key attributes like release time, department, and region, which will play an important role in subsequent time series analysis and regional comparison. Finally, the preprocessed text data is standardized to ensure data quality and consistency, laying a solid foundation for subsequent semantic analysis. The quality of the preprocessing link directly affects the accuracy of subsequent analysis. Therefore, the system adopts strict data quality control measures in this link to ensure the high quality of the input data.
[0027] S1 specifically includes: S1.1: Use a preset large language model to evaluate the scientific and technological literature text data, and obtain a complexity score, where the complexity score is used to indicate the professionalism, cross-disciplinarity, and structural complexity of the scientific and technological literature text data; As the basis of the intelligent data filtering mechanism, it aims to comprehensively quantify the complexity of literature texts. The system designs a multi-dimensional complexity evaluation system, focusing on three core dimensions: professionalism, cross-disciplinarity, and structural complexity. Professionalism evaluates the depth of the use of professional terms and scientific concepts in the literature. Literature with high professionalism usually contains a large number of domain-specific terms and esoteric theoretical content; cross-disciplinarity evaluates the breadth of the literature involving multiple disciplines or technical fields. Literature with high cross-disciplinarity often integrates the knowledge and methods of multiple disciplines; structural complexity evaluates the complexity of the organizational structure, logical relationships, and expression methods of the literature. Literature with high structural complexity may contain multi-layer nested clauses, complex logical dependencies, or non-linear content organization. Through the in-depth semantic understanding of the full text of the literature, the large language model can accurately capture these complexity features and generate a comprehensive complexity score. This deep learning-based evaluation method can more accurately understand the inherent complexity of the text compared to traditional keyword- or rule-based methods, providing a more reliable decision-making basis for subsequent processing.
[0028] S1.1.1: Collect training texts with complexity annotations and fine-tune the pre-set initial large language model to obtain a large language model; The construction of the training dataset is a key link in this step. The system collects a large number of scientific and technological literature samples and invites domain experts to annotate the complexity of these samples. The annotation process considers multiple dimensions, such as the density of professional terms, the depth of concepts, the complexity of the logical structure, and the degree of cross-disciplinarity in the literature. These annotated literature samples are used to fine-tune pre-trained large language models, such as BERT, GPT, etc. The fine-tuning process uses supervised learning methods to optimize the model parameters by minimizing the difference between the predicted complexity score and the expert annotation score. At the same time, to enhance the generalization ability of the model, the training data covers scientific and technological literature from different periods, different fields, and different types, ensuring that the model can adapt to the complexity evaluation needs of various literature texts. After sufficient training and verification, the system obtains a high-performance large language model dedicated to the complexity evaluation of scientific and technological literature, providing a powerful tool for subsequent complexity scoring.
[0029] In the fine-tuning training process of the large language model, the following technical solutions are specifically adopted: Training data construction: Select 10,000 literatures with different complexities from the public scientific and technological literature library, and have 5 domain experts annotate the complexity according to a 1-10 scale to form a training set with a Kappa coefficient of 0.85 for annotation consistency.
[0030] Model Architecture: Based on a pre-trained language model with a 12-layer Transformer structure, it contains 12 attention heads, the hidden layer dimension is 768, and the total number of parameters is approximately 110 million. To adapt to the characteristics of scientific and technological literature, 3,000 professional terms are added to the vocabulary.
[0031] Training Parameter Settings: The Adam optimizer is used, the learning rate is set to 3e-5, a linear learning rate warm-up and decay strategy is used, the batch size is 16, and 4 epochs are trained. The loss function uses a combination of mean squared error (MSE) and cross-entropy to optimize both classification accuracy and scoring precision simultaneously.
[0032] Validation Method: 5-fold cross-validation is used, and mean absolute error (MAE) and accuracy are used as evaluation metrics, achieving an MAE of 0.72 and a complexity classification accuracy of 85% on the validation set.
[0033] S1.1.2: Use the large language model to calculate the analysis features of the scientific and technological literature text data, and the analysis features include professional vocabulary density, domain coverage, and logical hierarchy depth; These features are the basis for complexity scoring, including key indicators such as professional vocabulary density, domain coverage, and logical hierarchy depth. Professional vocabulary density is calculated by identifying the proportion of professional terms, scientific and technological concepts, and academic vocabulary in the text, reflecting the professional depth of the literature. The system utilizes the powerful semantic understanding ability of the large language model to accurately identify professional terms in various fields and can effectively grasp emerging scientific and technological concepts. Domain coverage measures the breadth of the disciplines or technical fields covered by the literature. The system determines the interdisciplinary characteristics of the literature by analyzing the frequency and distribution of different domain concepts in the text. For highly cross-integrated literature, the system will give a higher complexity assessment. Logical hierarchy depth is an important indicator for evaluating the structural complexity of the literature. The system analyzes features such as the hierarchical structure, clause relationships, and logical dependencies in the literature to quantify the organizational complexity of the literature. In addition, the system also analyzes auxiliary features such as the syntactic complexity, semantic density, and reasoning depth of the literature to comprehensively capture the complex characteristics of the literature. These analysis features together constitute a multi-dimensional feature space, providing a rich information basis for accurately evaluating the complexity of the literature.
[0034] S1.1.3: Calculate the text complexity score based on the analysis features, and generate an evaluation basis description based on the text complexity score; The calculation process adopts a weighted comprehensive scoring method, assigns weights according to the importance of different features, and obtains the final complexity score. The setting of weights is based on the experience of domain experts and the statistical results of a large number of literature analyses, which can reflect the contribution degree of different features to complexity. To improve the interpretability of the scoring, the system also generates a detailed description of the evaluation basis, including the score of each dimension, key influencing factors, and typical feature examples. For example, for a literature rated as highly complex, the system may give the following description: "This literature contains a high density of artificial intelligence professional terms (35% of the vocabulary is professional terms), and at the same time involves the cross-content of three fields: computer science, statistics, and cognitive science. The literature structure presents a multi-layer nested relationship, and there are complex logical dependencies between clauses." Such a detailed evaluation basis helps users understand the source and meaning of the complexity score, enhancing the credibility and practicality of the evaluation results.
[0035] S1.1.4: Obtain the complexity score of the scientific and technological literature text data according to the description of the evaluation basis.
[0036] Taking into account the above analysis results, according to the preset scoring criteria, the literature complexity is quantified into specific scores, usually using a score range of 1-10 or a grading method of "low-medium-high". The complexity score not only reflects the overall complexity of the literature but also includes the refined evaluation results of each dimension, providing precise guidance for subsequent shunt processing. The system will classify the literature into two categories: simple literature and complex literature according to the scoring results. Simple literature usually has a low complexity score, straightforward content, and clear structure, and is suitable for lightweight processing; while complex literature has a high complexity score and may contain profound professional content, complex logical structures, or cross-field integrated knowledge, which requires more in-depth analysis. This complexity-based classification provides a scientific basis for the optimal allocation of system resources, enabling the system to concentrate more computing resources and analysis capabilities on those complex literatures that truly require in-depth processing, while quickly processing relatively simple literatures, improving the overall processing efficiency and analysis quality.
[0037] Through the above implementation, the system has completed the complexity evaluation of the scientific and technological literature text data, laying a foundation for subsequent strategic shunt processing. This complexity evaluation method based on large language models has higher accuracy and adaptability compared to traditional rule-based or simple feature-based methods, and can effectively process various types and fields of scientific and technological literature. At the same time, by generating a detailed description of the evaluation basis, the system enhances the interpretability and credibility of the complexity evaluation, enabling users to understand and verify the evaluation results. This high-quality complexity evaluation provides a reliable guarantee for subsequent resource optimization allocation and processing strategy selection, effectively improving the efficiency and quality of the entire literature analysis system.
[0038] S1.2: Classify the scientific and technological literature text data according to the complexity score to obtain a first type of text and a second type of text, where the first type of text is used to indicate that the complexity score is lower than a preset score threshold, and the second type of text is used to indicate that the complexity score is higher than the score threshold; The system sets a preset score threshold (which can be dynamically adjusted according to the system processing capacity and actual application requirements). Literature with a complexity score lower than the threshold is classified as the first type of text. This type of literature usually has clear semantics, a clear structure, and a moderate degree of professionalism, and is suitable for lightweight processing. Literature with a complexity score higher than the threshold is classified as the second type of text. This type of literature may contain profound professional content, complex structural organizations, or cross-disciplinary knowledge integration, and requires more in-depth analysis. In actual implementation, considering the multi-dimensional characteristics of the complexity score, the system may adopt a multi-threshold classification method, that is, set thresholds according to different dimensions of complexity. Only when the literature is lower than the corresponding thresholds in all key dimensions will it be classified as the first type of text. This detailed classification strategy ensures the effective identification of complex literature and avoids misjudgments that may be caused by simple classification. At the same time, the system records the complexity details and classification basis of each piece of literature, providing a reference for subsequent processing and facilitating the continuous optimization of the classification strategy.
[0039] S1.3: Perform lightweight analysis on the first type of text and in-depth analysis on the second type of text to obtain preliminary processing results; It is the core implementation link of the intelligent data filtering mechanism. Through a differential processing strategy, it realizes the optimal allocation of computing resources and the improvement of processing efficiency. For the first type of text (low complexity), the system adopts lightweight processing methods, such as basic text cleaning, simple semantic analysis, and conventional feature extraction. These methods have a small computational burden and a fast processing speed, and are suitable for batch processing of relatively simple literature texts. Specifically, the system may adopt a simplified language model, a shallow neural network, or an efficient statistical analysis method to quickly extract the core content and topic features of the text. For the second type of text (high complexity), the system starts the in-depth analysis mode and applies more powerful analysis tools and algorithms, such as a large language model with full parameters, deep semantic parsing, and complex logical structure analysis. The in-depth analysis mode invests more computing resources and processing time, and can more accurately understand the connotation and structure of complex literature, and capture subtle semantic features and logical relationships. The system may also apply specialized analysis strategies for different types of complex literature (such as interdisciplinary literature, high-professional literature, complex structure literature) to further improve the pertinence and efficiency of processing. Through this differential processing, the system significantly improves the overall processing efficiency while ensuring the analysis quality, and realizes the reasonable allocation of computing resources.
[0040] S1.4: Obtain the target text data according to the preliminary processing results.
[0041] Complete the preprocessing process of scientific and technological literature text data. Integrate and standardize the results of lightweight analysis and in-depth analysis to ensure that subsequent analysis can receive input data with a unified format and consistent quality. First, the system conducts a quality check on the processing results of the two types of texts, evaluating the integrity of key information extraction, the accuracy of semantic understanding, and the rationality of structural analysis to ensure that the processing results meet the predetermined quality standards. If it is found that the processing results of some literature do not meet the requirements, the system will mark them as needing to be reprocessed, and may adjust the processing strategy or parameters, or convert the text originally classified as the first type into the second type for more in-depth analysis. Second, the system standardizes the processing results that meet the quality standards, including operations such as unified format, feature regularization, and metadata supplementation, to ensure the consistency of the data form. Finally, the system organizes the standardized data into a target text data set, which contains various information such as text content, structural information, semantic features, and complexity attributes, providing high-quality input for subsequent semantic encoding and topic clustering. Through this series of processes, the system not only completes the basic preprocessing of literature text data, but also improves the overall processing efficiency through strategic diversion and differential processing, while ensuring the quality consistency of the processing results.
[0042] The intelligent data filtering mechanism based on LLM of the present invention has successfully achieved the efficient preprocessing of scientific and technological literature text data. The core advantages of this mechanism are as follows: First, through the accurate assessment of the literature complexity by the large language model, it provides a scientific decision-making basis for subsequent processing; Second, by classifying the literature into different complexity categories, it realizes the strategic allocation of processing resources; Third, through the differential processing strategies for different categories of literature, it significantly improves the processing efficiency while ensuring the analysis quality; Fourth, through strict quality control and standardization processing, it ensures the consistency and reliability of the preprocessing results. This innovative preprocessing method enables the system to more effectively handle large-scale and heterogeneous scientific and technological literature text data, laying a solid foundation for subsequent in-depth analysis.
[0043] S2: Use a pre-trained semantic analysis model to encode the target text data to obtain a document semantic vector; The system uses a pre-trained Chinese BERT model as the semantic encoder. Through pre-training on a vast amount of Chinese corpora, this model already has strong semantic understanding capabilities. Considering the professional characteristics of scientific and technological literature, the system performs additional fine-tuning on the BERT model to enable it to better understand the technical terms and expressions in scientific and technological literature. For long literature texts, the system adopts a segmented processing strategy, divides the text into segments of appropriate lengths, obtains semantic representations separately and then integrates them, thus cleverly solving the input length limitation problem of the BERT model when dealing with long texts. In this way, the system can generate vector representations that express rich semantic information for each literature document. These vectors can capture the deep semantics and context relationships of the literature text, providing high-quality feature representations for subsequent topic clustering. Compared with traditional methods such as the bag-of-words model or Word2Vec, the BERT semantic vectors can more accurately capture the meanings of words in specific contexts, understand the semantic essence of the literature content, and thus effectively solve the semantic understanding challenges brought about by the heterogeneity and specialization of the literature content.
[0044] S3: Perform topic clustering on the target text data according to the document semantic vectors, and extract multiple topics and the topic representations of each topic; The system performs topic clustering on the target text data based on the document semantic vectors, extracting multiple topics and their representations. This step is the core part of the present invention and adopts a topic mining method based on BERTopic. First, the system uses the UMAP (Uniform Manifold Approximation and Projection) algorithm to perform dimensionality reduction on the high-dimensional semantic vectors, mapping the BERT vectors of hundreds or even thousands of dimensions into a low-dimensional space (usually 2 - 10 dimensions), while preserving the semantic relationships between documents. The UMAP algorithm is based on Riemannian geometry and algebraic topology theory, and can maximize the preservation of the local and global structure of the data while reducing the dimension. Compared with traditional dimensionality reduction algorithms such as t-SNE, it has higher computational efficiency and better preserves the semantic distance relationship between data points. In the reduced-dimensional vector space, the system applies the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) clustering algorithm to perform density clustering on the documents. HDBSCAN is a hierarchical extended version of the DBSCAN algorithm, which can identify clusters of arbitrary shapes and effectively handle noisy data, and is particularly suitable for the irregular characteristics of the topic distribution in literature texts. Through clustering, the system groups semantically similar documents to form a preliminary topic structure. Subsequently, the system calculates the eigenvalue of the words in each cluster, mainly using the innovative c-TF-IDF (class-based Term Frequency-Inverse Document Frequency) method, treating each cluster as a "virtual document" and calculating the importance of the words in this cluster relative to other clusters. Finally, the system uses the Maximal Marginal Relevance (MMR) algorithm to optimize the keyword set, reducing the redundancy between keywords while ensuring a high correlation between the keywords and the topic, thereby forming a topic representation that can accurately represent the topic content and cover different aspects of the topic.
[0045] The core parameter settings and optimization strategies of the BERTopic model are as follows: UMAP dimensionality reduction parameters: n_neighbors is set to 15 to control the size of the local neighborhood; min_dist is set to 0.1 to balance the local and global structures; n_components is set to 5 to determine the dimension of the vector after dimensionality reduction. The parameters are optimized through grid search on different literature sets to maximize the balance between clustering quality and computational efficiency.
[0046] HDBSCAN clustering parameters: min_cluster_size is set to 10 to ensure that each topic contains sufficient samples; min_samples is set to 5 to control the clustering stability; cluster_selection_epsilon is set to 0.5 to optimize the recognition of small clusters. The system evaluates the clustering quality by combining the silhouette coefficient and the DB index to achieve dynamic adjustment of the parameters.
[0047] c-TF-IDF optimization: The word frequency calculation uses smoothing to reduce the bias of long texts; the IDF calculation uses logarithmic scaling to balance the weights of common and rare words; the top 20 highest-weight words are used for keyword extraction, and the λ parameter of the MMR algorithm is set to 0.6 to balance relevance and diversity.
[0048] S4: Hierarchically cluster the topics according to the topic representations to form a hierarchical topic structure; The system constructs a similarity matrix between topics based on the semantic similarity of topic keywords and the distribution of documents among different topics. The system calculates the cosine similarity or Jaccard similarity coefficient between sets of topic keywords, and at the same time considers the proportion of documents shared between different topics, and forms a comprehensive topic similarity matrix through weighted combination. Subsequently, the system applies a hierarchical clustering algorithm (such as Ward's method) to hierarchically cluster the topics. Ward's method selects the two clusters with the smallest increase in the within-group sum of squares after merging at each step of the clustering process, tending to generate clusters of similar sizes, which is particularly suitable for organizing the topics of scientific and technological literature. The clustering result can be represented as a dendrogram (i.e., a tree diagram). By setting an appropriate threshold, the system cuts out a multi-level topic structure from the tree diagram. Based on the hierarchical clustering result, the system defines a multi-level classification system for the topics of scientific and technological literature, usually including first-level topics (such as macro directions like "scientific and technological work", "scientific and technological enterprises", "scientific and technological finance", etc.), second-level topics (such as specific fields like "project management", "enterprise incubator", etc.) and third-level topics (such as specific literature measures). Finally, the system optimizes and adjusts the hierarchical topic system in combination with the knowledge of domain experts to ensure the scientificity and practicality of topic classification, which may include merging topics with similar semantics, splitting topics with overly broad content, or readjusting the hierarchical attribution of certain topics. In this way, the system finally constructs a multi-level topic structure that can grasp both the macro direction and the in-depth specific measures.
[0049] S5: Divide the target text data into multiple time windows, analyze the change characteristics of the topics in each time window, and construct a topic evolution time series based on the change characteristics, where the change characteristics include frequency of occurrence and intensity change; The system divides the target text data into multiple time windows according to the publication time, analyzes the changing characteristics of the themes in each time window, and constructs a time series of theme evolution. First, the system divides the text data into a series of consecutive time windows according to the literature publication date. The division of the time windows can be flexibly set according to research needs, such as annually, quarterly, or in units of five-year plan cycles. Within each time window, the system independently applies the BERTopic model for theme mining and performs a series of processing steps such as document embedding, dimensionality reduction, clustering, and theme representation. This independent modeling method based on time windows can capture the unique literature theme structure of each period without being interfered by data from other periods. Subsequently, the system calculates the occurrence frequency and intensity of each theme in different time windows. The theme frequency is measured by calculating the proportion of the number of documents belonging to the theme to the total number of documents in the window, while the theme intensity is quantified by a weighted combination of factors such as the number of documents of the theme, the document length, and the level of the publishing institution. These frequency and intensity data constitute the time series of theme evolution, reflecting the rise and fall of the attention of different literature themes. In addition, the system also introduces the Dynamic Topic Model (DTM) technology to analyze the semantic changes of the theme content over time. DTM assumes that there is a certain continuity of the same theme in adjacent periods, but allows the word distribution of the theme to gradually change. In this way, the system can track the evolution trend of the keywords within the theme, identify new words and decaying words, and reveal the subtle changes in the literature language and concerns. Finally, the system constructs a complete theme evolution map by calculating the similarity matrix of themes between adjacent time windows and tracking the processes of theme continuation, differentiation, and integration.
[0050] S6: Predict the trend of the change in the intensity of the theme according to the time series of theme evolution.
[0051] The system first uses the topic intensity time series data obtained from the time evolution analysis and applies traditional time series analysis methods for preliminary trend prediction. Commonly used methods include the Autoregressive Integrated Moving Average model (ARIMA) and the exponential smoothing method. The ARIMA model constructs a mathematical model to predict future values by analyzing the autocorrelation, difference stationarity, and moving average characteristics of the time series; while the exponential smoothing method performs weighted averaging on historical data, with the largest weight for the most recent data and the weights decaying exponentially over time. These methods are particularly suitable for capturing the linear trends and seasonal fluctuations of topic intensity. However, considering the non-linear characteristics of the evolution of scientific and technological literature, the system also combines deep learning methods to establish a more complex prediction model. The Long Short-Term Memory network (LSTM), as a special type of recurrent neural network, can learn long-term dependencies and is particularly suitable for processing time series data; while the Transformer model uses the attention mechanism to process sequence data in parallel and capture the associations between different time points. These deep learning models learn the complex patterns of topic evolution through the training of a large amount of historical data and provide more accurate non-linear predictions. In addition to predicting the intensity changes of existing topics, the system also identifies emerging topics and fading topics through the analysis of the topic evolution map, and predicts the future focus of attention of the literature. Finally, the system integrates external factors (such as the dynamics of scientific and technological development, changes in the international literature environment, major social events, etc.) to correct and adjust the prediction results, improving the accuracy and reliability of the prediction. Through multi-dimensional analysis and correction, the system can ultimately provide a comprehensive and accurate prediction of the trends of scientific and technological literature, providing forward-looking reference information for decision-makers.
[0052] In topic trend prediction, the system adopts the following deep learning model architectures and training strategies: LSTM network structure: It contains 2 layers of bidirectional LSTM, with 128 hidden units in each layer, a dropout rate of 0.3, and Batch Normalization is connected after the input layer to accelerate training. The time window length is set to 8, and the intensity changes of the topic in the next 4 time windows are predicted.
[0053] Transformer structure: It uses 4 layers of Transformer encoders, 8 attention heads, the dimension of the feed-forward network is 512, and the positional encoding uses sine and cosine functions, with a global receptive field to capture long-term dependencies.
[0054] Model training and evaluation: The training data is divided into a training set and a test set at a ratio of 80%:20%. An early stopping strategy is adopted to avoid overfitting. The evaluation metrics include the Root Mean Square Error (RMSE) and the Mean Absolute Percentage Error (MAPE). The system implements model integration, combines the prediction results of the ARIMA and deep learning models, and generates the final prediction through weighted averaging, significantly improving the prediction accuracy.
[0055] S7: Based on the changing trend of the theme intensity, construct a knowledge graph that includes literature entities, influence paths, and effect quantification indicators; The system constructs a domain knowledge graph for the impact analysis of scientific and technological literature as the knowledge basis for subsequent reasoning analysis. This knowledge graph consists of three main components: the literature entity network, the scientific and technological innovation ecological network, and the literature-influence causal chain network. The literature entity network includes entities such as literature documents, publishing institutions, implementation objects, key stakeholders, etc., and relationships such as "publication", "supervision", "implementation", etc. between them, reflecting the basic information and management structure of the literature. The scientific and technological innovation ecological network includes entities such as scientific research institutions, enterprise entities, innovation projects, scientific and technological fields, etc., and relationships such as collaboration, competition, and resource flow between them, reflecting the scientific and technological innovation environment in which the literature plays a role. The literature-influence causal chain network records the causal relationships between historical literature and the observed impacts, including information such as the paths of direct and indirect impacts, time delays, and impact intensities, providing an empirical basis for literature impact reasoning. The system constructs this knowledge graph by integrating multi-source data, and the data sources include literature texts, scientific and technological project databases, scientific research output statistics, enterprise innovation survey data, etc. To ensure the accuracy and integrity of the knowledge graph, the system adopts a semi-automated knowledge extraction method and combines expert review for verification. During the construction of the knowledge graph, special attention is paid to capturing the association patterns between literature themes and specific impacts, and these association patterns will become an important basis for subsequent reasoning analysis.
[0056] S8: Use the knowledge graph to establish a multi-level analysis framework; The framework includes four main levels: the direct impact level, the system response level, the innovation outcome level, and the socioeconomic level. The direct impact level analyzes the direct effects of the literature on specific objects, such as financial support, tax incentives, regulatory requirements, etc., and the direct changes in the behavior of the objects caused by these effects. This level focuses on the immediate effects after the implementation of the literature, which can usually be directly observed through government fund allocation data, tax exemption statistics, etc. The system response level analyzes the responses of various elements within the innovation system to the literature and their interactions, including resource reallocation, organizational structure adjustment, behavior pattern changes, etc. This level focuses on the internal dynamic adjustment process of the innovation system and reflects the conduction mechanism guided by the literature. The innovation outcome level analyzes the ultimate impact of the literature on scientific and technological innovation outcomes, including quantitative indicators such as R & D investment, patent output, technological breakthroughs, product innovation, etc., reflecting the actual promotion effect of the literature on innovation activities. The socioeconomic level analyzes the impact of the literature on a broader socioeconomic level, including aspects such as industrial structure, employment changes, economic growth, sustainable development, etc., reflecting the long-term comprehensive benefits of the literature. In this multi-level framework, the system defines clear evaluation indicators and measurement criteria for each level, forming a structured literature impact evaluation system. Through this hierarchical analysis framework, the system can evaluate the impact path and effect of the literature from different dimensions and depths, effectively addressing the challenges of complexity and multi-dimensionality in literature impact evaluation.
[0057] S9: Establish a retrieval and reasoning system according to the multi-level analysis framework, and the retrieval and reasoning system includes an inference chain from the literature to the impact; The system implements an innovative Reason-while-Retrieve (RwR) framework, which organically combines the retrieval and reasoning processes, enabling the reasoning results to be based on both logical deduction and factual evidence. First, the system decomposes the complex literature impact analysis problem into a series of sub-problems, forming a problem tree, where there are logical dependencies among the sub-problems. For example, analyzing the impact of a certain innovation incentive literature can be decomposed into sub-problems such as R & D investment, talent attraction, and technological cooperation. Subsequently, the system constructs a progressive reasoning chain based on a large language model, with each reasoning step corresponding to a sub-problem. The large language model generates reasoning hypotheses or intermediate conclusions based on the current sub-problem and the existing information. This progressive reasoning supports the analysis of complex causal relationships and is suitable for the multi-path conduction of literature impact. In each reasoning step, the system retrieves the most relevant knowledge from the knowledge graph according to the current sub-problem and the reasoning state. The retrieval uses a hybrid strategy, combining semantic similarity search, relationship path query, and reasoning relevance scoring to ensure that the retrieved knowledge is both relevant and useful. The system integrates the retrieved knowledge with the reasoning process of the large language model to verify or adjust the reasoning hypothesis. If the retrieved evidence supports the reasoning hypothesis, the credibility of the hypothesis is enhanced; if there is a conflict, re-reasoning or adjusting the hypothesis is triggered. Throughout the reasoning process, the system clearly tracks and quantifies the sources of uncertainty in the reasoning chain, including knowledge gaps, evidence conflicts, reasoning jumps, etc., to provide a confidence assessment for the final conclusion. Through this interactive reasoning chain construction, the system can conduct evidence-based and traceable in-depth analysis of complex literature impacts, greatly improving the reliability and interpretability of the reasoning results.
[0058] The technical implementation details of the RwR framework are as follows: Knowledge graph construction: Use the Neo4j graph database to store the literature impact knowledge network, which contains approximately 50,000 entity nodes and 200,000 relationship edges. Entity extraction adopts a method that combines named entity recognition (NER) and distant supervision, and relationship extraction uses a BERT-based relationship classification model with an F1 score of 0.83.
[0059] Implementation of the progressive reasoning chain: Build a reasoning framework based on a large language model with 16 billion parameters, and design reasoning templates through few-shot prompting engineering to implement 9 basic reasoning modes, including conditional reasoning, comparative reasoning, and counterfactual reasoning, etc. The temperature parameter 0.3 is used in the reasoning process to maintain the certainty and coherence of the reasoning.
[0060] Retrieval-reasoning fusion mechanism: Design a bidirectional attention mechanism to enhance each other between the retrieval results and the reasoning process. The retrieval uses a hybrid retrieval strategy, combining BM25 and vector similarity, and pays attention to exact matching and semantic relevance. The reasoning fusion adopts a confidence-weighted method to automatically adjust the reasoning weights according to the evidence support degree.
[0061] Uncertainty quantification: Based on the Bayesian framework, calculate the confidence interval for each inference step, use Monte Carlo sampling analysis to evaluate the robustness of the prediction, and quantify the knowledge gap through information entropy to provide users with a transparent reliability assessment.
[0062] Among them, the inference chain is established through the following steps, including: S9.1: Based on the multi-level analysis framework, decompose the impact analysis task into multiple subtasks with logical dependencies; Task decomposition is the key first step in dealing with complex inference problems. The system adopts a structured task decomposition method to split the overall impact analysis problem into a series of subtasks with clear boundaries and dependencies. First, based on the previously constructed four-level impact analysis framework (direct impact layer, system response layer, innovation achievement layer, and socio-economic layer), the system determines the main dimensions of the impact analysis. Within each dimension, the system is further refined into specific subtasks. For example, analyzing the impact of innovation literature on a company's R & D investment can be decomposed into the following subtasks: What R & D resource support does the literature provide for the company, what is the company's response mechanism to this support, how are R & D resources transformed into R & D investment growth, and what is the relationship between the investment growth and innovation output, etc. These subtasks form a directed acyclic graph (DAG), where each node in the graph represents a subtask, and the edges represent the dependencies between subtasks. Through this task decomposition, the system transforms the complex impact analysis problem into a series of manageable sub-problems, each with clear inputs, outputs, and evaluation criteria. Task decomposition not only reduces the complexity of a single inference step but also clarifies the logical path of the inference, providing a clear guiding framework for the subsequent inference process.
[0063] S9.2: Use the large language model to process the subtasks to obtain a progressive inference chain; The system sequentially inputs the decomposed subtasks into a pre-trained large language model and uses its powerful reasoning ability to generate answers for each subtask. The processing process adopts a "context accumulation" strategy, that is, the processing of subsequent subtasks takes the results of previous subtasks as context input to ensure the coherence and consistency of the reasoning process. To improve the reasoning quality, the system adopts "reasoning prompt engineering" technology, designs a special prompt template for each type of subtask, and guides the model to generate normalized and structured reasoning results. In addition, the system also implements a "multi-step thinking" mechanism, requiring the model to first analyze the problem, clarify the thinking, then gradually deduce, and finally summarize the conclusion. This explicit thinking process helps to reduce logical jumps and errors in reasoning. To handle the uncertainty in reasoning, the system adopts a "confidence annotation" method, requiring the model to annotate the confidence level for each reasoning step and clearly indicate the assumptions and limitations in the reasoning. Through these technologies, the system generates a progressive reasoning chain composed of multiple reasoning nodes, and each node contains the analysis process, reasoning result and confidence evaluation of the subtask. This reasoning chain represents the preliminary reasoning path from the literature to the impact, but it is only based on the general knowledge of the language model and needs to be verified and enhanced through fact retrieval.
[0064] S9.3: Retrieve the associated data related to each node of the progressive reasoning chain in the knowledge graph; To provide factual basis for the reasoning results of the language model, through the retrieval function of the knowledge graph, the reasoning is connected with the actual data. The system constructs multi-level retrieval queries for each node in the reasoning chain. First, the system extracts key entities and relationships from the reasoning node, such as literature name, implementation object, impact type, etc., as the basic retrieval conditions. Then, the system expands the retrieval scope according to the reasoning content and context, including similar literatures, related fields, similar impact cases, etc. In the retrieval process, the system adopts a hybrid retrieval strategy, combining exact matching and semantic similarity search to ensure the comprehensiveness and relevance of the retrieval results. For the retrieved results, the system conducts relevance scoring and deduplication processing, and preferentially retains the data that is most relevant to the current reasoning node and has the highest evidence value. The system's retrieval is not limited to the direct relationships in the knowledge graph, but also supports multi-hop queries, and can discover indirect but meaningful associated evidence. For example, if the reasoning involves the impact of a certain literature on an enterprise's R & D, the system will not only retrieve the direct literature-enterprise relationship, but also search for multi-hop paths such as literature-fund-enterprise or literature-talent-enterprise to capture complex impact mechanisms. Through this in-depth retrieval, the system collects relevant factual data for each reasoning node and provides an objective basis for reasoning verification.
[0065] S9.4: Combine the associated data and the reasoning process indicated by the large language model to verify and correct the reasoning results and generate a reasoning chain.
[0066] This is the core part of the RwR framework, which realizes the deep integration of reasoning and retrieval. First, the system matches the retrieved associated data with the corresponding reasoning nodes and analyzes the consistency between the data and the reasoning. For each reasoning node, the system calculates the "evidence support degree" to quantify the degree of support of the factual data for the reasoning conclusion. Based on the evaluation of the evidence support degree, the system classifies the reasoning results: for the reasoning fully supported by evidence, the system retains the original conclusion and adds specific factual basis to improve the credibility of the conclusion; for the reasoning partially supported by evidence, the system modifies or qualifies the conclusion to make it more compliant with the constraints of the factual data; for the reasoning with insufficient or conflicting evidence, the system marks it as "to be verified" and tries to resolve the conflict by additional retrieval or adjustment of the reasoning path. When dealing with evidence conflicts, the system adopts the "evidence weight" strategy, considering the authority, timeliness, and direct relevance of the data sources to resolve the conflicts between different pieces of evidence. After this verification and correction process, the system generates a complete reasoning chain that is both based on logical reasoning and supported by facts, with clear source annotations for each reasoning step: whether it is reasoning based on the language model, direct evidence of factual data, or a comprehensive judgment of both. This transparent source annotation makes the reasoning process traceable and verifiable, greatly enhancing the reliability and interpretability of the reasoning results.
[0067] The above constructs an innovative knowledge reasoning system that tightly combines reasoning and retrieval. Compared with traditional pure language model reasoning or simple information retrieval, this system has significant advantages: First, through task decomposition and progressive reasoning, the system effectively handles complex impact analysis problems, making the reasoning process clearer and more controllable; Second, the system combines the reasoning ability of large language models with the factual basis of knowledge graphs, retaining both the flexibility and generalization ability of the language model and avoiding its "hallucination" problem, improving the accuracy and reliability of reasoning; Third, the reasoning process of the system is transparent and interpretable, with a clear reasoning path and factual basis for each conclusion, facilitating user understanding and verification; Finally, the system supports interactive reasoning and can dynamically adjust the reasoning path according to new information or questions, with strong adaptability and scalability. This retrieval reasoning system based on the RwR framework provides powerful intelligent support for the impact analysis of scientific and technological literature, can extract key information from a large amount of literature and data, construct a reliable causal reasoning chain, effectively answer complex questions of "why" and "how", and provide in-depth insights and basis for scientific and technological decision-making.
[0068] S10: Adjust the effect quantification index, and perform multi-scenario prediction by combining the adjusted effect quantification index and the retrieval reasoning system to obtain the development trends under different conditions; Based on the inference results of the RwR framework, the system further conducts scenario simulations and sensitivity analyses of the literature's impact. First, by adjusting the key parameters of the literature (such as the scale of funds, the proportion of tax incentives, the entry threshold, etc.), the system simulates the potential impact differences of different literature design schemes. This parameter adjustment simulation helps decision-makers understand how subtle changes in the literature design affect the final effect, providing a quantitative reference for literature optimization. Second, under different external environment assumptions (such as economic growth rate, international competition situation, technological development stage, etc.), the system evaluates the robustness and adaptability of the literature's impact. By changing the external condition parameters, the system can identify the sensitivity of the literature's effect to the environment, helping to design a more flexible and adaptable literature plan. In addition, the system also simulates the changing trend of the literature's impact over time, including short-term effects, medium-term adjustments, and long-term impacts, revealing the time characteristics of the literature's impact. This time-dynamic simulation particularly focuses on the persistence and decay characteristics of the literature's effect, contributing to the design of a reasonable literature implementation cycle and update mechanism. Finally, the system conducts factor sensitivity analysis to identify the key factors and conditions that are most sensitive to the literature's impact, providing targeted suggestions for literature optimization. The system presents the simulation and analysis results in a visual manner, including impact path diagrams, sensitivity heatmaps, scenario comparison diagrams, etc., enabling decision-makers to intuitively understand the impact characteristics and optimization space of the literature plan. Through this multi-scenario prediction analysis, the system provides rich decision-making support information for literature formulators, helping to design more effective and targeted scientific and technological literature.
[0069] Among them, as Figure 2 shown, S2 specifically includes: S2.1: Segment the target text data to obtain text segments of a specified length; This addresses the limitation of pre-trained language models in processing long texts. Pre-trained models such as BERT usually have input length limitations (generally 512 tokens), while scientific and technological literature is often long and cannot be directly input into the model. To solve this problem, the system adopts an intelligent segmentation strategy to divide the long text into segments of appropriate length. The segmentation process takes into account the semantic integrity and preferentially cuts at natural paragraph boundaries or sentence boundaries to avoid splitting complete semantic units. For particularly long paragraphs, the system uses a sliding window technique and sets an appropriate overlapping area (usually 15 - 20% of the text length) to ensure the coherence of context information. For text parts containing special structures (such as tables, chart descriptions, formulas, etc.), the system adopts special processing rules to preserve their structural characteristics or convert them into a form that the model can understand. Through this intelligent segmentation processing, the system converts the original long text into a series of text segments with moderate length and complete semantics, providing suitable input units for subsequent semantic encoding.
[0070] S2.2: Obtain the semantic representation of the text fragment using the semantic analysis model; Use pre-trained large language models (such as BERT, RoBERTa, etc.) as semantic encoders. These models have already possessed strong semantic understanding capabilities through pre-training on a vast amount of text. To better adapt to the field of scientific and technological literature, the system also performs domain adaptation fine-tuning on the basic model, using a large amount of scientific and technological literature data to perform additional training on the model to enhance its understanding ability of professional terms and scientific and technological expressions. During the encoding process, the system inputs the text fragment into the model to obtain a deep semantic representation. Specifically, the system extracts the hidden state vectors of the last few layers of the model, and these vectors contain rich semantic information and context relationships. For BERT-like models, the system usually uses the [CLS] token vector of the last layer as the semantic representation of the entire fragment, or averages the vector representations of all tokens. These high-dimensional vectors (usually 768-dimensional or 1024-dimensional) capture the deep semantic features of the text, including multiple levels such as word meaning, syntactic relationship, and topic information, providing a high-quality semantic basis for subsequent document representation.
[0071] S2.3: Integrate the semantic representation of the text fragment into a document semantic vector, and the document semantic vector is used to indicate the semantic information and context relationship of the target text data.
[0072] This step solves the problem of how to synthesize the local representations of multiple fragments into the global representation of the entire document. For this purpose, the system designs a hierarchical semantic integration method. First, for each text fragment, the system obtains its semantic vector representation, and these vectors capture the semantic information at the fragment level. Then, the system adopts a weighted integration mechanism to combine these fragment vectors into the overall representation at the document level. The integration process considers multiple factors: fragment position (fragments at the beginning and end of the document may contain more important overview information), fragment content importance (evaluated by indicators such as keyword density), and semantic relevance between fragments. The system implements an integration algorithm based on the attention mechanism to automatically learn the importance weights of different fragments and give key fragments higher influence. In addition, the system also retains the mapping relationship between the fragment-level vector and the document-level vector, facilitating the tracing of the specific source of key information during subsequent analysis. The finally generated document semantic vector is a high-dimensional dense vector, comprehensively representing the semantic content and internal structural relationship of the document, providing an ideal input form for subsequent topic mining.
[0073] After completing the semantic representation, the system enters the topic clustering stage to identify the topic distribution in the scientific and technological literature based on the document semantic vector. First, since the semantic vectors generated by BERT encoding are usually of high dimension (such as 768 dimensions), directly performing clustering in the high-dimensional space may face the problem of "curse of dimensionality", affecting the clustering effect.
[0074] Among them, as Figure 3 shown, S3 specifically includes: S3.1: Use a preset dimensionality reduction model to perform dimensionality reduction on the document semantic vector to obtain a low-dimensional semantic vector; Use the UMAP (Uniform Manifold Approximation and Projection) algorithm as the main dimensionality reduction tool. The UMAP algorithm is based on Riemannian geometry and algebraic topology theories and can maximize the preservation of the local and global structures of data while reducing the dimension. Compared with traditional dimensionality reduction methods (such as PCA, t-SNE), UMAP has the advantages of high computational efficiency and strong scalability and is especially suitable for processing large-scale literature data. In practical applications, the system usually reduces high-dimensional vectors to a low-dimensional space of 5-15 dimensions, and this dimension range can balance information preservation and computational efficiency. The selection of dimensionality reduction parameters takes into account the data scale and distribution characteristics, and the system adopts a dynamic parameter adjustment strategy to automatically optimize the parameter settings according to the actual data distribution to ensure that the dimensionality reduction results not only retain key information but also facilitate subsequent processing.
[0075] S3.2: Based on the low-dimensional semantic vector, use a preset clustering model to perform density clustering on the target text data to form multiple themes; Select HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) as the main clustering algorithm. HDBSCAN is a hierarchical extended version of the DBSCAN algorithm and has several key advantages: it can identify clusters of arbitrary shapes, has good robustness to noisy data, can automatically determine the optimal number of clusters, and has strong ability to process clusters with different densities. These characteristics make HDBSCAN particularly suitable for the needs of scientific literature theme clustering because the theme distribution of scientific literature is usually not uniform, there are core themes and marginal themes, and the density difference is significant. During the clustering process, the system optimizes the key parameters of HDBSCAN: the minimum cluster size parameter is dynamically adjusted according to the total amount of literature to ensure that the identified themes are representative enough; the core distance parameter is adaptively set based on the local density distribution of the data to balance the clustering granularity and coverage. Through this optimized density clustering process, the system can naturally identify theme groups with semantic consistency from the literature data, and each group represents a relatively independent literature theme.
[0076] S3.3: Calculate the word feature values of the words in each theme, and extract keywords based on the word feature values. The word feature values include word frequency and inverse document frequency; To accurately represent the core content of each topic, the system innovatively applies the c-TF-IDF (class-based Term Frequency-Inverse Document Frequency) method. Different from traditional TF-IDF, c-TF-IDF treats each topic as a "virtual document" and calculates the importance of a term in that topic relative to other topics. Specifically, the TF (term frequency) part calculates the frequency of a term in a specific topic, reflecting the contribution of the term to that topic; the IDF (inverse document frequency) part calculates the distribution breadth of a term across all topics, giving higher weights to terms with strong topic specificity (appearing in only a few topics). Through c-TF-IDF calculation, the system generates a sorted list of all terms in each topic, and terms with high weights can usually accurately reflect the core content and unique features of the topic. The system further combines the semantic representation information of terms. By calculating the similarity between the word vector and the topic center vector, it captures keywords with strong semantic relevance but low frequency, enhancing the semantic accuracy of topic representation.
[0077] S3.4: Optimize the keywords using a preset maximum similarity matching algorithm to obtain the topic representation of the topic.
[0078] Optimize the keywords using a preset maximum similarity matching (MMR, Maximal Marginal Relevance) algorithm to obtain the topic representation of the topic. The MMR algorithm is a selection method that balances relevance and diversity. When selecting keywords, it considers both the relevance of the term to the topic and the diversity of the selected term set. Specifically, the system defines an objective function that combines relevance and diversity: MMR(w) = λ·sim(w,t) - (1-λ)·maxsim(w,S), where w is the candidate term, t is the topic center, S is the selected keyword set, sim represents the similarity function, and λ is the balance parameter (usually set to 0.5 - 0.7). The system iteratively selects keywords in descending order of MMR values until the predetermined number of keywords (usually 10 - 20) is reached. This method ensures that the selected keyword set is not only highly relevant to the topic but also covers different aspects of the topic, avoiding the problems of content duplication or one-sided topic representation. Finally, each topic is represented by a set of optimized keywords, and these keywords together constitute the semantic feature representation of the topic, providing a basis for subsequent topic hierarchy analysis and evolution analysis.
[0079] Among them, S3 also includes: S3.5: Based on the document semantic vector, establish a scoring index system including literature level classification, publication time, and citation frequency; This indicator system comprehensively considers multiple important dimensions of the literature: the hierarchical classification reflects the official attributes and authority of the literature, such as national, provincial and ministerial, local, etc.; the release time reflects the timeliness of the literature, and recently released literature may have higher reference value; the citation frequency reflects the influence and recognition of the literature, and highly cited literature usually represents important literature directions or research focuses. The system integrates these indicators into a unified importance score, assigns a weight value to each piece of literature, and these weights will play an important role in the subsequent clustering process.
[0080] S3.6: Construct a cost-sensitive decision tree based on the document semantic vector and the scoring indicator system, and the cost-sensitive decision tree is used to indicate increasing the assigned weight for samples whose importance reaches a preset importance threshold; The system regards the mis-clustering of literature as a kind of "cost", and the mis-clustering cost of important literature is higher. Through the cost-sensitive decision tree, the system increases the assigned weight for samples whose importance reaches the preset importance threshold, ensuring that these important literatures are processed more accurately during the clustering process. The construction process of the decision tree considers the semantic features and importance indicators of the literature, and learns decision rules that can effectively distinguish different types of literature by minimizing the weighted error rate.
[0081] S3.8: Extract classification rules from the cost-sensitive decision tree, and the classification rules include feature thresholds and classification paths; Each path from the root node to the leaf node of the decision tree represents a classification rule, which contains a series of feature judgment conditions (feature thresholds) and the final classification result. The system extracts these rules to form a rule set, and each rule describes the feature pattern and the belonging category of a certain type of literature. These rules are not only used for subsequent clustering optimization, but also provide an interpretable basis for literature classification, enabling the system to clearly explain why a certain piece of literature is classified into a specific topic. The rule extraction process also includes rule refinement and optimization, merging similar rules, removing redundant conditions, and adjusting the threshold range to make the rule set more refined and comprehensive.
[0082] S3.8: Optimize the initial center point distribution of the topic clustering according to the classification rules, and perform topic clustering on the target text data according to the optimization result.
[0083] Traditional density clustering algorithms (such as HDBSCAN) may not fully consider the influence of important documents when dealing with documents of different importance. To address this issue, the system innovatively applies the classification rules of cost-sensitive decision trees to the clustering initialization process. Specifically, the system uses the classification rules to identify a set of important documents, and with these documents as the core, optimizes the initial density estimation and core point selection of the clustering algorithm. This method ensures that important documents can become the "seed points" of clustering, guiding the formation of clustering boundaries and improving the recognition accuracy of the themes represented by important documents. After completing the initial optimization, the system executes an improved version of the HDBSCAN algorithm to organize all documents into semantically consistent theme clusters while maintaining the clustering accuracy of important documents. Finally, the system obtains an optimized clustering result that takes into account both semantic similarity and document importance.
[0084] It realizes the conversion process from text data to theme representation. The innovation of this process lies in: combining the deep semantic understanding ability of BERT and the flexible grouping ability of density clustering to achieve accurate theme recognition of scientific and technological literature; introducing the c-TF-IDF and MMR algorithms to obtain relevant and diverse theme representations; innovatively applying the idea of cost-sensitive learning to improve the clustering accuracy of important documents. These innovation points together constitute an efficient and accurate scientific and technological literature theme mining framework, laying a solid foundation for subsequent hierarchical analysis and evolutionary prediction.
[0085] Among them, as Figure 4 shown, S4 specifically includes: S4.1: Calculate the semantic similarity based on the theme representation, and construct a similarity matrix between the themes according to the theme representation and the semantic similarity; The system constructs a topic vector representation for each topic, which is based on two key pieces of information: one is the keyword set of the topic, and the other is the central point of the document semantic vectors belonging to that topic. For the keyword set, the system performs a weighted average of the word vectors of each keyword (with the weight being the c-TF-IDF value of the word) to generate a keyword vector representation; for the document set, the system calculates the weighted average of all the document semantic vectors under that topic (with the weight being the importance score of the document) to generate a document vector representation. Then, the system fuses these two vector representations into a unified topic vector for subsequent similarity calculation. In the similarity calculation step, the system adopts a multi-angle similarity measurement method. In addition to the commonly used cosine similarity, the system also calculates the Jaccard similarity coefficient of the topic keyword set and the overlap degree of the topic document distribution. These three similarities are integrated into a comprehensive similarity score to more comprehensively reflect the association strength between topics. Finally, the system organizes the similarity values between all topic pairs into a similarity matrix, which is a symmetric matrix, and each element represents the semantic similarity degree between the topics corresponding to the row and column. This matrix provides a key input for subsequent hierarchical clustering and determines which topics should be grouped into higher-level categories.
[0086] S4.2: Perform hierarchical clustering on the topics using a preset hierarchical clustering model to obtain a hierarchical system represented as a tree structure, where the hierarchical system includes multiple first-level topics and multiple second-level topics corresponding to the first-level topics; The present invention selects the Ward hierarchical clustering method. This method considers minimizing the increase in within-group variance when merging clusters, tends to generate clusters of similar sizes, and is suitable for the organizational requirements of scientific and technological literature topics. The hierarchical clustering process is a bottom-up iterative merging process: initially, each topic is regarded as an independent cluster; then, in each iteration step, the system finds the two clusters with the highest similarity and merges them; this process continues until all topics are merged into one cluster or a preset stop condition is reached. The clustering process can be visualized as a dendrogram (i.e., a hierarchical clustering tree), where the leaf nodes of the tree are the original topics, the internal nodes represent higher-level topic categories, and the root node contains all topics. The system obtains a multi-level topic classification system by "cutting" at appropriate positions in the dendrogram. The selection of the cutting position comprehensively considers factors such as the number of clusters, the internal consistency of the clusters, and the discrimination between different clusters. In the present invention, the system usually constructs a two-level or three-level topic classification system, including multiple first-level topics (such as macro categories like "Technological Innovation", "Industrial Development", "Science and Technology Finance", etc.) and multiple second-level topics corresponding to the first-level topics (such as specific fields like "Basic Research", "Transformation of Scientific and Technological Achievements" under "Technological Innovation"). This hierarchical system provides a clear organizational structure for the topics of scientific and technological literature, helping to understand the subordinate relationship and correlation degree between topics.
[0087] S4.3: Optimize and adjust the hierarchical system to form a hierarchical structure of themes.
[0088] The automatically generated hierarchical system may have some unreasonable aspects and needs to be further optimized to improve its scientificity and practicality. The optimization process of the system includes multiple aspects: First, split or merge overly large or small clusters to maintain the balance of themes at each level; Second, based on the semantic relationships between themes, adjust the attribution relationships of some themes to ensure semantic consistency in classification; Third, utilize external domain knowledge bases (such as subject classification systems, technical field classification standards, etc.) to standardize the naming and organizational structure of themes and improve the professionalism and understandability of classification. In practical applications, the system also supports an expert intervention mode, allowing domain experts to review and modify the automatically generated hierarchical structure based on their own knowledge, further enhancing the practical value of theme classification. Finally, the system forms a multi-level classification system for the themes of scientific and technological literature. This system not only retains the objectivity driven by data but also takes into account the guiding role of professional knowledge, providing an effective semantic framework for the organization and retrieval of scientific and technological literature.
[0089] After completing the construction of the hierarchical structure of themes, the system enters the theme evolution analysis stage, focusing on the changing trends of themes over time. This analysis is of great significance for grasping the dynamics of scientific and technological development.
[0090] Among them, as Figure 5 shown, S5 specifically includes: S5.1: Divide the target text data in the time dimension according to the publication time of the target text data to obtain multiple consecutive time windows; The division of time windows is the basis of time series analysis. The system designs a flexible time division strategy according to research requirements and data characteristics. Usually, the system uses natural time units (such as monthly, quarterly, annual) or literature cycles (such as five-year plan cycles) as the basis for dividing time windows. For cases with a large amount of data, smaller time units (such as monthly or quarterly) can be selected to capture more detailed changes; for cases with a small amount of data or when focusing on long-term trends, larger time units (such as annual or multi-year) can be chosen. The system supports dynamic time window settings, allowing users to adjust the window size or sliding step according to specific needs. The setting of time windows should ensure both the statistical significance of the data volume within the window and the requirements of time resolution. In actual processing, the system first organizes all data in chronological order based on the publication time attribute of the literature, and then divides the data into a series of consecutive time periods according to the preset window parameters. The literature within each time period constitutes a data set for a time window. This division of time windows provides a basic time framework for subsequent theme evolution analysis.
[0091] S5.2: Use the preset topic mining model to perform topic mining on the target text data within the time window; To track the change of topics over time, the system needs to independently identify the topic structure in each time window. The present invention adopts two topic mining strategies: one is the independent modeling strategy, and the other is the incremental update strategy. In the independent modeling strategy, the system independently applies the BERTopic model to the data of each time window, and executes a full set of topic mining processes (including document semantic encoding, dimensionality reduction, clustering, and topic representation extraction). This method can capture the topic patterns unique to each period and is not affected by other periods, but may lead to the problem of inconsistent topic identification. In the incremental update strategy, the system processes the data of the new window incrementally based on the topic model of the previous time window, only updates the topic distribution and representation, while maintaining the continuity of topic identification. This method helps to maintain the consistency of topics and facilitates cross-time comparison, but may miss emerging topics or ignore significant changes in topic content. To take into account the advantages of these two strategies, the system adopts a hybrid method: perform independent modeling regularly (such as annually or for each literature cycle), while adopting incremental update during the intermediate period, and set up a new topic discovery mechanism to identify emerging topics in a timely manner during incremental update. Through this hybrid strategy, the system can effectively capture the dynamic changes of the topic structure while maintaining topic continuity.
[0092] S5.3: According to the topic mining results, calculate the topic features of each topic in different time windows, where the topic features include the occurrence frequency and intensity; Topic features are the basic indicators for quantifying the temporal changes of topics. Topic frequency reflects the popularity of a topic, usually measured by calculating the proportion of the number of documents belonging to the topic in the total number of documents in the window; topic intensity reflects the importance or attention of a topic. The system calculates the topic intensity by synthesizing multiple factors, including: the number of topic documents, the average length of the documents, the average importance score of the documents (based on the previously constructed scoring index system), the concentration of topic keywords in the documents, etc. The calculation of topic intensity takes into account the balance between the number of documents and the importance of the documents, and can more accurately reflect the actual influence of the topic. The system calculates these feature values for each topic in each time window and organizes the results into time series data to form the time profile of topic features. These time profiles are the data basis for subsequent evolution analysis and reflect the change patterns of topics over time.
[0093] S5.4: Use the preset dynamic topic model to analyze the topic features to obtain the change features of the topic, where the change features are used to indicate the semantic changes of the topic over time; In addition to changes in topic frequency and intensity, the content structure of topics also evolves over time. This semantic change reflects the subtle shift in topic focus. To capture this change, the system applies Dynamic Topic Models (DTM) technology to analyze the temporal changes in the distribution of words within a topic. The DTM model assumes a certain continuity of the same topic in adjacent periods, but allows the word distribution of the topic to gradually change. The system compares the keyword distributions of each topic in different time windows to identify newly added words, disappearing words, and words with significant changes. These lexical changes reflect the evolution trend of topic content, such as the emergence of new technology concepts, the fading out of old concepts, and changes in focus. The system also calculates the degree of semantic drift within a topic, that is, the distance between the semantic representations of the same topic at different time points, to quantify the speed and magnitude of topic changes. In addition, the system analyzes the mutual influence between topics, such as semantic penetration, topic differentiation, and integration between topics, to reveal the dynamic evolution law of the topic network. These change characteristics together constitute a multi-dimensional description of topic evolution, reflecting not only the change in "how much" but also the change in "what".
[0094] S5.5: Generate the time series of topic evolution based on the change characteristics.
[0095] Integrate the various change characteristics analyzed above into structured time series data as input for trend prediction. The system organizes multiple types of time series of topic evolution: the time series of topic popularity, recording the frequency and intensity of each topic at different time points; the time series of topic content change, recording the changes in the keyword distribution of topics; the time series of topic relationship change, recording the dynamic mutual influence and structural reorganization between topics. These time series are stored in various forms, including numerical time series (such as the numerical change in topic intensity), vector time series (such as the change in the semantic representation of topics over time), and structural time series (such as the evolution of topic network topology). The system also adds rich meta-information to the time series data, such as time granularity, data source, processing method, etc., for subsequent interpretation and use. All time series data are organized into a unified database to support multi-dimensional query and analysis, providing high-quality training data for the final trend prediction.
[0096] The present invention realizes a comprehensive analysis of the thematic structure and evolution law of scientific and technological literature. This analysis not only reveals the organizational structure of the themes, but also captures the changing patterns of the themes over time, providing an important basis for understanding the development trends of science and technology. In particular, through the hierarchical organization of themes and the analysis of temporal changes, the system establishes a panoramic view from micro-themes to macro-domains and from static structures to dynamic evolutions, which helps to grasp the overall trends and internal laws of scientific and technological development from multiple perspectives and at multiple levels. This structured and dynamic analysis method breaks through the limitations of traditional literature analysis and provides more systematic and in-depth knowledge support for scientific and technological decision-making.
[0097] The embodiment of the present application also provides a device for analyzing and predicting the thematic trends of scientific and technological literature, including: A data preprocessing module, configured to collect scientific and technological literature text data, preprocess the scientific and technological literature text data, and obtain target text data; A semantic encoding module, configured to encode the target text data by using a pre-trained semantic analysis model to obtain a document semantic vector; A thematic clustering module, configured to perform thematic clustering on the target text data according to the document semantic vector, and extract multiple themes and the thematic representations of each theme; A hierarchical analysis module, configured to perform hierarchical clustering on the themes according to the thematic representations to form a thematic hierarchical structure; A time evolution analysis module, configured to divide the target text data into multiple time windows, analyze the change characteristics of the themes in each time window, and construct a thematic evolution time series according to the change characteristics, where the change characteristics include the occurrence frequency and intensity change; A trend prediction module, configured to predict the intensity change trend of the themes according to the thematic evolution time series.
[0098] The embodiment of the present application also provides a computer device, which includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for analyzing and predicting the thematic trends of scientific and technological literature as described above.
[0099] The embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method for analyzing and predicting the thematic trends of scientific and technological literature as described above.
[0100] The embodiments of the present application also provide a computer program product, including computer instructions, which implement the steps of the method for analyzing and predicting the theme trends of the above-mentioned scientific and technological literature when executed by a processor.
[0101] The present application has the following technical effects: Innovatively combines the BERT semantic model with the topic model to construct a BERTopic scientific and technological literature analysis framework, which can capture the deep semantic relationships in the literature text more accurately than traditional topic models such as LDA; designs a multi-level topic classification system, organizes the topics through hierarchical clustering algorithms, and realizes an all-round analysis from the macroscopic literature orientation to the microscopic literature measures; proposes a method for analyzing the dynamic evolution of literature topics, and realizes the quantitative analysis of the evolution law of literature topics through time window division and dynamic topic modeling; develops a literature trend prediction model based on the topic evolution pattern, combines time series analysis and deep learning methods, and provides a forward-looking reference for future literature orientation.
[0102] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method for analyzing and predicting the theme trends of scientific and technological literature described in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0103] In addition, the embodiments of the present disclosure also provide a computer program product, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method for analyzing and predicting the theme trends of scientific and technological literature provided in any one of the above embodiments of the present disclosure. For details, please refer to the above method embodiments and will not be elaborated here.
[0104] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0105] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and apparatuses described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0106] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0107] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0108] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0109] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for analyzing and predicting the trend of scientific and technological literature, characterized in that: include: Collecting scientific and technological literature text data, and preprocessing the scientific and technological literature text data to obtain target text data; Encoding the target text data using a pre-trained semantic analysis model to obtain a document semantic vector; Performing topic clustering on the target text data according to the document semantic vector, extracting multiple topics and topic representation of each topic; hierarchically clustering the topics according to the topic representation to form a topic hierarchical structure; Divide the target text data into multiple time windows, analyze the change characteristics of the topic in each time window, and construct a topic evolution time series based on the change characteristics, wherein the change characteristics include occurrence frequency and intensity changes; According to the topic evolution time series, the intensity change trend of the topic is predicted.
2. The method according to claim 1, characterized in that The method of encoding the target text data using a pre-trained semantic analysis model to obtain a document semantic vector includes: Segmenting the target text data to obtain text segments of a specified length; Acquire the semantic representation of the text segment using the semantic analysis model; The semantic representations of the text segments are integrated into a document semantic vector, where the document semantic vector is used to indicate the semantic information and contextual relationship of the target text data.
3. The method according to claim 2, characterized in that The subject clustering of the target text data according to the document semantic vector to extract multiple subjects and subject representation of each subject includes: Using a preset dimensionality reduction model to perform dimensionality reduction processing on the document semantic vector to obtain a low-dimensional semantic vector; Based on the low-dimensional semantic vector, the target text data is density clustered using a preset clustering model to form multiple topics; Calculate the word feature values of the words in each topic, and extract keywords based on the word feature values, wherein the word feature values include word frequency and inverse document frequency; The keywords are optimized using a preset maximum similarity matching algorithm to obtain a topic representation of the topic.
4. The method according to claim 3, characterized in that The step of performing hierarchical clustering on the topics according to the topic representation to form a topic hierarchical structure includes: Calculate semantic similarity based on the topic representation, and construct a similarity matrix between the topics according to the topic representation and the semantic similarity; Performing hierarchical clustering on the topics using a preset hierarchical clustering model to obtain a hierarchical system represented as a tree structure, wherein the hierarchical system includes a plurality of first-level topics and a plurality of second-level topics corresponding to the first-level topics; The hierarchical system is optimized and adjusted to form a thematic hierarchical structure.
5. The method according to claim 4, characterized in that The step of dividing the target text data into multiple time windows, analyzing the change characteristics of the topics in each time window, and constructing a topic evolution time series according to the change characteristics includes: According to the release time of the target text data, the target text data is divided into time dimensions to obtain a plurality of continuous time windows; Using a preset topic mining model to perform topic mining on the target text data within the time window; According to the topic mining results, the topic features of each topic in different time windows are calculated, and the topic features include occurrence frequency and intensity; Analyzing the topic features using a preset dynamic topic model to obtain a change feature of the topic, wherein the change feature is used to indicate a semantic change of the topic over time; The theme evolution time series is generated according to the change characteristics.
6. The method according to claim 5, characterized in that The preprocessing of the scientific literature text data to obtain target text data includes: Using a preset large language model to evaluate the scientific and technological literature text data to obtain a complexity score, wherein the complexity score is used to indicate the professionalism, intersectionality and structural complexity of the scientific and technological literature text data; Classifying the scientific and technological literature text data according to the complexity score to obtain a first category of text and a second category of text, wherein the first category of text is used to indicate that the complexity score is lower than a preset score threshold, and the second category of text is used to indicate that the complexity score is higher than the score threshold; Performing a lightweight analysis on the first type of text and performing a deep analysis on the second type of text to obtain a preliminary processing result; According to the preliminary processing result, the target text data is obtained.
7. The method according to claim 6, characterized in that The use of a preset large language model to evaluate the scientific literature text data to obtain a complexity score includes: Collect training texts with complexity annotations, and fine-tune the preset initial large language model to obtain a large language model; Calculate the analysis features of the scientific literature text data using the large language model, wherein the analysis features include professional vocabulary density, field coverage, and logical level depth; Calculating a text complexity score based on the analysis features, and generating an evaluation basis description based on the text complexity score; According to the evaluation basis description, a complexity score of the scientific literature text data is obtained.
8. The method according to claim 7, characterized in that The subject clustering of the target text data according to the document semantic vector to extract multiple subjects and subject representations of each subject also includes: Based on the document semantic vector, a scoring index system including document grade classification, publication time and citation frequency is established; According to the document semantic vector and the scoring index system, a cost-sensitive decision tree is constructed, wherein the cost-sensitive decision tree is used to indicate the weights given to samples whose importance reaches a preset importance threshold; Extracting classification rules from the cost-sensitive decision tree, wherein the classification rules include feature thresholds and classification paths; The initial center point distribution of the topic clustering is optimized according to the classification rule, and the target text data is subject-clustered according to the optimization result.
9. The method according to claim 8, characterized in that After predicting the intensity change trend of the topic according to the topic evolution time series, the method further includes: Based on the changing trend of the topic intensity, a knowledge graph including literature entities, impact paths and effect quantitative indicators is constructed; Establishing a multi-level analysis framework using the knowledge graph; Establishing a retrieval reasoning system according to the multi-level analysis framework, wherein the retrieval reasoning system includes a reasoning chain from literature to impact; Adjusting the effect quantification index, combining the adjusted effect quantification index with the retrieval reasoning system to perform multi-scenario prediction, and obtaining development trends under different conditions; The reasoning chain is established through the following steps, including: Based on the multi-level analysis framework, the impact analysis task is decomposed into a plurality of subtasks with logical dependencies; Processing the subtasks using the large language model to obtain a progressive reasoning chain; Retrieving associated data related to each node of the progressive reasoning chain in the knowledge graph; The associated data and the reasoning process indicated by the large language model are combined to verify and correct the reasoning result and generate a reasoning chain.
10. A device for analyzing and predicting the trend of scientific and technological literature, characterized in that: include: A data preprocessing module is used to collect scientific and technological literature text data, preprocess the scientific and technological literature text data, and obtain target text data; A semantic encoding module, used to encode the target text data using a pre-trained semantic analysis model to obtain a document semantic vector; A topic clustering module, used to perform topic clustering on the target text data according to the document semantic vector, and extract multiple topics and topic representations of each topic; A hierarchical analysis module, used for hierarchically clustering the topics according to the topic representation to form a topic hierarchical structure; A time evolution analysis module, used to divide the target text data into multiple time windows, analyze the change characteristics of the topic in each time window, and construct a topic evolution time series based on the change characteristics, wherein the change characteristics include the frequency of occurrence and the change of intensity; The trend prediction module is used to predict the intensity change trend of the topic based on the topic evolution time series.
Citation Information
Patent Citations
Subject topic evolution reasoning method combining time lag calculation in science and technology intelligence analysis
CN111046167A
Dynamic knowledge hotspot evolution and trend analysis method
CN111694930A
Scientific and technological theme evolution stage prediction method and system based on multi-graph representation
CN119691159A
System and method for automatically generating systematic reviews of a scientific field
US20110295903A1
Cited By
Artificial intelligence-based literature structured extraction method and system
CN120336416A
Data set quality evaluation method and device, equipment, medium and program product
CN120372325A
Hot topic monitoring method and system based on improved HDBSCAN clustering algorithm
CN120429486A
Industrial document intelligent retrieval method and system based on large model
CN120744081A
Theme recognition method and system for large-scale text data and readable medium
CN120745647A