A Method and System for Predicting the Evolution Stages of Science and Technology Themes Based on Multigraph Representation
By constructing a multi-graph representation and graph neural network model, the problem of excessive analysis of the particle size in the evolution of science and technology themes is solved, and accurate prediction of the evolution process of science and technology themes is achieved, and the development trends and phased changes of science and technology themes can be more accurately tracked.
Patent Information
- Application Number
- CN202411746825.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The analysis particle size of the evolution of science and technology themes in the prior art is too coarse, making it difficult to accurately predict the evolution process of science and technology themes.
By constructing multiple graph representations, we obtain the keywords of science and technology themes, their word frequency sequences and citation relationships in scientific and technological literature, use graph neural network models for feature extraction and modeling, and combine semantic distances and citation relationships to predict the evolution stage of science and technology themes.
Accurate prediction of the evolution process of scientific and technological themes is achieved, and the development trends and phased changes of scientific and technological themes are more accurately tracked.
Smart Images

Figure CN119691159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prediction of the evolution of scientific and technological themes, and particularly to a method and system for predicting the evolution stage of scientific and technological themes based on a multi-graph representation. Background Art
[0002] In the process of the evolution and development of scientific and technological themes over time, phased changes such as emergence, growth, integration, and decline will occur. Achieving accurate discrimination and prediction of the evolution stage of scientific and technological themes will help to reveal the dynamic process of scientific knowledge production and dissemination, track the frontiers of research hotspots and the development trends of emerging technologies in the field, and thus more comprehensively judge the development trend of scientific and technological innovation in the field.
[0003] Tracking the evolution of scientific and technological themes is a long and slow process. However, in the current big data era, with the increasing frequency of data interaction and the continuous complexity of the structure and relationship of scientific knowledge, it is difficult to fully grasp the context of theme evolution based on single-dimensional theme evolution research. Summary of the Invention
[0004] By providing a method and system for predicting the evolution stage of scientific and technological themes based on a multi-graph representation, the present invention solves the technical problem of too coarse analysis granularity in the research on the evolution of scientific and technological themes in the prior art, and realizes accurate prediction of the evolution process of scientific and technological themes.
[0005] The present invention provides a method for predicting the evolution stage of scientific and technological themes based on a multi-graph representation, including:
[0006] Obtaining the scientific and technological theme keywords in scientific and technological literature;
[0007] Obtaining the word frequency sequence of the scientific and technological theme keywords;
[0008] Through the formula Calculate to obtain the semantic distance a between the scientific and technological theme keyword and the scientific and technological theme keyword ij ; where is the word vector of the scientific and technological theme keyword within k years, is the word vector of the scientific and technological theme keyword within k years;
[0009] Construct the adjacency matrix of the semantic features S of the keywords within k years where n is the number of all scientific and technological theme keywords within k years;
[0010] Obtaining the citation relationship between each of the scientific and technological theme keywords, and obtaining the citation frequency between each pair of the scientific and technological theme keywords Among them, count_num is the keyword of the scientific and technological theme within k years The citation frequency of ;
[0011] Construct an adjacency matrix of the citation relationship of each scientific and technological theme keyword
[0012] Construct a multi-graph G=(V={V w ,X},E={E s ,E c}); where V w represents the set of scientific and technological theme keyword nodes, X represents the set of word frequency sequences of scientific and technological theme keyword nodes, in represents the word frequency sequence T corresponding to each scientific and technological theme keyword node, w represents the scientific and technological theme keyword node belonging to the set of multi-graph nodes, and E s represents the adjacency matrix The set of semantic distance edges between scientific and technological theme keywords in, and E c represents the adjacency matrix The set of citation relationship edges between scientific and technological theme keywords in;
[0013] Input the multi-graph G into the trained prediction model of the theme evolution stage, and use GCN to perform graph convolution operation on the adjacency matrix and obtain the modeled features of each scientific and technological theme keyword ; Use GAT to model the adjacency matrix based on the self-attention mechanism and obtain the modeled features of each scientific and technological theme keyword ; Concatenate and fuse the word frequency sequence of the scientific and technological theme keyword with the feature The feature to obtain the fused feature Input the feature into the MLP layer to predict the scientific and technological theme evolution stage, and obtain the final prediction result
[0014] Specifically, the obtaining of the scientific and technological theme keywords in the scientific and technological literature includes:
[0015] Combine the title and abstract corpus of the scientific and technological literature as the theme corpus set of the scientific and technological literature;
[0016] Query and match the preset keywords in the theme corpus set, and use the keywords obtained by the query and match as the scientific and technological theme keywords in the scientific and technological literature.
[0017] Specifically, obtaining the citation relationships between the scientific and technological theme keywords includes:
[0018] Obtaining the citation relationships between scientific and technological documents according to the detailed information of each scientific and technological document, and obtaining the citation relationships between the scientific and technological theme keywords through relationship mapping based on the citation relationships between the scientific and technological documents.
[0019] Specifically, after obtaining the frequency sequences of the scientific and technological theme keywords, it further includes:
[0020] Regarding the frequency sequences of the scientific and technological theme keywords as data points to calculate the shape distance to construct an undirected weighted graph, and obtaining the adjacency matrix A;
[0021] Normalizing the adjacency matrix A to obtain the similarity matrix W between graph vertices;
[0022] Adding the elements of each column of the similarity matrix W and placing them on the diagonal position to form a diagonal matrix, obtaining the weighted degree matrix D;
[0023] Obtaining the Laplacian matrix L according to the similarity matrix W and the weighted degree matrix D, and performing eigenvalue decomposition;
[0024] Taking the eigenvectors corresponding to the first λ minimum eigenvalues of the Laplacian matrix L to form the eigenmatrix H;
[0025] Clustering the eigenmatrix H to obtain the clustering labels of the corresponding frequency sequences;
[0026] Inputting the frequency sequences of the scientific and technological theme keywords after clustering labels into the to-be-trained prediction model for the theme evolution stage until convergence or reaching the number of iterations, obtaining the trained prediction model for the theme evolution stage.
[0027] Specifically, using cross-entropy loss to calculate the loss of scientific and technological theme evolution stage prediction, the cross-entropy loss is
[0028] The present invention also provides a scientific and technological theme evolution stage prediction system based on a multi-graph representation, including:
[0029] A scientific and technological theme keyword acquisition module, used to acquire scientific and technological theme keywords in scientific and technological documents;
[0030] A frequency sequence acquisition module, used to acquire the frequency sequences of the scientific and technological theme keywords;
[0031] A semantic distance calculation module, used to calculate through the formula to obtain the semantic distance a between the scientific and technological theme keyword and the scientific and technological theme keyword ij ; where, is the word vector of the scientific and technological theme keywords within k years of is the word vector of the scientific and technological theme keywords within k years of;
[0032] Semantic feature adjacency matrix construction module, used to construct the adjacency matrix of the semantic features S of the keywords within k years where, n is the number of all scientific and technological theme keywords within k years;
[0033] Citation frequency acquisition module, used to obtain the citation relationship between each of the scientific and technological theme keywords, and obtain the citation frequency between each pair of the scientific and technological theme keywords where, count_num is the scientific and technological theme keyword within k years to the citation frequency of;
[0034] Citation relationship adjacency matrix construction module, used to construct the adjacency matrix of the citation relationships of each scientific and technological theme keyword
[0035] Multigraph construction module, used to construct a multigraph G=(V={V w ,X},E={E s ,E c}); where, V w represents the set of scientific and technological theme keyword nodes, X represents the set of word frequency sequences of the scientific and technological theme keyword nodes, in represents the word frequency sequence T corresponding to each scientific and technological theme keyword node, w represents the scientific and technological theme keyword node belonging to the set of multigraph nodes, E s represents the adjacency matrix in which the set of semantic distance edges between each scientific and technological theme keyword, E c represents the adjacency matrix in which the set of citation relationship edges between each scientific and technological theme keyword;
[0036] Scientific and technological theme evolution stage prediction module, used to input the multigraph G into a trained theme evolution stage prediction model, and use GCN to perform graph convolution operation on the adjacency matrix and obtain the modeled features of each scientific and technological theme keyword Model the features Use GAT to model the adjacency matrix based on the self-attention mechanism and obtain the modeled features of each scientific and technological theme keyword Model the features Combine the word frequency sequence of the scientific and technological theme keyword with the features The described features are spliced and fused to obtain the fused features The described features are input into the MLP layer for prediction in the scientific and technological theme evolution stage to obtain the final prediction result
[0037] Specifically, the scientific and technological theme keyword acquisition module includes:
[0038] A theme corpus construction unit for combining the titles and abstract corpora of the scientific and technological documents as the theme corpus of the scientific and technological documents;
[0039] A scientific and technological theme keyword query unit for querying and matching preset keywords in the theme corpus, and taking the keywords obtained by query and match as the scientific and technological theme keywords in the scientific and technological documents.
[0040] Specifically, the citation frequency acquisition module includes:
[0041] A citation relationship acquisition unit for obtaining the citation relationships between the scientific and technological documents according to the detailed information of each scientific and technological document, and obtaining the citation relationships between the scientific and technological theme keywords through relationship mapping according to the citation relationships between the scientific and technological documents;
[0042] A citation frequency acquisition unit for obtaining the citation frequencies between each pair of the scientific and technological theme keywords according to the citation relationships between the scientific and technological theme keywords where count_num is the scientific and technological theme keyword within k years The citation frequency.
[0043] Specifically, it further includes:
[0044] An adjacency matrix A construction module for regarding the word frequency sequences of the scientific and technological theme keywords as data points to calculate the shape distance to construct an undirected weighted graph, and obtaining the adjacency matrix A;
[0045] A similarity matrix W acquisition module for normalizing the adjacency matrix A to obtain the similarity matrix W between the graph vertices;
[0046] A weighted degree matrix D acquisition module for adding the elements of each column of the similarity matrix W and placing them on the diagonal position to form a diagonal matrix, and obtaining the weighted degree matrix D;
[0047] An eigenvalue decomposition module for obtaining the Laplacian matrix L according to the similarity matrix W and the weighted degree matrix D, and performing eigenvalue decomposition;
[0048] A feature matrix H obtaining module, configured to form a feature matrix H by taking eigenvectors corresponding to the first λ minimum eigenvalues of the Laplacian matrix L;
[0049] A word frequency sequence marking module, configured to cluster the feature matrix H to obtain clustering labels of corresponding word frequency sequences;
[0050] A model training module, configured to input the word frequency sequences of the technological theme keywords after clustering labels into a to-be-trained prediction model for the theme evolution stage until convergence or reaching the number of iterations, to obtain the trained prediction model for the theme evolution stage.
[0051] Specifically, cross-entropy loss is used to calculate the loss of the prediction of the technological theme evolution stage, and the cross-entropy loss is
[0052] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0053] First, perform theme matching on the technological theme keywords in the scientific and technological literature, screen out the technological theme keywords that can represent the theme, then obtain the word frequency sequences of the technological theme keywords, and extract the semantic association relationship features between the technological theme keywords and the citation relationship features between the technological theme keywords. Secondly, construct a multi-graph network of technological theme keywords and model the technological theme keywords representing the theme. Subsequently, for features of different dimensions, different graph neural network models are used for feature extraction and modeling. Finally, a prediction model is used to predict the technological theme evolution stage, solving the technical problem of too coarse analysis granularity in the research on technological theme evolution in the prior art, and realizing accurate prediction of the technological theme evolution process. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of a method for predicting the technological theme evolution stage based on multi-graph representation provided by an embodiment of the present invention;
[0055] Figure 2 It is a flowchart of a method for predicting the technological theme evolution stage based on multi-graph representation provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of the working principle of a prediction model for the theme evolution stage in a method for predicting the technological theme evolution stage based on multi-graph representation provided by an embodiment of the present invention;
[0057] Figure 4 It is a module diagram of a system for predicting the technological theme evolution stage based on multi-graph representation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Embodiments of the present invention provide a method and system for predicting the evolution stage of scientific and technological themes based on multi-graph representation, solving the technical problem of too coarse analysis granularity in the research on the evolution of scientific and technological themes in the prior art, and achieving accurate prediction of the evolution process of scientific and technological themes.
[0059] The technical solution in the embodiments of the present invention to solve the above technical problems has the following general idea:
[0060] As Figure 1 shown, the prediction of the evolution stage of scientific and technological themes based on multi-graph in the embodiments of the present invention specifically includes steps such as scientific and technological theme extraction, scientific and technological theme stage discrimination, scientific and technological theme feature extraction, multi-graph modeling, and prediction of the evolution stage of scientific and technological themes. First, extract the scientific and technological theme keywords that can represent the theme, and use a supervised learning model to judge the trend of the scientific and technological theme keywords after theme extraction and assign a trend label. Secondly, carry out the work of extracting the features of scientific and technological themes. Specifically: calculate the word frequency of scientific and technological theme keywords to obtain the intensity features of scientific and technological themes, and obtain the word embeddings of scientific and technological theme keywords to extract the content features of scientific and technological themes, and map based on the literature citation relationship to obtain the citation relationship of scientific and technological theme keywords to extract the structural features of scientific and technological themes. After obtaining the features in three dimensions, use the scientific and technological theme nodes and the features in three dimensions to perform multi-graph modeling of scientific and technological themes to achieve the representation of multi-dimensional information. Finally, use a graph representation learning model and a regression model to predict the scientific and technological theme keywords extracted in the multi-graph that represent scientific and technological themes, and obtain the trend of its next-stage evolution, so as to achieve the prediction of the evolution stage of scientific and technological themes.
[0061] Specifically, in the data preparation stage, it is necessary to collect the literature in the subject field in the scientific literature database as research data. First, data cleaning is realized by removing abnormal data in the original scientific literature text data (the abnormal data situations include undefined and messy characters, null data values, etc.). Secondly, it is necessary to carry out the work of extracting scientific and technological themes. Specifically, for each piece of literature, combine the title and abstract corpus of the literature as the research theme corpus set of the literature, and then match the scientific and technological theme keywords under the literature to check whether these keywords appear in the research theme corpus set of the literature, and screen out the scientific and technological theme keywords that have appeared in the research theme corpus set to represent the scientific and technological themes in the literature of this field, and obtain the set M of scientific and technological theme keywords in each year k k 。
[0062] To make the subsequent modeling of semantic features more accurate, the obtained technology theme keywords are processed with camel case naming. For example, the technology theme keyword "Information System" is compressed and converted to "InformationSystem", and the technology theme keyword "Social Network Analysis" is converted to "SocialNetworkAnalysis", etc.
[0063] In the stage of extracting theme strength features, the embodiments of the present invention extract the word frequency sequence of technology theme keywords in multi-dimensional time as a main feature T of the technology theme. Specifically, the time window is set to 10, and the technology theme keywords within the specific year k are selected for collection The word frequency sequence T=(t k-9 ,t k-8 ,t k-7 ,…,t k-1 ,t k ) on the ten-year time span is used as the strength feature of the technology theme to obtain rich time series strength information. Collecting the word frequency information over a longer time span helps to comprehensively represent the strength feature of the technology theme, making the discrimination and prediction in the theme evolution stage more accurate.
[0064] When extracting the theme content features, the specific process is as follows:
[0065] (1) Obtain the semantic word vector of the technology theme
[0066] When obtaining the word vector, it is first necessary to prepare a corpus. Taking the technology theme words in 2019 as an example, which is consistent with the time window of 10 for the theme strength, the embodiments of the present invention form the corpus of 2019 by combining the titles and abstracts of all scientific and technological literatures from 2010 to 2019. Then, specific word embedding operations are performed to calculate the semantic expression vector of the technology theme keywords in the context globally It should be noted that the minimum word frequency statistic MIN_WORD is set to 1 here, aiming to obtain all the global semantic information of the technology themes in the corpus. The formula for obtaining the word vector is as follows:
[0067]
[0068] (2) Construct a semantic content adjacency matrix
[0069] After obtaining the word embedding vectors of each technology theme keyword, in order to construct a multi-graph, it is necessary to obtain the semantic distance between every two technology theme keywords. Therefore, the cosine similarity between the word vectors of technology theme keywords is calculated to reflect the strength of the semantic relationship between them, and then the semantic distance adjacency matrix of technology theme keywords within the unit year k is constructed.
[0070] The specific extraction process of the theme citation feature is as follows:
[0071] (1) Find the citation relationship between scientific and technological literatures
[0072] First, data processing is performed on scientific and technological literatures to extract the citation relationship between literatures under the established year k. For example, taking 2019 as an example, find which literatures in the previous two years k-1 and k-2 are cited by the literatures in that year k, so as to obtain the citation relationship between literatures in a 3-year time window.
[0073] (2) Extend the citation relationship to between technology theme keywords
[0074] After obtaining the citation relationship between literatures within k years, this citation relationship is extended and mapped to technology theme keywords. Specifically, for example, when literature D1 cites literature D2, all technology theme keywords in literature D1 also have a citation relationship with all technology theme keywords in literature D2. This mapping relationship belongs to a "keyword Cartesian product mapping". Finally, count the citation frequency between every two technology theme keywords, that is where count_num is the technology theme keyword under year k for the citation frequency.
[0075] (3) Construct the citation structure adjacency matrix
[0076] After obtaining the citation relationship between technology theme keywords, the adjacency matrix of the citation frequency between technology theme keywords can be constructed
[0077] The semantic content feature S and citation structure feature C of the above-obtained technology theme keywords are actually based on the multi-graph structure between technology theme keywords, and two associated relationship graphs are modeled respectively in the content feature dimension and the structure feature dimension.
[0078] Furthermore, construct a multi-graph G with technology theme keywords as nodes and semantic distance and citation frequency as edges.
[0079] After obtaining the multiplex graph G, based on the dataset with trend labels, a graph neural network is used to perform feature learning on graph features of different dimensions. Specifically, GCN is used to learn the citation structure features of science and technology topic keywords, and GAT is used to learn the semantic content features of science and technology topic keywords. The newly learned feature representations and word frequency features are concatenated for the prediction of the evolution stage.
[0080] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0081] As Figure 2 shown, the method for predicting the evolution stage of science and technology topics based on multiplex graph representation provided by the embodiment of the present invention includes:
[0082] Step S110: Obtain the science and technology topic keywords in scientific and technological literature;
[0083] Specifically explaining this step, obtaining the science and technology topic keywords in scientific and technological literature includes:
[0084] Combining the titles and abstract corpora of scientific and technological literature as the topic corpus set of scientific and technological literature;
[0085] Query and match preset keywords in the topic corpus set, and use the keywords obtained by query and match as the science and technology topic keywords in scientific and technological literature.
[0086] Step S120: Obtain the word frequency sequence of science and technology topic keywords; The word frequency sequence of science and technology topic keywords is used as the theme intensity to directly measure the degree of attention of science and technology topics, and is a key quantitative indicator reflecting the evolution stage of science and technology topics.
[0087] Step S130: Calculate the semantic distance a between the science and technology topic keyword and the science and technology topic keyword through the formula; where, ij ; among them, is the word vector of the science and technology topic keyword within k years, is the word vector of the science and technology topic keyword within k years;
[0088] Step S140: Construct the adjacency matrix of the semantic feature S of keywords within k years; where n is the number of all science and technology topic keywords within k years; The semantic distance relationship between science and technology topic keywords is used as the content feature, which reflects the correlation relationship of the change of the theme evolution stage among text carriers.
[0089] Step S150: Obtain the citation relationships between the keywords of each scientific and technological theme, and get the citation frequencies between every two keywords of each scientific and technological theme where count_num is the keyword of the scientific and technological theme within k years For the citation frequency;
[0090] Specifically, obtaining the citation relationships between the keywords of each scientific and technological theme includes:
[0091] Obtain the citation relationships between each scientific and technological literature according to the detailed information of each scientific and technological literature, and obtain the citation relationships between the keywords of each scientific and technological theme through relationship mapping based on the citation relationships between each scientific and technological literature. Taking the keyword citation frequency based on literature citation between scientific and technological themes as a structural feature can reflect the interaction relationship between scientific knowledge.
[0092] Taking the keywords of the scientific and technological theme in 2019 as an example. After screening to obtain the keywords of the scientific and technological theme in 2019, query the detailed information of the scientific and technological literature to which the keywords of the scientific and technological theme in 2019 belong, obtain the DOI of the cited scientific and technological literature, and then query whether the time of these cited scientific and technological literatures is within the data collected in this field (1) and whether it is within 2017 - 2019 (within 3 years) (2). After meeting (1) and (2), through the logic of "Keyword—ArticleCitation—Keyword", the citation relationships between the keywords of each scientific and technological theme can be obtained.
[0093] Step S160: Construct the adjacency matrix of the citation relationships of each scientific and technological theme keyword
[0094] Step S170: Construct a multigraph G=(V={V w ,X},E={E s ,E c}); where V w represents the set of scientific and technological theme keyword nodes, X represents the set of word frequency sequences of scientific and technological theme keyword nodes, in represents the word frequency sequence T corresponding to each scientific and technological theme keyword node, w represents the scientific and technological theme keyword node belonging to the set of multigraph nodes, and E s represents the set of semantic distance edges between the keywords of each scientific and technological theme in the adjacency matrix , and E c represents the set of citation relationship edges between the keywords of each scientific and technological theme in the adjacency matrix ;
[0095] Step S180: Input the multi-graph G into the trained prediction model for the theme evolution stage, and use GCN to perform graph convolution operation on the adjacency matrix and obtain the keyword of each technology theme The modeled features The specific formula is Use GAT to model the adjacency matrix based on the self-attention mechanism and obtain the keyword of each technology theme The modeled features The specific formula is Concatenate and fuse the word frequency sequence of the technology theme keyword with the feature Feature to obtain the fused feature The specific formula is where represents the word frequency sequence of the technology theme keyword; input the feature into the MLP layer to predict the technology theme evolution stage, and obtain the final prediction result That is Namely Such as Figure 3 shown
[0096] Specifically, use cross-entropy loss to calculate the loss of the technology theme evolution stage prediction, and the cross-entropy loss is
[0097] The following specifically describes the training process of the theme evolution stage prediction model. After obtaining the word frequency sequence of the technology theme keyword, it further includes:
[0098] Regard the word frequency sequence of each technology theme keyword as a data point to calculate the shape distance to construct an undirected weighted graph, and obtain the adjacency matrix A;
[0099] Specifically, use the word frequency sequence data of each technology theme keyword as vertices, and use the dynamic time warping distance between each word frequency sequence as the edge weight to construct the adjacency matrix A.
[0100] Normalize the adjacency matrix A to obtain the similarity matrix W between graph vertices;
[0101] Add the elements of each column of the similarity matrix W and place them on the diagonal position to form a diagonal matrix to obtain the weighted degree matrix D;
[0102] According to the similarity matrix W and the weighted degree matrix D, obtain the Laplacian matrix L and perform eigenvalue decomposition;
[0103] Take the eigenvectors corresponding to the first λ smallest eigenvalues of the Laplacian matrix L to form the eigenmatrix H;
[0104] In this embodiment, the method for determining λ is as follows:
[0105] Cluster the Fiedler vectors of the Laplacian matrix L, observe the variation relationship between the number of clusters k and the sum of squared errors of the clusters, and determine the value range of the number of clusters k through the elbow method;
[0106] Set λ to three groups of values: k, k - 1, and k - 2. On the basis of ensuring that the selected features can distinguish the differences between clusters, select the smaller value of λ.
[0107] Cluster the feature matrix H to obtain the cluster labels of the corresponding word frequency sequences;
[0108] Input the word frequency sequences of the scientific and technological theme keywords after clustering labels into the prediction model of the theme evolution stage to be trained until convergence or the number of iterations is reached, and obtain the trained prediction model of the theme evolution stage.
[0109] During the training process of the model, the embodiment of the present invention sets four hyperparameters: the hidden layer, the learning rate, epoch, and stopstep. The specific numerical settings are shown in Table 1.
[0110] Table 1 Hyperparameter settings for model training
[0111] LearningRate Epoch EarlyStop HiddenDim 1e-2 10 5 10 1e-3 50 10 30 1e-4 100 20 50 — 200 — —
[0112] As Figure 4 shown, the scientific and technological theme evolution stage prediction system based on the multi-graph representation provided by the embodiment of the present invention includes:
[0113] A scientific and technological theme keyword acquisition module 100, configured to acquire scientific and technological theme keywords in scientific and technological literatures;
[0114] Specifically, the scientific and technological theme keyword acquisition module 100 includes:
[0115] A theme corpus construction unit, configured to combine the titles and abstract corpora of scientific and technological literatures as the theme corpus of scientific and technological literatures;
[0116] A scientific and technological theme keyword query unit, configured to query and match preset keywords in the theme corpus, and use the keywords obtained by the query and match as the scientific and technological theme keywords in scientific and technological literatures.
[0117] A word frequency sequence acquisition module 200, configured to acquire the word frequency sequences of scientific and technological theme keywords; use the word frequency sequences of scientific and technological theme keywords as the theme intensity to directly measure the degree of attention of scientific and technological themes, which is a key quantitative index reflecting the evolution stage of scientific and technological themes.
[0118] A semantic distance calculation module 300, configured to use the formula Calculate the scientific and technological theme keywords and the scientific and technological theme keywords The semantic distance a between them ij ; among them, is the word vector of the scientific and technological theme keywords within k years of, is the word vector of the scientific and technological theme keywords within k years of;
[0119] The semantic feature adjacency matrix construction module 400 is used to construct the adjacency matrix of the semantic features S of the keywords within k years Among them, n is the number of all scientific and technological theme keywords within k years; taking the semantic distance relationship between scientific and technological theme keywords as the content feature reflects the correlation relationship of the changes in the theme evolution stage among text carriers
[0120] The citation frequency acquisition module 500 is used to obtain the citation relationship between each scientific and technological theme keyword and get the citation frequency between each pair of scientific and technological theme keywords Among them, count_num is the scientific and technological theme keyword within k years to the citation frequency of;
[0121] Specifically, the citation frequency acquisition module 500 includes:
[0122] The citation relationship acquisition unit is used to obtain the citation relationship between each scientific and technological literature according to the detailed information of each scientific and technological literature, and obtain the citation relationship between each scientific and technological theme keyword through relationship mapping based on the citation relationship between each scientific and technological literature. Taking the keyword citation frequency based on literature citation between scientific and technological theme keywords as the structural feature can reflect the interaction relationship between scientific knowledge
[0123] The citation frequency acquisition unit is used to obtain the citation frequency between each pair of scientific and technological theme keywords according to the citation relationship between each scientific and technological theme keyword Among them, count_num is the scientific and technological theme keyword within k years to the citation frequency of
[0124] The citation relationship adjacency matrix construction module 600 is used to construct the citation relationship adjacency matrix of each scientific and technological theme keyword
[0125] The multi-graph construction module 700 is used to construct the multi-graph G=(V={V w ,X},E={E s ,E c}); among them, V wDenote the set of nodes of keywords for scientific and technological themes, and X denote the set of frequency sequences of nodes of keywords for scientific and technological themes. in denote the frequency sequence T corresponding to each node of keyword for scientific and technological theme, w denote the node of keyword for scientific and technological theme belonging to the set of nodes of the multi-graph, and E s denote the adjacency matrix the set of edges of semantic distances between keywords for scientific and technological themes in, and E c denote the adjacency matrix the set of edges of citation relationships between keywords for scientific and technological themes in;
[0126] The scientific and technological theme evolution stage prediction module 800 is used to input the multi-graph G into the trained theme evolution stage prediction model, and use GCN to perform graph convolution operation on the adjacency matrix and obtain the features of each keyword for scientific and technological theme after modeling The specific formula is Use GAT to model the adjacency matrix based on the self-attention mechanism and obtain the features of each keyword for scientific and technological theme after modeling The specific formula is Concatenate and fuse the frequency sequence of the keyword for scientific and technological theme with the features features to obtain the fused features The specific formula is where denote the keyword for scientific and technological theme 's frequency sequence; input the features into the MLP layer to predict the scientific and technological theme evolution stage, and obtain the final prediction result that is
[0127] Specifically, use cross-entropy loss to calculate the loss of scientific and technological theme evolution stage prediction, and the cross-entropy loss is
[0128] The training process of the theme evolution stage prediction model is specifically described below, and it also includes:
[0129] The adjacency matrix A construction module is used to construct an undirected weighted graph by calculating the shape distance with the frequency sequences of each keyword for scientific and technological theme as data points, and obtain the adjacency matrix A;
[0130] Specifically, the adjacency matrix A construction module is specifically used to construct the adjacency matrix A with the frequency sequence data of each keyword for scientific and technological theme as vertices and the dynamic time warping distance between each frequency sequence as edge weights.
[0131] A similar matrix W obtaining module, configured to normalize the adjacency matrix A to obtain a similar matrix W between graph vertices;
[0132] A weighted degree matrix D obtaining module, configured to add the elements of each column of the similar matrix W and place them on the diagonal positions to form a diagonal matrix, thereby obtaining the weighted degree matrix D;
[0133] An eigenvalue decomposition module, configured to obtain a Laplacian matrix L based on the similar matrix W and the weighted degree matrix D, and perform eigenvalue decomposition;
[0134] An eigenmatrix H obtaining module, configured to take the eigenvectors corresponding to the first λ smallest eigenvalues of the Laplacian matrix L to form an eigenmatrix H;
[0135] In this embodiment, the method for determining λ is as follows:
[0136] Cluster the Fiedler vector of the Laplacian matrix L, observe the variation relationship between the number of clusters k and the sum of squared errors of the clustering, and determine the value range of the number of clusters k through the elbow method;
[0137] Set λ to three groups of values: k, k - 1, and k - 2, and select a smaller value of λ on the basis of ensuring that the selected features can distinguish the differences between clusters.
[0138] A word frequency sequence marking module, configured to cluster the eigenmatrix H to obtain clustering labels of the corresponding word frequency sequences;
[0139] A model training module, configured to input the word frequency sequences of the science and technology theme keywords after clustering labels into a to-be-trained theme evolution stage prediction model for training until convergence or reaching the number of iterations, thereby obtaining a trained theme evolution stage prediction model.
[0140] In summary, to solve the technical problems of too coarse analysis granularity and failure to fully utilize multi-dimensional evolution correlation information in the process of science and technology theme evolution in the prior art, the embodiment of the present invention proposes a new method for predicting the science and technology theme evolution stage, which conducts comprehensive analysis from multiple dimensions of theme strength, theme content, and theme structure, so as to achieve all-round and multi-level capture of theme evolution, explore the multi-feature interaction in the process of science and technology theme evolution in the discipline field, and will help to accurately track the theme evolution process, detect the stage of theme evolution, and predict the future evolution stage, etc.
[0141] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0142] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0145] Details not described in the embodiments of the present invention are all well-known technologies to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for predicting the evolution stage of scientific and technological themes based on multi-graph representation, characterized in that, Including: Obtain the scientific and technological theme keywords in the scientific and technological literature; Obtain the word frequency sequence of the scientific and technological theme keywords; The semantic distance a between the scientific and technological theme keywords is calculated through the formula and the scientific and technological theme keyword is calculated through the formula ; where ij ; among them, is the word vector of the scientific and technological theme keyword within k years, is the word vector of the scientific and technological theme keyword within k years; Construct the adjacency matrix of the semantic features S of keywords within k years where n is the number of all scientific and technological theme keywords within k years; Obtain the citation relationships between the scientific and technological theme keywords, and obtain the citation frequencies between every two of the scientific and technological theme keywords where count_num is the scientific and technological theme keyword within k years For the citation frequency; Construct the adjacency matrix of the citation relationships of the keywords for each technology theme Construct a multi - graph \(G=(V = \{V w , X\}, E=\{E s , E c \}); where \(V w represents the set of science and technology theme keyword nodes, \(X\) represents the set of word frequency sequence sets of science and technology theme keyword nodes, in represents the word frequency sequence \(T\) corresponding to each science and technology theme keyword node, \(w\) represents the science and technology theme keyword node belonging to the multi - graph node set, \(E s represents the adjacency matrix the set of semantic distance edges between science and technology theme keywords in it, \(E c represents the adjacency matrix the set of citation relationship edges between science and technology theme keywords in it; Input the multiplex graph G into the trained prediction model for the theme evolution stage, and use GCN to perform graph convolution operation on the adjacency matrix to obtain keywords of each technology theme the modeled features Use GAT to model the adjacency matrix based on the self-attention mechanism to obtain keywords of each technology theme the modeled features Concatenate and fuse the word frequency sequence of the technology theme keywords with the features the features to obtain the fused features Input the features into the MLP layer to predict the technology theme evolution stage, and obtain the final prediction result 2. The method for predicting the evolution stage of scientific and technological themes based on the multi-graph representation according to claim 1, wherein The obtaining of the scientific and technological theme keywords in the scientific and technological literature includes: Combine the title and abstract corpus of the scientific and technological literature as the theme corpus set of the scientific and technological literature; Query and match the preset keywords in the theme corpus set, and use the keywords obtained by the query and match as the scientific and technological theme keywords in the scientific and technological literature.
3. The method for predicting the evolution stage of scientific and technological themes based on the multi-graph representation according to claim 1, wherein, The obtaining of the citation relationship between each of the scientific and technological theme keywords includes: Obtain the citation relationship between each scientific and technological literature according to the detailed information of each scientific and technological literature, and obtain the citation relationship between each scientific and technological theme keyword through relationship mapping according to the citation relationship between each scientific and technological literature.
4. The method for predicting the evolution stage of scientific and technological themes based on the multi-graph representation according to claim 1, wherein After obtaining the word frequency sequence of the scientific and technological theme keywords, it further includes: Regard the word frequency sequences of the scientific and technological theme keywords as data points to calculate the shape distance to construct an undirected weighted graph, and obtain the adjacency matrix A; Normalize the adjacency matrix A to obtain the similarity matrix W between graph vertices; Add the elements of each column of the similarity matrix W and place them on the diagonal position to form a diagonal matrix to obtain the weighted degree matrix D; Obtain the Laplacian matrix L according to the similarity matrix W and the weighted degree matrix D, and perform eigenvalue decomposition; Take the eigenvectors corresponding to the first λ smallest eigenvalues of the Laplacian matrix L to form the eigenmatrix H; Cluster the eigenmatrix H to obtain the clustering labels of the corresponding word frequency sequences; Input the word frequency sequences of the scientific and technological theme keywords after clustering labels into the to-be-trained prediction model of the theme evolution stage for training until convergence or reaching the number of iterations, and obtain the trained prediction model of the theme evolution stage.
5. The method for predicting the evolution stage of a scientific and technological theme based on a multi-graph representation according to any one of claims 1-4, characterized in that, Use the cross-entropy loss to calculate the loss of the prediction of the evolution stage of the technology topic, and the cross-entropy loss is 6. A prediction system for the evolution stage of scientific and technological themes based on multi-graph representation, characterized in that, Including: A scientific and technological theme keyword acquisition module for obtaining the scientific and technological theme keywords in the scientific and technological literature; A word frequency sequence acquisition module for obtaining the word frequency sequence of the scientific and technological theme keywords; The semantic distance calculation module is used to calculate, through the formula the semantic distance a between the science and technology topic keyword and the science and technology topic keyword ; where ij is the word vector of the science and technology topic keyword within k years, is the word vector of the science and technology topic keyword within k years; A semantic feature adjacency matrix construction module for constructing an adjacency matrix of the semantic features S of keywords within k years where n is the number of all scientific and technological theme keywords within k years; A citation frequency acquisition module is used to obtain the citation relationships between the scientific and technological theme keywords, and obtain the citation frequencies between every two of the scientific and technological theme keywords where count_num is the scientific and technological theme keyword within k years For the citation frequency; A citation relationship adjacency matrix construction module for constructing an adjacency matrix of citation relationships of keywords for each scientific and technological theme A multi-graph construction module for constructing a multi-graph G=(V={V w ,X},E={E s ,E c}); where V w represents the set of technology theme keyword nodes, X represents the set of word frequency sequence sets of technology theme keyword nodes, in represents the word frequency sequence T corresponding to each technology theme keyword node, w represents the technology theme keyword node belonging to the multi-graph node set, E s represents the adjacency matrix the set of semantic distance edges between technology theme keywords in, E c represents the adjacency matrix the set of reference relationship edges between technology theme keywords in; The scientific and technological theme evolution stage prediction module is used to input the multiplex graph G into the trained theme evolution stage prediction model, and use GCN to perform graph convolution operation on the adjacency matrix and obtain the scientific and technological theme keywords The modeled features Use GAT to model the adjacency matrix based on the self-attention mechanism and obtain the scientific and technological theme keywords The modeled features Concatenate and fuse the word frequency sequence of the scientific and technological theme keywords with the features The features to obtain the fused features Input the features into the MLP layer to predict the scientific and technological theme evolution stage, and obtain the final prediction result 7. The technology theme evolution stage prediction system based on the multiple graph representation according to claim 6, wherein The scientific and technological theme keyword acquisition module includes: A theme corpus set construction unit for combining the title and abstract corpus of the scientific and technological literature as the theme corpus set of the scientific and technological literature; A scientific and technological theme keyword query unit for querying and matching the preset keywords in the theme corpus set, and using the keywords obtained by the query and match as the scientific and technological theme keywords in the scientific and technological literature.
8. The technology theme evolution stage prediction system based on the multiple graph representation according to claim 6, characterized in that The citation frequency acquisition module includes: A citation relationship acquisition unit for obtaining the citation relationship between each scientific and technological literature according to the detailed information of each scientific and technological literature, and obtaining the citation relationship between each scientific and technological theme keyword through relationship mapping according to the citation relationship between each scientific and technological literature; A citation frequency acquisition unit, configured to obtain the citation frequencies between any two of the technology theme keywords according to the citation relationships between the technology theme keywords where count_num is a technology theme keyword within k years For the citation frequency 9. The technology theme evolution stage prediction system based on the multi-graph representation according to claim 6, wherein It further includes: An adjacency matrix A construction module for regarding the word frequency sequences of the scientific and technological theme keywords as data points to calculate the shape distance to construct an undirected weighted graph, and obtaining the adjacency matrix A; A similarity matrix W obtaining module for normalizing the adjacency matrix A to obtain the similarity matrix W between graph vertices; A weighted degree matrix D obtaining module for adding the elements of each column of the similarity matrix W and placing them on the diagonal position to form a diagonal matrix to obtain the weighted degree matrix D; An eigenvalue decomposition module, configured to obtain a Laplacian matrix L based on the similarity matrix W and the weighted degree matrix D, and perform eigenvalue decomposition; An eigenmatrix H obtaining module, configured to form an eigenmatrix H by taking eigenvectors corresponding to the first λ minimum eigenvalues of the Laplacian matrix L; A word frequency sequence labeling module, configured to cluster the eigenmatrix H to obtain clustering labels of corresponding word frequency sequences; A model training module, configured to input the word frequency sequence of the scientific and technological theme keywords after clustering labels into a to-be-trained prediction model for the theme evolution stage until convergence or reaching the number of iterations, to obtain the trained prediction model for the theme evolution stage.
10. The prediction system for the evolution stage of scientific and technological themes based on the multi-graph representation according to any one of claims 6-9, characterized in that The cross-entropy loss is used to calculate the loss of the prediction of the evolution stage of the technology theme, and the cross-entropy loss is
Citation Information
Patent Citations
Method and system for obtaining scientific knowledge discovery
CN115841110A
Literature-based geoscience research hotspot extraction and visualization method and system
CN117708333A