Scientific and technological achievement analysis and prediction method and system based on big data

By constructing a dynamic knowledge flow semantic model and a causal enhanced spatiotemporal graph, the problem of identifying dynamic evolution patterns and causal relationships in the analysis of scientific and technological achievements is solved, thereby improving the accuracy of trend prediction for scientific and technological achievements.

CN121524942APending Publication Date: 2026-02-13NANJING DATA ASSOCIATION
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511706800.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing methods for analyzing scientific and technological achievements are insufficient to capture the dynamic evolution of scientific and technological concepts over time and lack accurate identification of semantic associations and causal relationships, resulting in predictions that lag behind the actual trends in scientific and technological development.

Method used

A dynamic knowledge flow semantic model is constructed using a temporal graph convolutional network. By building a graph of relationships between scientific and technological concepts, feature vectors of knowledge flow are extracted to form a causal enhanced spatiotemporal graph. Causal tracing analysis is then performed using a spatiotemporal graph neural network to generate a trend prediction report of scientific and technological achievements.

Benefits of technology

It enables dynamic capture of the semantic evolution patterns of scientific and technological achievements and accurate identification of causal relationships, thereby improving the accuracy of predicting the development trend of scientific and technological achievements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524942A_ABST
    Figure CN121524942A_ABST
Patent Text Reader

Abstract

The invention discloses a scientific and technological achievement analysis and prediction method and system based on big data, and relates to the technical field of machine learning and big data analysis, and the method comprises the steps: collecting and preprocessing multi-source scientific and technological achievement semantic data, and constructing a scientific and technological concept relation graph; the method comprises the following steps: performing training by taking a time sequence diagram convolutional network as a basic framework and taking a scientific and technological concept relation graph as a training sample, constructing a dynamic knowledge flow semantic model, performing evolution feature extraction on the scientific and technological concept relation graph by utilizing the dynamic knowledge flow semantic model, and outputting a knowledge flow feature vector; and inputting the causal enhanced space-time diagram into a space-time diagram neural network, aggregating semantic association and causal relationships among the scientific and technological achievements in a space dimension, capturing a dynamic change mode of scientific and technological achievement characteristics in a time dimension, and outputting a scientific and technological concept time sequence predicted value sequence. According to the method, the causal enhancement space-time diagram is constructed, so that trend deduction and causal traceability analysis are carried out for the time dimension, and the accuracy of scientific and technological achievement development trend prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning and big data analysis, and particularly relates to a scientific and technological achievement analysis and prediction method and system based on big data. BACKGROUND

[0002] With the continuous acceleration of scientific and technological innovation and knowledge production, a large amount of scientific and technological achievements are continuously generated in the form of patents, academic papers and scientific research projects. How to use big data technology to comprehensively analyze multi-source heterogeneous scientific and technological achievement information has become an important research direction of scientific and technological management and scientific research decision-making. The existing scientific and technological achievement analysis method usually relies on big data mining and natural language processing technology, and constructs a scientific and technological knowledge graph or a scientific research hotspot evolution model through text feature extraction, theme clustering and trend fitting, etc., to identify the scientific and technological frontiers and innovation activity patterns.

[0003] The traditional scientific and technological achievement analysis method focuses on static feature mining, and is difficult to capture the dynamic evolution law of scientific and technological concepts in time series, resulting in that the prediction result lags behind the real trend of scientific and technological development; the existing method often only relies on co-occurrence or similarity indicators when dealing with the complex semantic association and causal relationship between scientific and technological concepts, and lacks the ability of causal explanation and trend tracing of semantic evolution process, which limits the application depth in scientific and technological achievement evolution prediction and trend insight. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a scientific and technological achievement analysis and prediction method based on big data to solve the problems that the semantic evolution law of scientific and technological achievements is difficult to be dynamically captured and the causal association relationship is difficult to be accurately identified.

[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a big data-based scientific and technological achievement analysis and prediction method, which comprises collecting and pre-processing multi-source scientific and technological achievement semantic data, and constructing a scientific and technological concept relationship graph; taking a time series graph convolution network as a basic framework, training based on the scientific and technological concept relationship graph as a training sample, constructing a dynamic knowledge flow semantic model, and using the dynamic knowledge flow semantic model to extract evolution features of the scientific and technological concept relationship graph, and outputting a knowledge flow feature vector; based on the knowledge flow feature vector, establishing new connections between nodes of the scientific and technological concept relationship graph, and giving the connections a causal direction and weight according to the co-occurrence relationship and reference relationship between the nodes, to form a causal enhanced space-time graph; inputting the causal enhanced space-time graph into a space-time graph neural network, aggregating semantic associations and causal relationships between scientific and technological achievements in the spatial dimension, and capturing dynamic change patterns of scientific and technological achievement features in the time dimension, and outputting a scientific and technological concept time series prediction value sequence; based on the causal enhanced space-time graph, performing causal tracing analysis on the scientific and technological concept time series prediction value sequence, identifying scientific and technological concepts and causal relationships that cause trend changes, and generating a scientific and technological achievement trend prediction report.

[0007] As a preferred scheme of the big data-based scientific and technological achievement analysis and prediction method, the multi-source scientific and technological achievement semantic data comprises patent data, academic paper data and scientific research project data. The pre-processing comprises data cleaning, entity recognition and text normalization.

[0008] As a preferred scheme of the big data-based scientific and technological achievement analysis and prediction method, the construction of the scientific and technological concept relationship graph comprises the following steps. Performing co-occurrence statistical analysis on the pre-processed multi-source scientific and technological achievement semantic data, calculating co-occurrence probabilities between scientific and technological concepts, and generating candidate scientific and technological concept association pairs; Determining semantic relationship types of the candidate scientific and technological concept association pairs according to a preset semantic rule library, and converting them into scientific and technological concept triples; Performing vectorization representation processing on the scientific and technological concept triples, and generating a scientific and technological concept relationship vector set; Taking the scientific and technological concepts as nodes and the semantic relationships as edges, and combining them with the scientific and technological concept relationship vector set, a scientific and technological concept relationship graph is constructed.

[0009] As a preferred scheme of the big data-based scientific and technological achievement analysis and prediction method, the construction of the dynamic knowledge flow semantic model comprises the following steps. Extracting scientific and technological concept node features, semantic relationship structures and time attributes from the scientific and technological concept relationship graph, and generating a feature mapping matrix; Inputting the feature mapping matrix into the time series graph convolution network, initializing parameters of the time series graph convolution network and performing multiple rounds of training, and obtaining a space-time feature parameter set; The time-space feature parameter set is used to optimize the time sequence diagram convolution network structure, and a dynamic knowledge flow semantic model is formed.

[0010] As a preferred scheme of the scientific and technological achievement analysis and prediction method based on big data, the output knowledge flow feature vector is outputted by the following steps, The scientific and technological concept relationship graph is inputted into the dynamic knowledge flow semantic model, the change trend and migration path features of the scientific and technological concept nodes in the time dimension are extracted, and a node-level evolution feature set is generated; The semantic evolution correlation degree between the scientific and technological concepts is calculated according to the node-level evolution feature set, the evolution information of each scientific and technological concept node is aggregated, and the knowledge flow feature vector is outputted.

[0011] As a preferred scheme of the scientific and technological achievement analysis and prediction method based on big data, the following steps are taken to form the causal enhanced time-space graph, Based on the knowledge flow feature vector, the feature matching analysis of the scientific and technological concept nodes in the scientific and technological concept relationship graph is performed, and the potential semantic contact node pairs are identified; New connection edges are established between the potential semantic contact node pairs, and the co-occurrence frequency and reference strength of each new connection edge are calculated to generate a node association information set; The causal direction of the connection edge is determined according to the time attribute and semantic relationship structure of the scientific and technological concept nodes in the node association information set, and the causal weight is obtained through a causal probability calculation method; The new connection edges and the corresponding causal weights are fused with the scientific and technological concept relationship graph to form the causal enhanced time-space graph.

[0012] As a preferred scheme of the scientific and technological achievement analysis and prediction method based on big data, the following steps are taken to output the scientific and technological concept time sequence prediction value sequence, The causal enhanced time-space graph is inputted into the time-space graph neural network, the semantic relationship and causal features between the scientific and technological concept nodes are calculated by the spatial convolution operation, and the spatial dependence feature vector is obtained; The spatial dependence feature vector is extracted by the time convolution operation, the dynamic evolution mode of each scientific and technological concept node feature is analyzed, and the time sequence node state vector is outputted; The time sequence node state vector is mapped to the prediction target dimension through the full connection layer of the time-space graph neural network, and the scientific and technological concept time sequence prediction value sequence is obtained.

[0013] As a preferred scheme of the scientific and technological achievement analysis and prediction method based on big data, the following steps are taken to perform causal tracing analysis on the scientific and technological concept time sequence prediction value sequence based on the causal enhanced time-space graph, Nodes of scientific and technological concepts that exceed a preset threshold for change in time series prediction values ​​are selected from the scientific and technological concept time series and marked as key scientific and technological concept nodes. Extract all upstream causal edges and source nodes connected to key technology concept nodes from the causal enhanced spatiotemporal graph to form a local causal origin subgraph; The upstream causal path is identified by performing a reverse path search on the local causal origination subgraph using a graph traversal algorithm.

[0014] As a preferred embodiment of the big data-based scientific and technological achievement analysis and prediction method of the present invention, the steps for generating the scientific and technological achievement trend prediction report are as follows: Calculate the causal contribution of each upstream causal path, filter out upstream causal paths that exceed the preset contribution threshold, and generate a core causal path list. The key scientific and technological concept nodes, core causal path lists, and corresponding causal contribution values ​​are organized in natural language to generate a scientific and technological achievement trend prediction report.

[0015] Secondly, this invention provides a big data-based system for analyzing and predicting scientific and technological achievements, comprising: The data acquisition module is used to collect semantic data of multi-source scientific and technological achievements, perform preprocessing, and construct a scientific and technological concept relationship diagram. The model building module is used to build a dynamic knowledge flow semantic model based on a temporal graph convolutional network and trained with a technology concept relationship graph as training samples. The dynamic knowledge flow semantic model is then used to extract evolutionary features from the technology concept relationship graph and output a knowledge flow feature vector. The causal graph construction module is used to establish new connections between nodes in the science and technology concept relationship graph based on knowledge flow feature vectors, and to assign causal direction and weight to the connections according to the co-occurrence relationship and reference relationship between nodes, forming a causal enhanced spatiotemporal graph; The trend prediction module is used to input the causal enhanced spatiotemporal graph into the spatiotemporal graph neural network, aggregate the semantic associations and causal relationships between scientific and technological achievements in the spatial dimension, capture the dynamic change patterns of the characteristics of scientific and technological achievements in the time dimension, and output the time-series prediction value sequence of scientific and technological concepts. The causal tracing module is used to perform causal tracing analysis on the time-series predicted value sequence of scientific and technological concepts based on the causal enhanced spatiotemporal graph, identify the scientific and technological concepts and causal relationships that lead to trend changes, and generate a scientific and technological achievement trend prediction report.

[0016] The application has the beneficial effects that: by taking the time sequence diagram convolution network as a basic framework, training by taking the science and technology concept relationship diagram as a training sample, constructing a dynamic knowledge flow semantic model, the dynamic evolution law of the science and technology achievement semantic change over time is captured, and the potential semantic flow and migration path among the science and technology concepts are mined; by constructing a causal enhancement space-time graph, trend deduction and causal tracing analysis are realized in the time dimension, and the accuracy of the science and technology achievement development trend prediction is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Fig. 1 The flowchart of the science and technology achievement analysis and prediction method based on big data.

[0019] Fig. 2 The schematic diagram of the science and technology achievement analysis and prediction system based on big data.

[0020] Fig. 3 The flowchart of the dynamic knowledge flow semantic model construction.

[0021] Fig. 4 The flowchart of the formation of the causal enhancement space-time graph. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below with reference to the drawings of the specification.

[0023] In the following description, many specific details are set forth in order to provide a thorough understanding of the application, but the application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the application, therefore the application is not limited to the specific embodiments disclosed below.

[0024] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or selective embodiment that excludes other embodiments.

[0025] REFERENCE Figs. 1-4 For one embodiment of the application, the embodiment provides a science and technology achievement analysis and prediction method based on big data, comprising the following steps: S1, collect multi-source scientific and technological achievement semantic data and preprocess, construct scientific and technological concept relation graph.

[0026] The multi-source scientific and technological achievement semantic data includes patent data, academic paper data and scientific research project data.

[0027] It should be noted that the patent data refers to the text data of the claim, specification and abstract containing the detailed description of the technical solution extracted from the patent documents disclosed by the patent offices of various countries, which is obtained by searching and downloading in batches through the patent database; the academic paper data refers to the text data of the title, abstract and full text containing the research background, method, result and conclusion extracted from academic journals and conference papers, which is obtained through the API interface of the academic database (such as CNKI); the scientific research project data refers to the text data containing the research target, technical route and innovation point extracted from the project report, interim check report and final report disclosed by the government funding agencies and enterprise R&D departments, which is obtained through information scraping of the official project publicity platform.

[0028] The preprocessing includes data cleaning, entity recognition and text normalization.

[0029] It should be noted that data cleaning refers to denoising, deduplication and format standardization processing of redundant fields, repeated records and abnormal characters in multi-source scientific and technological achievement semantic data to ensure data integrity and consistency; entity recognition refers to the positioning and classification of scientific and technological concept entities (such as technical terms, methods and materials) in multi-source scientific and technological achievement semantic data after data cleaning, which converts multi-source scientific and technological achievement semantic data into structured entities with clear type labels; text normalization refers to word shape restoration, sentence segmentation, word segmentation and stop word filtering of multi-source scientific and technological achievement semantic data from different sources, eliminating language differences and expression redundancy.

[0030] The co-occurrence statistical analysis is performed on the preprocessed multi-source scientific and technological achievement semantic data, and the co-occurrence probability between scientific and technological concepts is calculated to generate candidate scientific and technological concept association pairs.

[0031] Further, the preprocessed multi-source scientific and technological achievement semantic data is traversed, and two scientific and technological concepts appearing in the same sentence, the same paragraph and the same document are counted as one co-occurrence event, and the co-occurrence times of each pair of scientific and technological concepts are accumulated; the co-occurrence times of each pair of scientific and technological concepts are normalized with the total number of occurrences of the corresponding scientific and technological concepts in the multi-source scientific and technological achievement semantic data, and the co-occurrence probability of each pair of scientific and technological concepts is calculated; scientific and technological concept pairs with a co-occurrence probability greater than a preset co-occurrence probability threshold are selected, and the corresponding scientific and technological concept identifiers and co-occurrence probabilities are recorded to generate candidate scientific and technological concept association pairs.

[0032] It should be noted that the co-occurrence probability threshold is set according to the average co-occurrence probability distribution and standard deviation statistics of the scientific and technological concept in the multi-source scientific and technological achievement semantic data, and is used to screen the limited value of the semantic association strength between scientific and technological concepts. The exemplary value range is 0.2-0.5. Below 0.2 will cause a large number of low correlation concepts to be misjudged as associated, causing noise interference and semantic confusion, and above 0.5 can only retain a few high-frequency co-occurrence concept pairs, resulting in insufficient semantic association coverage.

[0033] According to the preset semantic rule library, the semantic relationship type of the candidate scientific and technological concept association pair is determined, and is converted into a scientific and technological concept triple.

[0034] Further, the preset semantic rule library is called to perform dependency analysis on the syntactic structure containing the candidate scientific and technological concept association pair in the preprocessed multi-source scientific and technological achievement semantic data: the semantic role labeling is performed on the verb phrase, preposition phrase and functional word, the corresponding semantic relationship type is matched for each candidate scientific and technological concept association pair according to the pre-defined mode of the semantic rule library about “improvement relationship”, “reference relationship”, “dependence relationship” and “containment relationship”; when the matching is successful, the former scientific and technological concept in the candidate scientific and technological concept association pair is taken as the starting concept, the latter scientific and technological concept is taken as the target concept, and the matched semantic relationship type is taken as the relationship identifier, and is assembled into a scientific and technological concept triple.

[0035] It should be noted that the semantic rule library is a rule set for determining the semantic relationship type of the candidate scientific and technological concept association pair, which is based on the typical sentence patterns, dependency syntax structures and semantic collocation patterns in large-scale patent texts, academic papers and scientific research project documents that have been labeled, and is formed by sorting and inducing the rule entries of the corresponding trigger words, syntactic templates and concept position patterns of different semantic relationships.

[0036] The scientific and technological concept triple is processed by vectorization representation to generate a set of scientific and technological concept relationship vectors.

[0037] Further, the starting scientific and technological concept, the semantic relationship type and the target scientific and technological concept in each scientific and technological concept triple are encoded in a fixed order, and word segmentation and word frequency statistics operations are performed, and the words with a frequency higher than a preset frequency threshold are retained and converted into numerical indexes; according to the index position, an equal-length numerical vector is generated for each starting scientific and technological concept, semantic relationship type and target scientific and technological concept, and is spliced into a single scientific and technological concept triple vector by vector splicing; all single scientific and technological concept triple vectors are sorted according to the numerical index to generate a set of scientific and technological concept relationship vectors.

[0038] It should be noted that the preset frequency threshold is a limited value for screening low-frequency invalid words, which is set according to the word frequency distribution statistical results and quantile analysis of all words in the preprocessed multi-source scientific and technological achievement semantic data. The exemplary value range is 3-10. When the value is higher than 10, some low-frequency but domain characteristic words will be mistakenly deleted, resulting in semantic feature loss. When the value is lower than 3, a large number of noise words are retained, causing vector dimension redundancy and reducing the subsequent feature calculation accuracy.

[0039] The scientific and technological concept is taken as a node, the semantic relationship is taken as an edge, and the scientific and technological concept relationship vector set is combined to construct a scientific and technological concept relationship graph.

[0040] Further, each starting scientific and technological concept and target scientific and technological concept in the scientific and technological concept triple is created as a graph node, and the semantic relationship type is taken as the edge type connecting the starting scientific and technological concept node and the target scientific and technological concept node. Meanwhile, according to the numerical value vector corresponding to each scientific and technological concept triple in the scientific and technological concept relationship vector set, the numerical value information in the numerical value vector is hung on the corresponding edge as an edge attribute, which is used to represent the semantic association strength and relationship characteristics between scientific and technological concepts. All scientific and technological concept nodes and semantic relationship edges with attributes are connected according to the triple correspondence to form a scientific and technological concept relationship graph.

[0041] S2, based on the time sequence graph convolution network as the infrastructure, the scientific and technological concept relationship graph is taken as the training sample for training, a dynamic knowledge flow semantic model is constructed, and the dynamic knowledge flow semantic model is used for evolution feature extraction of the scientific and technological concept relationship graph, and a knowledge flow feature vector is output.

[0042] The scientific and technological concept node features, semantic relationship structure and time attribute are extracted from the scientific and technological concept relationship graph to generate a feature mapping matrix.

[0043] Further, the connection number of each scientific and technological concept node in the scientific and technological concept relationship graph, the number of times of being cited by other scientific and technological concept nodes, the number of times of citing other scientific and technological concept nodes, and the total number of adjacent nodes are counted, and the average semantic association strength of the scientific and technological concept node is obtained by calculating the semantic relationship weight average value between each scientific and technological concept node and the connected nodes. The connection number of the scientific and technological concept node, the number of times of being cited, the number of times of citing, the number of adjacent nodes and the average semantic association strength are integrated to generate a node structure feature sequence. The connection direction, relationship type and appearance time label of each semantic relationship edge are read, and the values in the same time period in the node structure feature sequence are aggregated and arranged. The values are numerically encoded according to the unified field order. The encoding results are arranged in the order of the unique identification of the scientific and technological concept node to generate a feature mapping matrix.

[0044] The feature mapping matrix is input into the time sequence graph convolution network, the time sequence graph convolution network parameters are initialized, and multiple rounds of training are performed to obtain a set of spatiotemporal feature parameters.

[0045] Further, an adjacency matrix is generated according to the connection relationship of the technology concept nodes in the technology concept relationship graph, and the node adjacency weight is initialized according to the weight distribution of the semantic relationship edge in the adjacency matrix; the time step is set according to the interval of the appearance time label in the technology concept relationship graph, and the convolution kernel size is set in combination with the dimension of the feature mapping matrix; the feature mapping matrix is divided into continuous time segments in the order of the appearance time label, and is input into the network in turn for training, in each round of training, the weighted aggregation value of the features of adjacent technology concept nodes is calculated through the spatial convolution layer, and the continuity of the feature change between different time segments is calculated through the time convolution layer; the node adjacency weight, the time step and the convolution kernel parameters are iteratively optimized and updated through the back propagation algorithm until the error converges, and the spatiotemporal feature parameter set is output.

[0046] The spatiotemporal feature parameter set is used to optimize the structure of the time series graph convolution network to form a dynamic knowledge flow semantic model.

[0047] Further, according to the spatiotemporal feature parameter set, the node adjacency weight, the convolution kernel size and the time step parameters in the time series graph convolution network are adjusted correspondingly; by calculating the stability index and the gradient change rate of each parameter in the spatiotemporal feature parameter set, the network layer with slow convergence in the training process is identified and the learning rate is adjusted correspondingly; in the spatial dimension, the connection range of the convolution layer is reconstructed according to the distribution of the node adjacency weight to enhance the feature capture ability of the local high correlation nodes, and in the time dimension, the sampling interval of the time convolution layer is adjusted according to the dynamic change trend of the time step to improve the sensitivity to long-term semantic changes; after the structure optimization is completed, the modified time series graph convolution network is executed for verification operation to confirm that the change trend of the output feature in the spatial and time dimensions remains stable, and the dynamic knowledge flow semantic model is output.

[0048] The technology concept relationship graph is input into the dynamic knowledge flow semantic model to extract the change trend and migration path feature of the technology concept nodes in the time dimension to generate a node-level evolution feature set.

[0049] Further, the technology concept relationship graph is input into the dynamic knowledge flow semantic model, the feature mapping matrix values between each technology concept node and the adjacent nodes are calculated in the spatial convolution layer, the feature values of the adjacent nodes are weighted and summed through the node adjacency weight to obtain the spatial aggregation result of each technology concept node; in the time convolution layer, the spatial aggregation results of the continuous time segments are convolved with the time step as a sliding window to calculate the direction, amplitude and duration period of the value change in the technology concept node feature mapping matrix; the spatial aggregation results and the time change trend of each technology concept node in all time segments are jointly encoded to generate a node feature record containing the technology concept node identifier, the time index, the feature change amount and the migration direction; all node feature records are integrated in chronological order to generate a node-level evolution feature set.

[0050] According to the node-level evolution feature set, the semantic evolution correlation degree between the science and technology concepts is calculated, the evolution information of each science and technology concept node is aggregated, and a knowledge flow feature vector is output.

[0051] Further, the science and technology concept nodes in the node-level evolution feature set are paired two by two, and the feature change sequence of each pair of science and technology concept nodes in the time dimension is extracted. At the same time, the root mean square of the difference value of the feature change sequence of each pair of science and technology concept nodes in the same time window is calculated in a sliding time window unit, and the correlation index between the time sequences is obtained. The proportion of the consistent change direction of each pair of science and technology concept nodes in the adjacent time window is calculated, and the correlation index and the proportion of the consistent change direction are combined into a semantic evolution correlation degree value by weighted average. According to the semantic evolution correlation degree value, a semantic evolution correlation matrix is constructed according to the numbering order of the science and technology concept nodes, and the semantic evolution correlation degree value is normalized by minimum-maximum to standardize the numerical range to the interval [0, 1]. Based on the row weighted sum of each science and technology concept node in the semantic evolution correlation matrix, the overall semantic evolution strength of each science and technology concept node and all other science and technology concept nodes is calculated, and the time change trend quantity in the node-level evolution feature set is weighted and fused to generate a node comprehensive evolution feature. The node comprehensive evolution features of all science and technology concept nodes are arranged and spliced according to the numbering order of the science and technology concept nodes, and a knowledge flow feature vector is output.

[0052] It should be noted that when calculating the feature change sequence of the science and technology concept node in the time dimension, the sliding time window of the time convolution layer is used as a unit to perform principal component analysis on the spatial aggregation results of the science and technology concept node in the window. The projection value in the first principal component direction is extracted as the feature change value, and the feature change sequence of the science and technology concept node in the time dimension is integrated and obtained. The first principal component is used to represent the dominant trend of the numerical change in the window, so that the feature change value can reflect the change direction and the change amplitude at the same time, thereby ensuring the consistency and comparability of the feature change value of each science and technology concept node in the time dimension in the calculation of the semantic evolution correlation degree.

[0053] The expression for calculating the semantic evolution correlation degree is: ; Wherein, represents the semantic evolution correlation degree between the science and technology concept node and the science and technology concept node ; represents the weighted coefficient of the feature change difference item; represents the weighted coefficient of the change direction consistency item; represents the total number of time windows for analysis; represents the science and technology concept node at time step Characteristic change values; Indicates at time step At that time, the concept node of science and technology Characteristic change values; This indicates the technology concept node across all time windows. With technology concept nodes Number of time windows with consistent direction of change and total number of time windows The ratio; and It is an index of technology concept nodes; It is a time step index.

[0054] It should be noted that, and This is achieved by calculating the correlation coefficient between the root mean square difference of feature changes and the proportion of changes in the same direction for all pairs of technology concept nodes in the node-level evolution feature set, and the fitting result of the semantic evolution trend. The correlation coefficient is then subjected to min-max normalization, mapping the values ​​to the [0,1] interval. Finally, an exponential normalization function is used to normalize the result, ensuring that the two weighted coefficients satisfy... =1, which is used to reflect the relative importance of the feature change difference term and the direction consistency term in the semantic evolution correlation degree calculation.

[0055] S3. Based on the knowledge flow feature vector, new connections are established between nodes in the science and technology concept relationship graph, and the causal direction and weight of the connections are assigned according to the co-occurrence relationship and reference relationship between nodes, forming a causal enhanced spatiotemporal graph.

[0056] Based on knowledge flow feature vectors, feature matching analysis is performed on the nodes of science and technology concepts in the science and technology concept relationship graph to identify potential semantic connection node pairs.

[0057] Furthermore, the feature vectors of the knowledge flow are matched with the feature mapping matrix values ​​corresponding to each technology concept node in the technology concept relationship graph. By calculating the cosine similarity and temporal evolution direction similarity of the feature vectors between different technology concept nodes, the semantic similarity index of each technology concept node pair is obtained. The cosine similarity and temporal evolution direction similarity of the feature vectors are fused into a semantic comprehensive similarity by weighted averaging, and compared with a preset semantic comprehensive similarity threshold. When the semantic comprehensive similarity is greater than the semantic comprehensive similarity threshold, it is determined that the technology concept node pair has a potential semantic connection. Technology concept node pairs that meet the semantic comprehensive similarity threshold condition are recorded as potential semantic connection node pairs, and a list of potential semantic connection node pairs is output.

[0058] It should be noted that the semantic comprehensive similarity threshold is set based on the mean and standard deviation statistics of the semantic comprehensive similarity distribution of all pairs of scientific concept nodes, and is used to filter the limited value of the semantic connection strength between scientific concept nodes. The exemplary value range is 0.6-0.8. If it is lower than 0.6, it is easy to misjudge the nodes with low semantic correlation degree as potential contact node pairs, resulting in an increase in semantic noise. If it is higher than 0.8, only a small number of high-correlation node pairs are retained, which is easy to cause insufficient semantic coverage.

[0059] New connection edges are established between the potential semantic contact node pairs, and the co-occurrence frequency and citation strength of each new connection edge are calculated to generate a node association information set.

[0060] Further, according to the list of potential semantic contact node pairs, new connection edges are created for each pair of scientific concept nodes in the scientific concept relationship graph; the number of literature articles and the time distribution of the potential semantic contact node pairs appearing in the preprocessed multi-source scientific achievement semantic data are counted, the co-occurrence frequency is calculated, the number of citation events related to the node pair is extracted, the ratio of the number of citations to the number of citation sources is calculated to obtain the citation strength; the scientific concept node identifier, co-occurrence frequency, citation strength and appearance time label corresponding to each new connection edge are combined and encoded to form a structured record; the structured records are integrated in order of scientific concept node identifier to generate a node association information set.

[0061] The causal direction of the connection edge is determined according to the time attribute and semantic relationship structure of the scientific concept nodes in the node association information set, and the causal weight is obtained through a causal probability calculation method.

[0062] Further, according to the time attribute in the node association information set, the order of two scientific concept nodes in the time dimension is determined. When the appearance time of scientific concept node A is earlier than that of scientific concept node B, and the co-occurrence frequency of the two nodes shows a monotonous increasing trend with time, the connection direction is set to point from scientific concept node A to scientific concept node B; when the time distribution of scientific concept node A and scientific concept node B alternates, the causal order is determined according to the difference direction of the citation strength of the two; the co-occurrence frequency, citation strength and time difference of each connection edge in the node association information set are normalized, and the causal probability value of scientific concept node A to scientific concept node B is calculated as the causal weight using the logistic regression method, which is recorded in the corresponding connection edge attribute.

[0063] The expression for calculating the causal weight is: ; Wherein, represents the causal weight from scientific concept node to scientific concept node ; and respectively represent two different technology concept node identities in the causality enhanced spatio-temporal graph; represent a technology concept node co-occurrence frequency of the technology concept node ; represent a technology concept node ; represent a technology concept node ; represent a technology concept node ; is a weighting coefficient of the co-occurrence frequency feature quantity; is a weighting coefficient of the reference strength feature quantity; is a weighting coefficient of the time difference feature quantity; is the base of the natural logarithm.

[0064] It should be noted that , and are weighting coefficients obtained by calculating the Pearson correlation coefficients of the co-occurrence frequency, reference strength and time difference of all connection edges in the node association information set and the consistency of the causal direction, and are normalized by proportion to obtain the weight ratio, which satisfies , used to reflect the relative importance of the three feature quantities in the calculation of the causal weight.

[0065] The new connection edge and the corresponding causal weight are fused with the technology concept relationship graph to form the causality enhanced spatio-temporal graph.

[0066] Further, the new connection edge is indexed and matched according to the node identity, the matched new connection edge is inserted into the gap between the corresponding technology concept nodes, when the direction of the original semantic relationship edge is consistent with that of the new connection edge, the original semantic relationship edge is retained and the edge weight is updated to the weighted average value of the original semantic relationship edge and the new connection edge, when the directions are opposite, the connection edge corresponding to the causal direction is retained and the causal weight is covered, and the causality enhanced spatio-temporal graph containing the technology concept node, the semantic relationship edge, the causal direction and the causal weight is output.

[0067] S4, input the causality enhanced spatio-temporal graph into the spatio-temporal graph neural network, aggregate the semantic association and causal relationship between the technology achievements in the spatial dimension, and capture the dynamic change pattern of the technology achievement characteristics in the time dimension, and output the technology concept time sequence prediction value sequence.

[0068] It should be explained that the spatio-temporal graph neural network is pre-trained, and in the specific operation, the causal enhanced spatio-temporal graph and the corresponding historical trend reference sequence are divided into a training set and a validation set according to a fixed ratio (8:2); in the training set, the feature vector of each science and technology concept node in the causal enhanced spatio-temporal graph is subjected to standard normalization operation, and the characteristic value is scaled to the interval [0, 1] to ensure that the dimensions of different dimensions are consistent; the normalized causal enhanced spatio-temporal graph is input into the spatio-temporal graph neural network, the error between the predicted value output by the spatio-temporal graph neural network and the historical trend reference sequence is calculated through the AdamW optimization algorithm and the smooth L1 loss function, and the spatial convolution layer and the time convolution layer parameters of the spatio-temporal graph neural network are updated layer by layer by using the back propagation mechanism; forward inference is regularly performed on the validation set, the validation error change is monitored, when the validation error tends to converge and the decline amplitude is no longer significant in continuous multiple iterations, the training is terminated, and the trained spatio-temporal graph neural network is output.

[0069] The causal enhanced spatio-temporal graph is input into the spatio-temporal graph neural network, the semantic relationship and the causal feature between the science and technology concept nodes are calculated by neighborhood aggregation through spatial convolution operation, and the spatial dependence feature vector is obtained.

[0070] Further, the causal enhanced spatio-temporal graph is input into the spatio-temporal graph neural network, and the adjacency matrix is constructed according to the node connection relationship in the causal enhanced spatio-temporal graph, and the knowledge flow feature vector of each science and technology concept node is taken as the input feature; in the spatial convolution layer, the adjacent nodes of each science and technology concept node are weighted and aggregated by taking the adjacency matrix as the weight template, the semantic relationship strength and the causal weight between the science and technology concept nodes are jointly applied to the aggregation function, and the local feature vector containing the interaction information of multiple nodes is obtained; the nonlinear activation and normalization processing are performed on the local feature vector to suppress the influence of abnormal feature values and enhance the stability of feature expression, and the spatial dependence feature vector containing the local feature dependence information of all science and technology concept nodes in the spatial dimension is output.

[0071] The spatial dependence feature vector is subjected to time series feature extraction through time convolution operation, the dynamic evolution mode of the feature of each science and technology concept node is analyzed, and the time series node state vector is output.

[0072] Further, the space-dependent feature vector is input into the spatio-temporal graph neural network, in the time convolution layer, indexed by the causally enhanced time label recorded in the spatio-temporal graph, and the sliding convolution calculation is performed in time sequence, in each time window, the change gradient of the technology concept node feature between adjacent time segments is calculated, and the change direction, change amplitude and change period features are extracted; the feature change results of each technology concept node in different time windows are aggregated by time weighting to reflect the dynamic evolution trend in the time dimension; through the continuous calculation process of the convolution kernel sliding, the feature change pattern of the technology concept node in the long-term and short-term time scale is captured, and the time sequence node state vector containing the time evolution state of each technology concept node is output.

[0073] The time sequence node state vector is mapped to the prediction target dimension through the full connection layer of the spatio-temporal graph neural network, and the technology concept time sequence prediction value sequence is obtained.

[0074] Further, the time sequence node state vector is input into the spatio-temporal graph neural network, in the full connection layer, the time sequence node state vector is linearly mapped and feature compressed, and the multi-dimensional time state information is mapped to the prediction target dimension; in the mapping process, the technology concept node identification in the causally enhanced spatio-temporal graph is kept consistent with the order of the technology concept node, so that the output result is one-to-one corresponding to the input technology concept node; the output values of each technology concept node in the prediction target dimension are calculated through forward propagation, and the non-linear expression ability is enhanced by using the activation function; the output values of all technology concept nodes are integrated in time label order to obtain the technology concept time sequence prediction value sequence.

[0075] It should be noted that the prediction target dimension refers to the numerical space dimension in the output layer of the spatio-temporal graph neural network for representing the future state or change trend of the technology concept node, which is used to map the time sequence node state vector to a quantifiable technology concept time sequence prediction result.

[0076] S5, based on the causally enhanced spatio-temporal graph, the technology concept time sequence prediction value sequence is causally traced, the technology concept causing the trend change and the causal relationship are identified, and a technology achievement trend prediction report is generated.

[0077] The technology concept nodes exceeding the preset change amplitude threshold are selected from the technology concept time sequence prediction value sequence and marked as key technology concept nodes.

[0078] Further, the technology concept time sequence prediction value sequence is traversed, the prediction value difference between each technology concept node in consecutive time segments is extracted, and the prediction value change amplitude is calculated; the prediction value change amplitude is compared with the preset change amplitude threshold, when the prediction value change amplitude of the technology concept node exceeds the change amplitude threshold, it is determined that the trend of the technology concept node fluctuates significantly, and it is marked as a key technology concept node.

[0079] It should be noted that the change range threshold is set based on the mean and standard deviation statistics of the prediction value change range distribution of all nodes in the technology concept time series prediction value sequence, which is used to identify the limiting value of the trend fluctuation degree of the technology concept node in the time dimension. The exemplary value range is 0.15-0.30. When it is lower than 0.15, it is easy to misjudge the slight value fluctuation as a trend change, resulting in too many key technology concept nodes and noise interference. When it is higher than 0.30, some technology concept nodes that actually exist significant changes may be ignored, causing trend identification omission.

[0080] All upstream causal edges and source nodes connected to the key technology concept node are extracted from the causal enhanced space-time graph to form a local causal tracing subgraph.

[0081] Further, according to the key technology concept node list, all connection edges pointing to the key technology concept node are searched in the causal enhanced space-time graph, and the source technology concept node corresponding to the starting point of the connection edge is extracted. The direct upstream causal edge of each key technology concept node and the corresponding source technology concept node set are extracted and stored in the form of an adjacency list. For the case where there is a multi-layer causal dependence relationship, the source technology concept node is taken as a new search starting point, and the upstream causal edge of the source technology concept node and the earlier appearing source technology concept node are recursively extracted until there is no causal edge pointing to the source technology concept node. The key technology concept node, the upstream and all source technology concept nodes are sorted according to the time label sequence to generate a local causal tracing subgraph.

[0082] The local causal tracing subgraph is searched by a graph traversal algorithm to identify the upstream causal path.

[0083] Further, the local causal tracing subgraph is executed by a depth-first traversal algorithm with the key technology concept node as the starting point. Each causal connection edge extending from the key technology concept node to the upstream node is traced in turn, and the key technology concept node sequence and the causal direction of the causal connection edge in the traversal path are recorded. In the traversal process, if a repeated key technology concept node is encountered, the search of this branch is terminated to avoid path circulation. When a source key technology concept node without an upstream node is searched, the key technology concept node sequence and the connection edge sequence that have been traversed are combined and recorded as a complete upstream causal path.

[0084] The causal contribution degree of each upstream causal path is calculated, and the upstream causal path exceeding the preset contribution degree threshold is screened to generate a core causal path list.

[0085] Furthermore, the causal contribution of each upstream causal path is calculated. The causal weights of each causal connection edge in the upstream causal path are weighted and summed with the reciprocal of the time difference between nodes to obtain the comprehensive influence intensity of the upstream causal path on the trend changes of key scientific and technological concept nodes. The calculated causal contribution is compared with a preset contribution threshold. When the causal contribution is higher than the contribution threshold, the upstream causal path is determined to be a core causal path. All upstream causal paths that meet the conditions are sorted from high to low according to their causal contribution to generate a list of core causal paths.

[0086] It should be noted that the contribution threshold is a limited value set based on the statistical results of the mean and standard deviation of the causal contribution distribution of all upstream causal paths. It is used to screen the significance of the influence of upstream causal paths on trend changes. An exemplary value range is 0.4 to 0.7. When it is below 0.4, low-impact paths will be misjudged as core causal paths, resulting in redundant results and increased noise. When it is above 0.7, only a small number of high-contribution paths are retained, which may miss medium-contribution paths that have a real impact on the formation of the trend.

[0087] The key scientific and technological concept nodes, core causal path lists, and corresponding causal contribution values ​​are organized in natural language to generate a scientific and technological achievement trend prediction report.

[0088] Furthermore, based on the list of key scientific and technological concept nodes and the list of core causal paths, the predicted value changes, time tags, and corresponding upstream core causal paths of the key scientific and technological concept nodes are matched accordingly. The node names, causal weights, and causal directions in each core causal path are semantically transformed, and the relationships of the core causal paths are expressed in natural language sentences. The expressed natural language sentences are combined according to the time sequence and causal chain hierarchy to generate a scientific and technological achievement trend prediction report that includes an explanation of the trend changes of key scientific and technological concept nodes, an analysis of upstream causal paths, and an explanation of causal contribution.

[0089] The embodiment also provides a big data-based scientific and technological achievement analysis and prediction system, comprising a data acquisition module, a model construction module, a causal diagram construction module, a trend prediction module and a causal tracing module; the data acquisition module is used for acquiring multi-source scientific and technological achievement semantic data, pre-processing the multi-source scientific and technological achievement semantic data, and constructing a scientific and technological concept relationship diagram; the model construction module is used for taking a time series diagram convolution network as a basic framework, taking the scientific and technological concept relationship diagram as a training sample for training, constructing a dynamic knowledge flow semantic model, and using the dynamic knowledge flow semantic model to extract evolution features of the scientific and technological concept relationship diagram and output a knowledge flow feature vector; the causal diagram construction module is used for establishing new connections between nodes of the scientific and technological concept relationship diagram based on the knowledge flow feature vector, giving the connections causal directions and weights according to co-occurrence relationships and reference relationships between the nodes, and forming a causal enhanced space-time diagram; the trend prediction module is used for inputting the causal enhanced space-time diagram into a space-time diagram neural network, aggregating semantic associations and causal relationships between scientific and technological achievements in a spatial dimension, capturing dynamic change patterns of scientific and technological achievement features in a time dimension, and outputting a scientific and technological concept time series prediction value sequence; and the causal tracing module is used for performing causal tracing analysis on the scientific and technological concept time series prediction value sequence based on the causal enhanced space-time diagram, identifying scientific and technological concepts and causal relationships that cause trend changes, and generating a scientific and technological achievement trend prediction report.

[0090] To sum up, the application captures the dynamic evolution law of scientific and technological achievement semantics changing over time, and mines potential semantic flow and migration paths between scientific and technological concepts by taking a time series diagram convolution network as a basic framework, taking a scientific and technological concept relationship diagram as a training sample for training, and constructing a dynamic knowledge flow semantic model.

[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application but not limit the application, and although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all should be covered in the scope of the claims of the application.

Claims

1. A method for analyzing and predicting scientific and technological achievements based on big data, characterized in that: include, Collect semantic data of scientific and technological achievements from multiple sources and preprocess them to construct a graph of scientific and technological concept relationships; Based on a temporal graph convolutional network, a dynamic knowledge flow semantic model is constructed using a technology concept relationship graph as training samples. The dynamic knowledge flow semantic model is then used to extract evolutionary features from the technology concept relationship graph and output a knowledge flow feature vector. Based on knowledge flow feature vectors, new connections are established between nodes in the science and technology concept relationship graph, and causal directions and weights are assigned to the connections according to the co-occurrence and reference relationships between nodes, forming a causal-enhanced spatiotemporal graph; By inputting the causal enhanced spatiotemporal graph into the spatiotemporal graph neural network, the semantic associations and causal relationships between scientific and technological achievements are aggregated in the spatial dimension, and the dynamic change patterns of the characteristics of scientific and technological achievements are captured in the temporal dimension, and the time-series predicted value sequence of scientific and technological concepts is output. Based on the causal enhanced spatiotemporal graph, a causal source analysis is performed on the time series prediction value sequence of scientific and technological concepts to identify the scientific and technological concepts and causal relationships that lead to trend changes, and a scientific and technological achievement trend prediction report is generated.

2. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The multi-source scientific and technological achievement semantic data includes patent data, academic paper data, and scientific research project data; The preprocessing includes data cleaning, entity recognition, and text normalization.

3. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for constructing the technology concept relationship diagram are as follows: Co-occurrence statistical analysis is performed on the preprocessed multi-source scientific and technological achievement semantic data to calculate the co-occurrence probability between scientific and technological concepts and generate candidate scientific and technological concept association pairs. Based on a pre-defined semantic rule base, the semantic relationship type of candidate technology concept association pairs is determined and transformed into technology concept triples; Vectorize the triples of science and technology concepts to generate a set of vector vectors representing the relationships between science and technology concepts. By using technological concepts as nodes, semantic relationships as edges, and combining them with a set of technological concept relationship vectors, a technological concept relationship graph is constructed.

4. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for constructing a dynamic knowledge flow semantic model are as follows: Extract the features of technology concept nodes, semantic relationship structure and temporal attributes from the technology concept relationship graph to generate a feature mapping matrix; The feature mapping matrix is ​​input into the temporal graph convolutional network, the parameters of the temporal graph convolutional network are initialized and multiple rounds of training are performed to obtain the spatiotemporal feature parameter set; By optimizing the temporal graph convolutional network structure using spatiotemporal feature parameter sets, a dynamic knowledge flow semantic model is formed.

5. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for outputting the knowledge flow feature vector are as follows: The relationship graph of scientific and technological concepts is input into the dynamic knowledge flow semantic model to extract the changing trends and migration path features of scientific and technological concept nodes in the time dimension, and generate a node-level evolution feature set. The semantic evolution correlation between scientific and technological concepts is calculated based on the node-level evolution feature set. The evolution information of each scientific and technological concept node is aggregated, and the knowledge flow feature vector is output.

6. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for forming the causal enhanced spatiotemporal graph are as follows: Based on knowledge flow feature vectors, feature matching analysis is performed on the nodes of science and technology concepts in the science and technology concept relationship graph to identify potential semantic connection node pairs. New connection edges are established between potentially semantically related node pairs, and the co-occurrence frequency and reference strength of each new connection edge are calculated to generate a node association information set; The causal direction of the connecting edges is determined based on the temporal attributes and semantic relationship structure of the technology concept nodes in the node association information set, and the causal weight is obtained through the causal probability calculation method. The new connecting edges and their corresponding causal weights are integrated with the technology concept relationship graph to form a causal-enhanced spatiotemporal graph.

7. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for outputting the time-series predicted value sequence of the technological concept are as follows: By inputting the causal enhanced spatiotemporal graph into the spatiotemporal graph neural network, the semantic relationships and causal features between the nodes of scientific and technological concepts are aggregated and calculated through spatial convolution operations to obtain spatial dependency feature vectors. Temporal features are extracted from spatially dependent feature vectors through temporal convolution operations, the dynamic evolution pattern of the features of each technological concept node is analyzed, and the temporal node state vector is output. By mapping the state vectors of time-series nodes to the dimension of the prediction target through the fully connected layer of the spatiotemporal graph neural network, a sequence of time-series predicted values ​​of scientific and technological concepts can be obtained.

8. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for conducting causal source analysis on the time-series predicted value series of scientific and technological concepts based on causal enhanced spatiotemporal graphs are as follows: Nodes of scientific and technological concepts that exceed a preset threshold for change in time series prediction values ​​are selected from the scientific and technological concept time series and marked as key scientific and technological concept nodes. Extract all upstream causal edges and source nodes connected to key technological concept nodes from the causal enhanced spatiotemporal graph to form a local causal origin subgraph; The upstream causal path is identified by performing a reverse path search on the local causal origination subgraph using a graph traversal algorithm.

9. The method for analyzing and predicting scientific and technological achievements based on big data as described in claim 1, characterized in that: The steps for generating the technology achievement trend prediction report are as follows: Calculate the causal contribution of each upstream causal path, filter out upstream causal paths that exceed the preset contribution threshold, and generate a core causal path list. The key scientific and technological concept nodes, core causal path lists, and corresponding causal contribution values ​​are organized in natural language to generate a scientific and technological achievement trend prediction report.

10. A big data-based system for analyzing and predicting scientific and technological achievements, based on the big data-based method for analyzing and predicting scientific and technological achievements as described in any one of claims 1 to 9, characterized in that: include, The data acquisition module is used to collect semantic data of multi-source scientific and technological achievements and perform preprocessing to construct a scientific and technological concept relationship diagram; The model building module is used to build a dynamic knowledge flow semantic model based on a temporal graph convolutional network and trained with a technology concept relationship graph as training samples. The dynamic knowledge flow semantic model is then used to extract evolutionary features from the technology concept relationship graph and output a knowledge flow feature vector. The causal graph construction module is used to establish new connections between nodes in the science and technology concept relationship graph based on knowledge flow feature vectors, and to assign causal direction and weight to the connections according to the co-occurrence relationship and reference relationship between nodes, forming a causal enhanced spatiotemporal graph; The trend prediction module is used to input the causal enhanced spatiotemporal graph into the spatiotemporal graph neural network, aggregate the semantic associations and causal relationships between scientific and technological achievements in the spatial dimension, capture the dynamic change patterns of the characteristics of scientific and technological achievements in the time dimension, and output the time-series prediction value sequence of scientific and technological concepts. The causal tracing module is used to perform causal tracing analysis on the time-series predicted value sequence of scientific and technological concepts based on the causal enhanced spatiotemporal graph, identify the scientific and technological concepts and causal relationships that lead to trend changes, and generate a scientific and technological achievement trend prediction report.

Citation Information

Cited By

  • Scientific field-oriented technology association dynamic tracking method

    CN121920358A

  • A scientific field-oriented technology association dynamic tracking method

    CN121920358B