Hotspot technology mining method and system based on scientific and technical literature data
By constructing a network of technical elements based on papers and patent documents, dividing them by time slices, and calculating and predicting their importance, the problem of incomplete data in existing methods is solved. This enables accurate identification of emerging technologies and early discovery of potential integration directions, improving the accuracy and foresight of hot technology mining.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中铁科学研究院集团有限公司
- Filing Date
- 2026-03-10
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for identifying hot technologies fail to fully integrate academic papers and patent literature, lack in-depth analysis of the relationships between technological elements, and lack scientific quantitative indicators. This results in highly subjective findings that are difficult to meet the requirements of accuracy and objectivity.
By acquiring a collection of scientific and technological literature data containing papers and patents, technical phrases are extracted and semantically normalized to construct a current technical element network of technical means and technical objects. The network is divided by time slices and time-series importance indicators are calculated to identify emerging technology clusters. Dynamic graph representation learning and prediction of future related edge formation are then performed to generate a list of emerging technology themes and integration directions.
It clearly demonstrates the development trajectory of technology, accurately identifies emerging technology fields, and discovers the direction of technological integration in advance, thereby improving the accuracy, comprehensiveness, and foresight of hot technology discovery.
Smart Images

Figure CN121833933A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scientific and technological information mining technology, and more specifically, to a method and system for mining hot technologies based on scientific and technological literature data. Background Technology
[0002] In today's rapidly evolving technological landscape, timely and accurate identification of cutting-edge technologies is crucial for research institutions to formulate strategic plans, for enterprises to grasp market trends, and for governments to guide industrial upgrading. Scientific literature, as an important carrier of scientific activities and achievements, contains a wealth of technical information and is a vital data source for identifying these cutting-edge technologies.
[0003] Traditional methods for identifying hot technologies suffer from the following limitations. Firstly, some methods focus only on a single type of scientific and technological literature, such as analyzing only academic papers or patent documents. This results in insufficiently comprehensive technical information and an inability to fully present the overall development trend of the technology. Academic papers emphasize theoretical research and academic discussion, while patent documents focus more on technological innovation and practical application; combining both is necessary to comprehensively reflect the full picture of the technology. Secondly, existing methods often lack in-depth analysis of the relationships between technical elements when processing scientific and technological literature data. Most methods simply perform statistical analysis of keywords in the literature, failing to construct a network of connections between technical means and technical objects, and thus failing to accurately grasp the internal structure and evolutionary laws of the technology. Furthermore, the lack of scientific and systematic quantitative indicators and methods for assessing the vitality of technological evolution and identifying emerging technologies leads to a high degree of subjectivity in the mining results, making it difficult to meet the requirements of accuracy and objectivity in practical applications. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, the present invention provides a method for mining hot technologies based on scientific and technological literature data, the method comprising:
[0005] Acquire a set of scientific and technological literature data, which includes a subset of academic paper data and a subset of patent data. The subset of academic paper data includes abstract text units and full-text text units. The subset of patent data includes abstract text units and claims text units. The scientific and technological literature data set is subjected to technical phrase extraction and semantic normalization processing to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components. The current technical element network includes multiple technical means entity nodes, multiple technical object entity nodes, and a set of associated edges connecting the technical means entity nodes and the technical object entity nodes. The current technical element network is divided into time slices according to a preset time slice division rule to generate a time-series technical element network segment set, which contains multiple technical element network segment units arranged in chronological order. The time-series network fragment set of technical elements is processed by calculating the node temporal importance index and the edge temporal importance index. Technical community structures whose comprehensive scores of technical evolution vitality exceed a preset vitality threshold are identified as emerging technology clusters, and a set of emerging technology clusters is generated. The time-series network fragment set of technical elements is subjected to dynamic graph representation learning and future association edge formation prediction processing to generate a potential technology combination prediction set containing predicted association edges. Each predicted association edge in the potential technology combination prediction set is subjected to technical field crossover degree evaluation processing and novelty evaluation processing to generate a predicted association edge evaluation result set. Based on the node growth trend of the technical means entity nodes and technical object entity nodes in the emerging technology cluster set, and combined with the technical field crossover assessment results and novelty assessment results in the predicted association edge assessment result set, an emerging technology topic list containing multiple emerging technology topic identifiers and a technology integration direction list containing multiple technology integration direction identifiers are generated. The entries in the technology integration direction list are sorted based on their technical field crossover score and novelty comprehensive score.
[0006] Furthermore, this invention also provides a hotspot technology mining system based on scientific and technological literature data, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned hotspot technology mining method based on scientific and technological literature data by executing the machine-executable instructions.
[0007] Based on the above, by acquiring a collection of scientific and technological literature data including papers and patents, and performing technical phrase extraction and semantic normalization, a current technology element network with technical means and technical objects as core elements is constructed. Dividing this network into preset time slices and performing time-series processing clearly demonstrates the changes of technical elements over time, helping to grasp the development trajectory of technology. By calculating the temporal importance indicators of nodes and edges in the time-seriesd technology element network segments, emerging technology clusters are scientifically identified, generating a set of emerging technology clusters and accurately identifying emerging technology fields with development potential. Dynamic graph representation learning and future edge formation prediction are performed on the time-seriesd technology element network segments to generate a potential technology combination prediction set. Furthermore, the degree of technological crossover and novelty of this set are evaluated, enabling the early discovery of possible technology integration directions and innovation points. Finally, by combining the node growth trend of emerging technology clusters and the evaluation results of predicted edges, a list of emerging technology topics and a list of technology integration directions are generated, improving the accuracy, comprehensiveness, and forward-looking nature of hot technology mining. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the execution flow of the hot technology mining method based on scientific and technological literature data provided in the embodiments of the present invention.
[0009] Figure 2 This is a schematic diagram of exemplary hardware and software components of the hotspot technology mining system based on scientific and technological literature data provided in an embodiment of the present invention. Detailed Implementation
[0010] Figure 1 This is a flowchart illustrating a hotspot technology mining method based on scientific and technological literature data, provided in one embodiment of the present invention. A detailed description follows.
[0011] Step S110: Obtain a set of scientific and technological literature data, which includes a subset of academic paper data and a subset of patent data. The subset of academic paper data includes abstract text units and full-text text units, and the subset of patent data includes abstract text units and claims text units.
[0012] In this embodiment, scientific literature data mining in the field of artificial intelligence is used as the application scenario. Relevant literature data is systematically selected from multiple academic and patent databases. Academic databases include top conferences and authoritative journals in the field of computer science, ensuring coverage of cutting-edge research results in this field. Patent databases cover patent literature published through official channels, especially patents related to artificial intelligence algorithms and hardware architecture. The abstract text unit of a paper typically consists of 300-500 words, including research background, technical methods, core results, and a summary of conclusions; the full text text unit of a paper includes a complete introduction, related work, method description, experimental design, results analysis, and discussion, ranging in length from several thousand to tens of thousands of words. The patent abstract text unit is approximately 200-300 words, briefly describing the technical field of the patent, the technical problem to be solved, the key points of the technical solution, and the beneficial effects; the patent claim text unit precisely defines the scope of patent protection through independent claims and dependent claims, including the combination relationship of technical features and limiting conditions. During data acquisition, sensitive information such as author email addresses and institutional internal codes in the document metadata was anonymized, retaining only publicly available technical information such as document titles, publication dates, abstracts, full text, and claims. Simultaneously, the acquired text data underwent format standardization, converting documents in different formats such as PDF, XML, and HTML into plain text, removing non-text elements such as formulas, figures, and references to ensure consistency in subsequent processing. Through these processes, a collection of scientific and technological literature data was constructed.
[0013] Step S120: Perform technical phrase extraction and semantic normalization on the scientific and technological literature data set to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components. The current technical element network includes multiple technical means entity nodes, multiple technical object entity nodes, and a set of associated edges connecting the technical means entity nodes and the technical object entity nodes.
[0014] In the application scenario of artificial intelligence, the acquired collection of scientific and technological literature data is processed. First, sentence boundary detection is performed on different types of text units (paper abstracts, full texts, patent abstracts, patent claims) to segment continuous text into semantically complete sentence units. Next, part-of-speech tagging and dependency parsing are performed on each sentence unit to identify nouns, verbs, adjectives, and their grammatical relationships. Based on preset technical phrase composition patterns (verb-noun combinations, adjective-noun combinations, noun-noun complex combinations), candidate technical phrases are extracted from the sentences. Subsequently, stop word filtering is used to remove common technical terms, resulting in a preliminary set of technical phrases. Structural analysis is performed on these phrases to distinguish between two types: technical means (such as "convolutional neural network training methods") and technical objects (such as "image classification datasets"), and semantic normalization is performed on each type—the phrases are converted into vector representations through a word embedding model, semantic similarity is calculated, and clustering is performed to map different phrases expressing the same meaning to the same entity node. Finally, based on the co-occurrence relationships and dependency paths of phrases in sentences, relational edges between entities are constructed. Edge weights are adjusted based on global co-occurrence frequency, forming a technical element network containing entity nodes and relational edges. In this network, technical means entity nodes and technical object entity nodes are connected by directed edges, with edge weights reflecting the closeness of their association.
[0015] Step S121: Perform sentence boundary detection processing on the paper abstract text unit, the paper full text text unit, the patent abstract text unit, and the patent claims text unit respectively to obtain multiple sentence units corresponding to each text unit.
[0016] For various text units in the field of artificial intelligence, a sentence boundary detection method integrating punctuation rules and machine learning models is adopted. First, preliminary segmentation is performed based on punctuation marks (period, question mark, exclamation mark) to divide the text into candidate sentence units. For text containing abbreviations (such as "eg", "ie", "CNN"), decimal points, ellipses, and other special symbols, a domain-specific abbreviation list is constructed for identification to avoid incorrect segmentation. For example, in the sentence "Using a CNN model to process image data (see Fig. 1).", the abbreviation list identifies "CNN" and "Fig." as special markers, determining the period after "data" as the sentence boundary. For patent claim text units, the special format of numbered items (such as "1. A…", "2. According to claim 1…") needs to be processed. Regular expressions are used to match the claim numbers to ensure that each claim item is an independent sentence unit. Simultaneously, a BERT-based sentence boundary detection model is used for secondary verification. This model is trained on a large-scale labeled corpus and can learn the contextual features of sentence boundaries. The preliminary segmentation results are input into the sentence boundary detection model. The model outputs the probability that each punctuation mark represents a sentence boundary. When the probability exceeds a preset threshold, it is confirmed as a boundary. Through this process, the text units of a paper abstract are segmented into 5-10 sentence units, the text units of the full paper are segmented into dozens to hundreds of sentence units, the text units of a patent abstract are segmented into 3-5 sentence units, and the text units of a patent claim are segmented into multiple sentence units according to the claim items, ensuring that each sentence unit is semantically complete and has accurate boundaries.
[0017] Step S122: Perform part-of-speech tagging and dependency parsing on the multiple sentence units to generate a part-of-speech tag sequence and dependency tree structure corresponding to each sentence unit.
[0018] The segmented sentence units are processed using a Transformer-based part-of-speech tagging model. This model takes the word sequence of the sentence as input and outputs a part-of-speech tag for each word, including 36 general part-of-speech tags such as noun (NN), verb (VB), adjective (JJ), adverb (RB), and preposition (IN), as well as AI-specific tags (such as technical term tags (TECH)). For example, after tagging the sentence "Transformer model performs excellently in natural language processing tasks," "Transformer" is labeled as NN-TECH, "model" as NN, "in" as IN, "natural language processing" as NN-TECH, "task" as NN, "in" as IN, "performance" as VB, and "excellent" as JJ. After part-of-speech tagging, a dependency parsing model based on graph neural networks is used to construct a dependency tree. This part-of-speech tagging model captures long-distance dependencies between words through a multi-head attention mechanism, outputting a tree structure containing 40 types of dependency relations, including subject-verb (nsubj), verb-object (dobj), modifier-head (amod), and prepositional (prep). For example, in the sentence above, "performance" (VB) is the predicate (nsubj) of the subject "model" (NN), "task" (NN) forms a prepositional relation (prep) with "performance" through the preposition "in" (IN), and "excellent" (JJ) is the predicate nominative (acomp) of "performance". The part-of-speech tag sequence is stored in list form, with each element being a (word, tag) pair; the dependency tree is stored in adjacency list form, with each node containing a word index, dependency relation type, and parent node index. Through the above processing, each sentence unit is converted into a structured grammatical representation.
[0019] Step S123: Based on the part-of-speech tag sequence and the dependency tree structure, extract noun phrase combinations that conform to the preset technical phrase composition pattern from each sentence unit. The preset technical phrase composition pattern includes verb and noun combination patterns, adjective and noun combination patterns, and noun and noun compound combination patterns.
[0020] Based on the generated part-of-speech tag sequence and dependency tree structure, technical phrases are extracted through a combination of rule matching and grammatical pattern recognition. For verb-noun combination patterns, the verb (VB) in the sentence and its corresponding direct object (dobj) or prepositional object (pobj) are first identified. For example, if a path "verb → dobj → noun phrase" exists in the dependency tree, then "verb + noun phrase" is extracted as the technical phrase. In the sentence "using the backpropagation algorithm to optimize neural network parameters," the direct object of "using" (VB) is "algorithm" (NN), and its modifier is "backpropagation" (NN-TECH), so "using the backpropagation algorithm" is extracted. For adjective-noun combination patterns, the attributive relation (amod) of adjective (JJ) modifying noun (NN) is identified, and the "adjective + noun" structure is extracted. For example, in "efficient convolutional neural network architecture," "efficient" (JJ) modifies "architecture" (NN), and "convolutional neural network" (NN-TECH) is included as a nested modifier, so "efficient convolutional neural network architecture" is extracted. For compound noun combinations, we identify consecutive noun sequences (NN chains) connected by parallel or compound relationships. For example, in "deep learning framework performance evaluation metrics," "deep learning" (NN-TECH) modifies "framework" (NN), and "performance evaluation" (NN) modifies "metrics" (NN), forming a multi-layered nested noun compound structure, which is then fully extracted as "deep learning framework performance evaluation metrics." During the extraction process, we recursively traverse the dependency tree to collect all noun phrase combinations that conform to the above pattern, ensuring the completeness and accuracy of the technical phrases.
[0021] Step S124: Perform stop word filtering on all extracted noun phrase combinations to remove general technical terms that appear in the preset technical stop word list, thereby obtaining a candidate technical phrase set. The preset technical stop word list contains general technical terms for methods, systems, devices, and equipment.
[0022] A pre-defined technical stop word list containing common technical terms in the field of artificial intelligence was constructed. This list was generated through statistical analysis of multiple documents in the field and includes 200 high-frequency common terms such as "method," "system," "device," "equipment," "module," "unit," "process," "step," "mechanism," and "scheme." Word-by-word matching was performed on the extracted noun phrase combinations. If a phrase contained a word from the stop word list, it was removed. For example, in "image recognition method based on deep learning," "method" is a stop word; removing it yields "image recognition based on deep learning." In "neural network training system and device," "system" and "device" are stop words; removing them yields "neural network training." For cases where the stop word is located in the middle of a phrase, such as "research on convolutional neural network implementation device and method," removing "device" and "method" yields "research on convolutional neural network implementation." During the filtering process, the core modifying relationships of the phrases were preserved to ensure that the phrases still retain their complete technical meaning after removing stop words. After filtering, the remaining phrases were deduplicated, and identical phrases were merged to obtain a set of candidate technical phrases. The phrases in this candidate technology phrase set have had redundant and generic terms removed, retaining only the core expressions with specific technical meanings.
[0023] Step S125: Perform phrase structure analysis on each candidate technical phrase in the candidate technical phrase set, identify the core word component and modifier component in each candidate technical phrase, classify the candidate technical phrase into technical means phrase type or technical object phrase type according to the core word component and the modifier component, and generate a set of technical phrase units carrying phrase type tags, the phrase type tags including technical means phrase tags and technical object phrase tags.
[0024] For each phrase in the candidate technical phrase set, phrase structure analysis is performed using a core word identification algorithm based on dependency syntax. First, the noun core (usually the last noun) in the phrase is determined through part-of-speech tag sequences. Then, its modifiers (adjectives, nouns, word segments, etc.) are recursively searched. For example, in the phrase "text classification model based on attention mechanism," the core word is "model" (NN), and the modifiers include "based on attention mechanism" (prepositional phrase) and "text classification" (noun phrase). Based on the semantic category of the core word, the phrase is divided into technical means or technical objects: if the core word represents a method, algorithm, tool, or technology for achieving a specific function (such as "algorithm," "model," "method," "technology," "framework," "system"), it is marked as a technical means phrase; if the core word represents the object, goal, product, or scenario of the technology application (such as "data," "image," "task," "problem," "dataset," "scenario"), it is marked as a technical object phrase. For example, the core word of "convolutional neural network training algorithm" is "algorithm," and it is marked as a technical means phrase; the core word of "medical image segmentation dataset" is "dataset," and it is marked as a technical object phrase. For complex phrases containing multiple core words, the primary core word is determined by analyzing the modification relationships. For example, in "machine translation methods and systems based on Transformer", the primary core word is "method", which is marked as a technical means phrase. Each phrase unit is assigned the label "technical means phrase" or "technical object phrase", forming a set of technical phrase units carrying type labels.
[0025] Step S126: Perform semantic normalization on all technical phrase units carrying technical means phrase tags in the technical phrase unit set, map different technical phrase units expressing the same technical means meaning to the same technical means entity node, and assign a technical means entity identifier to each technical means entity node.
[0026] Step S1261: Obtain all technical phrase units carrying technical means phrase tags in the technical phrase unit set. Each technical means phrase unit contains the text content of the technical means phrase and the context sentence unit of the technical means phrase in the original document.
[0027] From the set of technical phrase units, all units labeled "technical means phrases" are selected. Each unit contains two parts of data: the technical means phrase text (e.g., "backpropagation algorithm," "BP algorithm," "error backpropagation training method") and the context sentence of the phrase in the original literature (e.g., "In the process of neural network training, the backpropagation algorithm is used to adjust weight parameters"). The context sentence is used to assist semantic understanding, especially to provide contextual information when the phrase is ambiguous. The above technical means phrase units are sorted by frequency of occurrence, and phrase units that appear at least three times are retained to form an initial technical means phrase pool, ensuring the efficiency and accuracy of subsequent processing.
[0028] Step S1262: Input the text content of each technical means phrase unit into the pre-trained word embedding model, and convert the text content of each technical means phrase unit into a technical means phrase embedding vector through the pre-trained word embedding model.
[0029] A BERT model pre-trained on a large-scale corpus of artificial intelligence (including abstracts of multiple papers and patent texts) is used as the word embedding model. The technical phrase text is segmented (e.g., "backpropagation algorithm" is segmented into "backward," "propagation," and "algorithm") and input into the last hidden state of the BERT model. A weighted average of the word embedding vectors (with attention weights) is performed to obtain phrase-level embedding vectors. The embedding vectors are 768-dimensional, with each dimension corresponding to a feature direction in the semantic space. For example, the embedding vectors of "backpropagation algorithm" and "BP algorithm" are semantically closer together, while they are farther away from the embedding vector of "convolutional neural network." Through this process, each technical phrase unit is converted into a fixed-dimensional numerical vector.
[0030] Step S1263: Perform pairwise semantic similarity calculation on all technical means phrase units in the set of technical means phrase units, calculate the cosine similarity between any two technical means phrase embedding vectors, and obtain the semantic similarity matrix.
[0031] Pairwise cosine similarity is calculated for the embedding vectors of technical phrases using the formula: Cosine similarity = Dot product of vector A and vector B / (Modulus of vector A × Modulus of vector B). The result ranges from -1 to 1, with values closer to 1 indicating greater semantic similarity. The similarity values of all phrase pairs are organized into an N×N semantic similarity matrix (N being the number of technical phrase units). The element in the i-th row and j-th column of the matrix represents the similarity between the i-th and j-th phrases. For example, the similarity between "backpropagation algorithm" and "BP algorithm" is 0.92, with "gradient descent algorithm" being 0.65, and with "convolutional neural network" being 0.32. This matrix visually demonstrates the degree of semantic association between phrases.
[0032] Step S1264: Perform hierarchical clustering on the semantic similarity matrix, and divide the technical means phrase units with semantic similarity exceeding the preset clustering similarity threshold into the same technical means phrase cluster. Each technical means phrase cluster contains multiple technical means phrase units that express similar technical means meanings.
[0033] A cohesive hierarchical clustering algorithm is used to process the semantic similarity matrix. Initially, each phrase is an independent cluster. Then, the two clusters with the highest similarity are iteratively merged until the similarity between all clusters is below a preset threshold (experimentally determined to be 0.85). During clustering, the similarity between clusters is calculated using the group average method (the average of the similarity of all phrase pairs in two clusters). For example, "backpropagation algorithm," "BP algorithm," and "error backpropagation method" all have similarities higher than 0.85 and are merged into one cluster; "stochastic gradient descent," "SGD algorithm," and "stochastic gradient descent optimization method" are merged into another cluster. Each cluster contains 2-15 phrase units, which semantically express the same or highly similar technical means. Through clustering, a large number of technical means phrase units are summarized into a smaller number of phrase clusters, preparing for entity node mapping.
[0034] Step S1265: Select the most frequently occurring technical means phrase unit from each technical means phrase cluster as the representative technical means phrase of the technical means phrase cluster, and use the text content of the representative technical means phrase as the node name of the technical means entity node corresponding to the technical means phrase cluster.
[0035] For each technical means phrase cluster, the frequency of all phrase units is statistically analyzed, and the phrase with the highest frequency is selected as the representative phrase. For example, in a cluster containing "backpropagation algorithm" (occurring 120 times), "BP algorithm" (occurring 85 times), and "error backpropagation method" (occurring 45 times), "backpropagation algorithm" has the highest frequency and is selected as the representative phrase. If multiple phrases have the same frequency, the most standardized expression is determined through manual review as the representative phrase (e.g., "convolutional neural network" is selected instead of "CNN network"). The text content of the representative phrase will serve as the name of the technical means entity node, ensuring the accuracy and universality of the node name.
[0036] Step S1266: Create a technical means entity node for each technical means phrase cluster, assign a unique technical means entity identifier to the technical means entity node, and associate and map all technical means phrase units within the technical means phrase cluster with the technical means entity identifier.
[0037] For each technical phrase cluster, an entity node object is created, containing attributes such as node name (representative phrase), unique identifier (e.g., "TMI-0001", "TMI-0002", where TMI represents the technical entity), a list of all phrase units within the cluster, and the total frequency of phrase occurrences. A many-to-one mapping relationship is established between the text content of all technical phrase units within the cluster and this entity identifier. For example, "TMI-0001" corresponds to phrase units such as "backpropagation algorithm," "BP algorithm," and "error backpropagation method." This mapping relationship is stored in a relational database, supporting quick retrieval of the corresponding entity node from the phrase unit.
[0038] Step S1267: Store the context sentence unit corresponding to each technical means phrase unit in the original document as the context information of the technical means entity node for subsequent functional description of the technical means entity node.
[0039] The context sentences of all phrase units within a technical means phrase cluster are collected, deduplicated, and used as the context information for the entity node of that technical means. For example, the context information of "TMI-0001" (backpropagation algorithm) includes sentences such as "In neural network training, the backpropagation algorithm updates weights by calculating gradients" and "The BP algorithm is the core optimization method in deep learning." This context information is stored in the "contexts" attribute of the entity node. Subsequently, a functional description of the technical means entity can be generated using a text summarization algorithm to assist in the summarization of technical themes.
[0040] Step S127: Perform semantic normalization on all technical phrase units carrying technical object phrase tags in the technical phrase unit set, map different technical phrase units expressing the same technical object meaning to the same technical object entity node, and assign a technical object entity identifier to each technical object entity node.
[0041] Step S1271: Obtain all technical phrase units carrying technical object phrase tags in the technical phrase unit set. Each technical object phrase unit contains the text content of the technical object phrase and the context sentence unit of the technical object phrase in the original document.
[0042] From the set of technical phrase units, all units labeled "technical object phrases" are selected. Each unit contains the technical object phrase text (e.g., "image data", "image dataset", "image sample set") and the corresponding context sentence (e.g., "model training requires a large amount of labeled image data"). Similarly, phrase units that appear more than or equal to 3 times are retained to form a technical object phrase pool.
[0043] Step S1272: Input the text content of each technical object phrase unit into a pre-trained word embedding model, and convert the text content of each technical object phrase unit into a technical object phrase embedding vector through the pre-trained word embedding model.
[0044] Using the same BERT pre-trained model as the technical object phrases, the text of the technical object phrases is converted into 768-dimensional embedding vectors. For example, the embedding vectors of "image data" and "picture dataset" are close in semantic space, reflecting that they refer to the same technical object.
[0045] Step S1273: Perform pairwise semantic similarity calculation on all technical object phrase units in the set of technical object phrase units, calculate the cosine similarity between any two technical object phrase embedding vectors, and obtain the semantic similarity matrix.
[0046] The calculation method is the same as that for technical phrases, generating an N×N semantic similarity matrix that reflects the degree of semantic association between technical object phrases.
[0047] Step S1274: Perform hierarchical clustering on the semantic similarity matrix, and divide the technical object phrase units with semantic similarity exceeding the preset clustering similarity threshold into the same technical object phrase cluster. Each technical object phrase cluster contains multiple technical object phrase units that express the meaning of similar technical objects.
[0048] The same hierarchical clustering algorithm was used, with a clustering threshold set to 0.80 (the semantic differences between technical object phrases are usually slightly greater than those between technical means phrases). For example, "image data", "image dataset", and "image sample set" were merged into one cluster, while "text corpus", "text dataset", and "text sample" were merged into another cluster.
[0049] Step S1275: Select the most frequently occurring technical object phrase unit from each technical object phrase cluster as the representative technical object phrase of the technical object phrase cluster, and use the text content of the representative technical object phrase as the node name of the technical object entity node corresponding to the technical object phrase cluster.
[0050] The frequency of phrases within a cluster is statistically analyzed, and the most frequent phrase is selected as the representative phrase. For example, "image data" is used as the representative name of a cluster containing phrases such as "image data" and "image dataset".
[0051] Step S1276: Create a technical object entity node for each technical object phrase cluster, assign a unique technical object entity identifier to the technical object entity node, and associate and map all technical object phrase units within the technical object phrase cluster with the technical object entity identifier.
[0052] Create technical object entity nodes, assign unique identifiers (such as "TOI-0001" and "TOI-0002", where TOI represents a technical object entity), and establish a many-to-one mapping relationship between phrase units and entity identifiers.
[0053] Step S1277: Store the context sentence unit corresponding to each technical object phrase unit in the original document as the context information of the technical object entity node for subsequent functional description of the technical object entity node.
[0054] Collect the context sentences of phrases within the cluster and use them as the "contexts" attribute of the technical object entity node to help understand the specific meaning and application scenarios of the entity.
[0055] Step S128: Extract the co-occurrence relationship between the technical phrase unit carrying the technical means phrase tag and the technical phrase unit carrying the technical object phrase tag from each sentence unit. Based on the co-occurrence relationship and the direct dependency path in the dependency relationship tree structure, construct the initial association edge between the technical means entity node and the technical object entity node, and assign an initial association strength value to each initial association edge.
[0056] For each sentence unit, identify the co-occurring technical means phrase units and technical object phrase units. For example, in the sentence "using convolutional neural networks to process image data," "convolutional neural network" (technical means) and "image data" (technical object) co-occur. Analyze the direct dependency paths between the two using a dependency tree. If the technical means phrase forms a verb-object relationship with the technical object phrase through a verb (e.g., "process" → dobj → "image data," "use" → nsubj → "convolutional neural network"), then a direct association is determined. Create an initial association edge for each co-occurring technical means entity node and technical object entity node, with the initial association strength value set to the co-occurrence frequency of that phrase pair in the current sentence (initially 1). If the same entity pair appears in multiple sentences, the strength value is accumulated. For example, if "convolutional neural network" and "image data" co-occur in 10 sentences, the initial association strength value is 10. The direction of the association edge points from technical means to technical object, forming a directed edge.
[0057] Step S129: Based on the global co-occurrence frequency of each technical means entity node and each technical object entity node in the scientific and technological literature data set, the initial association strength value is weighted and adjusted to generate an adjusted association strength value, and the adjusted association strength value is used as the edge weight value of each association edge in the association edge set.
[0058] The global co-occurrence frequency (i.e., the total number of sentence units containing the entity pair) of statistical technique entity nodes and technical object entity nodes in the entire scientific and technological literature dataset is used to adjust the initial association strength value (intra-sentence co-occurrence frequency) by multiplying it by the logarithmic function of the global co-occurrence frequency (log(1 + global co-occurrence frequency)). For example, if the initial strength value for "convolutional neural network" and "image data" is 10, and the global co-occurrence frequency is 100, then the adjusted strength value is 10 × log(1 + 100) ≈ 10 × 4.6 = 46. This adjustment gives higher edge weights to entity pairs that frequently co-occur in multiple documents, reflecting a closer technical association between them. The adjusted edge weight values are stored in the "weight" attribute of the associated edges, forming a set of associated edges.
[0059] Step S1210: Aggregate all technical means entity nodes, all technical object entity nodes, and all associated edges carrying edge weight values to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components.
[0060] The set of entity nodes for technical means, the set of entity nodes for technical objects, and the set of associated edges are integrated into a directed weighted graph structure, namely the current technical element network. This current technical element network is stored in the form of an adjacency list: each node contains attributes such as a unique identifier, node type (technical means / technical object), name, and context information; each edge contains attributes such as a source node identifier, a target node identifier, a weight value, and co-occurrence frequency. The network is stored as a JSON file, where the node list (nodes) contains all entity nodes, and the edge list (edges) contains all associated edges. For example, the node list contains {"id":"TMI-0001", "type":"technical means", "name":"backpropagation algorithm", ...}, and the edge list contains {"source":"TMI-0001", "target":"TOI-0001", "weight":46, ...}. This current technical element network comprehensively reflects the relationships between technical means and technical objects in the field of artificial intelligence.
[0061] Step S130: The current technical element network is divided into time slices according to a preset time slice division rule to generate a time-series-based set of technical element network segments. The time-series-based set of technical element network segments contains multiple technical element network segment units arranged in chronological order.
[0062] Step S131: Obtain the publication time information corresponding to each paper and each patent document in the scientific and technological literature data set. The publication time information includes year information and month information.
[0063] The publication date of each document is extracted from the metadata of the scientific literature database. Papers typically include a "publication date" (e.g., "2020-06-15"), while patents include a "publication date" (e.g., "2019-11-20"). The time information is parsed into year and month (e.g., June 2020) and stored as a tuple in (year, month) format. For documents without a clearly specified month, January is defaulted to; documents with only a year are not included in the time series analysis to ensure the accuracy of the time information.
[0064] Step S132: Based on the publicly disclosed time information, map all technical means entity nodes and all technical object entity nodes onto the corresponding time coordinate axis. Each entity node carries the timestamp information of its first appearance and the time sequence information of each appearance.
[0065] For each entity node, collect the publication dates of all documents containing its corresponding phrase unit, extract the first appearance timestamp (earliest (year, month)) and the sequence of all appearance times (a list of (year, month) sorted by time). For example, the first appearance of "Transformer model" was in June 2017, and the sequence of appearance times is [(2017, 6), (2018, 3), (2018, 10), ...]. Store this time information as the "first_appearance" and "appearance_times" attributes of the entity node.
[0066] Step S133: Map each associated edge in the current technical element network to the corresponding time coordinate axis. Each associated edge carries the timestamp information of its first establishment based on the co-occurrence relationship and the time point sequence information of each co-occurrence.
[0067] For each associated edge (technical means entity → technical object entity), the publication dates of all documents containing the co-occurring phrases of that entity are collected, and the first co-occurrence timestamp and the sequence of all co-occurrence time points are extracted. For example, the first co-occurrence time of the associated edge between "Transformer model" and "natural language processing" is June 2017, and the sequence of co-occurrence time points is [(2017, 6), (2018, 5), ...]. The time information is stored as the "first_cooccurrence" and "cooccurrence_times" attributes of the associated edge.
[0068] Step S134: Divide the time coordinate axis evenly into multiple continuous time segment units according to the preset time slice length parameter.
[0069] The preset time slice length is 1 year (12 months), with the earliest year of the scientific literature database as the start year and the latest year of the database as the end year (e.g., 2010-2023). The time axis is divided into 14 time slice units, each corresponding to a calendar year (e.g., 2010, 2011, ..., 2023). The time slice unit is represented as (start year, end year), such as (2020, 2020) representing the 2020 time slice.
[0070] Step S135: For each time segment unit, select all technical means entity nodes and all technical object entity nodes that have appeared in the current technical element network to form a node subset of the time segment unit.
[0071] For each time segment (e.g., 2020), iterate through the "appearance_times" attribute of all entity nodes. If a (year, month) (e.g., any month in 2020) exists in the time sequence belonging to that time segment, then include that entity node in the node subset. For example, "Transformer model" has a record in 2020, therefore it is included in the 2020 node subset. The node subset contains technical means entity nodes and technical object entity nodes, each labeled with its type.
[0072] Step S136: For each time segment unit, select all associated edges established or existing within the current technical element network to form a subset of edges for the time segment unit. Each selected associated edge retains its edge weight value within the time segment unit.
[0073] For each time segment, iterate through the "cooccurrence_times" attribute of all associated edges. If a time point belonging to that time segment exists in the co-occurrence time point sequence, then include that associated edge in the edge subset. Simultaneously, calculate the edge weight value of that associated edge within that time segment: calculate the co-occurrence frequency within that time segment and multiply it by the logarithmic function of the global co-occurrence frequency (consistent with step S129). For example, if "Transformer model" and "Natural Language Processing" co-occurred 5 times in 2020, with a global co-occurrence frequency of 100, then the edge weight value within that time segment is 5 × log(1 + 100) ≈ 23.
[0074] Step S137: Based on the node subset and edge subset of each time segment unit, construct the technology element network segment unit corresponding to that time segment unit. The technology element network segment unit includes active technology means entity nodes, technology object entity nodes, and associated edges connecting the entity nodes within that time segment unit.
[0075] Step S1371: Obtain the node subset corresponding to each time segment unit. The node subset includes multiple technical means entity nodes and multiple technical object entity nodes. Each entity node carries the number of times the entity node appears in the time segment unit and the time of its first appearance.
[0076] Extract the number of times each entity node appears in the current time segment (i.e., the number of time points in "appearance_times" belonging to this time segment) and the first appearance time (global first appearance time) from the node subset. For example, the "Transformer model" appears 8 times in 2020, and its first appearance time is June 2017.
[0077] Step S1372: Obtain the edge subset corresponding to each time segment unit. The edge subset contains multiple associated edges. Each associated edge carries the edge weight value of the associated edge in the time segment unit and the co-occurrence information of the associated edge in the time segment unit.
[0078] Extract the edge weight value (calculated in step S136) and co-occurrence count (number of co-occurrence time points in the current time segment) of each associated edge from the edge subset.
[0079] Step S1373: Based on each technical means entity node and each technical object entity node in the node subset, create a corresponding entity node object in the technical element network segment unit, and set a node attribute field for each entity node object. The node attribute field includes an entity node identifier field, an entity node type field, a number of occurrences field, and a first occurrence time field.
[0080] Create an object for each entity node, with attribute fields including: entity node identifier (e.g., "TMI-0001"), entity node type ("technical means" or "technical object"), occurrence count (number of occurrences within the current time segment), and first occurrence time (global first occurrence time). For example, the attributes of the "Transformer model" node in the 2020 network segment are {"id":"TMI-0002", "type":"technical means", "occurrences":8", "first_appearance":(2017, 6)}.
[0081] Step S1374: Based on each associated edge in the edge subset, create a corresponding associated edge object in the technical element network segment unit. The associated edge object connects the specified technical means entity node and technical object entity node in the edge subset. Set an edge attribute field for each associated edge object. The edge attribute field includes an edge weight value field and a co-occurrence count field.
[0082] Create an object for each associated edge, connecting the source node (technical means) and the target node (technical object). The attribute fields include the edge weight value (weight within the current time segment) and the co-occurrence count (co-occurrence count within the current time segment). For example, in the 2020 network segment, the associated edge attribute for "Transformer model → Natural Language Processing" is {"source":"TMI-0002", "target":"TOI-0003", "weight":23", "cooccurrences":5}.
[0083] Step S1375: Normalize the occurrence count field of all entity node objects in the network segment unit of the technical element, divide the occurrence count by the total occurrence count of all entity nodes in the time segment unit, and obtain the relative occurrence frequency of each entity node in the time segment unit.
[0084] Calculate the total number of occurrences of all entity nodes within the current time segment. The relative frequency of each node is calculated as: node occurrence count / total occurrence count. For example, if the total occurrence count in 2020 is 1000, and the "Transformer model" appears 8 times, the relative frequency is 0.008. This value serves as an initial measure of node importance.
[0085] Step S1376: Normalize the edge weight value field of all associated edge objects in the network segment unit of the technical element, divide the edge weight value by the total edge weight value of all associated edges in the time segment unit, and obtain the relative association strength of each associated edge in the time segment unit.
[0086] Calculate the sum of the weights of all related edges within the current time segment. The relative correlation strength of each edge = edge weight value / total weight value. For example, if the total weight value in 2020 is 5000, and a certain related edge has a weight of 23, the relative correlation strength is 0.0046.
[0087] Step S1377: Use the relative frequency of occurrence as the initial value of the node importance of the entity node object in the technical element network segment unit, and use the relative correlation strength as the initial value of the edge importance of the associated edge object in the technical element network segment unit to complete the construction of the technical element network segment unit corresponding to the time segment unit.
[0088] The relative frequency of occurrence is assigned to the "importance" attribute of the node, and the relative association strength is assigned to the "importance" attribute of the edge. At this point, the network segment unit corresponding to each time segment unit is completed, containing a list of nodes, a list of edges, and their respective importance attributes.
[0089] Step S138: Arrange the technical element network segment units corresponding to all time segment units in the order of the time segment units to generate a time-seriesd technical element network segment set. Each technical element network segment unit in the time-seriesd technical element network segment set corresponds to a specific time window.
[0090] The network segments of technological elements, corresponding to 14 time segments (2010-2023), are arranged chronologically to form a time-series list. Each network segment contains information about its corresponding time window (e.g., "2020"). This collection of technological element network segments is stored as an ordered JSON array, with each element containing the network segment data for each time segment. This collection allows for a visual observation of the evolution of the technological element network over time.
[0091] Example Implementation Section: Step S140: Perform node temporal importance index calculation and edge temporal importance index calculation on the time-seriesd set of technology element network segments, determine the technology community structure with a comprehensive score of technology evolution vitality exceeding a preset vitality threshold as an emerging technology cluster, and generate an emerging technology cluster set.
[0092] In this embodiment, within the context of artificial intelligence, a time-series-based set of network segments of technological elements is processed. The calculation of node temporal importance indicators aims to assess the changes in the importance of entity nodes across different time segments, while the calculation of edge temporal importance indicators focuses on the evolution of the weights of associated edges over time. The technological community structure is a tightly connected group of nodes in the network; by calculating its comprehensive score of technological evolution vitality, emerging technology clusters with development potential can be identified. First, for each entity node, the node degree time series is extracted from the network segment set, the node degree change rate is calculated and smoothed to obtain the node influence growth rate feature. For associated edges, the edge weight time series is extracted, and the edge weight change rate and edge formation acceleration feature are calculated. Then, community detection is performed on the network segment units for each time segment to identify the technological community structure and track its evolution process, extracting stability change features. Finally, based on the node influence growth rate, edge formation acceleration, and community stability change features, the comprehensive score of the technological evolution vitality of the technological community structure is calculated, and communities exceeding a preset vitality threshold are identified as emerging technology clusters.
[0093] Step S141: For each technical means entity node and each technical object entity node, extract the node degree value of the entity node in each time segment unit from the time-seriesd technical element network segment set. The node degree value represents the number of associated edges connected to the entity node, and generate the node degree time series of the entity node.
[0094] From the time-seriesd set of network segments of technical elements, each technical means entity node and technical object entity node is traversed. For each entity node, the corresponding technical element network segment unit for each time segment is examined sequentially, and the number of associated edges connected to the node in that time segment unit is counted; this number is the node degree value. The node degree values of each time segment unit are arranged in chronological order to form the node degree time series of that entity node. For example, if a technical means entity node has a node degree value of A in time segment unit 1, a node degree value of B in time segment unit 2, and a node degree value of C in time segment unit 3, then its node degree time series is [A, B, C]. This time series reflects the changes in the breadth of association of the entity node in different time segments.
[0095] Step S142: Based on the node degree time series, calculate the change in node degree value of the entity node between adjacent time segment units, and divide the change in node degree value by the length of the time segment unit to obtain the node degree change rate of the entity node in each time segment unit.
[0096] Based on the node degree time series of entity nodes, the change in node degree values between adjacent time segment units is calculated. For the node degree values of the i-th and (i-1)-th time segment units in the time series, the change is the difference between the i-th and (i-1)-th node degree values. The length of the time segment unit is a preset fixed time interval, such as one year. The node degree change rate is obtained by dividing the change in node degree value by the length of the time segment unit. For example, if the node degree time series is [A, B, C], the adjacent changes are BA and CB respectively, and the time segment unit length is D, then the node degree change rates are (BA) / D and (CB) / D respectively. The node degree change rate reflects the rate of change of the node's association breadth.
[0097] Step S143: Perform sliding window averaging on the node degree change rate sequence to eliminate the influence of short-term fluctuations, generate a smoothed node degree change rate sequence, and identify the continuous rising phase from the smoothed node degree change rate sequence. Use the cumulative increase of the node degree change rate within the continuous rising phase as the node influence growth rate feature of the entity node.
[0098] The sliding window average method is used to process the node degree change rate sequence. The size of the sliding window is a preset fixed value, such as E time segment units. For each element in the change rate sequence, the average value of a total of E elements before and after it is taken as the smoothed value, generating a smoothed node degree change rate sequence. From the smoothed sequence, identify the stage where the change rate continuously increases within multiple consecutive time segment units, that is, the continuous rising stage. Add up all the node degree change rates within this stage to obtain the cumulative increase, which is used as the node influence growth rate feature of the entity node. For example, the smoothed change rate sequence is [F, G, H], where F < G < H, the continuous rising stage is these three time segment units, and the cumulative increase is F + G + H.
[0099] Step S144: For each associated edge, extract the edge weight value of this associated edge in each time segment unit from the time - serialized set of technical element network segments, generating the edge weight time series of this associated edge.
[0100] For each associated edge, traverse the time - serialized set of technical element network segments. In the network segment unit corresponding to each time segment unit, find the edge weight value of this associated edge. Arrange the above - mentioned edge weight values in the order of time segment units to form an edge weight time series. For example, the edge weight value of a certain associated edge in time segment unit 1 is I, in time segment unit 2 is J, and in time segment unit 3 is K, then its edge weight time series is [I, J, K]. This edge weight time series reflects the change of the association strength of the associated edge over time.
[0101] Step S145: According to the edge weight time series, calculate the change amount of the edge weight value between adjacent time segment units of this associated edge, and divide the change amount of the edge weight value by the length of the time segment unit to obtain the edge weight change rate of this associated edge, and take the difference between the edge weight change rates of adjacent time segment units as the edge formation acceleration feature of this associated edge.
[0102] According to the edge weight time series, calculate the change amount of the edge weight value between adjacent time segment units, that is, subtract the edge weight value of the previous time segment unit from the edge weight value of the subsequent time segment unit. Divide the change amount by the length of the time segment unit to obtain the edge weight change rate. Then, calculate the difference between adjacent edge weight change rates, that is, subtract the previous change rate from the subsequent change rate, and this difference is the edge formation acceleration feature. For example, the edge weight time series is [I, J, K], the adjacent change amounts are J - I and K - J, the length of the time segment unit is D, the edge weight change rates are (J - I) / D and (K - J) / D, and the edge formation acceleration feature is [(K - J) / D-(J - I) / D].
[0103] Step S146: Perform community detection processing on the network segment unit of the technology element corresponding to each time segment unit, and identify multiple technology community structures contained in each network segment unit of the technology element. The technology community structure is composed of closely connected technology means entity nodes and technology object entity nodes.
[0104] A community detection algorithm is used to process the technical element network segment units of each time segment. The algorithm analyzes the connections between nodes in the network and groups closely connected nodes into the same community. In this embodiment, a community detection algorithm based on modularity optimization is used. Modularity is an indicator of the quality of community partitioning; a higher value indicates a better partition. By iteratively optimizing the modularity, technical means entity nodes and technical object entity nodes in the technical element network segment units are divided into multiple technical community structures. The connections between nodes within each technical community structure are relatively dense, while the connections between communities are relatively sparse.
[0105] Step S147: Track the evolution of the same technical community structure in consecutive time segments, calculate the overlap of community members between adjacent time segments for each technical community structure, generate a stability change curve for the technical community structure based on the overlap of community members, and extract the stability change features of the community structure from the stability change curve. The stability change features of the community structure include the identifiers of the rapid expansion phase and the identifiers of the rapid contraction phase of community members.
[0106] For each technical community structure, tracking is performed within consecutive time-segment units. The overlap degree of community members in adjacent time-segment units is calculated, by dividing the number of nodes shared by the two time-segment units by the sum of the total number of nodes in both communities. Based on the overlap degree values of different time-segment units, stability change curves are plotted. Phases of rapid increase in the number of community members are identified from the curves and marked as rapid expansion phases; phases of rapid decrease in the number of community members are identified and marked as rapid shrinkage phases. These phase markers constitute the stability change characteristics of the community structure.
[0107] Step S148: For each technical community structure, calculate the corresponding node growth score, edge evolution score, and community stability score based on the node influence growth rate characteristics of the entity nodes it contains, the edge formation acceleration characteristics of the associated edges, and the community structure stability change characteristics; and weight the node growth score, edge evolution score, and community stability score to generate a comprehensive score of technical evolution vitality for each technical community structure.
[0108] For example, step S1481: For each technical community structure, calculate the average value of the node influence growth rate characteristics of all technical means entity nodes and technical object entity nodes within it, and use it as the node growth rate benchmark value of the technical community structure; compare and normalize the node growth rate benchmark value with the node growth rate benchmark values of all technical community structures in the same time segment unit to generate the node growth score of the technical community structure.
[0109] The node influence growth rate characteristics of all entity nodes within the technical community structure are collected, and their arithmetic mean is calculated to obtain a baseline value for the node growth rate. This baseline value is compared with the baseline values of the node growth rates of all other technical community structures within the same time segment, and normalized to map it to the [0, 1] interval to obtain a node growth score. The normalization process uses the min-max normalization method, i.e., (benchmark value of node growth rate - minimum value) / (maximum value - minimum value), where the minimum and maximum values are the minimum and maximum values of the baseline values of the node growth rates of all communities within the same time segment, respectively.
[0110] Step S1482: For each technical community structure, calculate the average value of the edge formation acceleration characteristics of all associated edges within it, and use it as the edge acceleration benchmark value of the technical community structure; compare and normalize the edge acceleration benchmark value with the edge acceleration benchmark values of all technical community structures in the same time segment unit to generate the edge evolution score of the technical community structure.
[0111] The arithmetic mean of the edge formation acceleration characteristics of all associated edges within the computational technology community structure is used as the baseline value for edge acceleration. Using the same min-max normalization method as the node growth score, the baseline value for edge acceleration is compared with and normalized to the baseline values for edge acceleration of other communities within the same time segment to obtain the edge evolution score.
[0112] Step S1483: For each technical community structure, based on the rapid expansion phase identifier and rapid shrinkage phase identifier of community members in its community structure stability change characteristics, calculate a scalar evaluation value reflecting its structural stability; compare and normalize the scalar evaluation value with the scalar evaluation values of all technical community structures in the same time segment unit to generate the community stability score of the technical community structure.
[0113] Based on the stage indicators in the characteristics of community structural stability changes, a scalar evaluation value is assigned to community stability. If the community is currently in a rapid expansion stage, the scalar evaluation value is M; if it is in a stable stage, the scalar evaluation value is N; and if it is in a rapid shrinkage stage, the scalar evaluation value is P, where M>N>P. This scalar evaluation value is compared with the scalar evaluation values of all communities within the same time segment, and a min-max normalization process is used to obtain the community stability score.
[0114] Step S1484: Multiply the node growth score, edge evolution score and community stability score by preset weight coefficients and then sum them to obtain the comprehensive score of the technological evolution vitality of each technical community structure.
[0115] The weighting coefficients for node growth score and edge evolution score are predefined as Q, R, and S, respectively, with Q + R + S = 1. The node growth score is multiplied by Q, the edge evolution score by R, and the community stability score by S, and then the three products are added together to obtain the comprehensive score for technological evolution vitality.
[0116] Step S149: The technology community structure whose comprehensive score of technological evolution vitality exceeds the preset vitality threshold is identified as an emerging technology cluster, and all technical means entity nodes and technical object entity nodes in the technology community structure are aggregated to generate an emerging technology cluster set. Each emerging technology cluster in the emerging technology cluster set corresponds to a technology community structure whose comprehensive score of technological evolution vitality exceeds the preset vitality threshold.
[0117] A vitality threshold T is preset, and technology community structures with a comprehensive technology evolution vitality score greater than T are identified as emerging technology clusters. All technical means entity nodes and technical object entity nodes contained in each emerging technology cluster are collected and aggregated to form an emerging technology cluster set. Each emerging technology cluster includes information such as a cluster identifier, a node list, and a comprehensive score.
[0118] Step S150: Perform dynamic graph representation learning processing and future association edge formation prediction processing on the time-seriesd set of network fragments of technical elements to generate a potential technology combination prediction set containing predicted association edges. Perform technical field crossover degree evaluation processing and novelty evaluation processing on each predicted association edge in the potential technology combination prediction set to generate a predicted association edge evaluation result set.
[0119] In the field of artificial intelligence, dynamic graph representation learning is applied to a time-series set of network fragments of technological elements to capture the network's dynamic evolutionary characteristics. Dynamic graph representation learning learns the evolutionary feature representations of entity nodes by combining graph structure information and time-series information. Based on these feature representations, future association edge formation is predicted, forecasting the possible future association edges between technological means entity nodes and technological object entity nodes, generating a set of potential technology combination predictions. Then, each predicted association edge undergoes a technological crossover assessment and a novelty assessment; the assessment results are used for subsequent analysis of technology integration directions.
[0120] Step S151: Construct a dynamic graph neural network model, which includes a graph convolutional network layer, a temporal modeling network layer, and a future edge prediction output layer. The graph convolutional network layer is used to learn the structural features of the technical element network segment unit within each time segment unit, and the temporal modeling network layer is used to learn the evolution law of the technical element network segment unit in the time dimension.
[0121] The dynamic graph neural network model consists of graph convolutional network layers, temporal modeling network layers, and a future edge prediction output layer. The graph convolutional network layers employ multi-layer graph convolution operations, with each layer containing multiple neurons, learning the structural features of nodes by aggregating features from neighboring nodes. The temporal modeling network layers use a gated recurrent unit structure, including input gates, forget gates, and output gates, to process time-series data and capture the network's evolution across different time segments. The future edge prediction output layer uses a fully connected network structure, outputting the probability of future edge formation between pairs of entity nodes. The model's input is a time-series-based set of network segments containing technical elements, and its output is a probability matrix for future edge formation.
[0122] Step S152: Input each technical element network segment unit in the time-seriesd technical element network segment set into the graph convolutional network layer, and perform neighbor feature aggregation processing on the technical means entity nodes and technical object entity nodes in each technical element network segment unit through the graph convolutional network layer to generate a set of node structure feature vectors corresponding to each time segment unit.
[0123] Each time-segment unit is represented as a graph-structured data unit, including a node feature matrix and an adjacency matrix. The node feature matrix contains attribute information of entity nodes, such as node type and frequency of occurrence. The adjacency matrix represents the relationships between nodes. The graph-structured data is input into graph convolutional network layers. Each layer of the graph convolutional network aggregates the neighbor features of a node through weighted summation, updating the node's feature representation. After multiple layers of graph convolutional operations, a set of node structural feature vectors corresponding to each time-segment unit is generated, with each vector reflecting the structural characteristics of the node in that time segment.
[0124] Step S153: Input the set of node structure feature vectors into the temporal modeling network layer in chronological order. The temporal modeling network layer uses a gated recurrent unit structure to perform time dependency modeling on the set of node structure feature vectors, capture the feature evolution trajectory of each entity node in continuous time slices, and generate the evolution feature representation vector corresponding to each entity node.
[0125] The set of node structural feature vectors from different time segments is arranged chronologically and input into the temporal modeling network layer. Gated recurrent units control the input of the current feature through input gates, forget gates determine how many historical features are retained, and output gates output the current hidden state. In this way, the temporal modeling network layer can capture the feature changes of entity nodes over consecutive time segments, generating an evolutionary feature representation vector. This evolutionary feature representation vector integrates the structural features and temporal evolution features of the node, more comprehensively reflecting the characteristics of the node.
[0126] Step S154: Concatenate the evolution feature representation vectors of all entity nodes to generate a global evolution feature representation matrix that can characterize the evolution mode of the entire set of technology element network segments.
[0127] The evolutionary feature representation vectors of each entity node are concatenated in the order of the nodes to form a two-dimensional matrix, namely the global evolutionary feature representation matrix. The rows of the global evolutionary feature representation matrix correspond to different entity nodes, and the columns correspond to the dimensions of the evolutionary feature representation vectors. This global evolutionary feature representation matrix integrates the evolutionary information of all entity nodes and can characterize the evolutionary pattern of the entire technological element network.
[0128] Step S155: Input the global evolution feature representation matrix into the future edge prediction output layer. The future edge prediction output layer calculates the probability of forming a related edge between each pair of technical means entity nodes and technical object entity nodes in the future time segment unit, and generates a probability matrix of future related edges between all entity node pairs.
[0129] The global evolutionary feature representation matrix is input into the future edge prediction output layer, which contains multiple fully connected layers and activation functions. The evolutionary features of entity nodes are processed through the fully connected layers to calculate the probability of forming an associated edge between each pair of technical means entity nodes and technical object entity nodes. The probability values are mapped to the [0, 1] interval using a sigmoid activation function, generating a future associated edge formation probability matrix. Rows in the matrix correspond to technical means entity nodes, columns correspond to technical object entity nodes, and matrix elements represent the probability of forming an associated edge between corresponding node pairs.
[0130] Step S156: Select entity node pairs whose probability values exceed a preset probability threshold from the future association edge formation probability matrix, use the entity node pairs as candidate entity node pairs, and generate a predicted association edge for each candidate entity node pair. Each predicted association edge includes the initiator entity node identifier, the receiver entity node identifier, and the future association edge formation probability value.
[0131] A predetermined probability threshold U is set, and entity node pairs with probability values greater than U are selected as candidate entity node pairs from the probability matrix of future associated edges. A predicted associated edge is generated for each candidate entity node pair. The attributes of the associated edge include the identifier of the initiating entity node (technical means entity node), the identifier of the receiving entity node (technical object entity node), and the probability value of future associated edge formation.
[0132] Step S157: Aggregate all candidate entity node pairs and their corresponding predicted association edges to generate a potential technology combination prediction set containing multiple predicted association edges.
[0133] All generated predictive edges are combined to form a potential technology combination prediction set. Each predictive edge in this potential technology combination prediction set represents a possible future technology combination.
[0134] Step S158: For each predicted associated edge in the potential technology combination prediction set, obtain the first entity node and the second entity node connected by the predicted associated edge, wherein the first entity node and the second entity node are respectively a technology means entity node or a technology object entity node.
[0135] Iterate through each predicted association edge in the potential technology combination prediction set, and extract the two entity nodes connected by the association edge, namely the first entity node and the second entity node. Among them, the first entity node is the technology means entity node, and the second entity node is the technology object entity node.
[0136] Step S159: Calculate the technical field distance between the first entity node and the second entity node based on the technical field classification label of the first entity node and the technical field classification label of the second entity node. The technical field distance is determined based on the hierarchical path length in the International Patent Classification System or the subject classification system.
[0137] Each entity node is pre-assigned a technical field classification label, which is determined based on the International Patent Classification (IPC) or a subject-specific classification system. For example, the IPC divides technical fields into multiple departments, major categories, and minor categories. The hierarchical path length of the technical fields to which two entity nodes belong in the classification system is calculated, which is the number of hierarchical nodes traversed from one technical field to another. This hierarchical path length is the technical field distance.
[0138] Step S1510: Map the distance of the technical field to the technical field intersection score interval, and generate the technical field intersection score of the predicted association edge. The technical field intersection score is positively correlated with the technical field distance.
[0139] The scoring interval for the degree of crossover in technical fields is set to [0, 1]. The distance between technical fields is mapped to this interval, with a higher crossover score indicating a greater distance between them. For example, using a linear mapping formula, a technical field distance of V is mapped to (V / maximum possible distance) to obtain the degree of crossover score.
[0140] Step S1511: Obtain all existing associated edges in the current technology element network to form an existing technology combination set, and determine whether the combination of the first entity node and the second entity node connected by the predicted associated edge appears in the existing technology combination set.
[0141] Extract all existing associated edges from the current technology element network and store the entity node pairs of these associated edges in the existing technology combination set. For the entity node pairs of the predicted associated edges, check whether they exist in the existing technology combination set.
[0142] Step S1512: If the combination of the first entity node and the second entity node does not appear in the existing technology combination set, then the predicted association edge is determined to have global novelty, and a global novelty identifier is assigned to the predicted association edge.
[0143] If an entity node pair does not appear in the existing set of technology combinations, it means that the technology combination has never appeared in the current network and has global novelty. A global novelty identifier, such as "GN", is assigned to it.
[0144] Step S1513: If the combination of the first entity node and the second entity node appears in the existing technology combination set, then the most recent occurrence time of the combination in the time-seriesd technology element network segment set is further determined. If the most recent occurrence time is more than a preset time interval threshold away from the current time segment unit, then the predicted association edge is determined to have revival novelty, and a revival novelty identifier is assigned to the predicted association edge.
[0145] If an entity node pair appears in the existing technology combination set, find its most recent occurrence time in the time-seriesd network segment set. Calculate the time interval between the most recent occurrence time and the current time segment unit. If this interval exceeds a preset time interval threshold W, the predicted association edge is determined to have revival novelty, and a revival novelty identifier, such as "RN", is assigned.
[0146] Step S1514: Based on the allocation of the global novelty identifier or the revival novelty identifier, and combined with the future association edge formation probability value of the predicted association edge, generate a comprehensive novelty score for the predicted association edge.
[0147] For predicted association edges with a global novelty identifier, the comprehensive novelty score = probability of future association edge formation × X, where X is the global novelty weight coefficient. For predicted association edges with a revitalization novelty identifier, the comprehensive novelty score = probability of future association edge formation × Y, where Y is the revitalization novelty weight coefficient, and X > Y. For association edges without novelty, the comprehensive novelty score = probability of future association edge formation × Z, where Z... <Y。
[0148] Step S1515: Combine the technical field crossover score, novelty comprehensive score, and future association edge formation probability value of each predicted association edge to generate the predicted association edge evaluation result corresponding to each predicted association edge. The predicted association edge evaluation result includes the technical field crossover score field, the novelty comprehensive score field, and the future association edge formation probability field.
[0149] The evaluation results for each predicted related edge are formed by combining three assessment metrics: the degree of cross-disciplinary collaboration score, the overall novelty score, and the probability of future related edge formation. The evaluation results are stored in structured data format, containing the corresponding three fields.
[0150] Step S1516: Aggregate the evaluation results of all predicted associated edges to generate a set of predicted associated edge evaluation results.
[0151] The evaluation results of all predicted associated edges are integrated into a set to form a predicted associated edge evaluation result set, which is used for subsequent analysis and ranking of technology integration directions.
[0152] Step S160: Based on the node growth trend of the technical means entity nodes and technical object entity nodes in the emerging technology cluster set, and combined with the technical field crossover assessment results and novelty assessment results in the predicted association edge assessment result set, generate an emerging technology topic list containing multiple emerging technology topic identifiers and a technology integration direction list containing multiple technology integration direction identifiers, wherein the entries in the technology integration direction list are sorted based on their technical field crossover score and novelty comprehensive score.
[0153] By combining information from the set of emerging technology clusters and the set of predicted association edge evaluation results, a list of emerging technology topics and a list of technology integration directions are generated. The list of emerging technology topics is determined based on the growth trend of entity nodes in the emerging technology clusters, while the list of technology integration directions is sorted according to the evaluation results of the predicted association edges. First, entity nodes are extracted from the set of emerging technology clusters, and their growth trends are analyzed to generate emerging technology topics. Then, association edges with high crossover and high novelty are selected from the set of predicted association edge evaluation results to generate technology integration directions, which are then sorted according to the evaluation results. Finally, these two lists are output.
[0154] Step S161: Extract all technical means entity nodes and technical object entity nodes contained in each emerging technology cluster from the emerging technology cluster set, obtain the node influence growth rate characteristics of each entity node, and perform weighted average processing on the node influence growth rate characteristics of all entity nodes in each emerging technology cluster to generate an overall growth trend score for each emerging technology cluster.
[0155] For each emerging technology cluster, the growth rate characteristics of the node influence of all technical means entity nodes and technical object entity nodes within it are collected. Each node is assigned a weight based on its importance within the cluster (e.g., node degree or frequency of occurrence), and then the node influence growth rate characteristics are weighted and averaged to obtain the overall growth trend score for that emerging technology cluster. The weights can be assigned as a proportion of node importance to total importance.
[0156] Step S162: Perform topic summarization processing on all technical means entity nodes and technical object entity nodes within each emerging technology cluster, extract the core entity node with the highest frequency from the entity nodes of each emerging technology cluster, and generate the emerging technology topic description text corresponding to the emerging technology cluster based on the technical phrase unit corresponding to the core entity node.
[0157] Frequency analysis is performed on entity nodes within emerging technology clusters to extract the top-frequency core entity nodes. Based on the technical phrase units corresponding to these core entity nodes and the relationships between them, a descriptive text for the emerging technology topic is generated. The descriptive text should concisely and clearly summarize the core technical content of the technology cluster.
[0158] Step S163: Assign an emerging technology topic identifier to each emerging technology cluster, and associate and store the emerging technology topic identifier, the emerging technology topic description text, and the overall growth trend score to generate an emerging technology topic record unit.
[0159] Assign a unique emerging technology topic identifier, such as "ET-Z," to each emerging technology cluster. Link the identifier, topic description text, and overall growth trend score to form an emerging technology topic record unit, which facilitates subsequent storage and retrieval.
[0160] Step S164: Sort all emerging technology clusters and their corresponding emerging technology topic record units from high to low according to the overall growth trend score, and generate an emerging technology topic list containing multiple emerging technology topic identifiers.
[0161] The emerging technology topic recording units are sorted from high to low according to their overall growth trend scores to form an emerging technology topic list. Each entry in the list includes an emerging technology topic identifier, topic description text, and overall growth trend score.
[0162] Step S165: Select predicted related edges from the predicted related edge evaluation result set that have a technical field crossover score exceeding a preset crossover threshold and a novelty comprehensive score exceeding a preset novelty threshold, and form a set of filtered predicted related edges.
[0163] The preset threshold for crossover is AA, and the threshold for novelty is BB. Predicted related edges with a crossover score greater than AA and a comprehensive novelty score greater than BB are selected from the predicted related edge evaluation result set, forming a filtered set of predicted related edges.
[0164] Step S166: Parse each predicted association edge in the filtered predicted association edge set, obtain the first entity node and the second entity node connected by each predicted association edge, and generate the technology fusion direction description text corresponding to the predicted association edge based on the technology phrase unit corresponding to the first entity node and the second entity node.
[0165] After parsing and filtering, each associated edge in the predicted associated edge set is analyzed to obtain the technical means entity nodes and technical object entity nodes connected to it. Based on the technical phrase units corresponding to these two nodes and their technical meanings, a technical integration direction description text is generated, describing the potential application directions of this technical combination.
[0166] Step S167: Assign a technology fusion direction identifier to each predicted association edge in the filtered predicted association edge set, and associate and store the technology fusion direction identifier, the technology fusion direction description text, the technology field crossover score, and the novelty comprehensive score to generate a technology fusion direction record unit.
[0167] Assign a unique technology fusion direction identifier, such as "TF-ZZ", to each predicted association edge. Store the identifier, fusion direction description text, technology field crossover score, and novelty comprehensive score together to form a technology fusion direction record unit.
[0168] Step S168: Sort the technology fusion direction record units corresponding to all predicted associated edges in the filtered predicted associated edge set from high to low according to the weighted sum of the technical field crossover score and the novelty comprehensive score, and generate a technology fusion direction list containing multiple technology fusion direction identifiers.
[0169] The technology convergence direction recording unit is calculated by weighting the score for the degree of cross-disciplinary collaboration and the comprehensive score for novelty, with weights CC and DD, respectively, and CC+DD=1. Recording units are then sorted from highest to lowest weighted sum to generate a list of technology convergence directions.
[0170] Step S169: Link and output the emerging technology topic list and the technology integration direction list. Each emerging technology topic identifier in the emerging technology topic list corresponds to a technology cluster whose comprehensive score of technological evolution vitality exceeds a preset vitality threshold. Each technology integration direction identifier in the technology integration direction list corresponds to a predicted association edge whose technical field crossover score exceeds a preset crossover threshold and whose comprehensive novelty score exceeds a preset novelty threshold.
[0171] The list of emerging technology topics and the list of technology convergence directions are output in a structured form, such as through tables or JSON format. It is ensured that each emerging technology topic identifier and technology convergence direction identifier corresponds to a relevant technology cluster and predicted association edge, facilitating user viewing and analysis. In an exemplary embodiment, a hot technology mining system based on scientific literature data is provided. This system can be a terminal, server, etc., and its internal structure diagram can be as follows: Figure 2As shown, this hotspot technology mining system based on scientific and technological literature data includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a hotspot technology mining method based on scientific and technological literature data. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the shell of the hot technology mining system based on scientific and technological literature data, or an external keyboard, touchpad, or mouse, etc.
[0172] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for mining hot technologies based on scientific and technological literature data, characterized in that, The method includes: A set of scientific and technological literature data is obtained, and technical phrase extraction and semantic normalization processing are performed on the set of scientific and technological literature data to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components. The current technology element network is divided into time slices according to a preset time slice division rule to generate a time-series-based set of technology element network segments. The time-series network fragment set of technical elements is processed by calculating the temporal importance index of nodes and the temporal importance index of edges. Technical community structures whose comprehensive scores of technical evolution vitality exceed a preset vitality threshold are identified as emerging technology clusters, and a set of emerging technology clusters is generated. The time-series network fragment set of technical elements is subjected to dynamic graph representation learning and future association edge formation prediction processing to generate a potential technology combination prediction set containing predicted association edges. Each predicted association edge in the potential technology combination prediction set is subjected to technical field crossover degree evaluation processing and novelty evaluation processing to generate a predicted association edge evaluation result set. Based on the node growth trend of the technical means entity nodes and technical object entity nodes in the emerging technology cluster set, and combined with the technical field crossover evaluation results and novelty evaluation results in the predicted association edge evaluation result set, an emerging technology theme list containing multiple emerging technology theme identifiers and a technology integration direction list containing multiple technology integration direction identifiers are generated.
2. The method for hotspot technology mining based on scientific and technological literature data according to claim 1, characterized in that, The scientific and technological literature data set includes a subset of academic paper data and a subset of patent literature data. The subset of academic paper data includes abstract text units and full-text text units. The subset of patent literature data includes abstract text units and claims text units. The process of extracting technical phrases and normalizing their semantics from the scientific literature dataset to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components includes: Sentence boundary detection processing is performed on the paper abstract text unit, the paper full text text unit, the patent abstract text unit, and the patent claims text unit to obtain multiple sentence units corresponding to each text unit; The multiple sentence units are subjected to part-of-speech tagging and dependency parsing to generate a part-of-speech tag sequence and dependency tree structure for each sentence unit. Based on the part-of-speech tag sequence and the dependency tree structure, noun phrase combinations that conform to the preset technical phrase composition pattern are extracted from each sentence unit. The preset technical phrase composition pattern includes verb and noun combination pattern, adjective and noun combination pattern, and noun and noun compound combination pattern. All extracted noun phrase combinations are subjected to stop word filtering to remove general technical terms that appear in a preset technical stop word list, resulting in a candidate technical phrase set. The preset technical stop word list contains general technical terms for methods, systems, devices, and equipment. Each candidate technical phrase in the candidate technical phrase set is subjected to phrase structure analysis processing to identify the core word component and the modifier component in each candidate technical phrase. Based on the core word component and the modifier component, the candidate technical phrase is divided into technical means phrase type or technical object phrase type, and a set of technical phrase units carrying phrase type labels is generated. The phrase type labels include technical means phrase labels and technical object phrase labels. Semantic normalization is performed on all technical phrase units carrying technical means phrase tags in the set of technical phrase units. Different technical phrase units expressing the same technical means meaning are mapped to the same technical means entity node, and a technical means entity identifier is assigned to each technical means entity node. Semantic normalization is performed on all technical phrase units carrying technical object phrase tags in the set of technical phrase units. Different technical phrase units expressing the same technical object meaning are mapped to the same technical object entity node, and a technical object entity identifier is assigned to each technical object entity node. Extract the co-occurrence relationship between the technical phrase unit carrying the technical means phrase tag and the technical phrase unit carrying the technical object phrase tag from each sentence unit. Based on the co-occurrence relationship and the direct dependency path in the dependency relationship tree structure, construct the initial association edge between the technical means entity node and the technical object entity node, and assign an initial association strength value to each initial association edge. Based on the global co-occurrence frequency of each technical means entity node and each technical object entity node in the scientific and technological literature data set, the initial association strength value is weighted and adjusted to generate an adjusted association strength value, and the adjusted association strength value is used as the edge weight value of each association edge in the association edge set. All technical means entity nodes, all technical object entity nodes, and all associated edges carrying edge weights are aggregated to generate a current technical element network with technical means entity nodes and technical object entity nodes as core components.
3. The method for hotspot technology mining based on scientific and technological literature data according to claim 1, characterized in that, The step of dividing the current technology element network into time slices according to a preset time slice division rule to generate a time-series-based set of technology element network segments includes: Obtain the publication time information for each paper and each patent document in the scientific and technological literature data set, wherein the publication time information includes year information and month information; Based on the publicly disclosed time information, all technical means entity nodes and all technical object entity nodes are mapped to the corresponding time coordinate axis. Each entity node carries the timestamp information of its first appearance and the time sequence information of each appearance. Each associated edge in the current technology element network is mapped to the corresponding time coordinate axis, and each associated edge carries the timestamp information of its first establishment based on the co-occurrence relationship and the time point sequence information of each co-occurrence. According to the preset time slice length parameter, the time coordinate axis is evenly divided into multiple continuous time slice units; For each time segment unit, all technical means entity nodes and all technical object entity nodes that have appeared in the current technical element network are selected to form a node subset of the time segment unit. For each time segment unit, all associated edges established or existing within the current technical element network are selected to form a subset of edges for that time segment unit. Each selected associated edge retains its edge weight value within that time segment unit. Based on the node subset and edge subset of each time segment unit, a technology element network segment unit corresponding to that time segment unit is constructed. The technology element network segment unit includes active technology means entity nodes, technology object entity nodes, and associated edges connecting the entity nodes within that time segment unit. All the technical element network segment units corresponding to the time segment units are arranged in chronological order to generate a time-seriesd set of technical element network segments. Each technical element network segment unit in the time-seriesd set of technical element network segments corresponds to a specific time window.
4. The method for hotspot technology mining based on scientific and technological literature data according to claim 1, characterized in that, The process of calculating the temporal importance index of nodes and the temporal importance index of edges in the time-seriesd set of network segments of technical elements, and determining the technical community structure with a comprehensive score of technical evolution vitality exceeding a preset vitality threshold as an emerging technology cluster, generates a set of emerging technology clusters, including: For each technical means entity node and each technical object entity node, extract the node degree value of the entity node in each time segment unit from the time-seriesd set of technical element network segments. The node degree value represents the number of associated edges connected to the entity node, and generate the node degree time series of the entity node. Based on the node degree time series, calculate the change in node degree value of the entity node between adjacent time segments, and divide the change in node degree value by the length of the time segment to obtain the node degree change rate of the entity node in each time segment. The node degree change rate sequence is subjected to sliding window averaging to eliminate the impact of short-term fluctuations and generate a smoothed node degree change rate sequence. A continuous rising phase is identified from the smoothed node degree change rate sequence, and the cumulative increase in the node degree change rate within the continuous rising phase is taken as the node influence growth rate feature of the entity node. For each associated edge, the edge weight value of the associated edge in each time segment unit is extracted from the time-seriesd set of network segments of technical elements, and the edge weight time series of the associated edge is generated. Based on the edge weight time series, the change in edge weight value of the associated edge between adjacent time segments is calculated, and the change in edge weight value is divided by the length of the time segment to obtain the edge weight change rate of the associated edge. The difference between the edge weight change rates of adjacent time segments is used as the edge formation acceleration feature of the associated edge. Community detection processing is performed on the network segment unit of the technology element corresponding to each time segment unit to identify multiple technology community structures contained in each technology element network segment unit. The technology community structure is composed of closely connected technology means entity nodes and technology object entity nodes. The evolution of the same technical community structure is tracked in consecutive time segments. The overlap of community members of each technical community structure between adjacent time segments is calculated. The stability change curve of the technical community structure is generated based on the overlap of community members. The stability change features of the community structure are extracted from the stability change curve. The stability change features of the community structure include the identifier of the rapid expansion stage of community members and the identifier of the rapid shrinkage stage of community members. For each technical community structure, based on the node influence growth rate characteristics of its constituent entity nodes, the edge formation acceleration characteristics of its associated edges, and the stability change characteristics of its community structure, the corresponding node growth score, edge evolution score, and community stability score are calculated. The node growth score, edge evolution score, and community stability score are then weighted and combined to generate a comprehensive score for the technical evolution vitality of each technical community structure. Technology community structures whose comprehensive score of technological evolution vitality exceeds a preset vitality threshold are identified as emerging technology clusters. All technical means entity nodes and technical object entity nodes in the technology community structure are aggregated to generate an emerging technology cluster set. Each emerging technology cluster in the emerging technology cluster set corresponds to a technology community structure whose comprehensive score of technological evolution vitality exceeds a preset vitality threshold.
5. The method for hotspot technology mining based on scientific and technological literature data according to claim 1, characterized in that, The process of performing dynamic graph representation learning and future edge formation prediction on the time-seriesd set of network fragments of technical elements to generate a potential technical combination prediction set containing predicted edges includes: A dynamic graph neural network model is constructed, which includes a graph convolutional network layer, a temporal modeling network layer, and a future edge prediction output layer. The graph convolutional network layer is used to learn the structural features of the technical element network segment unit within each time segment unit, and the temporal modeling network layer is used to learn the evolution law of the technical element network segment unit in the time dimension. Each technical element network segment unit in the time-seriesd technical element network segment set is input into the graph convolutional network layer. The graph convolutional network layer performs neighbor feature aggregation processing on the technical means entity nodes and technical object entity nodes in each technical element network segment unit to generate a set of node structure feature vectors corresponding to each time segment unit. The set of node structure feature vectors is input into the temporal modeling network layer in chronological order. The temporal modeling network layer uses a gated recurrent unit structure to perform time dependency modeling on the set of node structure feature vectors, captures the feature evolution trajectory of each entity node in continuous time slices, and generates the evolution feature representation vector corresponding to each entity node. The evolution feature representation vectors of all entity nodes are concatenated to generate a global evolution feature representation matrix that can characterize the evolution mode of the entire set of network segments of technical elements; The global evolution feature representation matrix is input into the future edge prediction output layer. The future edge prediction output layer calculates the probability of forming an associated edge between each pair of technical means entity nodes and technical object entity nodes in the future time segment unit, and generates a probability matrix of future associated edges between all entity node pairs. Entity node pairs whose probability values exceed a preset probability threshold are selected from the probability matrix of future associated edges. These entity node pairs are then used as candidate entity node pairs. A predicted associated edge is generated for each candidate entity node pair. Each predicted associated edge includes the identifier of the initiating entity node, the identifier of the receiving entity node, and the probability value of future associated edge formation. Aggregate all candidate entity node pairs and their corresponding predicted association edges to generate a potential technology combination prediction set containing multiple predicted association edges.
6. The method for hotspot technology mining based on scientific and technological literature data according to claim 5, characterized in that, The process of evaluating the degree of technical field crossover and novelty of each predicted edge in the potential technology combination prediction set generates a set of predicted edge evaluation results, including: For each predicted associated edge in the potential technology combination prediction set, obtain the first entity node and the second entity node connected by the predicted associated edge, where the first entity node and the second entity node are respectively a technology means entity node or a technology object entity node; Based on the technical field classification label of the first entity node and the technical field classification label of the second entity node, the technical field distance between the first entity node and the second entity node is calculated. The technical field distance is determined based on the hierarchical path length in the International Patent Classification System or the subject classification system. The distance of the technical field is mapped to the technical field intersection score interval to generate the technical field intersection score of the predicted association edge. The technical field intersection score is positively correlated with the technical field distance. Obtain all existing associated edges in the current technology element network to form an existing technology combination set, and determine whether the combination of the first entity node and the second entity node connected by the predicted associated edge appears in the existing technology combination set. If the combination of the first entity node and the second entity node does not appear in the existing technology combination set, then the predicted association edge is determined to have global novelty, and a global novelty identifier is assigned to the predicted association edge. If the combination of the first entity node and the second entity node appears in the existing technology combination set, then the most recent occurrence time of the combination in the time-seriesd technology element network segment set is further determined. If the most recent occurrence time is more than a preset time interval threshold away from the current time segment unit, then the predicted association edge is determined to have revival novelty, and a revival novelty identifier is assigned to the predicted association edge. Based on the allocation of the global novelty identifier or the revival novelty identifier, and combined with the probability value of the future associated edge formation of the predicted associated edge, a comprehensive novelty score for the predicted associated edge is generated. The technical field crossover score, novelty comprehensive score, and future association edge formation probability value of each predicted association edge are combined to generate the predicted association edge evaluation result for each predicted association edge. The predicted association edge evaluation result includes the technical field crossover score field, the novelty comprehensive score field, and the future association edge formation probability field. The evaluation results of all predicted associated edges are aggregated to generate a set of predicted associated edge evaluation results.
7. The method for hotspot technology mining based on scientific and technological literature data according to claim 1, characterized in that, The process involves generating an emerging technology topic list containing multiple emerging technology topic identifiers and a technology integration direction list containing multiple technology integration direction identifiers based on the node growth trends of the technical means entity nodes and technical object entity nodes in the emerging technology cluster set, combined with the technical field crossover assessment results and novelty assessment results in the predicted association edge assessment result set. The entries in the technology integration direction list are sorted based on their technical field crossover score and novelty comprehensive score, including: Extract all technical means entity nodes and technical object entity nodes contained in each emerging technology cluster from the set of emerging technology clusters, obtain the node influence growth rate characteristics of each entity node, and perform weighted average processing on the node influence growth rate characteristics of all entity nodes in each emerging technology cluster to generate an overall growth trend score for each emerging technology cluster. Thematic summarization is performed on all technical means entity nodes and technical object entity nodes within each emerging technology cluster. The core entity node with the highest frequency of occurrence is extracted from the entity nodes of each emerging technology cluster, and the emerging technology theme description text corresponding to the core entity node is generated based on the technical phrase unit corresponding to the core entity node. Each emerging technology cluster is assigned an emerging technology topic identifier, and the emerging technology topic identifier, the emerging technology topic description text, and the overall growth trend score are associated and stored to generate an emerging technology topic record unit. Sort all emerging technology clusters and their corresponding emerging technology theme record units according to their overall growth trend scores from high to low, and generate an emerging technology theme list containing multiple emerging technology theme identifiers; From the set of predicted associated edges evaluation results, predicted associated edges with technical field crossover scores exceeding a preset crossover threshold and novelty comprehensive scores exceeding a preset novelty threshold are selected to form a set of filtered predicted associated edges. Each predicted associated edge in the filtered predicted associated edge set is parsed to obtain the first entity node and the second entity node connected to each predicted associated edge, and the technology fusion direction description text corresponding to the predicted associated edge is generated according to the technology phrase unit corresponding to the first entity node and the second entity node. A technology fusion direction identifier is assigned to each predicted association edge in the filtered predicted association edge set, and the technology fusion direction identifier, the technology fusion direction description text, the technical field crossover score, and the novelty comprehensive score are associated and stored to generate a technology fusion direction record unit. The technology fusion direction record units corresponding to all predicted associated edges in the filtered predicted associated edge set are sorted from high to low according to the weighted sum of the technical field crossover score and the novelty comprehensive score, generating a technology fusion direction list containing multiple technology fusion direction identifiers. The emerging technology topic list and the technology integration direction list are associated and output. Each emerging technology topic identifier in the emerging technology topic list corresponds to a technology cluster whose comprehensive score of technological evolution vitality exceeds a preset vitality threshold. Each technology integration direction identifier in the technology integration direction list corresponds to a predicted association edge whose technical field crossover score exceeds a preset crossover threshold and whose comprehensive novelty score exceeds a preset novelty threshold.
8. The method for hotspot technology mining based on scientific and technological literature data according to claim 2, characterized in that, The step involves semantically normalizing all technical phrase units carrying technical means phrase tags in the technical phrase unit set, mapping different technical phrase units expressing the same technical means meaning to the same technical means entity node, and assigning a technical means entity identifier to each technical means entity node, including: Obtain all technical phrase units carrying technical means phrase tags in the technical phrase unit set. Each technical means phrase unit contains the text content of the technical means phrase and the context sentence unit of the technical means phrase in the original document. The text content of each technical means phrase unit is input into a pre-trained word embedding model, and the text content of each technical means phrase unit is converted into a technical means phrase embedding vector through the pre-trained word embedding model. Perform pairwise semantic similarity calculations on all technical means phrase units in the set of technical means phrase units, calculate the cosine similarity between any two technical means phrase embedding vectors, and obtain the semantic similarity matrix; Hierarchical clustering is performed on the semantic similarity matrix to divide the technical means phrase units with semantic similarity exceeding a preset clustering similarity threshold into the same technical means phrase cluster. Each technical means phrase cluster contains multiple technical means phrase units that express similar technical means meanings. Select the most frequently occurring technical phrase unit from each technical phrase cluster as the representative technical phrase of the technical phrase cluster, and use the text content of the representative technical phrase as the node name of the technical entity node corresponding to the technical phrase cluster. Create a technical means entity node for each technical means phrase cluster and assign a unique technical means entity identifier to the technical means entity node. Associate and map all technical means phrase units within the technical means phrase cluster with the technical means entity identifier. The context sentence unit corresponding to each technical means phrase unit in the original document is stored as the context information of the technical means entity node, and used for subsequent functional description of the technical means entity node.
9. The method for hotspot technology mining based on scientific and technological literature data according to claim 3, characterized in that, The step of constructing a network segment unit of technical elements corresponding to each time segment unit based on the node subset and edge subset of each time segment unit includes: Obtain a node subset corresponding to each time segment unit. The node subset contains multiple technical means entity nodes and multiple technical object entity nodes. Each entity node carries information on the number of times the entity node appears in the time segment unit and the information on the time of its first appearance. Obtain the edge subset corresponding to each time segment unit. The edge subset contains multiple associated edges. Each associated edge carries the edge weight value of the associated edge in the time segment unit and the co-occurrence information of the associated edge in the time segment unit. Based on each technical means entity node and each technical object entity node in the node subset, a corresponding entity node object is created in the technical element network segment unit, and a node attribute field is set for each entity node object. The node attribute field includes an entity node identifier field, an entity node type field, a number of occurrences field, and a first occurrence time field. For each associated edge in the edge subset, a corresponding associated edge object is created in the technical element network segment unit. The associated edge object connects the specified technical means entity node and technical object entity node in the edge subset. An edge attribute field is set for each associated edge object. The edge attribute field includes an edge weight value field and a co-occurrence count field. The occurrence count field of all entity node objects in the network segment unit of the technical element is normalized, and the occurrence count is divided by the total occurrence count of all entity nodes in the time segment unit to obtain the relative occurrence frequency of each entity node in the time segment unit. The edge weight value field of all associated edge objects in the network segment unit of the technical element is normalized, and the edge weight value is divided by the total edge weight value of all associated edges in the time segment unit to obtain the relative association strength of each associated edge in the time segment unit. The relative frequency of occurrence is used as the initial value of the node importance of the entity node object in the technical element network segment unit, and the relative correlation strength is used as the initial value of the edge importance of the associated edge object in the technical element network segment unit, thus completing the construction of the technical element network segment unit corresponding to the time segment unit.
10. A hotspot technology mining system based on scientific and technological literature data, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the hotspot technology mining method based on scientific and technological literature data according to any one of claims 1 to 9 by executing the machine-executable instructions.
Citation Information
Patent Citations
Science and technology frontier research focus analysis method and device based on national fund project mining
CN113761313A
Method and apparatus for choosing a stock portfolio, based on patent indicators
US6175824B1
FM model based method and apparatus for predicting medical hot spot, and computer device
WO2021139271A1
Cited By
Cross-clustering patent technology evolution path analysis method, system and terminal
CN122048594A