A method and system for generating a technology evolution path and a technology roadmap based on a large language model driven technology

By using a technology evolution path generation method and knowledge graph driven by a large language model, the problems of data utilization and heterogeneity integration in existing technology roadmaps are solved, enabling a comprehensive update and forward-looking improvement of the technology roadmap, and providing more accurate decision support.

CN122432705APending Publication Date: 2026-07-21HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-04-03
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing methods for drawing up technology roadmaps rely on expert experience, have limited data sources, struggle to fully utilize massive amounts of global data, suffer from serious subjective biases, ignore non-consensus knowledge, and face difficulties in integrating heterogeneous multi-source data, resulting in shortcomings in the comprehensiveness, timeliness, and accuracy of the roadmaps.

Method used

A technology evolution path generation method driven by a large language model is adopted. By extracting technology nodes from multi-source heterogeneous data through clustering and topic modeling, and combining knowledge graphs for link prediction, the technology roadmap is dynamically supplemented. Consensus and non-consensus knowledge are integrated to achieve a comprehensive update of the technology roadmap.

Benefits of technology

It significantly enhances the comprehensiveness, timeliness, and forward-looking nature of the technology roadmap, reflects the historical depth of the technological evolution process, provides more precise support for science and technology policy and strategic decision-making, and has good scalability and automation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432705A_ABST
    Figure CN122432705A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of production management, and specifically relates to a technical evolution path and a technical roadmap generation method and system based on a large language model driving. The method comprises the following steps: automatically mining technical topics and technical evolution paths from patent literature data by using feature extraction, cluster analysis, topic modeling and large model topic identification technology; identifying entities and relationships through a large model to extract consensus technical nodes and development paths from technical reports, and extracting key information such as non-consensus technical nodes, development suggestions and risk assessment from unstructured texts such as expert public statements; constructing a unified structure knowledge graph, fusing and storing technical nodes and development directions extracted from technical evolution paths, original roadmaps, consensus and non-consensus knowledge, and establishing potential associations between new and old technical nodes based on a large model driven knowledge graph link prediction method, so as to realize dynamic expansion of the technical roadmap and probabilistic evaluation of future development directions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of production management technology, specifically relating to a method and system for generating technology evolution paths and technology roadmaps based on a large language model. Background Technology

[0002] Technology roadmaps, as a key tool for showcasing the development paths and evolution trends of various technologies within a technological field, are widely used in areas such as technology forecasting, strategic planning, policy formulation, and resource allocation. However, existing technology roadmap drawing methods face numerous challenges and shortcomings in practical applications. First, while expert-based roadmap drawing methods rely on intensive discussions and experience among experts, they have significant limitations: on the one hand, the limited amount of historical data available to experts prevents them from fully utilizing the vast amounts of global data, such as patents, literature, and industry reports, leading to limitations in the comprehensiveness and forward-looking nature of technology roadmaps. On the other hand, experts are often influenced by personal experience, background knowledge, and industry perspectives when analyzing and drawing roadmaps, and this subjective bias affects the objectivity and accuracy of the final roadmap. Furthermore, traditional expert roadmaps typically only include consensus opinions reached by the expert group, neglecting non-consensus knowledge such as differing views, technological disputes, and potential risk assessments within the expert group, which also provide important references for technology forecasting and decision-making.

[0003] With the continuous development of data-driven methods, multi-source data fusion has gradually become an effective means of technology roadmap creation, alleviating to some extent the limitations of traditional expert-dependent methods in terms of data breadth and timeliness. However, existing data-driven methods still face significant challenges. Current mainstream practices mainly rely on publicly available scientific and technological literature such as papers and patents. While this has achieved breakthroughs in data scale and effectively compensated for the limited reference materials in expert discussions, its information sources remain relatively singular, making it difficult to support the construction of rich and diverse technology roadmaps. It is particularly noteworthy that other types of unstructured data that have important supplementary value to technology roadmaps, such as global engineering frontier reports, industry white papers, policy documents published by authoritative institutions, and non-consensus knowledge containing individual expert judgments (including different viewpoints on technology paths, potential alternatives, early warnings of development bottlenecks, and risk assessments), are often overlooked or not systematically integrated in existing methods.

[0004] Furthermore, the aforementioned multi-source data exhibit significant heterogeneity in text format, semantic granularity, and expression style. Existing technologies lack a unified semantic alignment and knowledge fusion mechanism, making it difficult to efficiently and accurately integrate such heterogeneous information into the overall structure of the technology roadmap. Meanwhile, current deep learning-based technology roadmap generation methods, such as those using Generative Adversarial Networks (GANs) or BERT models, are cumbersome, have long training cycles, are inefficient, and require large amounts of labeled data to train the models. The processes of multi-source data fusion, semantic analysis, and link prediction remain complex, limiting the speed and accuracy of technology roadmap generation. Therefore, to improve the comprehensiveness, timeliness, and accuracy of technology roadmaps, it is necessary to explore new technological approaches to address these issues and optimize the technology roadmap construction process. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for generating technology evolution paths and technology roadmaps based on a large language model, so as to achieve a comprehensive update of the technology roadmap in terms of supplementing historical paths, expanding cutting-edge directions, and probabilistically assessing future trends.

[0006] To achieve the above objectives, the first aspect of this invention provides a method for generating technology evolution paths based on a large language model, comprising the following steps: S1. Set a time interval according to the target technical field and obtain relevant patent literature data; preprocess the patent literature data and then extract the structured text containing key information; S2. Use a large language model to extract context-aware features from the structured text to obtain semantic vectors; S3. The semantic vectors are grouped using a clustering algorithm to form several technology clusters; S4. Perform topic modeling operations for each technology cluster to extract the set of keywords for that technology cluster; S5. Perform refined discrimination and summarization on the keyword set to generate technical nodes with clear technical meaning; S6. Calculate the correlation between technology nodes in different time intervals using similarity and organize them into a time series form of technology evolution path.

[0007] Furthermore, in step S1, the preprocessing includes denoising, cleaning, and structured preprocessing; the key information includes title, abstract, keywords, classification information, and time information; in step S2, the semantic vector is used to characterize the semantic features of the document in terms of technical content.

[0008] 3. The method for generating technology evolution paths based on a large language model according to claim 1, characterized in that, in step S3, the clustering algorithm adopts K-means, HDBSCAN, spectral clustering or hierarchical clustering; each technology cluster corresponds to a set of technology topics; In step S4, the topic modeling operation uses LDA, BERTopic, or NMF methods. The input of the topic modeling operation is the structured text of the patent document data corresponding to each technology cluster, and the output is the set of topic terms and the weight corresponding to each topic term.

[0009] Furthermore, in step S5, the refined discrimination and induction includes: inputting the set of keywords, keyword weights, and structured text of the corresponding patent document data for each technology cluster into the large language model; guiding the large language model to inductively summarize the technical topics represented by each cluster through prompting engineering; and outputting the technical topic name, technical topic definition, core technical features, and representative terms, with the technical topic name serving as the technical node.

[0010] Furthermore, step S6 specifically includes: sorting the technical topic names according to time information, and using cosine similarity to calculate the succession relationship between technical topics in adjacent time periods. When the cosine similarity between the technical topic name in the later time period and the technical topic name in the previous time period is higher than a preset threshold, it is determined that there is an evolutionary relationship between the two, and a directed relationship edge is established. If a technical topic name and multiple predecessor technical topic names simultaneously meet the similarity condition, multiple candidate evolutionary edges are retained, thereby forming a technical evolution path with a branching structure.

[0011] A second aspect of this invention provides a method for dynamically generating technology roadmaps based on large models and knowledge graphs, comprising the following steps: Step 1: Extract technical nodes, technical development relationships, and expert evaluation information from technical report data and expert opinion data; Step 2: Store the original technology roadmap, the technology evolution path obtained by any of the above-mentioned technology evolution path generation methods, and the extraction results of Step 1 in a unified manner, and establish the relationships between entities to obtain a knowledge graph; Step 3: Use a large language model to reason about the knowledge graph, dynamically supplement the technology development path, and finally update the technology roadmap.

[0012] Furthermore, step 1 specifically includes: performing text cleaning, segmentation, sentence segmentation, time annotation, source annotation, and terminology standardization on the technical report data and expert opinion data to obtain technical report text fragments and expert opinion text fragments; A knowledge graph ontology of the technology roadmap is constructed. Then, under the constraints of the knowledge graph ontology, the text fragments of the technology report type and the text fragments of the expert opinion type are respectively input into the large language model. Knowledge extraction is performed according to the preset prompt words to obtain the extraction results of the technology report type and the extraction results of the expert opinion type.

[0013] Furthermore, the knowledge graph ontology includes entity types and relationship types; the entity types include technology nodes, application scenarios, and expert evaluations; the relationship types include evolutionary relationships between technology nodes, recommended development and risk relationships between expert evaluations and technology nodes, and applicable relationships between technology nodes and application scenarios. The extracted results of the technical report category include technical nodes, mainstream development directions, key technical bottlenecks, and explicit development relationships between technologies. The extracted results of the expert opinion category include assessments of technical potential and warnings of technical risks.

[0014] Furthermore, step 2 specifically includes: using the aforementioned technology evolution path and original technology roadmap as the basic framework of the knowledge graph, supplementing it with the consensus and non-consensus knowledge provided by the extraction results, performing knowledge fusion, and storing the fused knowledge graph in a graph database; Step 3 specifically includes: firstly, preliminary screening of the entity types in the knowledge graph to obtain candidate node pairs; then, inputting the candidate node pairs and their neighborhood context into the large model for combination; judging whether there is a reasonable development or evolution relationship between any two technical nodes; and outputting the connection reason and the probability score of the path's future realization; the neighborhood context includes the entity nodes and relationships adjacent to the candidate node pairs in the knowledge graph. The predicted new paths, which include probability assessments, are added to the existing technology roadmap to form a multi-level technology roadmap that includes main paths, branch paths, and alternative paths. At the same time, the future development potential of different technology directions is quantitatively ranked based on the probability scores of the paths.

[0015] A third aspect of this invention provides a dynamic technology roadmap generation system based on large models and knowledge graphs, comprising: The data acquisition module is used to acquire multi-source heterogeneous data in the target technology field; the multi-source heterogeneous data includes patent document data, technical report data, expert opinion data, and original technology roadmap data; The technology evolution path mining module is used to obtain the technology evolution path according to any of the technology evolution path generation methods described above; The knowledge mining module is used to extract technical nodes, technical development relationships, and expert evaluation information from the technical report data and expert opinion data. The knowledge graph construction module is used to uniformly store the technology evolution path, the results extracted by the knowledge mining module and the original technology roadmap, and establish relationships between entities to obtain a knowledge graph. The link prediction and path update module is used to reason about the knowledge graph using a large language model, dynamically supplement the technology development path, and ultimately update the technology roadmap.

[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages: 1. The technology evolution path generation method based on large language model provided by this invention can effectively reflect the historical development of a specific technology field through clustering and topic recognition, provide supplementary content for the early stage of existing technology roadmaps, and significantly enhance their ability to depict the technology evolution process and their sense of historical depth.

[0017] 2. The technology roadmap dynamic generation method provided by this invention integrates multi-source data, especially non-consensus data of expert knowledge, and supplements the roadmap with experts' cutting-edge suggestions on technology; it introduces a link prediction mechanism to dynamically connect new and old technology nodes, thereby realizing a comprehensive update of the technology roadmap in terms of supplementing historical paths, expanding cutting-edge directions, and probabilistic assessment of future trends. It can significantly improve the comprehensiveness, timeliness, and forward-looking nature of the technology roadmap, and provide more accurate support for science and technology policy formulation and strategic decision-making.

[0018] 3. This invention breaks through the limitations of traditional expert discussion models in terms of data breadth, subjective bias, and update efficiency, and realizes the ability to automatically extract the laws of technological evolution from massive heterogeneous data; by distinguishing and integrating consensus and non-consensus knowledge, it not only retains the reliability of mainstream technical paths, but also incorporates key non-consensus information such as expert cutting-edge suggestions and potential risks, making the technology roadmap more strategically insightful and valuable for decision-making.

[0019] 4. This invention constructs a knowledge graph with a unified structure as an intermediate representation layer, which effectively solves the problem of semantic heterogeneity of multi-source data and provides a computable and scalable knowledge infrastructure for subsequent intelligent reasoning.

[0020] 5. The entire framework of this invention has good scalability and automation capabilities, supports continuous learning and incremental updates, and is suitable for the construction and maintenance of roadmaps in various technical fields, providing efficient, accurate and comprehensive intelligent decision support for science and technology policy formulation, industrial strategic layout and scientific research resource allocation. Attached Figure Description

[0021] Figure 1 This is a dynamic evaluation framework diagram of the technology roadmap based on large models and knowledge graphs in this embodiment of the invention.

[0022] Figure 2 This is a flowchart of the dynamic evaluation method for the technology roadmap based on large models and knowledge graphs in this embodiment of the invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0024] Example 1 To address the limitations of traditional manual technology roadmap generation, such as the constraints of expert time and the difficulty in systematically processing massive amounts of global scientific and technological data, this embodiment provides a technology evolution path generation method driven by a large language model, comprising the following steps: S1. Set a time interval according to the target technical field, and obtain relevant patent literature data from patent databases and academic paper databases; preprocess the patent literature data, and then extract the structured text containing key information. S2. Use a large language model to extract context-aware features from the structured text to obtain semantic vectors; S3. The semantic vectors are grouped using a clustering algorithm to form several technology clusters; S4. Perform topic modeling operations for each technology cluster to extract the set of keywords for that technology cluster; S5. Perform refined discrimination and summarization on the keyword set to generate technical nodes with clear technical meaning; S6. Calculate the correlation between technology nodes in different time intervals using similarity and organize them into a time series form of technology evolution path.

[0025] Specifically, 1. Data preprocessing Paper and patent data from Web of Science and Derwent databases were cleaned and structured, removing irrelevant text and duplicate records, and extracting fields such as title, abstract, keywords, classification information, and time information. These fields were then concatenated into a unified input text, which served as the input for subsequent feature extraction. These fields cover key information such as technical terms, R&D entities, and application scenarios.

[0026] 2. Feature Extraction The pre-processed papers and patent texts are input into a pre-trained language model to perform semantic encoding on each document, obtaining a corresponding high-dimensional semantic vector representation. The pre-trained language model can be a language model with text representation capabilities or a text embedding model. The semantic vectors are used to represent the semantic features of the documents in terms of technical content (e.g., key technical features and their potential associations) and serve as input for subsequent clustering analysis.

[0027] In practice, the title, abstract, and keywords can be concatenated in the format of "title + [separator] + abstract + [separator] + keywords," and then input into a pre-trained language model for forward encoding. The token-level hidden states output by the model are then pooled, for example using mean pooling, first-token pooling, or weighted pooling, to obtain a fixed-dimensional semantic vector for the documents. To improve the consistency of vector distribution among documents from different sources, L2 normalization can be performed after the vector output to facilitate subsequent similarity calculations based on Euclidean distance or cosine similarity.

[0028] 3. Clustering The obtained semantic vectors of the documents are clustered to identify potential technological sub-directions. Clustering methods can include K-means, HDBSCAN, spectral clustering, or hierarchical clustering; the following explanation uses K-means as an example.

[0029] When using K-means, the first step is to set the number of clusters, K. The value of K can be preset based on the expected number of sub-directions in the target technology field, or it can be selected in conjunction with the elbow method, silhouette coefficient, or Davies-Bouldin index. Then, using the semantic vectors of the documents as input, each document is iteratively assigned to the nearest cluster center, and the cluster centers are continuously updated until the clustering results converge or the preset number of iterations is reached. After clustering, each cluster corresponds to a set of potential technology topics. To improve clustering stability, K-means can be randomly initialized multiple times, and the result with the smallest sum of squares within each cluster can be selected as the final clustering result.

[0030] 4. Topic Modeling For each cluster, topic modeling is performed to extract a representative set of keywords within that cluster. Topic modeling methods can include LDA, BERTopic, NMF, or other methods suitable for text topic discovery; LDA will be used as an example below.

[0031] When using LDA, the text within each cluster is first segmented, stop words removed, and word frequency counted to construct a bag-of-words representation. Then, the number of topics is set, and the LDA model is trained based on the co-occurrence distribution of terms in the documents, yielding a set of high-probability terms for each topic and the probability distribution of each document across topics. For each cluster, the topic with the highest probability is selected as the core topic, and the terms with the highest probability values ​​within it are extracted as representative topic words for that cluster's core topic, resulting in a topic word set. To ensure the interpretability of the topic words, a minimum word frequency threshold can be set, and overly generalized high-frequency words within the domain can be removed.

[0032] 5. Topic Recognition The set of keywords, keyword weights (and corresponding term probabilities), and representative papers or patent texts corresponding to each cluster are input into the large language model. Through prompting engineering, the model is guided to summarize the technical themes represented by the cluster, outputting standardized technical theme names and their brief definitions. These theme names are then written as technical entities into the subsequent knowledge graph. This step aims to summarize the keyword set into clear and nameable technical themes, thereby achieving a mapping from document clustering to technical nodes.

[0033] In practice, the input information to the large language model includes a set of topic terms and the weights of each topic term; the output includes the technical topic name, technical topic definition, core technical features, and representative terms. The prompt word templates used for topic recognition can be found in the corresponding section of Table 2. These prompts determine whether the topic constitutes an independent, nameable technical node and output a standardized technical name and brief definition.

[0034] For example, for a cluster of keywords including "solid electrolyte," "lithium dendrite suppression," and "all-solid-state battery packaging," the large model can summarize "all-solid-state lithium battery manufacturing technology" as a technology node and mark its core features. These technology nodes, sorted by timestamp, constitute a technological evolution path reflecting the historical development of the field, which can serve as supplementary content for the early stages of the technology roadmap.

[0035] 6. Theme Evolution Path Identification After identifying the technical topics, they are sorted according to time information, and similarity calculations are used to identify the connections between technical topics in different time periods. Specifically, cosine similarity can be used to calculate the semantic similarity between technical topics in different time periods. When the similarity between a technical topic in a later time period and a technical topic in a previous time period is higher than a preset threshold, an evolutionary association is determined to exist between them, and a directed relation edge is established. If a technical topic and multiple predecessor topics simultaneously meet the conditions, multiple candidate evolutionary edges are retained, thus forming a technical evolution path with a branching structure.

[0036] In practice, technical topics can be grouped by year or preset time window, and the average document vector of the same technical topic can be used as the stage representation vector of that topic within that time period. Then, cosine similarity is calculated pairwise for topic nodes in adjacent time windows, and an evolutionary relationship is established while satisfying temporal order constraints and similarity threshold constraints. The similarity threshold can be set according to the semantic dispersion of different technical fields.

[0037] This evolutionary path can effectively reflect the historical development of a specific technology field, provide supplementary content for the early stages of existing technology roadmaps, and significantly enhance their ability to depict the process of technological evolution and their sense of historical depth.

[0038] Example 2 A method for dynamically generating technology roadmaps based on large models and knowledge graphs, such as Figure 1 and 2 As shown, it includes: Step 1, extracting technical nodes, technical development relationships, and expert evaluation information from technical report data and expert opinion data; Step 2: Store the original technology roadmap, the technology evolution path obtained by the technology evolution path generation method described in Example 1, and the extraction results of Step 1 in a unified manner, and establish the relationships between entities to obtain a knowledge graph; Step 3: Use a large language model to reason about the knowledge graph, dynamically supplement the technology development path, and finally update the technology roadmap.

[0039] This embodiment uses massive amounts of scientific and technological literature and patent data as the basic input for technology evolution analysis. It utilizes a large model to perform deep semantic understanding, feature extraction, and fusion of multi-source heterogeneous texts. It constructs a unified knowledge graph structure covering consensus and non-consensus knowledge, and introduces a link prediction mechanism on this basis to dynamically connect new and old technology nodes, thereby achieving a comprehensive update of the technology roadmap in terms of supplementing historical paths, expanding cutting-edge directions, and probabilistically assessing future trends.

[0040] To address the common problem that existing data-driven roadmap creation methods neglect non-consensus cutting-edge knowledge and lack the ability to integrate diverse expert perspectives, this invention designs a dual supplementary mechanism that integrates consensus and non-consensus knowledge. Consensus knowledge is primarily supplemented by authoritative technical reports (such as the *Global Engineering Frontiers Report*), reviews in high-impact journals, and white papers published by standardization organizations. Non-consensus knowledge is supplemented by technical insights expressed by experts in informal channels such as public speeches, social media, academic forums, and risk assessment documents. This includes differing assessments of the potential of emerging technologies, questions about the feasibility of technical pathways, warnings of potential technical risks, and suggestions for alternative technical solutions.

[0041] First, in step 1, the authoritative technical reports and expert knowledge data are preprocessed, including text cleaning, segmentation, sentence segmentation, time annotation, source annotation, and terminology standardization, so that the original text is transformed into text fragments suitable for knowledge extraction.

[0042] Then, the knowledge graph ontology is constructed. A technology roadmap knowledge graph ontology is pre-constructed to define the entity types and relationship types involved in subsequent extraction. The knowledge graph ontology is shown in Table 1. The entity types include technology entities, application scenarios, and expert evaluation entities; the relationship types include "evolution" relationships between technologies, "recommended development" and "risk-existing" relationships between expert evaluation nodes and technology nodes, and "applicable" relationships between technology nodes and application scenario nodes. This ontology is used to unify the extraction criteria from different data sources, enabling subsequent results to directly enter the same knowledge graph structure.

[0043] Table 1. Knowledge Graph Ontology Constructed for Technology Roadmap Update Framework

[0044] Under ontology constraints, knowledge extraction is performed using a large language model. For authoritative technical report data, preprocessed text fragments are input into the large language model. Based on preset prompts, widely recognized key technologies, key development directions, cutting-edge technologies, and explicit development relationships between technologies are identified, and the extracted results are considered consensus knowledge. For expert knowledge data, preprocessed text fragments are input into the large language model. Based on preset prompts, opinion-based information regarding technological potential assessments and risk warnings is identified, and the extracted results are considered non-consensus knowledge. For example, if an expert points out that "perovskite photovoltaics, although highly efficient, have questionable long-term stability and may be overtaken by tandem silicon-based technology," this opinion will be parsed into two non-consensus knowledge edges: "Perovskite photovoltaics → Expert assessment: Insufficient stability" and "Tandem silicon-based photovoltaics ← Expert assessment: Potential alternative direction," and stored in the knowledge graph. Preset prompt templates can be found in the corresponding section of Table 2.

[0045] This invention employs a large-scale prompt engineering strategy for unstructured text to guide models in accurately identifying implicit yet forward-looking viewpoints and transforming them into structured technical nodes or attribute tags, which are then injected into corresponding positions within the knowledge graph. This mechanism not only enriches the content dimensions of the technology roadmap but also preserves the uncertainties and diversity inherent in technological development, providing decision-makers with a more comprehensive perspective on risk-opportunity trade-offs.

[0046] For raw technology roadmap data, due to its limited volume and the fact that it already presents technological development trends in the form of charts, phase divisions, and node connections, technical entities and relationships can be extracted manually. Specifically, technical theme nodes and technological development relationships are directly extracted from existing raw technology roadmaps, and the extracted results serve as the basic framework for subsequent knowledge graph construction.

[0047] Further, in step 2, knowledge fusion and knowledge graph construction are carried out. After obtaining the results of technology evolution path extraction, original technology roadmap extraction, authoritative technology report extraction, and expert knowledge extraction, the knowledge fusion and knowledge graph construction stage begins. This stage uses the technology evolution path and original technology roadmap as the basic framework, supplemented by consensus and non-consensus knowledge provided by authoritative technology reports and expert knowledge, to achieve unified organization of multi-source knowledge.

[0048] Knowledge fusion addresses inconsistencies in terminology, entity granularity, and relation types across different data sources. Entities and relations extracted from various sources are mapped to a unified ontology, and a large language model is used to perform entity normalization, alias merging, semantic disambiguation, and relation alignment. For technical entities with the same semantics but different expressions, they are merged into the same knowledge graph node; for relations with consistent semantics but different expressions, they are mapped to a unified relation type. This knowledge graph supports standardized storage and association of knowledge from various sources, including technology evolution paths, original roadmaps, consensus reports, and non-consensus statements. This not only achieves the organic integration of multi-source knowledge but also provides a structured foundation for subsequent path reasoning and dynamic updates.

[0049] In practice, candidate entity pairs can be initially screened based on string similarity, keyword overlap, or semantic vector cosine similarity. Then, the candidate entity names, definitions, and contextual descriptions are input into a large language model, which determines whether they belong to the same technical entity, hierarchical technical entities, or related but different entities. When determined to be the same technical entity, entity normalization and merging are performed, retaining names from different sources as aliases. The prompt word templates used for knowledge fusion can be found in the corresponding section of Table 2.

[0050] The merged knowledge graph can be stored in a graph database. Technical entity nodes can record attributes such as name, definition, first appearance time, and source type; relationship edges can record attributes such as relationship type, source, time information, and confidence level; expert evaluation entities can record information such as the proposing entity, opinion bias, and evidence text.

[0051] In step 3, after the knowledge graph is constructed and integrated, link prediction and path updating are performed. This module is used to discover new associations in the knowledge graph that are not yet explicitly labeled but may be valid in technical logic, expanding the single-line technology development path in the traditional roadmap into a multi-directional development path, thereby achieving further expansion and updating of the existing technology roadmap.

[0052] First, candidate node pairs are generated based on existing technical entities, relationships, application scenarios, and expert-evaluated entities in the knowledge graph. Then, the candidate technical entities, along with their surrounding risk factors and application scenario nodes, are input into the large model, enabling it to gain a comprehensive understanding of both the technical nodes themselves and their surrounding nodes. Based on its embedded cross-domain knowledge and contextual understanding capabilities, the large model outputs connection reasons and probability values, thereby increasing the interpretability of link prediction reasoning and reducing the "black box effect."

[0053] In practice, candidate node pairs are initially screened based on semantic similarity, shared neighboring nodes, shared risk factors, or shared application scenarios to reduce the number of candidate combinations that the large language model needs to judge. Then, the screened candidate node pairs and their neighborhood context are input into the large language model for relational reasoning, and the model outputs the connection reason and a probability score for the future realization of the path (i.e., probabilistic evaluation). For example, when there are two nodes, "quantum dot display" and "Micro-LED," in the graph, the model can infer a "technology evolution" relationship between them based on information such as material compatibility, manufacturing process evolution trends, and industry investment trends, and provide the corresponding evolution probability. The prompt word templates used for link prediction can be found in the corresponding section of Table 2.

[0054] Finally, the new relationship edges obtained from the link prediction are added to the original technology roadmap structure to obtain the updated technology roadmap. Specifically, this includes adding the predicted new paths, which incorporate probability assessments, to the original roadmap structure, forming a multi-layered technology development landscape including main paths, branch paths, and alternative paths. Simultaneously, based on the probability scores of the paths, the future development potential of different technology directions is quantitatively ranked, assisting decision-makers in identifying high-value tracks, avoiding technological pitfalls, or establishing strategic reserves. Furthermore, this framework provides an incremental dynamic evaluation mechanism—when new scientific literature, patents, or expert opinions continuously flow in, the system can perform a new round of knowledge extraction, graph fusion, and link prediction processes to ensure that the technology roadmap always remains timely and cutting-edge. This embodiment uses the technology evolution path and the original technology roadmap as the updated basic framework, with authoritative technical reports and expert knowledge providing supplementary information as the basic framework, and the link prediction results further expanding the potential development directions in the original roadmap. The resulting updated technology roadmap not only reflects existing explicit development paths, but also embodies mainstream consensus, expert disagreements, and possible future development branches, thereby improving the timeliness, comprehensiveness, and forward-looking nature of the technology roadmap.

[0055] Table 2 Examples of Prompt Word Templates

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] Example 3 This invention provides a dynamic technology roadmap generation system based on large models and knowledge graphs, such as... Figure 1 As shown, it includes a data acquisition module, a technology evolution path mining module, a knowledge mining module, a knowledge graph construction module, and a link prediction and path update module.

[0062] The data acquisition module is used to acquire multi-source heterogeneous data in the target technology field; the multi-source heterogeneous data includes patent literature data, authoritative technical report data, expert opinion data, and original technology roadmap data; The technology evolution path mining module is used to extract technology nodes and identify technology evolution paths from massive amounts of papers and patent data (patent literature data); The knowledge mining module is used to extract technical nodes, technical development relationships, expert suggestions on technical development, or risk assessment judgments from the technical report data and expert opinion data. The knowledge graph construction module is used to uniformly store the technology evolution path, the results extracted by the knowledge mining module and the original technology roadmap, and establish relationships between entities to obtain a knowledge graph. The link prediction and path update module uses a large language model to reason about the knowledge graph, dynamically supplementing the technology development path, and ultimately achieving efficient and comprehensive updates to the technology roadmap.

[0063] The patent literature data includes both academic papers and patent data. Academic papers can be obtained from the Web of Science database, while patent data can be obtained from the Derwent database. Authoritative technical report data comes from publications such as *Global Engineering Frontiers*, *National Science and Technology Innovation Plan*, and other authoritative reports, planning documents, and white papers that reflect mainstream technological judgments and key development directions. Expert knowledge data is obtained through online searches, and its sources include textual materials from expert conference speeches, forum presentations, policy hearing records, and other informal public channels. Original technology roadmap data consists of existing technology roadmap charts, textual descriptions, and related materials. Figure 1 Generally, existing technology roadmaps are developed by manually collecting literature, convening experts, holding multiple rounds of intensive discussions, and reaching a consensus. However, this approach cannot fully utilize the massive amounts of data available globally, requires extensive expert consultation, is difficult to update, and ignores non-consensus knowledge. Therefore, this invention uses this data as one of the data sources for knowledge graph construction, enabling the construction and updating of new technology roadmaps.

[0064] Furthermore, a technology evolution path analysis based on a large language model is conducted. The technology evolution path module is used to identify technology topics and form technology evolution paths from paper and patent data. This module includes six steps: data preprocessing, feature extraction, clustering, topic modeling, topic identification, and topic evolution path identification. The specific process adopts the technology evolution path generation method based on a large language model driven by Example 1.

[0065] Furthermore, a dual extraction of consensus and non-consensus knowledge is performed. The consensus and non-consensus knowledge extraction module is used to extract technical entities, technical relationships, and expert evaluation information from authoritative technical reports and expert knowledge data to supplement consensus judgments and divergent viewpoints that are difficult to reflect in traditional technology roadmaps.

[0066] Further, knowledge fusion and knowledge graph construction are carried out. After obtaining the output results of the technology evolution path module, the extraction results of the original technology roadmap, the extraction results of authoritative technical reports, and the extraction results of expert knowledge, the knowledge fusion and knowledge graph construction stage begins. This stage uses the technology evolution path and the original technology roadmap as the basic framework, supplemented by consensus and non-consensus knowledge provided by authoritative technical reports and expert knowledge, to achieve a unified organization of multi-source knowledge.

[0067] Finally, the new relationship edges obtained from the link prediction are added to the original technology roadmap structure to obtain the updated technology roadmap. The technology evolution path and the original technology roadmap serve as the updated basic framework, while authoritative technical reports and expert knowledge provide supplementary information. The link prediction results further expand the potential development directions in the original roadmap. The resulting updated technology roadmap not only reflects existing explicit development paths but also embodies mainstream consensus, expert disagreements, and possible future development branches, thereby improving the timeliness, comprehensiveness, and forward-looking nature of the technology roadmap.

[0068] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating technology evolution paths based on a large language model, characterized in that, Includes the following steps: S1. Set a time interval according to the target technical field and obtain relevant patent literature data; preprocess the patent literature data and then extract the structured text containing key information; S2. Use a large language model to extract context-aware features from the structured text to obtain semantic vectors; S3. The semantic vectors are grouped using a clustering algorithm to form several technology clusters; S4. Perform topic modeling operations for each technology cluster to extract the set of keywords for that technology cluster; S5. Perform refined discrimination and summarization on the keyword set to generate technical nodes with clear technical meaning; S6. Calculate the correlation between technology nodes in different time intervals using similarity and organize them into a time series form of technology evolution path.

2. The method for generating technology evolution paths based on a large language model according to claim 1, characterized in that, In step S1, the preprocessing includes denoising, cleaning, and structured preprocessing; the key information includes title, abstract, keywords, classification information, and time information; in step S2, the semantic vector is used to characterize the semantic features of the document in terms of technical content.

3. The method for generating technology evolution paths based on a large language model according to claim 1, characterized in that, In step S3, the clustering algorithm used is K-means, HDBSCAN, spectral clustering, or hierarchical clustering; each technology cluster corresponds to a set of technology topics. In step S4, the topic modeling operation uses LDA, BERTopic, or NMF methods. The input of the topic modeling operation is the structured text of the patent document data corresponding to each technology cluster, and the output is the set of topic terms and the weight corresponding to each topic term.

4. The method for generating technology evolution paths based on a large language model according to claim 1, characterized in that, In step S5, the refined discrimination and induction includes: inputting the set of keywords, keyword weights and structured text of the corresponding patent document data for each technology cluster into the large language model, guiding the large language model to summarize the technical topics represented by each cluster through prompting engineering, and outputting the technical topic name, technical topic definition, core technical features and representative terms, and using the technical topic name as the technical node.

5. The method for generating technology evolution paths based on a large language model according to claim 1, characterized in that, Step S6 specifically includes: sorting the technical topic names according to time information, and using cosine similarity to calculate the succession relationship between technical topics in adjacent time periods. When the cosine similarity between the technical topic name in the later time period and the technical topic name in the previous time period is higher than a preset threshold, it is determined that there is an evolutionary relationship between the two, and a directed relationship edge is established. If a technical topic name and multiple predecessor technical topic names simultaneously meet the similarity condition, multiple candidate evolutionary edges are retained, thereby forming a technical evolution path with a branching structure.

6. A method for dynamically generating technology roadmaps based on large models and knowledge graphs, characterized in that, Includes the following steps: Step 1: Extract technical nodes, technical development relationships, and expert evaluation information from technical report data and expert opinion data; Step 2: Store the original technology roadmap, the technology evolution path obtained by the technology evolution path generation method according to any one of claims 1-5, and the extraction results of step 1 in a unified manner, and establish the relationships between entities to obtain a knowledge graph; Step 3: Use a large language model to reason about the knowledge graph, dynamically supplement the technology development path, and finally update the technology roadmap.

7. The method for dynamically generating technology roadmaps based on large models and knowledge graphs according to claim 6, characterized in that, Step 1 specifically includes: performing text cleaning, segmentation, sentence segmentation, time annotation, source annotation, and terminology standardization on the technical report data and expert opinion data to obtain technical report text fragments and expert opinion text fragments; A knowledge graph ontology of the technology roadmap is constructed. Then, under the constraints of the knowledge graph ontology, the text fragments of the technology report type and the text fragments of the expert opinion type are respectively input into the large language model. Knowledge extraction is performed according to the preset prompt words to obtain the extraction results of the technology report type and the extraction results of the expert opinion type.

8. The method for dynamically generating technology roadmaps based on large models and knowledge graphs according to claim 7, characterized in that, The knowledge graph ontology includes entity types and relationship types; the entity types include technology nodes, application scenarios, and expert evaluations; the relationship types include evolutionary relationships between technology nodes, recommended development and risk relationships between expert evaluations and technology nodes, and applicable relationships between technology nodes and application scenarios. The extracted results of the technical report category include technical nodes, mainstream development directions, key technical bottlenecks, and explicit development relationships between technologies. The extracted results of the expert opinion category include assessments of technical potential and warnings of technical risks.

9. The method for dynamically generating technology roadmaps based on large models and knowledge graphs according to claim 7, characterized in that, Step 2 specifically includes: using the technology evolution path and the original technology roadmap as the basic framework of the knowledge graph, supplementing it with the consensus and non-consensus knowledge provided by the extraction results, performing knowledge fusion, and storing the fused knowledge graph in a graph database; Step 3 specifically includes: firstly, preliminary screening of the entity types in the knowledge graph to obtain candidate node pairs; then, inputting the candidate node pairs and their neighborhood context into the large model for combination; judging whether there is a reasonable development or evolution relationship between any two technical nodes; and outputting the connection reason and the probability score of the path's future realization; the neighborhood context includes the entity nodes and relationships adjacent to the candidate node pairs in the knowledge graph. The predicted new paths, which include probability assessments, are added to the existing technology roadmap to form a multi-level technology roadmap that includes main paths, branch paths, and alternative paths. At the same time, the future development potential of different technology directions is quantitatively ranked based on the probability scores of the paths.

10. A dynamic generation system for technology roadmaps based on large models and knowledge graphs, characterized in that, include: The data acquisition module is used to acquire multi-source heterogeneous data in the target technology field; the multi-source heterogeneous data includes patent document data, technical report data, expert opinion data, and original technology roadmap data; A technology evolution path mining module is used to obtain a technology evolution path using the technology evolution path generation method according to any one of claims 1-4; The knowledge mining module is used to extract technical nodes, technical development relationships, and expert evaluation information from the technical report data and expert opinion data. The knowledge graph construction module is used to uniformly store the technology evolution path, the results extracted by the knowledge mining module and the original technology roadmap, and establish relationships between entities to obtain a knowledge graph. The link prediction and path update module is used to reason about the knowledge graph using a large language model, dynamically supplement the technology development path, and ultimately update the technology roadmap.