Energy system subversive technology identification method, equipment and medium

By combining knowledge graphs and reinforcement learning models, a method for identifying disruptive technologies in energy systems is constructed, which solves the problems of low identification efficiency and insufficient accuracy in existing technologies, and achieves efficient and accurate identification and continuous optimization of disruptive technologies.

CN120996046APending Publication Date: 2025-11-21STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Patent Information

Application Number
CN202511076743.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately identify disruptive technologies in energy systems. Traditional methods are inefficient and susceptible to subjective influences, while data mining methods fail to deeply understand the underlying principles and complex relationships of technologies, leading to misjudgments or omissions of potential disruptive technologies.

Method used

This paper adopts a method combining knowledge graph and reinforcement learning model. By acquiring relevant text data of energy system, extracting technical keywords as nodes, constructing association relationships and setting edge weights, and using intelligent agents to explore and learn in the knowledge graph, it identifies technologies with disruptive potential, designs a disruptive technology identification reward, exploration reward and path efficiency reward mechanism, and optimizes the knowledge graph structure.

Benefits of technology

It has achieved accurate identification of disruptive technologies, improved identification efficiency and accuracy, can adapt to the needs of large-scale data processing and analysis, continuously optimizes the performance of knowledge graphs, and adapts to the rapid changes in the field of energy technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996046A_ABST
    Figure CN120996046A_ABST
Patent Text Reader

Abstract

The invention relates to an energy system subversive technology identification method, equipment and a medium. The method comprises the following steps: acquiring related text data of an energy system, extracting technical keywords as nodes, acquiring an association relationship between the nodes to construct edges and set edge types and edge weights, forming a knowledge graph and dynamically updating the knowledge graph; the dynamically updated knowledge graph serves as an environment, exploration and learning are conducted in the knowledge graph through a reinforcement learning model, a technical list with subversive potential is recognized, the state of an intelligent agent is the current position of the intelligent agent in the knowledge graph and related information of an accessed node, and the state of the intelligent agent is the current position of the intelligent agent in the knowledge graph and the related information of the accessed node. The action is an operation performed by the intelligent agent in the knowledge graph, and the rewards comprise a subversive technology identification reward, an exploration reward and a path efficiency reward; and in the process of exploring the intelligent agent, optimizing and adjusting the knowledge graph according to the exploring behavior of the intelligent agent. Compared with the prior art, the method has the advantages that the subversive technology can be efficiently and accurately identified, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy system technology, and in particular to a method, device and medium for identifying disruptive technologies in energy systems. Background Technology

[0002] In the energy sector, disruptive technologies refer to those that can significantly impact existing energy systems, triggering fundamental changes in energy production, transmission, storage, or consumption. These technologies are typically characterized by high innovation, significant development potential, and the ability to replace or surpass traditional technologies. Examples include early replacements of firewood with coal and coal with oil, as well as the recent rise of renewable energy technologies and smart grid technologies. Their emergence often significantly improves energy efficiency, reduces energy costs, decreases environmental pollution, or opens up entirely new energy application areas, thereby driving the transformation, upgrading, and sustainable development of energy systems.

[0003] Currently, the main methods for identifying disruptive technologies in energy systems include:

[0004] Traditional literature analysis relies primarily on manual screening and analysis of a large volume of energy technology literature to identify potentially disruptive technologies. However, with the deepening of research in the energy field and the explosive growth of technical literature, manual analysis is inefficient and susceptible to subjective factors, making it difficult to grasp all relevant technical information in a timely and comprehensive manner, and easily overlooking some potentially disruptive emerging technologies.

[0005] Expert evaluation method: This method relies on the experience and judgment of domain experts to assess the disruptive potential of technologies. However, expert evaluation suffers from strong subjectivity and difficulty in ensuring consistency of results. Furthermore, experts may lack sufficient understanding of emerging or interdisciplinary technologies, leading to inaccurate assessments. In addition, the expert evaluation process is typically time-consuming and cannot quickly respond to rapid changes in the energy technology field.

[0006] Data mining: This method uses data mining techniques to extract information and patterns related to energy technologies from massive amounts of data to identify disruptive technologies. CN118261362A discloses a method for analyzing and predicting the development trend map of new energy technologies, which includes the following steps: Step 1: Data collection: Collect relevant data on new energy technologies, including historical data, policy data, market data, etc., to establish a new energy technology database; Step 2: Data mining and analysis: Perform data mining and analysis on the new energy technology database to extract key indicators and trend characteristics of new energy technology development; By constructing a new energy technology development trend map, a comprehensive and systematic understanding of the current status and trends of new energy technologies can be achieved, avoiding problems such as inaccurate data and strong subjectivity in traditional methods. Furthermore, based on the application of data mining and machine learning technologies, accurate prediction and analysis of new energy technology development trends can be achieved, providing a scientific basis and support for relevant decision-making and planning. However, although this method involves data mining, it only predicts development trends and cannot identify disruptive technologies. Moreover, this method struggles to deeply understand the underlying principles, application scenarios, and complex relationships with other technologies, easily leading to misjudgments or omissions of potentially disruptive technologies. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art by providing a method, device and medium for identifying disruptive technologies in energy systems.

[0008] The objective of this invention can be achieved through the following technical solutions:

[0009] According to a first aspect of the present invention, a method for identifying disruptive technologies in an energy system is provided, the method comprising the following steps:

[0010] Acquire text data related to energy systems, extract technical keywords as nodes, determine node attributes, obtain the relationships between technical keywords to construct edges, set edge types according to the type of relationships, set edge weights according to the degree of relevance of relationships, form a knowledge graph and update it dynamically;

[0011] Using a dynamically updated knowledge graph as the environment, reinforcement learning models are employed to explore and learn within the knowledge graph, identifying a list of technologies with disruptive potential. The agent's state comprises its current position within the knowledge graph and information about visited nodes. Actions are the operations performed by the agent within the knowledge graph, including selecting neighboring nodes, lingering and collecting neighboring information, and jumping to new nodes. Rewards include rewards for identifying disruptive technologies, exploration rewards, and path efficiency rewards. During the agent's exploration process, the knowledge graph is optimized and adjusted based on its exploration behavior.

[0012] As a preferred technical solution, the edge types include semantic association edges, application scenario edges, technology evolution edges, and social attention edges. The semantic association edges represent the semantic association relationship between technical keywords, the application scenario edges represent the synergistic / substitution relationship between technical keywords in different application scenarios, the technology evolution edges represent the technological development / improvement relationship between technical keywords, and the social attention edges represent the degree of social attention relationship between technical keywords.

[0013] As a preferred technical solution, the edge weights are specifically set as follows:

[0014] For semantically related edges, the co-occurrence frequency of technical keywords and the semantic strength of co-occurring statements are extracted from the text data, and the edge weight of the semantically related edge is determined based on the co-occurrence frequency and semantic strength.

[0015] For application scenario edges, the frequency of common occurrence of technical keywords in different application scenarios is calculated to determine the edge weight of the application scenario edge.

[0016] For technology evolution edges, the edge weights are calculated based on patent citation frequency, patent citation timeliness, document co-occurrence frequency, and knowledge inheritance relationship.

[0017] For the social attention edge, the edge weight is determined based on the co-occurrence frequency of technical keywords in social media, user interaction, and user sentiment, as well as the co-occurrence frequency, reporting influence, and reporting depth of technical keywords in news reports.

[0018] As a preferred technical solution, the node attributes include technical description, number of patents, technical breakthroughs, market size, cost-effectiveness, reliability indicators, and environmental protection indicators.

[0019] As a preferred technical solution, the state of the intelligent agent includes:

[0020] Current location node information: including the technical keywords and node attributes corresponding to the current node;

[0021] Visited Node List: Records the sequence of nodes that the agent has visited;

[0022] Local knowledge graph structure: contains the partial knowledge graph structure and edge weight information of the neighboring nodes of the current node.

[0023] As a preferred technical solution, the rewards include:

[0024] A reward for identifying disruptive technologies is obtained by weighted summation of multiple disruptive indicators and their weights, wherein the weights are determined based on the confidence level of the acquired disruptive indicators.

[0025] The exploration reward is determined based on the exploration level of a region and the technological novelty of the accessed node. The knowledge graph is divided into different regions. The exploration level of each region is defined as the ratio of the number of accessed nodes in that region to the total number of nodes. The technological novelty index is defined as the ratio of the technological novelty evaluation value of the accessed node to the time when the agent accesses the node plus 1. The exploration reward is negatively correlated with the exploration level of the region and positively correlated with the technological novelty index.

[0026] The path efficiency reward is set based on the path length from the starting node to the current node and a dynamic adjustment coefficient. The dynamic adjustment coefficient is determined based on the ratio of the value of the visited nodes to a preset value threshold. When the current node is the target node, the path efficiency reward is set to a positive preset value.

[0027] As a preferred technical solution, the disruptive indicators include technological innovation indicators, technology market potential indicators, technology performance indicators, and social impact indicators. The technological innovation indicators include the patent number growth rate and technology breakthrough score. The technology market potential indicators include the market growth rate and investment enthusiasm index. The technology performance indicators include the performance improvement rate and cost reduction rate. The social impact indicators include the rate of increase in public attention and the intensity of policy support.

[0028] As a preferred technical solution, the optimization and adjustment of the knowledge graph based on the agent's exploration behavior during the agent's exploration process specifically includes:

[0029] Edge weight adjustment based on agent access frequency: During the agent's exploration process, the number of times the agent visits a node through each edge is recorded to obtain the edge access frequency; the original edge weight is adjusted based on the edge access frequency, and the higher the edge access frequency, the greater the edge weight after the jump.

[0030] Edge weight adjustment based on agent feedback: The feedback value of each edge is calculated based on the agent's feedback. The feedback value of 1 indicates positive feedback and 0 indicates negative feedback. The original edge weights are adjusted based on the feedback value.

[0031] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0032] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] 1. This invention, by combining knowledge graphs and reinforcement learning models, can accurately identify a list of technologies with disruptive potential. The reinforcement learning model can learn the complex relationships and potential patterns between technologies within the knowledge graph, thereby more accurately identifying disruptive technologies that may have a significant impact on the energy system. Furthermore, it can efficiently explore and identify technologies within a vast energy technology knowledge graph, saving time and effort compared to manual screening and analysis. The autonomous exploration and learning capabilities of the intelligent agent enable the identification process to quickly converge on technologies with potential, improving identification efficiency.

[0035] 2. This invention uses a dynamically updated knowledge graph as the environment and utilizes a reinforcement learning model to explore and learn within the knowledge graph. It can autonomously select adjacent nodes, stay and collect information, and jump to new nodes within the knowledge graph. This exploration capability enables the agent to efficiently traverse the knowledge graph and effectively identify a list of technologies with disruptive potential.

[0036] 3. This invention designs a reward mechanism that includes rewards for identifying disruptive technologies, exploration rewards, and path efficiency rewards. These reward mechanisms can effectively guide the behavior of the agent, making it pay more attention to technologies with disruptive potential during the exploration process, while encouraging the agent to explore unknown areas and improve path efficiency, thereby improving the accuracy and efficiency of disruptive technology identification.

[0037] 4. This invention optimizes and adjusts the knowledge graph based on the agent's exploration behavior during the exploration process. For example, it adjusts edge weights based on the frequency of edge visits and feedback information. This allows the knowledge graph to automatically optimize its structure and content according to the actual exploration situation, further improving its quality and effectiveness. By continuously optimizing and adjusting based on the agent's exploration behavior, the knowledge graph can continuously improve its performance. This self-optimization capability enables the knowledge graph to continuously adapt to new data and new exploration needs during long-term use, maintaining its advantage in identifying disruptive technologies.

[0038] 5. The method of the present invention has good scalability. The knowledge graph construction and update mechanism and the reinforcement learning model optimization process can adapt to the needs of large-scale data processing and analysis, and can be expanded with the continuous development of the energy technology field and the increase of data volume. Attached Figure Description

[0039] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0041] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. “Multiple” in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0042] Example 1

[0043] This embodiment provides a method for identifying disruptive technologies in energy systems, such as... Figure 1 As shown, the method includes the following steps:

[0044] Step 1) Obtain text data related to the energy system, extract technical keywords as nodes, determine node attributes, obtain the relationships between technical keywords to construct edges, set edge types according to the type of relationships, set edge weights according to the degree of relevance of relationships, form a knowledge graph and update it dynamically.

[0045] In this embodiment, the acquired energy system-related text data includes, but is not limited to, patent data, literature data, public reports, and online discussion data, which can be crawled using web scraping technology. The collected text data is cleaned to remove irrelevant information, garbled characters, duplicate content, etc., the text content is converted into a unified encoding format, and the text is standardized, such as unifying capitalization, removing redundant spaces and punctuation marks, etc.

[0046] Step 11) Extract nodes

[0047] In one embodiment, the TF-IDF algorithm can be used to extract technical keywords as nodes in a knowledge graph. Specifically, the term frequency (TF) and inverse text frequency (IDF) of each word in the text are calculated. The TF-IDF value reflects the importance of a word in the text. The frequency of each word in the text is counted, and high-frequency words are selected as candidate keywords based on a set frequency threshold. It should be noted that simple term frequency statistics may filter out some common stop words, so a stop word list is needed for filtering. This process is common practice for those skilled in the art, and will not be elaborated further in this embodiment. Natural language processing methods can also be used in other embodiments, and this embodiment does not limit the means of extracting technical keywords.

[0048] Step 12) Determine node attributes

[0049] In one embodiment, node attributes include technical description, number of patents, technical breakthroughs, market size, cost-effectiveness, reliability indicators, and environmental indicators.

[0050] Technical Description: This section provides a detailed explanation of the basic principles, methods, or processes of the technology. For example, wind power generation technology uses wind to drive the blades of a wind turbine to rotate, thereby driving a generator to produce electricity; solar photovoltaic technology is based on the photoelectric effect of semiconductor materials, directly converting sunlight into electrical energy.

[0051] Number of patents: This indicates the number of patent applications or grants related to the technology. A large number of patents usually signifies strong innovation and high technological barriers. For example, a novel battery technology with over 100 patents indicates significant innovation in this technological field.

[0052] Technological breakthroughs: These reflect the innovative points and breakthroughs in the theory or practice of the technology. For example, a certain hydrogen energy technology has overcome the high energy consumption problem in the traditional hydrogen production process, realizing a green and efficient hydrogen production method.

[0053] Market size: refers to the scale of application and market demand for the technology. For example, the global market size for wind power technology reached [X] billion RMB in 2023, reflecting the market acceptance and application of the technology.

[0054] Cost-effectiveness: This measures the cost advantages and benefits of a technology, including the ratio of construction costs, operating costs, and maintenance costs to energy output or benefits. For example, the cost-effectiveness of wind power technology continues to improve as the technology matures and its scale expands, gradually enhancing its competitiveness in the energy market.

[0055] Reliability metrics, such as Mean Time Between Failures (MTBF) and failure rate, are used to evaluate the stability and reliability of a technology. For example, a longer MTBF for an energy storage technology indicates higher reliability and greater stability when applied in energy systems.

[0056] Environmental indicators: These measure the environmental impact of technology, including carbon emission intensity, renewable resource utilization rate, energy conversion efficiency, and waste emissions.

[0057] These node attributes can provide reinforcement learning agents with rich information, helping them to better understand and evaluate technical keywords, thereby more effectively exploring knowledge graphs and identifying technologies with disruptive potential.

[0058] Step 13) Construct edges

[0059] In this embodiment, the edge types include semantic association edges, application scenario edges, technology evolution edges, and social concern edges, which are constructed by identifying semantic association relationships, application scenario relationships, technology evolution relationships, and social concern relationships between nodes.

[0060] Semantic association edges represent the semantic relationships between technical keywords. Natural language processing techniques can be used to analyze text data and uncover these semantic connections. For example, if two technical keywords are frequently mentioned in the same document and are related in syntactic structure, such as "Technology A improves the efficiency of Technology B," then a semantic association edge can be constructed between them.

[0061] Application scenario edges represent the synergistic / substitutional relationships of technical keywords in different application scenarios. Application scenario edges are constructed based on the frequency of occurrence of technologies in different application scenarios. For example, if both technologies A and B frequently appear in the application scenario of offshore wind farms, an application scenario edge is constructed between them, indicating that they may have synergistic or substitutional relationships in this scenario.

[0062] A technology evolution edge represents the technological development / improvement relationship between technological keywords, indicating that one technology is a development or improvement of another. For example, "polycrystalline silicon solar cell technology" is an improved evolution of "monocrystalline silicon solar cell technology," and there is a technology evolution edge between them.

[0063] The social attention edge represents the relationship of social attention between key technological terms, reflecting the public's level of attention and opinion towards different technologies. For example, if the public is highly concerned about the environmental advantages of electric vehicle technology and also highly concerned about the pollution problems of traditional fuel vehicle technology, a social attention edge will be formed between the two.

[0064] Step 14) Set edge weights

[0065] For semantically related edges, the co-occurrence frequency of technical keywords and the semantic strength of co-occurring sentences are extracted from the text data. The edge weight of the semantically related edge is determined based on the co-occurrence frequency and semantic strength. If two technical keywords co-occur multiple times in the same document and their semantic relationship in sentences is close (e.g., in causal sentences or modifying sentences), the weight is increased. For example, if "smart grid technology" and "distributed energy access technology" co-occur in multiple documents and frequently appear in sentences such as "smart grid technology makes distributed energy access more efficient," it indicates that their semantic relationship is strong, and the weight of the semantically related edge is high.

[0066] For application scenario edges, the frequency of co-occurrence of technical keywords in different application scenarios is calculated to determine the edge weight of the application scenario edge. If two technical keywords frequently co-occur in a certain application scenario and play an important role in that scenario, having a key impact on the scenario's operational effectiveness, then the application scenario edge weight is increased. For example, in the offshore wind farm application scenario, if "offshore wind power generation technology" and "offshore wind power installation and maintenance technology" are indispensable and widely used in multiple offshore wind power projects, then the application scenario edge weight between them is high.

[0067] For technology evolution edges, the edge weights are calculated based on patent citation frequency, patent citation timeliness, document co-occurrence frequency, and knowledge transfer relationships. Patent citation frequency refers to the number of times a patent for technology A is directly cited by a related patent for technology B. If a patent for technology B cites a patent for technology A multiple times, this indicates that technology A has a significant impact on the evolution of technology B, and the edge weight can be increased accordingly. Patent citation timeliness is weighted by the citation time; recent citations may have a greater impact on technology evolution. An exponential decay function can be used to measure the timeliness of citations. For example, for citations within the past 5 years, the time weights are set as follows: 0.5 for citations in the 1st year (the most recent year), 0.4 for the 2nd year, 0.3 for the 3rd year, 0.2 for the 4th year, and 0.1 for the 5th year. If patent B cites patent A 10 times in year 1 and 5 times in year 3, the weighted citation count is 10 × 0.5 + 5 × 0.3 = 6.5, and the total citation count is 15. The weight of the technology evolution edge can be adjusted to 6.5 / 15 = 0.43. For document co-occurrence frequency, the number of times technology A and technology B are mentioned simultaneously in academic literature is counted. For example, if technology A and technology B co-occur in 100 documents within a certain period, and technology A appears in a total of 200 documents, the weight of the technology evolution edge can be initially set to 100 / 200 = 0.5. To support knowledge transfer relationships, natural language processing techniques are used to mine the influence of technology A on technology B described in the literature. For example, if the literature explicitly states that technology B is an improvement on technology A, the weight of the technology evolution edge can be increased. If 30 out of 100 co-occurring documents explicitly mention this knowledge transfer relationship, the weight can be increased by 0.15 from the previous 0.5, reaching 0.65.

[0068] For the social attention edge, the edge weight is determined based on the co-occurrence frequency of technical keywords in social media, user interaction, and user sentiment, as well as the co-occurrence frequency, impact, and depth of technical keywords in news reports. For the co-occurrence frequency of technical keywords: the number of times related keywords of technology A and technology B appear simultaneously in social media platforms or news reports is counted. For example, if keywords of technology A and technology B appear together 1000 times on Weibo in a certain month, and keyword A appears a total of 5000 times, the social attention edge weight can be initially set as 1000 / 5000 = 0.2. For user interaction, this includes likes, comments, and reposts. If users frequently interact with content that contains both technology A and technology B, it indicates a high level of social attention to their association. For example, in these 1000 co-occurrences, the average number of likes per post is 100, the number of comments is 50, and the number of shares is 20. The average number of likes for similar content on the platform is 50, the number of comments is 20, and the number of shares is 10. Therefore, the weight of the social attention edge can be increased by 0.1 from the previous 0.2, reaching 0.3. Regarding user sentiment, sentiment analysis techniques are used to determine users' sentiment towards technologies A and B. If users mostly give positive evaluations when mentioning the two technologies and believe they have good synergy or application prospects, then the weight can be increased. For example, after sentiment analysis, if 70% of the content in the 1000 co-occurrences shows positive sentiment towards the two technologies, the weight can be increased by another 0.1, reaching 0.4. Regarding the impact of media coverage, if the co-occurrence reports of technologies A and B come from authoritative media, such as well-known industry magazines and mainstream news websites, and these reports are widely reprinted, the weight can be increased. Regarding the depth of reporting, if the reporting focuses on the synergistic application of technology A and technology B, and their social impact, and the content provides an in-depth analysis of the relationship between them, then the weight can be further increased. For example, if 20 out of 50 co-occurring reports provide an in-depth analysis of the prospects for the synergistic application of the two technologies, the weight can be increased by 0.1, reaching 0.5.

[0069] By using the above methods, nodes, node attributes, edges, and edge weights are determined to construct a knowledge graph. This knowledge graph is dynamically updated. Data sources for the knowledge graph, such as databases, literature, and web pages, are continuously monitored to promptly detect data updates. For example, when new energy technology patents or academic papers are detected, key information is extracted from this new data and the knowledge graph is updated. Alternatively, key events can be set to trigger knowledge graph updates, such as the release of new technologies or policy changes. When an event occurs, the relevant parts of the knowledge graph are immediately updated to ensure the environment reflects the latest information. Simultaneously, the updated knowledge graph is fed back to the reinforcement learning agent, enabling it to make decisions based on the latest information. To avoid ambiguity in the purpose of this application, this embodiment does not provide a detailed description of the specific updating methods of the knowledge graph.

[0070] Step 2) Using a dynamically updated knowledge graph as the environment, a reinforcement learning model is used to explore and learn within the knowledge graph to identify a list of technologies with disruptive potential.

[0071] Intelligent Agent: Represents the entity that explores the knowledge graph, with the goal of identifying technologies with disruptive potential. The intelligent agent can traverse nodes and collect information within the knowledge graph.

[0072] The key elements related to reinforcement learning are set as follows:

[0073] 1. State: The agent's current position in the knowledge graph and information about the nodes it has visited, including:

[0074] Current location node information: including the technical keywords and node attributes corresponding to the current node;

[0075] List of visited nodes: Records the sequence of nodes that the agent has visited, helping the agent avoid repetitive visits and understand the explored paths;

[0076] Local knowledge graph structure: It contains the partial knowledge graph structure and edge weight information of the neighboring nodes of the current node, providing the agent with local environmental perception.

[0077] 2. Action: The operations performed by the agent in the knowledge graph, including:

[0078] Choosing neighboring nodes: The agent can choose to move from its current position node to a neighboring node along an edge. The selection of neighboring nodes is based on the edge connections of the current position node.

[0079] Stop and collect neighboring information: The agent can choose to stop at the current location node to further collect and analyze the detailed information of the node, which may consume a certain number of steps or time units.

[0080] Jumping to a new node: An agent can choose to jump to a new node in the knowledge graph that has not yet been visited to explore unknown areas. This action usually has a certain degree of randomness, but it may also be based on certain strategies (such as the popularity of the node, its relevance to the current goal, etc.).

[0081] 3. Reward: Used to measure the benefit an agent receives after taking a certain action in a given state, including:

[0082] Disruptive Technology Identification Reward: A substantial positive reward is given when an agent successfully identifies a technological node with disruptive potential. The reward is calculated by a weighted sum of multiple disruptive indicators and their weights, where the weights are determined based on the confidence level of the acquired disruptive indicators. The formula is as follows:

[0083]

[0084] Where R1 is the reward for identifying disruptive technologies, N is the number of disruptive indicators, and w i S represents the weight corresponding to index i. i Let be the score of indicator i.

[0085] Exploration Reward: Based on the exploration level of a region and the technological novelty of the accessed nodes, and combining the technological novelty of the nodes with the access order of the agents, agents that access novel technology nodes earlier are given additional rewards to encourage agents to prioritize the exploration of emerging technology regions in the knowledge graph. Specifically, the knowledge graph is divided into different regions, and the exploration level of each region is defined as the ratio of the number of nodes already visited in that region to the total number of nodes. The technological novelty index is defined as the ratio of the technological novelty evaluation value of the accessed node to the time when the agent accesses that node plus 1. The exploration reward is negatively correlated with the exploration level of the region and positively correlated with the technological novelty index. In one embodiment, the exploration reward is expressed as:

[0086]

[0087] Among them, V i X represents the exploration level of the region where node i is located. i The novelty evaluation value of node i represents the technological innovation of the technology, which can be determined by analyzing the publication time of the latest research and the patent application time of the technology. The larger the value, the more novel the technology. t represents the time step (time step) when the agent visits the node. α is a weighting coefficient used to balance the influence of access frequency and novelty factors on the reward. In this way, the agent can obtain more rewards when visiting novel nodes, and the earlier the novel node is visited, the higher the reward.

[0088] Path efficiency reward: This is set based on the path length from the starting node to the current node and a dynamic adjustment coefficient. Dynamic path evaluation is introduced, taking into account the agent's real-time feedback during exploration. This allows the agent to dynamically adjust the path efficiency reward based on the value of visited nodes, enhancing the agent's adaptive learning ability for efficient paths. The dynamic adjustment coefficient is determined based on the ratio of the value of visited nodes to a preset value threshold. When the current node is the target node, the path efficiency reward is set to a positive preset value. The path efficiency reward is expressed as:

[0089]

[0090] Where p is the path length from the starting node to the current node, and β is a dynamic adjustment coefficient, ranging from 0 to 1. This way, when the agent accumulates sufficient value along the path, it can receive a higher reward even if the path is long; conversely, the penalty is reduced, guiding the agent to comprehensively consider both path length and node value.

[0091] The total reward R = R1 + R2 + R3.

[0092] 4. Environment: This refers to the constructed knowledge graph of the energy system, containing information such as technical keyword nodes, node attributes, edges, and edge weights. The environment is dynamically updated based on the actions of the agent and provides feedback to the agent.

[0093] When new energy technology information emerges (such as new patents, research results, market dynamics, etc.), the node attributes and edge weights in the knowledge graph should be updated in a timely manner. At the same time, based on the agent's exploratory behavior, such as the agent's frequent visits to certain nodes or high attention to certain technology combinations, the knowledge graph can also be optimized and adjusted, for example, by adjusting the edge weights to reflect the agent's new understanding of the technological correlations.

[0094] The environment provides reward signals and new state information to the agent based on its actions and state changes. For example, when the agent chooses to move to an adjacent node, the environment provides the agent with the state information of that node and calculates the corresponding reward value according to the reward function, feeding it back to the agent.

[0095] In this embodiment, during the agent's exploration process, the knowledge graph is optimized and adjusted based on its exploration behavior, specifically including:

[0096] ① Edge weight adjustment based on agent access frequency: During the agent's exploration process, the number of times the agent visits a node through each edge is recorded to obtain the edge access frequency f; the original edge weights are adjusted based on the edge access frequency, with a higher edge access frequency resulting in a larger edge weight after the jump. In one embodiment, the adjustment method is as follows:

[0097] W new =W old +γ×f

[0098] Among them, W new W old Here, γ represents the edge weights before and after the update, γ is the learning rate which controls the speed of weight updates, and f is the edge access frequency.

[0099] ② Edge weight adjustment based on agent feedback: Calculate the feedback value for each edge based on the agent's feedback. A feedback value of 1 indicates positive feedback, and 0 indicates negative feedback. Adjust the original edge weights based on the feedback values. In one embodiment, the adjustment method is as follows:

[0100] W new =W old +δ×feedback

[0101] Where δ is the learning rate, which controls the speed of weight updates, and feedback is the feedback value.

[0102] 5. Policy: The policy by which an agent chooses actions determines what action the agent should take in a given state.

[0103] Exploration Strategy: In the early stages of training, a random exploration strategy is adopted, allowing the agent to randomly select actions to fully explore different regions and paths of the knowledge graph. As training progresses, exploration strategies based on the learned value function or policy network are gradually introduced, such as the ε-greedy strategy. In most cases, the agent selects the current optimal action, but still conducts random exploration with a certain probability to balance exploration and exploitation.

[0104] Utilizing a policy: Based on learned knowledge graph information and reward signals, the agent selects actions that maximize long-term rewards. For example, based on information about visited nodes and edge weights, it selects the neighboring nodes most likely to lead to nodes with disruptive potential technologies. A policy network, such as a Deep Q-Network (DQN), can be constructed to learn the optimal policy through continuous interaction with the environment.

[0105] Reinforcement learning algorithms perform the following steps:

[0106] 1. Initialization: Construct an initial energy system knowledge graph, including information such as nodes, edges, edge weights, and node attributes. Initialize the state of the agent, typically by placing it at the starting node of the knowledge graph (which can be a core technology node or a randomly selected node).

[0107] 2. State Evaluation: The agent evaluates the actionable actions based on the current state (including the current location node information, the list of visited nodes, and the local knowledge graph structure).

[0108] 3. Action Selection: Based on the current policy, the agent selects an action (select an adjacent node, stay and collect information, or jump to a new node).

[0109] 4. Environmental Feedback: The environment updates the knowledge graph based on the agent's actions (if necessary) and calculates the corresponding reward value to feed back to the agent. At the same time, the environment provides the agent with new state information (new location node information after the agent moves, updated list of visited nodes, and new local knowledge graph structure, etc.).

[0110] 5. Policy Update: The agent updates its policy based on the reward signals and new state information it receives, so as to continuously improve the decision-making process.

[0111] 6. Iterate through the loop: Repeat steps 2-5 until the termination condition is met (such as reaching the maximum number of iterations, the agent finding a certain number of disruptive technologies, or the knowledge graph exploration reaching a certain level of coverage).

[0112] 7. Output Results: Based on the agent's exploration results, output a list of technologies with disruptive potential. These technologies can be sorted according to the confidence level or reward value identified by the agent, providing a reference for energy technology decisions.

[0113] The above approach utilizes reinforcement learning to effectively explore knowledge graphs, uncover a list of technologies with disruptive potential, and provide strong support for the development and innovation of energy systems.

[0114] Example 2

[0115] This embodiment provides a specific implementation method for a disruptive indicator, based on embodiment 1.

[0116] In this embodiment, disruptive indicators include technological innovation indicators, technology market potential indicators, technological performance indicators, and social impact indicators.

[0117] (1) Technological innovation indicators

[0118] (11) Patent growth rate

[0119] Patent growth rate r 专利 This represents the average annual growth rate of the number of patents related to this technology over a specific period (e.g., 5 years). The formula is:

[0120]

[0121] Among them, P current P represents the current number of patents. past The initial number of patents in the past period is n, where n is the number of years.

[0122] (12) Technological Breakthrough Score

[0123] The technological breakthrough score is given by field experts based on the degree of breakthrough in key performance indicators, with a maximum score of 10 points. For example, if a battery technology achieves a major breakthrough in energy density, the score may be 8 points.

[0124] (2) Technology Market Potential Indicators

[0125] (21) Market growth rate

[0126] Calculate the compound annual growth rate (CAGR) of the technology's market over a specific period (e.g., 3-5 years) using market analysis reports or model forecasts. The formula is:

[0127]

[0128] Among them, M future M is the predicted future market size. current Let n be the current market size, and n be the number of years.

[0129] (22) Investment Popularity Index

[0130] The ratio of total venture capital investment in this technology sector over the past year to the historical average investment in this sector is used as the investment popularity index I. 投资 The formula is:

[0131]

[0132] Among them, I pastear The total investment over the past year, I history This represents the historical average total investment.

[0133] (3) Technical performance indicators

[0134] (31) Performance improvement ratio

[0135] Performance improvement ratio R 性能 This refers to the percentage improvement in key performance indicators (such as energy conversion efficiency and driving range) compared to existing mainstream technologies. The formula is:

[0136]

[0137] Among them, P new P represents the performance value of the new technology. old These are the old technical performance values.

[0138] (32) Cost reduction ratio

[0139] Cost reduction ratio R 成本 This represents the percentage reduction in cost per unit of product or service compared to existing technologies. The formula is:

[0140]

[0141] Among them, C old For the cost of old technology, C new Cost of new technologies.

[0142] (4) Social impact indicators

[0143] (41) Rate of increase in public attention r 关注度

[0144] The average monthly growth rate of public interest in the technology over the past year was calculated using data such as social media and news search trends. The formula is:

[0145]

[0146] Among them, H current Based on current search popularity, H past This reflects the search popularity from a year ago.

[0147] (42) Policy support intensity

[0148] The score is based on a comprehensive evaluation of the technology's level of policy support from national and local governments (such as subsidy amounts, tax incentives, and research project funding), with a maximum score of 10. For example, a hydrogen energy technology that receives substantial subsidies and multiple research project support could receive a score of 9.

[0149] Example 3

[0150] This embodiment, based on embodiment 1, provides a method for determining the feedback value in edge weight adjustment based on agent feedback.

[0151] The feedback value is the agent's evaluation of the edge, reflecting how the edge helps the agent's exploration. It is determined as follows:

[0152] 1) Successfully identified feedback value I

[0153] When an agent reaches a new node via an edge, if it identifies a disruptive technology, the feedback value I is 1; otherwise, I is 0. For example, if the agent reaches node B via edge AB and discovers a disruptive technology, the feedback value is 1, indicating that edge AB has high value.

[0154] 2) Explore the feedback value E of efficiency feedback

[0155] The feedback value is related to the efficiency with which the agent reaches a new node via an edge. If new information can be quickly obtained through an edge, the feedback value E is 1; if the edge leads to a loop or repeated visits, the feedback value E is negative (e.g., -1). For example, if the agent quickly reaches an unvisited node via edge CD, the feedback value is 1; if it enters a loop via edge EF, the feedback value is -1.

[0156] 3) Feedback value Q of path effectiveness feedback

[0157] The effectiveness of the path taken by the agent after traversing an edge is determined. If the edge leads the agent closer to a potentially disruptive technology, the feedback value Q is 1; if it deviates from the target, the feedback value Q is 0. For example, if the agent approaches the target region via edge GH, the feedback value is 1; if it deviates from the target via edge IJ, the feedback value is 0.

[0158] 4) Overall Feedback Value

[0159] The feedback value is calculated by combining multiple factors. For example:

[0160] feedback=μ1×I+μ2×E+μ3×Q

[0161] Where, μ i i = 1, 2, 3 are weighting coefficients, I is the feedback value of successful identification, E is the feedback value of exploration efficiency, and Q is the feedback value of path effectiveness.

[0162] The determination of feedback values ​​should be based on the agent's goals and task requirements. By designing a reasonable feedback mechanism, the agent can effectively explore the knowledge graph and identify technologies with disruptive potential.

[0163] Example 4

[0164] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0165] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0166] The processing unit performs the various methods and processes described above, such as steps 1)-2). For example, in some embodiments, steps 1)-2) may be implemented as a computer software program tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of steps 1)-2) described above may be performed. Alternatively, in other embodiments, the CPU may be configured to perform steps 1)-2) by any other suitable means (e.g., by means of firmware).

[0167] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0168] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0169] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0170] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for identifying disruptive technologies in an energy system, characterized in that, The method includes the following steps: Acquire text data related to energy systems, extract technical keywords as nodes, determine node attributes, obtain the relationships between technical keywords to construct edges, set edge types according to the type of relationships, set edge weights according to the degree of relevance of relationships, form a knowledge graph and update it dynamically; Using a dynamically updated knowledge graph as the environment, reinforcement learning models are employed to explore and learn within the knowledge graph, identifying a list of technologies with disruptive potential. The agent's state comprises its current position within the knowledge graph and information about visited nodes. Actions are the operations performed by the agent within the knowledge graph, including selecting neighboring nodes, lingering and collecting neighboring information, and jumping to new nodes. Rewards include rewards for identifying disruptive technologies, exploration rewards, and path efficiency rewards. During the agent's exploration process, the knowledge graph is optimized and adjusted based on its exploration behavior.

2. The method for identifying disruptive technologies in an energy system according to claim 1, characterized in that, The edge types include semantic association edges, application scenario edges, technology evolution edges, and social attention edges. The semantic association edges represent the semantic association between technical keywords, the application scenario edges represent the synergistic / substitutional relationship between technical keywords in different application scenarios, the technology evolution edges represent the technological development / improvement relationship between technical keywords, and the social attention edges represent the degree of social attention between technical keywords.

3. The method for identifying disruptive technologies in an energy system according to claim 2, characterized in that, The edge weights are specifically set as follows: For semantically related edges, the co-occurrence frequency of technical keywords and the semantic strength of co-occurring statements are extracted from the text data, and the edge weight of the semantically related edge is determined based on the co-occurrence frequency and semantic strength. For application scenario edges, the frequency of common occurrence of technical keywords in different application scenarios is calculated to determine the edge weight of the application scenario edge. For the technology evolution edge, the edge weight is calculated based on patent citation frequency, patent citation timeliness, document co-occurrence frequency, and knowledge inheritance relationship. For the social attention edge, the edge weight is determined based on the co-occurrence frequency of technical keywords in social media, user interaction, and user sentiment, as well as the co-occurrence frequency, reporting influence, and reporting depth of technical keywords in news reports.

4. The method for identifying disruptive technologies in an energy system according to claim 1, characterized in that, The node attributes include technical description, number of patents, technical breakthroughs, market size, cost-effectiveness, reliability indicators, and environmental indicators.

5. The method for identifying disruptive technologies in an energy system according to claim 1, characterized in that, The states of the intelligent agent include: Current location node information: including the technical keywords and node attributes corresponding to the current node; Visited Node List: Records the sequence of nodes that the agent has visited; Local knowledge graph structure: contains the partial knowledge graph structure and edge weight information of the neighboring nodes of the current node.

6. The method for identifying disruptive technologies in an energy system according to claim 1, characterized in that, The rewards include: A reward for identifying disruptive technologies is obtained by weighted summation of multiple disruptive indicators and their weights, wherein the weights are determined based on the confidence level of the acquired disruptive indicators. The exploration reward is determined based on the exploration level of a region and the technological novelty of the accessed node. The knowledge graph is divided into different regions. The exploration level of each region is defined as the ratio of the number of accessed nodes in that region to the total number of nodes. The technological novelty index is defined as the ratio of the technological novelty evaluation value of the accessed node to the time when the agent accesses the node plus 1. The exploration reward is negatively correlated with the exploration level of the region and positively correlated with the technological novelty index. The path efficiency reward is set based on the path length from the starting node to the current node and a dynamic adjustment coefficient. The dynamic adjustment coefficient is determined based on the ratio of the value of the visited nodes to a preset value threshold. When the current node is the target node, the path efficiency reward is set to a positive preset value.

7. The method for identifying disruptive technologies in an energy system according to claim 6, characterized in that, The disruptive indicators include technological innovation indicators, technology market potential indicators, technology performance indicators, and social impact indicators. The technological innovation indicators include the patent number growth rate and technology breakthrough score. The technology market potential indicators include the market growth rate and investment enthusiasm index. The technology performance indicators include the performance improvement rate and cost reduction rate. The social impact indicators include the rate of increase in public attention and the intensity of policy support.

8. The method for identifying disruptive technologies in an energy system according to claim 1, characterized in that, The optimization and adjustment of the knowledge graph based on the agent's exploration behavior during the agent's exploration process specifically includes: Edge weight adjustment based on agent access frequency: During the agent's exploration process, the number of times the agent visits a node through each edge is recorded to obtain the edge access frequency; the original edge weight is adjusted based on the edge access frequency, and the higher the edge access frequency, the greater the edge weight after the jump. Edge weight adjustment based on agent feedback: The feedback value of each edge is calculated based on the agent's feedback. The feedback value of 1 indicates positive feedback and 0 indicates negative feedback. The original edge weights are adjusted based on the feedback value.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • New energy technology development trend atlas analysis and prediction method

    CN118261362A

Cited By

  • Multi-agent task routing method and device based on causal atlas and related products

    CN121900891A

  • Multi-agent task routing methods, devices, and related products based on causal graphs

    CN121900891B