Technical permeability-containing energy field knowledge graph construction method and system
By constructing an energy domain knowledge graph based on semantic clustering and a large language model, the problems of automated analysis and dynamic prediction of multi-source data were solved, the quantitative assessment of technology penetration rate was realized, and the scientific nature and accuracy of decision-making in the energy sector were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUONENG ECONOMIC & TECH RES INST CO LTD
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies struggle to automate and scale up the analysis of multi-source data in the energy sector, lack dynamic forecasting capabilities, and are unable to effectively identify key inflection points in technology penetration, resulting in highly subjective decision-making and low reproducibility.
A standard technology node system is constructed by combining semantic clustering and large language model. The dynamic penetration rate of technology nodes is calculated by the Logistic main trend fitting algorithm, and a knowledge graph of the energy field containing technology penetration rate is generated by combining sliding window correction.
It enables quantitative evaluation of multi-source data, improves the foresight and objectivity of trend analysis, accurately identifies the initiation and saturation inflection points of technologies, and provides scientific data support.
Smart Images

Figure CN121980038A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of knowledge graph technology in the energy sector, and in particular to a method and system for constructing an energy sector knowledge graph that incorporates technology penetration rate. Background Technology
[0002] Driven by the dual-carbon strategy and high-quality development, energy technologies are exhibiting characteristics of multi-source concurrent development, shortened cycles, and cross-sectoral integration. Faced with the ever-increasing volume of multi-source data, including academic papers, patents, projects, and standards, how to systematically analyze technological trends and identify critical development periods has become a significant challenge for decision-making in the energy sector.
[0003] While current knowledge graph-based technology analysis can achieve knowledge association, it struggles to dynamically quantify the technology lifecycle and predict its evolution. Existing methods suffer from the following main problems: First, they rely excessively on expert experience, making automated and large-scale analysis difficult, and are highly subjective with low reproducibility; second, they lack a unified quantitative mechanism for integrating multi-source data, making it impossible to comprehensively assess the overall development level of technology; and third, they lack dynamic prediction capabilities based on time series, making it difficult to identify key inflection points in technology penetration and effectively support forward-looking decision-making.
[0004] Therefore, it is necessary to construct a knowledge graph method that can integrate multi-source data, quantify technology penetration rates, and achieve dynamic prediction in order to improve the objectivity, comprehensiveness, and foresight of the analysis. Summary of the Invention
[0005] This invention provides a method and system for constructing an energy field knowledge graph that includes technology penetration rate, aiming to solve the above-mentioned problems.
[0006] According to an embodiment of the present invention, a method for constructing an energy domain knowledge graph including technology penetration rate is provided, comprising: S1. Acquire multi-source unstructured text data in the energy field, construct a standard technology node system by combining semantic clustering and large language model naming, and calculate the time series fusion score of each technology node. S2. Based on the time series fusion score, a preset penetration rate calculation model is used, combined with the Logistic main trend fitting algorithm, to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage, thereby generating an energy field knowledge graph containing technology penetration rates.
[0007] According to an embodiment of the present invention, an energy domain knowledge graph construction system including technology penetration rate is provided, comprising: The technology node construction and quantification module is used to acquire multi-source unstructured text data in the energy field. It uses a combination of semantic clustering and large language model naming to construct a standard technology node system and calculates the time series fusion score of each technology node. The penetration rate calculation and trend prediction module is used to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage based on the time series fusion score, using a preset penetration rate calculation model and combined with the Logistic main trend fitting algorithm, thereby generating an energy field knowledge graph containing technology penetration rates.
[0008] The embodiments of the present invention have the following beneficial effects: This invention addresses the weaknesses of traditional energy technology assessment, such as strong subjectivity and the use of single indicators, by innovatively proposing a quantitative evaluation system based on multi-source data fusion. By introducing predicted peak values as a normalization benchmark, it effectively avoids premature peaking errors caused by relying solely on historical data, thus enhancing the foresight of trend analysis. Simultaneously, through a combination of dual benchmarks and industry averages for normalization, it examines both the vertical growth of the technology and its horizontal competitiveness within the industry. Furthermore, by incorporating Logistic fitting and sliding window correction, it can accurately identify the inflection points and saturation points of technologies, providing scientific and objective data support for strategic decision-making, investment allocation, and capacity layout in the energy sector. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart of a method for constructing an energy field knowledge graph including technology penetration rate, according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating a specific implementation of the energy domain knowledge graph construction method including technology penetration rate, according to an embodiment of the present invention. Figure 3 This is a detailed flowchart illustrating the construction of the technical node system and the calculation of the fusion score in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the principle of the penetration rate calculation model and trend prediction framework in an embodiment of the present invention. Figure 5 This is a structural diagram of an energy field knowledge graph construction system that includes technology penetration rate, according to an embodiment of the present invention. Detailed Implementation
[0011] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0012] Method Implementation Examples According to embodiments of the present invention, a method for constructing an energy domain knowledge graph including technology penetration rates is provided. Figure 1 This is a flowchart of a method for constructing an energy domain knowledge graph including technology penetration rates, according to an embodiment of the present invention. Figure 1 As shown, the method for constructing an energy domain knowledge graph including technology penetration rate in this embodiment of the invention specifically includes: S1. Acquire multi-source unstructured text data in the energy field, construct a standard technology node system by combining semantic clustering and large language model naming, and calculate the time series fusion score of each technology node. Data sources for multi-source unstructured text data include: academic paper databases, patent databases, scientific research project databases, and industry standard databases.
[0013] Figure 2 This is a flowchart illustrating the specific implementation of the energy domain knowledge graph construction method including technology penetration rate, according to an embodiment of the present invention. Figure 3 The following is a detailed flowchart illustrating the construction of the technology node system and the calculation of the fusion score in an embodiment of the present invention. First, the cleaned multi-source text data is semantically vectorized. Then, clustering algorithms are used to group results describing the same technical topic (e.g., "proton exchange membrane electrolysis of water" and "PEM hydrogen production") into the same cluster. Subsequently, a large language model (LLM) is called to standardize the naming of the cluster, generating a unique technology node name (e.g., "PEM electrolyzer"), and attaching related cross-source result data.
[0014] Then calculate the time of each technology node. Fusion score As a representative metric for measuring the popularity of a technology, the fusion score... Obtain it using the following formula: ; in, Indicates time, Indicates the data source type. This indicates the weight of the corresponding data source. express The number of results of this type at any given time.
[0015] If the four types of results are summed according to preset weights, it can be expressed by the following formula: ; in, express The number of such achievements per year This represents the weighting coefficient. In this embodiment, the weighting parameters are configured as follows: papers 0.3, patents 0.3, projects 0.2, and standards 0.2.
[0016] Furthermore, this embodiment of the invention also includes an outlier detection and smoothing mechanism. To eliminate the interference of data noise on trend judgment, this embodiment sets thresholds: an annual growth rate exceeding 200% is considered a sudden increase, and a decline rate exceeding 80% is considered a sharp decrease. For detected outliers, a three-period moving average method is used to replace the window mean, thereby improving the stability of the sequence.
[0017] The above steps transform unstructured data into a computable time series. For example, for the technology node "PEM electrolyzer," the system calculates its fusion score for 2024 as follows: .
[0018] S2. Based on the time series fusion score, a preset penetration rate calculation model is used, combined with the Logistic main trend fitting algorithm, to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage, thereby generating an energy domain knowledge graph containing technology penetration rates. Figure 4 This is a schematic diagram illustrating the principle of the penetration rate calculation model and trend prediction framework according to an embodiment of the present invention. Figure 4 It can be seen that S2 specifically includes: In this embodiment of the invention, to balance the historical performance and future potential of the technology, a dual-benchmark concept is introduced. A time series model (such as ARIMA) is used to analyze historical data. The sequence is modeled to predict the score for the next 3-5 periods, and the maximum value is taken as the predicted peak. The final normalized denominator Take historical peak The formula is as follows, comparing the larger value among the predicted peak values: ; At the same time, calculate the industry average of technology nodes in the same field. , as a horizontal comparison benchmark.
[0019] Next, the standardized penetration rate is calculated using a combined normalization algorithm. The algorithm uses weighting coefficients. The dynamic balance between its own growth potential and industry competitiveness is calculated using the following formula: ; in, Dynamic adjustments should be made based on the stage of the technology's life cycle: for emerging technologies... Value For growth-stage technologies, Value For mature technologies, Value .
[0020] Finally, in order to predict the long-term evolution trend of technology, the model in A Logistic S-shaped curve is fitted onto the sequence. This curve describes the nonlinear growth process of technology from its initial stage to saturation, and the formula is defined as follows: ; in This is the saturation value. For growth rate, This is the midpoint of a potential outbreak.
[0021] Based on the calculated permeability values, key inflection points are automatically identified and marked: when When the threshold first exceeds 30%, it is marked as the tipping point (entering the verification period); when it exceeds 60%, it is marked as the outbreak point (entering the expansion period); when it exceeds 80%, it is marked as the saturation point (entering the mainstream period).
[0022] Based on the aforementioned calculations of the dynamic penetration rate, development trend curve, and key inflection point annotations for each technology node, a knowledge graph is constructed using a graph database. Specifically, the standard technology nodes are used as entity nodes in the graph, with inter-technology relationships (such as upstream and downstream, substitution, and complementarity) as edges. The dynamic penetration rate sequence, S-shaped fitting curve parameters, and identified development stage information such as initiation points, outbreak points, and saturation points corresponding to each technology node are attached as attribute data to the corresponding entity nodes. This generates an energy domain knowledge graph that not only includes technology entities and static relationships but also deeply integrates dynamic quantitative indicators and trend prediction information of technology development.
[0023] To further illustrate the calculation process, let's take the "Hydrogen Energy: PEM Electrolyzer" technology node as an example: Assuming the current year is 2024, what is the fusion score of this node? The historical peak value is 70, while the industry average is 50. The predicted future peak value is 60.
[0024] Since the historical peak (70) is greater than the predicted peak (60), the final benchmark is... .
[0025] Assuming the technology is in its growth stage, weights are assigned. Substitute into the formula to calculate: ; Based on this, the system determines that the technology is in the expansion phase (between 60% and 80%) and predicts that it is in a rapid growth phase.
[0026] After completing the above calculations, the "PEM electrolyzer" is stored as an entity node in the graph database. Node attributes include its fusion score time series, the calculated penetration rate of 67.43%, its development stage "expansion phase," and growth parameters obtained from Logistic regression. Simultaneously, connections are established between this node and nodes in the hydrogen energy technology field, as well as other hydrogen production technology nodes such as alkaline electrolyzers and solid oxide electrolyzers, thus forming a knowledge graph with dynamic quantitative attributes. Through this graph, users can intuitively query and compare the development trends, penetration levels, and lifecycle stages of different technologies.
[0027] By employing the embodiments of the present invention, the following beneficial effects are achieved: This invention addresses the weaknesses of traditional energy technology assessment, such as strong subjectivity and the use of single indicators, by innovatively proposing a quantitative evaluation system based on multi-source data fusion. By introducing predicted peak values as a normalization benchmark, it effectively avoids premature peaking errors caused by relying solely on historical data, thus enhancing the foresight of trend analysis. Simultaneously, through a combination of dual benchmarks and industry averages for normalization, it examines both the vertical growth of the technology and its horizontal competitiveness within the industry. Furthermore, by incorporating Logistic fitting and sliding window correction, it can accurately identify the inflection points and saturation points of technologies, providing scientific and objective data support for strategic decision-making, investment allocation, and capacity layout in the energy sector.
[0028] System Implementation Examples According to embodiments of the present invention, an energy domain knowledge graph construction system incorporating technology penetration rates is provided. Figure 5 This is a structural diagram of an energy domain knowledge graph construction system including technology penetration rate, according to an embodiment of the present invention. Figure 5 As shown, the energy domain knowledge graph construction system including technology penetration rate in this embodiment of the invention includes: The technology node construction and quantification module is used to acquire multi-source unstructured text data in the energy field. It uses a combination of semantic clustering and large language model naming to construct a standard technology node system and calculates the time series fusion score of each technology node. The technical node construction and quantification module specifically includes: a data cleaning unit, a semantic clustering and naming unit, and a fusion score calculation unit; The data cleaning unit is used to preprocess the acquired multi-source unstructured text data, including deduplication, format standardization, key field extraction and cleaning, in order to eliminate noise and inconsistencies. The semantic clustering and naming unit is used to transform the cleaned text data into semantic vectors, and to use clustering algorithms to group semantically similar results into the same cluster. Then, a large language model is called to generate standardized technical node names for each cluster. The fusion score calculation unit is used to calculate the fusion score of each technical node at time t according to the preset data source weights and the following formula, thereby forming time series data.
[0029] ; in, Indicates time, Indicates the data source type. This indicates the weight of the corresponding data source. express The number of results of this type at any given time.
[0030] The penetration rate calculation and trend prediction module is used to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage based on the time series fusion score, using a preset penetration rate calculation model and combined with the Logistic main trend fitting algorithm, thereby generating an energy field knowledge graph containing technology penetration rates.
[0031] The penetration rate calculation and trend prediction module specifically includes: a dual-benchmark calculation unit, a combined normalization unit, and a Logistic trend fitting unit; The dual-benchmark calculation unit is used to calculate the industry average fusion score of technology nodes in the same field, and to predict the future peak fusion score of technology nodes using a time series prediction model. Combined with historical peak scores, the final normalized benchmark is determined. The combined normalization unit is used to calculate the standardized penetration rate P(t) of each technology node at time t using the following formula, based on the dynamically adjusted weighting coefficient α: ; The Logistic trend fitting unit is used to fit a Logistic S-shaped curve on the standardized penetration rate sequence P(t), and automatically identify the initiation point, outbreak point and saturation point of technological development according to preset stage thresholds (such as 30%, 60%, 80%).
[0032] This invention further includes: a graph generation and visualization module, connected to the penetration rate calculation and trend prediction module, used to perform structured storage of entities, relationships and attributes in a graph database based on the standard technology node system output by the technology node construction and quantification module, and the dynamic penetration rate, development trend curve and stage inflection point information of each technology node output by the penetration rate calculation and trend prediction module, and finally generate and output an energy field knowledge graph containing technology penetration rates, and provide a visual query and analysis interface.
[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing an energy domain knowledge graph that includes technology penetration rate, characterized in that... include: S1. Acquire multi-source unstructured text data in the energy field, construct a standard technology node system by combining semantic clustering and large language model naming, and calculate the time series fusion score of each technology node. S2. Based on the time series fusion score, a preset penetration rate calculation model is used, combined with the Logistic main trend fitting algorithm, to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage, thereby generating an energy field knowledge graph containing technology penetration rates.
2. The method according to claim 1, characterized in that, The specific steps for constructing the standard technology node system include: The multi-source unstructured text data is cleaned and its fields are standardized. Semantic clustering algorithms are used to group results describing the same technical topic into the same cluster; The clustered technical units are standardized and named using a large language model to form independent technical nodes.
3. The method according to claim 1, characterized in that, The fusion score Obtain it using the following formula: ; in, Indicates time, Indicates the data source type. This indicates the weight of the corresponding data source. express The number of results of this type at any given time.
4. The method according to claim 1, characterized in that, The preset permeability calculation model specifically includes: ; in, For penetration rate, For the current fusion score, This represents the industry average for technology nodes within the same field. The larger of the historical peak value and the predicted peak value. These are adjustable weighting coefficients.
5. The method according to claim 4, characterized in that, The larger of the historical peak value and the predicted peak value is obtained through the following method: ; in, This is the highest historical peak. The predicted peak value within a preset time period in the future.
6. The method according to claim 1, characterized in that, The Logistic main trend fitting algorithm specifically includes: In the standardized penetration rate sequence The formula for fitting an S-shaped curve is: ; in, This is the saturation value. For growth rate, This is the midpoint of a potential outbreak; When the actual penetration rate deviates from the fitted curve by more than a preset threshold, a sliding window averaging is initiated to smooth and correct the current time point and subsequent time points.
7. The method according to claim 4, characterized in that, The adjustable weighting coefficient is dynamically adjusted according to the lifecycle stage of the technology node.
8. The method according to claim 1, characterized in that, The identification of inflection points in the development stage specifically includes: comparing the calculated dynamic penetration rate with multiple preset stage thresholds to determine the development stage of the technology node.
9. The method according to claim 1, characterized in that, The data sources for the multi-source unstructured text data include: academic paper databases, patent databases, scientific research project databases, and industry standard databases.
10. A knowledge graph construction system for the energy field that includes technology penetration rate, characterized in that, include: The technology node construction and quantification module is used to acquire multi-source unstructured text data in the energy field. It uses a combination of semantic clustering and large language model naming to construct a standard technology node system and calculates the time series fusion score of each technology node. The penetration rate calculation and trend prediction module is used to calculate the dynamic penetration rate of technology nodes and identify inflection points in the development stage based on the time series fusion score, using a preset penetration rate calculation model and combined with the Logistic main trend fitting algorithm, thereby generating an energy field knowledge graph containing technology penetration rates.