Power technology innovation contribution degree evaluation method and system based on patent knowledge graph

By constructing a triple reliability scoring model and multi-dimensional indicators to screen high-quality data, and combining BP neural network to evaluate the innovation contribution of power patents, the quality problems of power patent data and the one-sidedness of evaluation are solved, and an accurate assessment of the innovation contribution of power technology is achieved.

CN120706520APending Publication Date: 2025-09-26ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510820216.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing technologies, entity linking errors and relationship mislabeling in power patent data lead to low-quality knowledge graph construction. Traditional methods are unable to comprehensively evaluate the contribution of power technology innovation and lack multi-dimensional comprehensive evaluation.

Method used

By constructing a triple reliability scoring model, combining local accuracy and local completeness indicators to screen high-quality data, combining technology diffusion intensity, network hubness and uniqueness indicators to conduct multi-dimensional evaluation, and using BP neural network to build a contribution evaluation model.

Benefits of technology

The construction quality of the power patent knowledge graph and the accuracy of the evaluation results have been significantly improved, the innovation contribution of power technology patents has been comprehensively quantified, and key patents and innovations have been identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706520A_ABST
    Figure CN120706520A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power technology innovation contribution degree evaluation method and system based on a patent knowledge graph, and relates to the technical field of electric power technology innovation contribution degree evaluation, and the method comprises the following steps: classifying electric power technology patent data based on a triple, and obtaining first triple data; obtaining local credibility data of the first triple data, constructing a triple reliability scoring model based on a BP neural network, and screening the first triple data to obtain second triple data; obtaining contribution degree evaluation features of the power patent knowledge graph constructed based on the second triple data; and constructing a contribution degree evaluation model based on a BP neural network according to the contribution degree evaluation features, sorting model outputs, and generating an innovation contribution degree evaluation table, thereby solving the problem of one-sided evaluation results of the innovation contribution degree of the power technology, and comprehensively quantifying the innovation contribution degree of the power technology patent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of evaluation system analysis, and more specifically, to a method and system for evaluating the contribution of electric power technology innovation based on patent knowledge graphs. Background Art

[0002] With the rapid development of power technology, the number of patents in the power sector has exploded worldwide. As an important carrier of technological innovation, the value assessment of patents is of great significance for technology investment decisions, R&D direction planning, and intellectual property management. However, traditional patent evaluation methods mainly rely on single indicators such as the number of citations and the size of patent families, which cannot fully reflect the technological innovation contribution of patents. Especially in the field of power technology, there are many technical branches (such as power generation, transmission, energy storage, smart grid, etc.), and the technical connections between each branch are complex. Traditional methods cannot effectively analyze the synergy and integration relationship between cross-domain technologies.

[0003] In recent years, knowledge graph technology, owing to its powerful association analysis capabilities, has gradually become a crucial tool for patent valuation. By constructing a knowledge graph for power patents, we can reveal the trajectory of technological development and the underlying relationships between patents, providing new insights for quantitatively assessing the contribution of technological innovation. However, existing technologies still face numerous challenges in practical application, such as low data quality, a single evaluation dimension, and insufficient algorithm adaptability, which hinder the accuracy and practicality of patent valuation.

[0004] For example, the invention patent announcement with announcement number CN117972113B discloses a method and system for predicting and evaluating patent authorization based on an attribute knowledge graph, including the following steps: performing patent data retrieval from a patent knowledge graph based on test patent text data, building a second knowledge graph based on the retrieved knowledge data, performing similarity analysis based on the test triple data and the retrieved knowledge data to obtain a first similarity, replacing the original triple data with the test triple data in the second knowledge graph to form a third knowledge graph, calculating the structural differences between the second and third knowledge graphs to obtain data relevance, and predicting and evaluating authorization based on the first similarity and the data relevance. Through the present invention, it is possible to perform similarity analysis of prior art based on the knowledge graph using technical entities, attributes, and relationship triples, and based on knowledge mining, analyze the relevance and interchangeability of the test patent with prior art, thereby achieving a scientific and accurate prediction of authorization results.

[0005] For example, the invention patent announcement with announcement number CN113901238A discloses a method and system for constructing a knowledge graph of city physical examination indicators, which includes the following steps: performing a first fusion of knowledge triples to obtain an indicator entity set, an indicator category entity set, and an indicator-belonging indicator category relationship set; performing a second fusion of the indicator category entity set to obtain a fused indicator category entity set; establishing an association relationship between each indicator entity in the indicator entity set; and constructing a city physical examination indicator knowledge graph through the indicator entity set, the fused indicator category entity set, the association relationship set between indicator entities, and the indicator-belonging indicator category relationship set. The present invention stores city physical examination indicators through a graph structure, thereby improving the efficiency of city physical examination indicator retrieval, facilitating indicator recommendation, and contributing to the development of city physical examination work; by simplifying the set of associated indicator pairs, redundant relationships between indicator entities are removed, greatly improving the efficiency of graph database relationship search.

[0006] The above disclosed technical solutions have at least the following technical problems:

[0007] There are problems such as entity linking errors and relationship mislabeling in the original patent data, which leads to low reliability of triple data, directly affecting the construction quality of the knowledge graph and the accuracy of subsequent analysis results. The lack of a systematic data quality assessment mechanism makes it difficult to effectively screen and repair low-quality triple data.

[0008] Traditional methods rely on a single indicator (such as the number of citations) and are unable to comprehensively evaluate the technological innovation contribution of patents from multiple dimensions such as technical relevance and knowledge diffusion paths. Existing technologies have not fully explored the network topology characteristics in the power patent knowledge graph, resulting in one-sided evaluation results.

[0009] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0010] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a method and system for evaluating the innovation contribution of electric power technology based on a patent knowledge graph. By constructing an electric power patent knowledge graph through triples, a multi-dimensional comprehensive evaluation of the innovation contribution of patent technology is performed to solve the problem of one-sided evaluation results of the innovation contribution of electric power technology.

[0011] To achieve the above object, the present invention provides the following technical solutions:

[0012] The invention discloses a method for evaluating the innovation contribution of electric power technology based on patent knowledge graph, comprising the following steps: classifying electric power technology patent data based on triples to obtain first triple data; obtaining local credibility data of the first triple data, constructing a triple reliability scoring model based on a BP neural network, screening the first triple data to obtain second triple data; obtaining contribution evaluation features of the electric power patent knowledge graph constructed based on the second triple data; constructing a contribution evaluation model based on a BP neural network according to the contribution evaluation features, and sorting the model output to generate an innovation contribution evaluation table.

[0013] In a preferred embodiment, the electric power technology patent data is classified based on triples to obtain first triple data, specifically: the first triple data includes several groups of triples, each triple has key fields of entity name, technology category and relationship; based on natural language processing technology, key entities and relationships between key entities in the electric power technology patent data are identified; based on the key entities and the relationships between key entities, several triples are combined to form structured data, namely the first triple data.

[0014] In a preferred embodiment, the local credibility data includes a local accuracy index; a reference array for each key field is preset; an actual extraction array of key fields for each group of triples is obtained; the reference array and the actual extraction array are subjected to similarity analysis to obtain the error of each key field; according to each key field error, the accuracy score of each key field is calculated based on a preset key field accuracy score model; and the accuracy scores of all key fields of each triple are averaged as the local accuracy index of each triple.

[0015] In a preferred embodiment, the local credibility data includes a local integrity indicator; a non-empty key field set of each triple is obtained based on the first triple data; a required key field set of each triple is preset; and the ratio of the non-empty key field set of the triple to the required field set of the triple is used as the local integrity ratio of each triple, as a local integrity indicator.

[0016] In a preferred embodiment, the contribution evaluation feature includes a technology diffusion intensity feature; each patent node and the relationship between each patent node are obtained based on the electric power patent knowledge graph, and each patent node is used as a row and column to construct an adjacency matrix, and the relationship between each patent node is used as the value of the corresponding element in the adjacency matrix; based on the adjacency matrix, the degree centrality value of the patent node is obtained by summing the values ​​of the corresponding elements of each patent node; normalization is performed based on the degree centrality values ​​of all patent nodes to finally obtain the technology diffusion intensity feature.

[0017] In a preferred embodiment, the contribution evaluation feature includes a technology network hub index; based on the electric power patent knowledge graph, each patent node and the relationship between each patent node are obtained, and an adjacency matrix is ​​constructed with the citing patent nodes as rows and the cited patent nodes as columns; based on the adjacency matrix analysis, the structural information of the patent citation relationship is obtained, and based on the pagearank method, the structural information of the patent citation relationship and the preset initial pagearank value of each patent node are analyzed to obtain the pagearank value; the pagearank value of each patent node is iteratively processed to obtain the pagearank value change before and after each pagearank value iteration; the pagearank value change is compared with the preset pagearank method convergence threshold, and when the pagearank value change is less than the preset pagearank method convergence threshold, the iteration is stopped to obtain the final pagearank value; the final pagearank values ​​of all patent nodes are sorted and normalized to obtain the technology network hub index.

[0018] In a preferred embodiment, the contribution evaluation features include a technical uniqueness score; based on the electric power patent knowledge graph, text vocabulary data is extracted from patent nodes and the relationship edges between nodes to form a patent corpus; based on the patent corpus, all unique words are screened out, and a fixed index position of each unique word is preset to form a vocabulary; based on all unique words in the vocabulary and in combination with the TF-IDF method, the TF-IDF value of each unique word is obtained; according to the TF-IDF value of each unique word and the fixed index position of each unique word, a patent node feature vector is constructed; according to the patent node feature vector and in combination with the cosine similarity method, patent node similarity data is obtained; based on the patent node feature vector and in combination with the cosine similarity method, normalization processing is performed on the patent node similarity data to obtain a technical uniqueness score.

[0019] In a preferred embodiment, a contribution evaluation model is constructed based on a BP neural network according to the contribution evaluation characteristics; the innovation contribution of each electric power technology patent is obtained through the contribution evaluation model; and the innovation contribution of the electric power technology patents is sorted in ascending order to generate an innovation contribution evaluation table.

[0020] The electric power technology innovation contribution evaluation system based on the patent knowledge graph includes a data acquisition module, a data screening module, a graph construction module and a contribution analysis module; the data acquisition module is used to classify the electric power technology patent data based on triples to obtain the first triple data; the data screening module is used to obtain the local credibility data of the first triple data, build a triple reliability scoring model based on the BP neural network, screen the first triple data, and obtain the second triple data; the graph construction module is used to obtain the contribution evaluation features of the electric power patent knowledge graph built based on the second triple data; the contribution analysis module is used to build a contribution evaluation model based on the BP neural network according to the contribution evaluation features, and sort the model output to generate an innovation contribution evaluation table.

[0021] The technical effects and advantages of the electric power technology innovation contribution evaluation method and system based on patent knowledge graph of the present invention are as follows:

[0022] 1. This paper constructs a triple reliability scoring model, combines local accuracy indicators and local completeness indicators, and systematically screens high-quality triple data, effectively solving problems such as entity link errors and relationship mislabeling in the original data, significantly improving the construction quality of the electric power patent knowledge graph and the accuracy of subsequent evaluation results.

[0023] 2. This invention breaks through the limitations of traditional single indicators by integrating a multi-dimensional comprehensive evaluation model that combines technology diffusion intensity characteristics, coreness index, and novelty index, deeply explores the network topology characteristics and semantic associations of the knowledge graph, and comprehensively quantifies the innovation contribution of power technology patents. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flow chart of the method for evaluating the contribution of electric power technology innovation based on patent knowledge graphs of the present invention.

[0025] Figure 2 This is a structural diagram of the electric power technology innovation contribution evaluation system based on the patent knowledge graph of the present invention. DETAILED DESCRIPTION

[0026] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0027] Example 1, Figure 1 The present invention provides a method for evaluating the contribution of electric power technology innovation based on patent knowledge graph, which includes the following steps:

[0028] S1, classifying the electric power technology patent data based on triples to obtain the first triplet data;

[0029] In this embodiment, the key entities and the relationships between key entities in the electric power technology patent data are identified based on natural language processing technology;

[0030] Based on the key entities and the relationships between the key entities, several triples are combined to form structured data, namely the first triple data.

[0031] It should be noted that the first triple data includes several groups of triples, each of which has key fields of entity name, technology category and relationship;

[0032] It should be noted that natural language processing (NLP) technology is an existing technology used to identify and extract key entities and relationships between entities in text;

[0033] It should be noted that power technology patent data refers to patent data related to power technology obtained through national patent databases, patent search tools, and patent data providers, including patent texts, applicants, inventors, patent classification numbers, and application dates;

[0034] S2, obtaining local credibility data of the first triplet data, building a triplet reliability scoring model based on a BP neural network, screening the first triplet data, and obtaining second triplet data;

[0035] Local credibility data includes local accuracy index and local completeness index;

[0036] The local accuracy index is used to measure the accuracy of key fields in triples and ensure the reliability of triple data, thereby providing a high-quality data foundation for the subsequent construction and analysis of the electric power patent knowledge graph.

[0037] The local integrity index is used to measure whether the triple contains all the required key field information, ensuring the integrity of the triple data, and thus providing a complete data foundation for the subsequent construction and analysis of the electric power patent knowledge graph.

[0038] The triplet reliability scoring model, constructed from local accuracy and completeness metrics, comprehensively assesses the accuracy and completeness of triplet data, screens for high-quality data, and improves the reliability of subsequent analysis and the construction of the power patent knowledge graph. By quantifying data quality and supporting automated data cleansing and repair, the model offers flexibility and scalability, optimizes data processing efficiency, and provides a reliable data foundation for technical decision-making.

[0039] In this embodiment, the local accuracy index and the local integrity index are obtained respectively, and a triple reliability scoring model is obtained through comprehensive weighted processing;

[0040] Obtain a comprehensive data quality score based on the triple reliability scoring model analysis;

[0041] Based on the comparison between the comprehensive data quality score and the preset triple data quality assessment threshold, traverse each triple in the first triple data, retain each triple with a comprehensive data quality score greater than or equal to the triple data quality assessment threshold, and form the second triple data;

[0042] It should be noted that the comprehensive data quality score is a quantitative score obtained after the quality of the triple data is evaluated through the triple reliability scoring model; it combines the local accuracy index and the local completeness index to measure the accuracy and completeness of the triple data.

[0043] In this embodiment, the specific method for obtaining the local accuracy indicator is as follows:

[0044] Preset reference array for each key field;

[0045] Get the actual extracted array of key fields for each triple;

[0046] Perform similarity analysis on the reference array and the actual extracted array to obtain the error of each key field;

[0047] According to the error of each key field, the accuracy score of each key field is calculated based on the preset key field accuracy score model;

[0048] The accuracy scores of all key fields of each triple are averaged as the local accuracy indicator of each triple.

[0049] It should be noted that the actual extracted array is the actual value of the key field extracted from the first triple data; for example, for the triple (Patent A, Technology Category, Power Generation Technology), the actual extracted array is the value of the "Technology Category" field as "Power Generation Technology";

[0050] It should be noted that the preset reference array is a standard value pre-set based on domain knowledge, historical data, or expert experience. For example, for the triple (Patent A, Technology Category, Power Generation Technology), based on knowledge in the power field, the preset technology category reference array can be ["Power Generation Technology", "Power Transmission Technology", "Energy Storage Technology"].

[0051] It should be noted that the following is an overview of the calculation:

[0052] The specific calculation formula for key field error is:

[0053] e x =||x t -x ref ||

[0054] The specific calculation formula for calculating the accuracy score of each key field based on the preset key field accuracy score model is:

[0055] A x (t) = exp(-λe x )

[0056] The specific calculation formula of the local accuracy index is:

[0057]

[0058] In the formula, x is the key field in the triple, e x is the key field error, x t The actual extraction array of the triple key field, x ref A is a reference array of preset key fields. x (t) is the accuracy score of the key field, t is the triplet label, λ is the decay feature, exp() is the exponential function, A(t) is the local accuracy index, and n is the number of key fields in a single triplet;

[0059] It should be noted that λ is the attenuation characteristic, λ>0, which is used to control the sensitivity of the accuracy score to the error and is set based on experience;

[0060] It should be noted that for categorical data (such as technology categories), e x The error should be 0 or 1 (match or mismatch); for numerical data (such as application date), e x The error should be the absolute difference between the actual value and the reference value.

[0061] In this embodiment, the specific method for obtaining the local integrity indicator is as follows:

[0062] Obtain a non-empty key field set for each triple based on the first triple data;

[0063] Preset the required key field set for each triple;

[0064] The ratio of the triple's non-empty key field set to the triple's required field set is taken as the local completeness ratio of each triple and used as the local completeness indicator.

[0065] It should be noted that the triplet non-empty key field set is formed by traversing all triples in the first triplet data and extracting all non-empty key fields;

[0066] It should be noted that the set of required key fields for triples is a predefined set of key fields based on the knowledge requirements and experience in the field of power technology. For example, for power patent data, the required field set may include "patent number", "inventor", "technology category", "application date", etc.

[0067] It should be noted that the specific calculation process involved in the above steps can be implemented with the help of existing mature technologies, so it will not be repeated here; the following is an overview of the calculation:

[0068] The specific calculation formula of the local integrity index is:

[0069]

[0070] Where C(t) is the local integrity index, F req is the set of required key fields for triples, F obt is the set of non-empty key fields of triples, and t is the triple label;

[0071] In this embodiment, the specific method for obtaining the triple reliability scoring model is as follows:

[0072] It should be noted that the specific calculation process involved in the above steps can be implemented with the help of existing mature technologies, so it will not be repeated here; the following is an overview of the calculation:

[0073] The specific calculation formula of the triple reliability scoring model is:

[0074] R(t)=ω A *A(t)+ω C *C(t)

[0075] Where R(t) is the output of the triple reliability scoring model, i.e., the comprehensive score of data quality, A(t) is the local accuracy index of triple t, C(t) is the local integrity index of triple t, ω A is the local accuracy indicator weight feature of triplet t, ω C is the weighted feature of the local integrity index of triplet t;

[0076] It should be noted that ω A and ω C Determined based on actual needs or through expert experience, A +ω C =1;

[0077] It should be noted that the output result of the triple reliability scoring model, that is, the comprehensive data quality score, is that the triples that are greater than or equal to the triple data quality assessment threshold are usable triples and are retained as part of the second triple data; the triples that are less than the triple data quality assessment threshold are invalid triples and are discarded.

[0078] Furthermore, the triple reliability scoring model has the following advantages for constructing the electric power patent knowledge graph:

[0079] High-quality triplets selected based on reliability scores can be used to calculate patent influence and technological innovation, improving the accuracy of patent evaluation results;

[0080] Filter low-quality triplets by comprehensively evaluating the local accuracy and local completeness of the triplets, screening out erroneous or noisy data, and ensuring that the relationships in the graph are authentic and reliable;

[0081] By reducing redundant information, the scoring model can detect repeated or contradictory triples and prevent invalid data from interfering with the construction and analysis of the knowledge graph.

[0082] S3, obtaining contribution evaluation features of the electric power patent knowledge graph constructed based on the second triple data;

[0083] Contribution evaluation characteristics include technology diffusion intensity characteristics, technology network hubness indicators, and technology uniqueness scores;

[0084] The technology diffusion intensity characteristic is an indicator used to measure the number of direct connections of a patent in the technology network, reflecting the influence of the patent in technology dissemination and innovation. It is a more intuitive indicator derived from the degree centrality value and is closer to the evaluation of the innovation contribution of power patents.

[0085] The Technology Network Hubness Index is an indicator used to measure the core position and importance of a patent in a technology network. It is based on the PageRank algorithm (webpage importance score) and reflects the influence and importance of a patent in the entire technology network by analyzing the citation relationships between patents.

[0086] The technical uniqueness score is an indicator used to measure the degree of difference between a patent and existing technologies. It reflects the innovativeness of the patent and evaluates its novelty by calculating the similarity between the patent and existing technologies.

[0087] In this embodiment, all entities in each triple in the second triple data are set as nodes, and the relationships between all entities are set as edges to construct an electric power patent knowledge graph. Based on the relationships between nodes and edges in the electric power patent knowledge graph, the technology diffusion intensity characteristics, technology network hub index, and technology uniqueness score, namely the second data, are obtained respectively;

[0088] It should be noted that the entities of the second triplet data include patent number, invention name, technical description, key technology, and citation;

[0089] It should be noted that the nodes of the electric power patent knowledge graph include patent nodes;

[0090] In this embodiment, the specific method for obtaining the technology diffusion intensity characteristics is as follows:

[0091] Based on the electric power patent knowledge graph, each patent node and the relationship between each patent node are obtained, and each patent node is used as a row and column to construct an adjacency matrix, and the relationship between each patent node is used as the value of the corresponding element in the adjacency matrix;

[0092] Based on the adjacency matrix, the degree centrality value of the patent node is obtained by summing the values ​​of the corresponding elements of each patent node;

[0093] Normalization is performed based on the degree centrality values ​​of all patent nodes to finally obtain the technology diffusion intensity characteristics.

[0094] It should be noted that the patents that play a key role in technology diffusion and innovation are judged by the technology diffusion intensity characteristics. The following is an overview of the calculation:

[0095] Sum the values ​​of the corresponding elements of each patent node to obtain the degree centrality value of the patent node. The specific calculation formula is:

[0096]

[0097] Based on the normalization of the degree centrality values ​​of all patent nodes, the specific calculation formula for the technology diffusion intensity characteristics is finally obtained:

[0098]

[0099] Where, DC i is the degree centrality value of patent i, B ij is an adjacency matrix element, representing the direct relationship between patent i and patent j, i is a patent node in the electric power patent knowledge graph, j is a patent node in the electric power patent knowledge graph, B is an adjacency matrix, representing the connection relationship between nodes in the network, DC′ i is the technology diffusion intensity feature, max(DC) is the degree centrality value of the patent with the largest degree centrality in the entire electric power patent knowledge graph, and min(DC) is the degree centrality value of the patent with the smallest degree centrality in the entire electric power patent knowledge graph.

[0100] It should be noted that if B ij =1, it means that there is a direct connection between i and j, otherwise B ij =0, which means there is no direct connection between i and j;

[0101] Furthermore, analyzing the characteristics of technology diffusion intensity has the following advantages for evaluating the contribution of power technology innovation:

[0102] Reflecting the influence of patents, the technology diffusion intensity characteristic measures the number of direct connections between patent nodes, that is, the number of other patents or technologies directly associated with the patent; patents with high technology diffusion intensity characteristics usually play an important role in technology dissemination and influence, and may represent a higher innovation contribution.

[0103] Identify key patents. By calculating the technology diffusion intensity characteristics, we can identify patents that are at the core of the power patent knowledge graph; these patents may have a significant driving effect on subsequent technological development and reflect their innovativeness.

[0104] In this embodiment, the specific method for obtaining the technical network hub index is as follows:

[0105] Based on the electric power patent knowledge graph, each patent node and its relationship is obtained, and an adjacency matrix is ​​constructed with the citing patent nodes as rows and the cited patent nodes as columns.

[0106] Obtaining structural information of patent citation relationships based on the adjacency matrix analysis, and analyzing the structural information of the patent citation relationships and the preset initial pagearank value of each patent node based on the pagearank method to obtain a pagearank value;

[0107] Iterate the pagearank value of each patent node to obtain the pagearank value change before and after each pagearank value iteration;

[0108] Compare the pagearank value change with the preset pagearank method convergence threshold, and when the pagearank value change is less than the preset pagearank method convergence threshold, stop iteration and obtain the final pagearank value;

[0109] The final pagearank values ​​of all patent nodes are sorted and normalized to obtain the technology network hub index.

[0110] It should be noted that the technological network hubness index is used to measure the technological influence of patents and evaluate the innovation contribution of power patents. The following is an overview of the calculation:

[0111] The specific calculation formula for obtaining the pagerank value is:

[0112]

[0113] The specific calculation formula for comparing the change in pagearank value with the preset pagearank method convergence threshold is:

[0114] max i |PR i (m) -PR i (m-1) |<θ

[0115] The specific calculation formula for obtaining the technology network hub index is:

[0116]

[0117] Where PR i is the pagearank value of patent node i, d is the damping feature, which is obtained based on empirical presets, N is the total number of nodes in the power patent knowledge graph, In(i) is the set of all patent nodes pointing to patent node i, L(j) is the out-degree of patent node j, that is, the number of patents cited by patent node j, PR i (m) is the pagerank value of patent node i in the mth iteration, θ is the preset convergence threshold, such as 0.0001, |PR i (m) -PR i (m-1) | is the difference between the pagearank value of patent node i in the mth iteration and the pagearank value of patent node i in the m-1th iteration, that is, the change in pagearank value before and after the iteration, PR′ i is the technology network hub index, min(PR) is the lowest patent node pagearank value in the entire electric power patent knowledge graph, and max(PR) is the highest patent node pagearank value in the entire electric power patent knowledge graph;

[0118] Furthermore, analyzing the technology network hub index has the following advantages for evaluating the contribution of power technology innovation:

[0119] Considering the structure of the electric power patent knowledge graph, the technical network hubness index reflects the importance of patents in the technical electric power patent knowledge graph by analyzing the citation relationship between patents.

[0120] Identify core technologies. Patents with higher technology network hubness indicators are usually located at the core of the technology power patent knowledge graph, and may represent key technologies or innovations with high influence.

[0121] Quantifying influence: By calculating the technology network hub index, we can quantify the impact of each patent on other patents, which helps to identify patents that have an important role in promoting technological development.

[0122] In this embodiment, the specific method for obtaining the technical uniqueness score is as follows:

[0123] Based on the electric power patent knowledge graph, text vocabulary data is extracted from patent nodes and the relationship edges between nodes to form a patent corpus;

[0124] Filter out all unique words based on the patent corpus and preset a fixed index position for each unique word to form a vocabulary;

[0125] Based on all unique words in the vocabulary and combined with the TF-IDF method, the TF-IDF value of each unique word is obtained;

[0126] Constructing a patent node feature vector based on the TF-IDF value of each unique word and the fixed index position of each unique word;

[0127] Obtain patent node similarity data based on patent node feature vectors and combined with the cosine similarity method;

[0128] Normalization is performed based on the similarity data of patent nodes to obtain a technology uniqueness score.

[0129] It should be noted that the overview of the calculations in the above steps is as follows:

[0130] The specific calculation formula of TF-IDF value is:

[0131] TF-IDF(t,d,D)=TF(t,d)*IDF(t,D)

[0132] The specific calculation formula of TF is:

[0133]

[0134] The specific calculation formula of IDF is:

[0135]

[0136] The specific calculation formula for patent node similarity data is:

[0137]

[0138] The specific calculation formula for the technology uniqueness score is:

[0139]

[0140] Where TF is term frequency, IDF is inverse document frequency, t is a word in the patent, TF(t,d) is the frequency of word t in patent d, f(t,d) is the number of times word t appears in d, nd is the total number of words in d, IDF(t,D) is a measure of the importance of word t in the entire patent set d, |D| is the total number of patents, |{d∈D:t∈d}| is the number of patents containing word t, and TS is the product of the number of words in t. ij is a vector and vector Patent node similarity data, is a vector and vector The dot product of is a vector The module length, is a vector The modulus length, TS′ i Score the uniqueness of the technology, is the average patent node similarity data of patent i, max(TS) is the patent node similarity data with the highest technical similarity in the entire electric power patent knowledge graph, and min(TS) is the patent node similarity data with the lowest technical similarity in the entire patent network;

[0141] It should be noted that TF-IDF (Term Frequency - Inverse Document Frequency) is an existing technology used to calculate the TF-IDF value of a vocabulary;

[0142] It should be noted that the patent node feature vector includes multiple dimensions, each dimension corresponds to a unique word in the vocabulary, and the dimension value is the TF-IDF value obtained by the TF-IDF method;

[0143] It should be noted that the vector is the patent node feature vector of patent i, vector is the patent node feature vector of patent j;

[0144] It should be noted that the value range of the technology uniqueness score is [0,1]. The larger the value of the technology uniqueness score, the higher the similarity between patents, which usually indicates low innovation. The smaller the value of the technology uniqueness score, the lower the similarity between patents, which may indicate higher novelty and innovation.

[0145] Furthermore, analyzing the technology uniqueness score has the following advantages for evaluating the contribution of power technology innovation:

[0146] To quantify novelty, a low technical uniqueness score (close to 0) generally indicates that the patent's technical description is significantly different from the prior art, implying a high degree of novelty and a greater degree of innovation contribution; whereas a high technical uniqueness score (close to 1) may indicate that the technical content is repetitive and the innovation is low, which can objectively quantify the difference between each patent and the prior art.

[0147] By scoring the technological uniqueness, we can quickly locate those innovations that are significantly different from existing technologies in large-scale patent data.

[0148] S4, construct a contribution evaluation model based on BP neural network according to the contribution evaluation characteristics, sort the model output, and generate an innovation contribution evaluation table.

[0149] In this embodiment, the specific method for obtaining the ranking results of the power technology innovation contribution evaluation is as follows:

[0150] Constructing a contribution evaluation model based on a BP neural network according to the contribution evaluation features;

[0151] Obtain the innovation contribution of each electric power technology patent through the contribution evaluation model;

[0152] The innovation contributions of power technology patents are sorted in ascending order to generate an innovation contribution evaluation table.

[0153] It should be noted that in the above steps, the following is an overview of the calculation:

[0154] The specific calculation formula of the contribution evaluation model is:

[0155]

[0156] Where, ICi is the innovation contribution of patent i to electric power technology, DC′ i is the technical diffusion intensity characteristic of patent i, PR′ i is the technology network hub index of patent i, TS′ i is the technical uniqueness score of patent i, α is the characteristic weight of technology diffusion intensity, is the weight of the technology network hub index, η is the weight of the technology uniqueness score, and

[0157] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0158] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0159] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0160] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0161] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0162] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph is characterized by: The steps include: Classify the electric power technology patent data based on triples to obtain the first triplet data; Obtaining local credibility data of the first triplet data, building a triplet reliability scoring model based on a BP neural network, screening the first triplet data, and obtaining the second triplet data; Obtaining contribution evaluation features of the electric power patent knowledge graph constructed based on the second triple data; According to the contribution evaluation characteristics, a contribution evaluation model is constructed based on the BP neural network, and the model output is sorted to generate an innovation contribution evaluation table.

2. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 1 is characterized in that: The electric power technology patent data is classified based on triples to obtain the first triplet data, which is specifically: The first triple data includes a plurality of triples, each triple having key fields of entity name, technology category and relationship; Identify key entities and relationships between key entities in power technology patent data based on natural language processing technology; Based on the key entities and the relationships between the key entities, several triples are combined to form structured data, namely the first triple data.

3. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 2 is characterized in that: The local credibility data includes a local accuracy index, and the local accuracy index is specifically: Preset reference array for each key field; Get the actual extracted array of key fields for each triple; Perform similarity analysis on the reference array and the actual extracted array to obtain the error of each key field; According to the error of each key field, the accuracy score of each key field is calculated based on the preset key field accuracy score model; The accuracy scores of all key fields of each triple are averaged as the local accuracy indicator of each triple.

4. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 3 is characterized in that: The local credibility data includes a local integrity indicator; the local integrity indicator is specifically: Obtain a non-empty key field set for each triple based on the first triple data; Preset the required key field set for each triple; The ratio of the triple's non-empty key field set to the triple's required field set is taken as the local completeness ratio of each triple and used as the local completeness indicator.

5. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 4 is characterized in that: The contribution evaluation features include technology diffusion intensity features, which are specifically: Based on the electric power patent knowledge graph, each patent node and the relationship between each patent node are obtained, and each patent node is used as a row and column to construct an adjacency matrix, and the relationship between each patent node is used as the value of the corresponding element in the adjacency matrix; Based on the adjacency matrix, the degree centrality value of the patent node is obtained by summing the values ​​of the corresponding elements of each patent node; Normalization is performed based on the degree centrality values ​​of all patent nodes to finally obtain the technology diffusion intensity characteristics.

6. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 5 is characterized in that: The contribution evaluation features include technical network hub indicators, which are specifically: Based on the electric power patent knowledge graph, each patent node and its relationship is obtained, and an adjacency matrix is ​​constructed with the citing patent nodes as rows and the cited patent nodes as columns. Obtaining structural information of patent citation relationships based on the adjacency matrix analysis, and analyzing the structural information of the patent citation relationships and the preset initial pagearank value of each patent node based on the pagearank method to obtain a pagearank value; Iterate the pagearank value of each patent node to obtain the pagearank value change before and after each pagearank value iteration; Compare the pagearank value change with the preset pagearank method convergence threshold, and when the pagearank value change is less than the preset pagearank method convergence threshold, stop iteration and obtain the final pagearank value; The final pagearank values ​​of all patent nodes are sorted and normalized to obtain the technology network hub index.

7. The method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to claim 6 is characterized in that: The contribution evaluation feature includes a technology uniqueness score, which is specifically: Based on the electric power patent knowledge graph, text vocabulary data is extracted from patent nodes and the relationship edges between nodes to form a patent corpus; Filter out all unique words based on the patent corpus and preset a fixed index position for each unique word to form a vocabulary; Based on all unique words in the vocabulary and combined with the TF-IDF method, the TF-IDF value of each unique word is obtained; Constructing a patent node feature vector based on the TF-IDF value of each unique word and the fixed index position of each unique word; Obtain patent node similarity data based on patent node feature vectors and combined with the cosine similarity method; Normalization is performed based on the similarity data of patent nodes to obtain a technology uniqueness score.

8. A system using the method for evaluating the contribution of electric power technology innovation based on patent knowledge graph according to any one of claims 1 to 7, characterized in that: It includes data acquisition module, data screening module, graph construction module and contribution analysis module; A data acquisition module, configured to classify the electric power technology patent data based on triples to obtain first triplet data; A data screening module is used to obtain local credibility data of the first triplet data, build a triplet reliability scoring model based on a BP neural network, screen the first triplet data, and obtain the second triplet data; A graph construction module, used to obtain contribution evaluation features of the electric power patent knowledge graph constructed based on the second triple data; The contribution analysis module is used to build a contribution evaluation model based on the BP neural network according to the contribution evaluation characteristics, sort the model output, and generate an innovation contribution evaluation table.

Citation Information

Patent Citations

  • Urban physical examination index knowledge graph construction method and system

    CN113901238A

  • Method and system for predicting and evaluating patent authorization based on attribute knowledge graph

    CN117972113B