Method, device and equipment for evaluating patent importance in patent combination
By building an undirected network of patent portfolio and applying a central algorithm, the problem of insufficient efficiency and accuracy of patent importance evaluation in the prior art is solved, and a scientific evaluation of the importance of patents in the patent portfolio is achieved, and management efficiency is improved.
Patent Information
- Application Number
- CN202411954438.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has difficulty effectively managing and evaluating patent importance in large patent portfolios, resulting in limited evaluation efficiency and accuracy.
By obtaining the target patent combination in the patent database, determining the target information of each patent document, building an undirected network, and using degree-centric, median or feature vector centrality algorithms to determine the importance of each patent document based on the network structure.
It has achieved a quantitative assessment of the importance of patents in the patent portfolio, improved patent management efficiency, and provided an important basis for the company's technical strategy and patent transfer negotiations.
Smart Images

Figure CN120067289A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly to a method, device, and equipment for evaluating the importance of patents in a patent portfolio. Background Art
[0002] With the increasingly fierce global technological competition, the patent portfolio owned by an enterprise has become an important indicator to measure its technological innovation ability and market competitiveness. However, in the face of thousands or even tens of thousands of patent assets, how to effectively manage and evaluate these patents, especially to determine which patents play a core role in the portfolio, is crucial for the long-term development of the enterprise.
[0003] Traditional methods for evaluating the importance of patents mostly rely on manually reading patent documents, analyzing patent citation relationships, evaluating the technical value and market potential of patents, etc. However, this method is not only time-consuming and laborious, but also difficult to comprehensively and objectively reflect the true importance of patents in the portfolio. In addition, with the rapid growth of the number of patents, the efficiency and accuracy of manual evaluation also face severe challenges.
[0004] In recent years, with the rapid development of big data and artificial intelligence technologies, more and more automated and intelligent tools have been introduced into patent analysis. For example, key information in patent documents can be extracted through text mining technology, technical associations between patents can be identified through clustering analysis, and the technical influence of patents can be evaluated through citation analysis. However, most of these tools and methods only stay at the level of preliminary screening and classification of patents, and are still insufficient for deeply mining important patents in the patent portfolio and giving quantitative evaluation results.
[0005] Therefore, there is an urgent need to provide a more reliable solution for evaluating the importance of patents in a patent portfolio. Summary of the Invention
[0006] The purpose of the present invention is to provide a method, device, and equipment for evaluating the importance of patents in a patent portfolio, which are used to determine the importance degree of each patent in the patent portfolio and improve the efficiency of patent management.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] In the first aspect, the present invention provides a method for evaluating the importance of patents in a patent portfolio, and the method includes:
[0009] Obtain a target patent portfolio from a patent database;
[0010] Determine the target information of each patent document in the target patent portfolio, where the target information at least includes the subject word, applicant, or classification number of the patent document;
[0011] Construct an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges;
[0012] Using a preset algorithm, based on the undirected network, determine the importance degree of each patent document in the target patent portfolio.
[0013] Optionally, when the target information is the subject word, determining the target information of each patent document in the target patent portfolio includes:
[0014] Preprocess the data in the target patent portfolio to obtain preprocessed patent document data;
[0015] Based on the preprocessed patent document data, construct a dictionary containing all vocabulary and a corpus composed of patent documents and corresponding word frequencies;
[0016] Set the LDA model parameters; the LDA model parameters at least include the number of subject words, the number of iterations, and the Dirichlet prior parameter of the LDA model;
[0017] Using the preprocessed patent document data, the dictionary, the corpus, and the LDA model parameters, train the LDA model;
[0018] According to the trained LDA model, calculate the probability distribution of each patent document on each subject word, and select the subject word with the highest probability as the subject word of each patent document in the target patent portfolio.
[0019] Optionally, the constructing an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges includes:
[0020] For any pair of patent documents in the target patent portfolio, determine the set of subject words of each patent document in the pair, and determine the same subject words shared by the pair of patent documents;
[0021] Regard each patent document as a network node. For any pair of patent documents and their same subject words, create an edge connecting the two nodes in the undirected network, and set the weight of the edge according to the number of the same subject words or the frequency of the same subject words appearing in the patent document;
[0022] Process all the patent documents in the target patent portfolio to form an undirected network containing all patent nodes and corresponding edges.
[0023] Optionally, the preset algorithm is at least the degree centrality algorithm, the betweenness centrality algorithm, or the eigenvector centrality algorithm.
[0024] Optionally, using the degree centrality algorithm, based on the undirected network, determine the importance of each patent document in the target patent portfolio, including:
[0025] Based on the undirected network, use the formula:
[0026] CD(N i )=∑{j=1}^{n}x{ij}(i≠j)
[0027] Determine the importance of each patent document in the target patent portfolio;
[0028] Where, CD(N i ) represents the degree centrality of node i in the undirected network, n represents the total number of nodes in the undirected network, x{ij} represents the relationship between node i and node j. If there is the same subject word between node i and node j, then x{ij} = 1; otherwise, x{ij} = 0.
[0029] Optionally, using the betweenness centrality algorithm, based on the undirected network, determine the importance of each patent document in the target patent portfolio, including:
[0030] Based on the undirected network, use the formula:
[0031]
[0032] Determine the importance of each patent document in the target patent portfolio;
[0033] Where, C b (v i ) represents the betweenness centrality of node i in the undirected network, g ij (v i ) is the number of paths passing through node v i among all the shortest paths from node i to node j, and g ij is the total number of shortest paths from node i to node j.
[0034] Optionally, using the eigenvector centrality algorithm, based on the undirected network, determine the importance of each patent document in the target patent portfolio, including:
[0035] Based on the undirected network, use the formula:
[0036]
[0037] Determine the importance of each patent document in the target patent portfolio;
[0038] Where, assume C e =(C e (v1 ), C e (v 2 ), …, C e (v n )) T is the central vector of all nodes in the undirected network, then λC e = A T C e , C e is the eigenvector of the adjacency matrix A T . λ is the eigenvalue. There are multiple solutions for λ. To ensure that the central value is greater than 0, the maximum λ is taken.
[0039] Optionally, after using the preset algorithm to determine the importance of each patent document in the target patent portfolio based on the undirected network, it further includes:
[0040] Adjust the display order of each patent document in the target patent portfolio according to the importance of each patent document in the target patent portfolio, so that the patent documents with higher importance are displayed first;
[0041] Or, assign identification information to patent documents with different importance levels according to the importance of each patent document in the target patent portfolio and display them; the identification information at least includes: color identification, graphic elements or icons;
[0042] Or, change the size scaling ratio of different patent documents in the target patent portfolio on the display interface according to the importance of each patent document in the target patent portfolio and display them; the higher the display ratio, the higher the importance of the patent document.
[0043] Compared with the prior art, the present invention provides a method for evaluating the importance of patents in a patent portfolio. By obtaining the target patent portfolio in the patent database, determining the target information of each patent document in the target patent portfolio, constructing an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges, and using a preset algorithm to determine the importance of each patent document in the target patent portfolio based on the undirected network. The present invention calculates the importance of patents in the patent portfolio through the node importance theory method in network relationship science, thereby judging the contribution degree of patents in the patent portfolio, which can improve the patent management efficiency and provide an important basis for the evaluation of the quality of patent portfolios, negotiations in the process of patent transfer, etc.
[0044] In a second aspect, the present invention provides an apparatus for evaluating the importance of patents in a patent portfolio. The apparatus includes:
[0045] A target patent portfolio acquisition module, configured to acquire a target patent portfolio in a patent database;
[0046] The subject term determination module is used to determine the target information of each patent document in the target patent portfolio, and the target information at least includes the subject term, the applicant or the classification number of the patent document;
[0047] The undirected network construction module is used to construct an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges;
[0048] The patent document importance determination module is used to determine the importance of each patent document in the target patent portfolio based on the undirected network by using a preset algorithm.
[0049] In a third aspect, the present invention provides a device for evaluating the importance of patents in a patent portfolio, and the device includes:
[0050] A memory, a processor, and a communication interface coupled to the processor; a computer program executable by the processor is stored on the memory; when the processor runs the computer program, it executes the above-mentioned method for evaluating the importance of patents in a patent portfolio.
[0051] The technical effects achieved by the device type solution provided in the second aspect and the device type solution provided in the third aspect are the same as those of the method type solution provided in the first aspect, and will not be elaborated here. Description of the Drawings
[0052] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0053] Figure 1 is a flowchart of a method for evaluating the importance of patents in a patent portfolio provided by the present invention;
[0054] Figure 2 is a schematic diagram of a simplified undirected network provided by the present invention;
[0055] Figure 3 is a schematic structural diagram of a device for evaluating the importance of patents in a patent portfolio provided by the present invention;
[0056] Figure 4 is a schematic structural diagram of a device for evaluating the importance of patents in a patent portfolio provided by the present invention. Detailed Embodiments
[0057] For the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0058] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0059] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or multiple.
[0060] Key node identification is an important research area in network relationship science, aiming to find those nodes that play important roles in information dissemination, influence diffusion, etc. from complex network structures. This process usually measures the importance of nodes based on a series of centrality metrics, including but not limited to degree centrality, betweenness centrality, closeness centrality, eigenvector centrality, etc. By identifying these key nodes, it can not only help us understand the basic structure and function of the network, but also provide guidance for various practical applications, such as optimizing information dissemination channels in social networks, enhancing the robustness and security of technical networks, and improving disease prevention and control strategies in medical networks.
[0061] A method, apparatus, and device for evaluating the importance of patents in a patent portfolio provided by the present invention regard a given patent portfolio as a network relationship, and each patent can be regarded as a node in this network relationship. The importance of patents in the patent portfolio is calculated by the node importance theory method in network relationship science. Thereby, the contribution degree of a patent in the patent portfolio can be judged, providing an important basis for the evaluation of the quality of the patent portfolio, the negotiation in the process of patent transfer, etc. Next, the solution provided by the embodiments of this specification will be described in conjunction with the accompanying drawings:
[0062] As Figure 1 shown, the process may include the following steps:
[0063] Step 110: Obtain a target patent portfolio from a patent database.
[0064] The value of a patent does not lie in a single patent, but in a group of patents with internal associations. The value of a patent portfolio is much greater than the sum of the values of all single patents within the patent portfolio. There are many means in the prior art for identifying patent portfolios. For example, a patent portfolio can be a patent portfolio formed manually based on expert experience, or a patent portfolio whose method is not manually constructed.
[0065] To obtain the target patent portfolio, the full text of the patents, claims, specifications, or abstract texts in the target patent portfolio can be obtained through a patent database. The patent database includes, but is not limited to, the EPO (European Patent Office), USPTO (United States Patent and Trademark Office), CNKI (China National Knowledge Infrastructure), etc.
[0066] Step 120: Determine the target information of each patent document in the target patent portfolio.
[0067] The target information may at least include the subject words, applicant, classification number, and inventors of the patent document, etc. In the embodiments of this specification, the subject words are taken as an example for the solution description. In actual applications, an undirected network can be constructed based on other dimensions (such as the applicant, classification number, and inventors, etc.) as edges, and the importance of each patent document can be calculated. Therefore, the solutions for evaluating the importance of each patent document with the relevant dimensions in the patent document as the edges of the undirected network and the patent document as the node all fall within the protection scope of the present invention.
[0068] The target patent portfolio includes multiple patent documents, and multiple subject words can be determined in each patent document.
[0069] When implementing step 120, the patent document can be initially analyzed to determine the overall structure of the patent document, identify key information, remove duplicate keywords and phrases, and combine words with similar meanings into one topic word. Synonyms and near-synonyms can also be considered as alternative topic words. Then, according to the content and technical field of the patent document, the relevance of each alternative topic word is evaluated, and the word that best matches the patent topic word and content is selected as the final topic word. The number of topic words should be determined according to the content and length of the patent document, ensuring both the comprehensiveness of the topic words and avoiding excessive redundant words. More specifically, the LDA model can be used to determine the target information of each patent document in the target patent portfolio, and the target information at least includes the topic words, applicant, or classification number of the patent document.
[0070] The LDA (Latent Dirichlet Allocation) model is a topic word model widely used in fields such as text mining, information retrieval, and natural language processing. The LDA model assumes that a document is composed of a mixture of multiple topic words, and each topic word is composed of a specific set of words. The goal of LDA is to automatically discover these latent topic words in the document and give the topic word distribution of each document.
[0071] Step 130: Taking patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges, construct an undirected network.
[0072] An undirected network means that the edge from any point i to point j in the network represents the same edge as the edge from point j to point i. This means that in the network, the edges have no direction and simply represent the existence of a connection relationship between two nodes.
[0073] For example: Let P1, P2, ……, Pn represent patent documents 1 to patent document n. The topic words included in P1 are Word1, Word2, Word3, Word4, Word5, the topic words in P2 are Word2, Word5, Word6, Word7, Word8, the topic words in P3 are Word3, Word4, Word6, Word7, Word8, and the topic words in Pn are Wordn1, Wordn2, Wordn3, Wordn4, Wordn5. It can be expressed as:
[0074] P1: [Word1, Word2, Word3, Word4, Word5]
[0075] P2: [Word2, Word6, Word7, Word8, Word9]
[0076] P3: [Word3, Word4, Word5, Word7, Word8]
[0077] ……
[0078] Pn: [Wordn1, Wordn2, Wordn3, Wordn4, Wordn5].
[0079] As Figure 2 shown, Figure 2 shows a simplified network composed of some nodes in an undirected network. Figure 2 In, P1, P2, P3, P4, and P5 represent 5 patent documents. P1 is connected to P2, P3, P4, and P5 respectively, indicating that document P1 has the same subject terms as P2, P3, P4, and P5; P2 is connected to P1 and P3 respectively, indicating that document P2 has the same subject terms as P3 and P1; P3 is connected to P2, P1, and P4 respectively, indicating that document P3 has the same subject terms as P2, P1, and P4; P4 is connected to P1, P3, and P5 respectively, indicating that document P4 has the same subject terms as P1, P3, and P5; P5 is connected to P1 and P4 respectively, indicating that document P5 has the same subject terms as P1 and P4. As for the numbers on the edges, they represent the weights of the edges, and this weight can represent the number of the same subject terms or the occurrence frequency of the same subject terms in the patent documents, etc.
[0080] Use a method of a series of centrality indicators in network relationship science to measure the importance of nodes, calculate the centrality indicators of each patent document, and obtain the importance of the patent document, specifically as in step 140.
[0081] Step 140: Adopt a preset algorithm, and based on the undirected network, determine the importance degree of each patent document in the target patent portfolio.
[0082] The preset algorithm can at least be a degree centrality algorithm, a betweenness centrality algorithm, or an eigenvector centrality algorithm.
[0083] Figure 1 In the method of, by obtaining the target patent portfolio in the patent database, determining the target information of each patent document in the target patent portfolio, using patents as nodes, and using the same target information of any two patent documents in the target patent portfolio as edges, constructing an undirected network, and adopting a preset algorithm, based on the undirected network, determining the importance degree of each patent document in the target patent portfolio. The present invention calculates the importance degree of patents in the patent portfolio through the node importance theory method in network relationship science, thereby judging the contribution degree of patents in the patent portfolio, which can improve the patent management efficiency and provide an important basis for the evaluation of the quality of the patent portfolio, the negotiation in the process of patent transfer, etc.
[0084] Based onFigure 1 For the method, some specific embodiments of the method are also provided in this specification, and the following will be described.
[0085] In step 120, when using the LDA model to determine the target information of each patent document in the target patent portfolio, where the target information at least includes the subject words, applicant or classification number of the patent document, for each patent document in the patent database, using the LDA model, by maximizing the probability distribution of the attribution of each word segment in each patent document to the subject word segment, the subject words of each patent document are obtained. Specifically, it may include:
[0086] Preprocess the data in the target patent portfolio to obtain preprocessed patent document data;
[0087] Construct a dictionary containing all vocabulary and a corpus composed of patent documents and corresponding word frequencies based on the preprocessed patent document data;
[0088] Set the LDA model parameters; the LDA model parameters at least include the number of subject words, the number of iterations, and the Dirichlet prior parameter of the LDA model;
[0089] Train the LDA model using the preprocessed patent document data, the dictionary, the corpus, and the LDA model parameters;
[0090] According to the trained LDA model, calculate the probability distribution of each patent document on each subject word, and select the subject word with the highest probability as the subject word of each patent document in the target patent portfolio.
[0091] When the above steps are specifically implemented, the LDA model is used to perform subject word modeling on the patent document to obtain the subject words of each patent document. The specific implementation steps may include:
[0092] Collect patent document data: Obtain the set of patent documents to be analyzed, and these documents can be publicly available patent specifications, patent abstracts, etc.
[0093] Data preprocessing: Remove irrelevant characters in the text, such as special symbols, numbers, spaces, redundant punctuation, etc., as well as common format markers in patent documents. For Chinese patent documents, word segmentation processing is required to break the sentences into individual words, and Chinese word segmentation tools such as jieba can be used for word segmentation. Remove some words that frequently appear in the text but are not helpful for subject word modeling, such as "de", "shi", "le", etc.
[0094] Model construction and training: Convert the preprocessed patent documents into a Bag of Words model or TF-IDF feature vectors, and construct a dictionary and a corpus. Each word in the dictionary has a unique ID, and each document in the corpus is a vector composed of word IDs and corresponding word frequencies.
[0095] Use the preprocessed patent document data and the set LDA model parameters to train the LDA model.
[0096] After training is completed, the high-frequency words under each topic word can be viewed, and these words can represent the main content of the topic word.
[0097] For each patent document, its probability distribution on each topic word can be obtained, so as to determine the main topic word of the document.
[0098] Through the above steps, valuable topic words can be extracted from patent documents, providing strong support for subsequent analysis and applications.
[0099] For step 130, it may specifically include:
[0100] For any pair of patent documents in the target patent portfolio, determine the set of topic words of each patent document in the pair, and determine the same topic words shared by the pair of patent documents;
[0101] Regard each patent document as a network node. For any pair of patent documents and their same topic words, create an edge connecting the two nodes in an undirected network. The weight of the edge is set according to the number of the same topic words or the occurrence frequency of the same topic words in the patent document.
[0102] Process all the patent documents in the target patent portfolio to form an undirected network containing all patent nodes and corresponding edges.
[0103] Specifically, when the above steps are implemented, the set of topic words of each patent document can be determined first; then in the target patent portfolio, according to the application date, publication date, citation relationship or other relevant criteria of the patent, determine adjacent pairs of patent documents; for any pair of patent documents, compare their respective sets of topic words, find the same topic words they share, and use them as the basis for constructing the edges of the undirected network; regard each patent document as a network node and assign a unique identifier to each node; for any pair of patent documents and their shared topic words, create an edge connecting the two nodes in the undirected network, and repeat the above steps until all pairs of patent documents are processed, forming an undirected network containing all patent nodes and corresponding edges.
[0104] In step 140, the preset algorithm may at least be a degree centrality algorithm, a betweenness centrality algorithm or a eigenvector centrality algorithm. Next, the implementation methods corresponding to the above algorithms may be described respectively:
[0105] Implementation method 1: using degree centrality algorithm
[0106] Implementation steps may include:
[0107] 1) Obtain patent documents in the patent portfolio through patent database;
[0108] 2) performing pre-processing operations such as word segmentation and stop word removal on the patent document, and then using the LDA model to determine the subject words of each patent document;
[0109] 3) Using patent documents as nodes and keywords that appear in both patents as edges in the network, an undirected network is constructed
[0110] 4) Use the degree centrality algorithm to measure the importance of nodes in complex networks.
[0111] Among them, degree centrality refers to the number of edges directly connected to a node in the network. The higher the degree centrality, the more direct connections the node has with other nodes, which means that it occupies a more central position in the network. Figure 2 , the number of edges of P1 is 4, the number of edges of P2 and P5 is 2, and the number of edges of P3 and P4 is 3. Therefore, in this patent portfolio, P1 is the most important, followed by P3 and P4, and P2 and P5 are the least important.
[0112] Further use the edge weights and as the degree centrality index to calculate, Figure 1 In the figure, P1 is 7, P2 is 2, P3 is 5, P4 is 4, and P5 is 2. The importance of each patent in the patent portfolio is P1>P3>P4>P2=P5.
[0113] Furthermore, using a degree centrality algorithm to determine the importance of each patent document in the target patent portfolio based on the undirected network may include:
[0114] Based on the undirected network, formula (1) is used:
[0115] CD(N i )=∑{j=1}^{n}x{ij}(i≠j) (1)
[0116] Determine the importance of each patent document in the target patent portfolio;
[0117] Among them, CD(N i) represents the degree centrality of node i in an undirected network, n represents the total number of nodes in the undirected network, and x{ij} represents the relationship between node i and node j. If there are the same topic words between node i and node j, then x{ij} = 1; otherwise, x{ij} = 0.
[0118] In the first embodiment above, for each node (i.e., patent document) in the undirected network, its degree centrality is calculated. Degree centrality refers to the number of edges directly connected to the node. In a network without weights, degree centrality is the number of neighbors of the node. If there are weights, degree centrality can be calculated as the sum of the weights of the edges connected to the node. According to the calculated degree centrality values, the patent documents in the target patent portfolio are sorted. The higher the degree centrality value of a patent document, the more connections it has in the network and the higher its importance in the target patent portfolio.
[0119] Embodiment 2: Betweenness centrality algorithm
[0120] The implementation steps may include:
[0121] 1) Obtain the patent documents in the patent portfolio through the patent database.
[0122] 2) Perform operations such as word segmentation and stop word removal on the patent text documents, and use the LDA model to determine the topic words of each patent document.
[0123] 3) Construct an undirected network with patent documents as nodes and the topic words that have appeared together in two patent documents as edges in the network.
[0124] 4) Use the betweenness centrality algorithm to measure the importance of nodes in the complex network.
[0125] Among them, betweenness centrality is an index to measure the importance of nodes in a network, which reflects the mediating role of a certain node in the network. Specifically, betweenness centrality measures the number of times a node appears on the shortest path between all node pairs. Nodes with high betweenness centrality mean that they play an important "bridge" role in the network and are the key paths for the flow of information, resources or other forms.
[0126] Furthermore, adopting the betweenness centrality algorithm, based on the undirected network, to determine the importance of each patent document in the target patent portfolio may include:
[0127] Based on the undirected network, use formula (2):
[0128]
[0129] Determine the importance of each patent document in the target patent portfolio;
[0130] Among them, C b (v i ) represents the betweenness centrality of node i in the undirected network, and g ij (v i ) is the number of paths passing through node v among all the shortest paths from node i to node j, while g i is the total number of shortest paths from node i to node j. ij
[0131] Using the betweenness centrality algorithm, calculate the betweenness value of each node in the network, that is, the proportion of the number of times the node appears in all the shortest paths. The higher the betweenness value of a patent, the stronger its ability to connect different subject words or technical fields as a bridge in the network, so its importance is higher.
[0132] Embodiment 3: Adopt the eigenvector centrality algorithm
[0133] The implementation steps may include:
[0134] 1) Obtain the patent documents in the patent portfolio through the patent database.
[0135] 2) Perform operations such as word segmentation and stop word removal on the patent text documents, and use the LDA model to determine the subject words of each patent document.
[0136] 3) Use the patent documents as nodes and the subject words that have appeared together in two patent documents as edges in the network to construct an undirected network.
[0137] 4) Use the eigenvector centrality algorithm to measure the importance of nodes in the complex network.
[0138] Furthermore, adopting the eigenvector centrality algorithm, based on the undirected network, to determine the importance of each patent document in the target patent portfolio may include:
[0139] Based on the undirected network, use the formula:
[0140]
[0141] Determine the importance of each patent document in the target patent portfolio;
[0142] Among them, assuming C e =(C e (v 1 ), C e (v 2 ), …, C e (v n )) T is the central vector of all nodes in the undirected network, then λCe = A T C e , C e is the eigenvector of the adjacency matrix A T , λ is the eigenvalue, and λ has multiple solutions. To ensure that the central value is greater than 0, the maximum λ is taken.
[0143] Finally, after determining the importance of each patent document in the patent portfolio, it may further include:
[0144] According to the importance of each patent document in the target patent portfolio, adjust the display order of each patent document in the target patent portfolio, and display the patent documents in descending order of importance, so that the patent documents with higher importance are displayed first or prominently;
[0145] Alternatively, according to the importance of each patent document in the target patent portfolio, assign identification information to the patent documents with different importance levels and display them; the identification information includes at least: color identification, graphic elements or icons;
[0146] Alternatively, according to the importance of each patent document in the target patent portfolio, change the size scaling ratio of different patent documents in the display interface and display them; the patent document with a larger display ratio has a higher importance level.
[0147] Based on the above embodiments, the technical solution provided by the present invention can achieve at least the following technical effects:
[0148] 1) By determining the importance of each patent in the patent portfolio through the technical solution provided by the present invention, an enterprise can allocate resources more precisely, improve the overall operation efficiency, and achieve better economic benefits.
[0149] 2) It helps the enterprise better evaluate the overall value of the patent portfolio and can provide an important basis for the enterprise to formulate a patent strategy. By evaluating the importance of each patent in the patent portfolio, the enterprise can identify potential technical fields and R & D directions, which helps the enterprise concentrate its efforts on technology R & D and innovation, and promotes the upgrading of products and technologies.
[0150] 3) Determining the importance of each patent in the patent portfolio helps the enterprise establish a more scientific and reasonable patent management system. The enterprise can classify and manage patents according to their importance, improve the efficiency and effectiveness of patent management, and reduce the costs and risks of patent management.
[0151] Based on the same idea, the present invention also provides a device for evaluating the importance of patents in a patent portfolio, as Figure 3 shown, the device may include:
[0152] A target patent portfolio acquisition module 310, configured to acquire a target patent portfolio from a patent database;
[0153] A subject term determination module 320, configured to determine target information of each patent document in the target patent portfolio, where the target information at least includes a subject term, an applicant, or a classification number of the patent document;
[0154] An undirected network construction module 330, configured to construct an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges;
[0155] A patent document importance determination module 340, configured to use a preset algorithm to determine the importance of each patent document in the target patent portfolio based on the undirected network.
[0156] Based on Figure 3 the device in
[0157] Optionally, when the target information is a subject term, the subject term determination module 320 may include:
[0158] A preprocessing unit, configured to preprocess the data in the target patent portfolio to obtain preprocessed patent document data;
[0159] A dictionary and corpus construction unit, configured to construct a dictionary containing all vocabulary and a corpus composed of patent documents and corresponding word frequencies based on the preprocessed patent document data;
[0160] An LDA model parameter setting unit, configured to set LDA model parameters; the LDA model parameters at least include the number of subject terms, the number of iterations, and the Dirichlet prior parameter of the LDA model;
[0161] An LDA model training unit, configured to train an LDA model by using the preprocessed patent document data, the dictionary, the corpus, and the LDA model parameters;
[0162] A subject term determination unit, configured to calculate the probability distribution of each patent document on each subject term according to the trained LDA model, and select the subject term with the highest probability as the subject term of each patent document in the target patent portfolio.
[0163] Optionally, the undirected network construction module 330 may include:
[0164] A same subject term determination unit, configured to, for any pair of patent documents in the target patent portfolio, determine the set of subject terms of each patent document in the pair of patent documents, and determine the same subject terms shared by the pair of patent documents;
[0165] A node and edge determination unit, which treats each patent document as a network node. For any pair of patent documents and their same subject terms, an edge connecting the two nodes is created in an undirected network, and the weight of the edge is set according to the number of same subject terms or the occurrence frequency of the same subject terms in the patent document;
[0166] An undirected network construction module, which processes all patent documents in the target patent portfolio to form an undirected network containing all patent nodes and corresponding edges.
[0167] Optionally, the preset algorithm can be at least the degree centrality algorithm, the betweenness centrality algorithm, or the eigenvector centrality algorithm.
[0168] Optionally, a first patent document importance determination unit is used for:
[0169] Based on the undirected network, using the formula:
[0170] CD(N i )=∑{j=1}^{n}x{ij}(i≠j)
[0171] Determine the importance of each patent document in the target patent portfolio;
[0172] Wherein, CD(N i ) represents the degree centrality of node i in the undirected network, n represents the total number of nodes in the undirected network, x{ij} represents the relationship between node i and node j. If there are the same subject terms between node i and node j, then x{ij} = 1; otherwise, x{ij} = 0.
[0173] Optionally, a second patent document importance determination unit is used for:
[0174] Based on the undirected network, using the formula:
[0175]
[0176] Determine the importance of each patent document in the target patent portfolio;
[0177] Wherein, C b (v i ) represents the betweenness centrality of node i in the undirected network, g ij (v i ) is the number of paths passing through node v among all the shortest paths from node i to node j i , and g ij is the total number of all shortest paths from node i to node j.
[0178] Optionally, a third patent document importance determination unit is used for:
[0179] Based on the undirected network, the formula:
[0180]
[0181] is used to determine the importance level of each patent document in the target patent portfolio;
[0182] where, assuming C e =(C e (v 1 ), C e (v 2 ), …, C e (v n )) T is the central vector of all nodes in the undirected network, then λC e = A T C e , C e is the eigenvector of the adjacency matrix A T , λ is the eigenvalue, and there are multiple solutions for λ. To ensure that the central value is greater than 0, the maximum λ is taken.
[0183] Optionally, the device may further include: a display module, configured to:
[0184] Adjust the display order of each patent document in the target patent portfolio according to the importance level of each patent document in the target patent portfolio, so that the patent documents with higher importance levels are displayed first;
[0185] Or, assign identification information to patent documents with different importance levels according to the importance level of each patent document in the target patent portfolio, and display them; the identification information includes at least: color identification, graphic elements or icons;
[0186] Or, change the size scaling ratio of different patent documents in the target patent portfolio on the display interface according to the importance level of each patent document in the target patent portfolio, and display them; the patent document with a larger display ratio has a higher importance level..
[0187] Based on the same idea, the embodiments of this specification also provide a device for evaluating the importance of patents in a patent portfolio. As Figure 4 shown, the device includes:
[0188] A memory, a processor, and a communication interface coupled to the processor; a computer program executable by the processor is stored on the memory; when the processor runs the computer program, it executes the foregoing method for evaluating the importance of patents in a patent portfolio.
[0189] As Figure 4As shown in the figure, the above-mentioned processor may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. The above-mentioned communication interface(s) may be one or more. The communication interface may use any device such as a transceiver for communicating with other devices or communication networks.
[0190] As Figure 4 shown in the figure, the above-mentioned terminal device may further include a communication line. The communication line may include a path for transmitting information between the above-mentioned components.
[0191] Optionally, as Figure 4 shown in the figure, the terminal device may further include a memory. A computer program that can be run by the processor is stored on the memory; when the processor runs the computer program, the method provided by the embodiment of the present invention is implemented.
[0192] In a specific implementation, as an embodiment, as Figure 4 shown in the figure, the processor may include one or more CPUs, such as Figure 4 CPU0 and CPU1 in
[0193] In a specific implementation, as an embodiment, as Figure 4 shown in the figure, the terminal device may include multiple processors, such as Figure 4 the processors in
[0194] The above mainly introduces the solution provided by the embodiment of the present invention from the perspective of the interaction between various modules. It can be understood that, in order to implement the above functions, each module includes the corresponding hardware structure and / or software unit for executing each function. Those skilled in the art should easily realize that the present invention can be implemented in the form of hardware or a combination of hardware and computer software by combining the units and algorithm steps of the examples described in the embodiments disclosed herein. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0195] Embodiments of the present invention may divide functional modules according to the above method examples. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0196] The processor in this specification may also have the function of a memory. The memory is used to store computer execution instructions for executing the solution of the present invention and is controlled by the processor to execute. The processor is used to execute the computer execution instructions stored in the memory, thereby implementing the method provided by the embodiments of the present invention.
[0197] The memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor through a communication line. The memory may also be integrated with the processor.
[0198] Optionally, the computer execution instructions in the embodiments of the present invention may also be referred to as application program code, and the embodiments of the present invention do not make specific limitations thereto.
[0199] The method disclosed in the embodiments of the present invention above can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with the ability to process signals. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0200] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0201] Although the present invention has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, the present specification and the drawings are merely exemplary illustrations of the invention defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for evaluating the importance of patents in a patent portfolio, characterized in that: include: Obtain target patent portfolios from patent databases; Determine target information of each patent document in the target patent portfolio, wherein the target information at least includes a subject word, an application subject or a classification number of the patent document; An undirected network is constructed with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges; A preset algorithm is used to determine the importance of each patent document in the target patent portfolio based on the undirected network.
2. The method for evaluating patent importance in a patent portfolio according to claim 1, characterized in that: When the target information is the keyword, determining the target information of each patent document in the target patent combination includes: Preprocessing the data in the target patent portfolio to obtain preprocessed patent document data; Based on the preprocessed patent document data, a dictionary containing all vocabulary and a corpus consisting of patent documents and corresponding word frequencies are constructed; Setting LDA model parameters; the LDA model parameters at least include the number of subject words, the number of iterations and the Dirichlet prior parameters of the LDA model; Training an LDA model using the preprocessed patent document data, the dictionary, the corpus, and the LDA model parameters; According to the trained LDA model, the probability distribution of each patent document on each subject word is calculated, and the subject word with the highest probability is selected as the subject word of each patent document in the target patent portfolio.
3. A patent importance assessment method in a patent portfolio according to claim 2, characterized in that: The undirected network is constructed by taking patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges, including: For any pair of patent documents in the target patent portfolio, determine a set of subject words for each patent document in any pair of patent documents, and determine the same subject words shared by any pair of patent documents; Each patent document is regarded as a network node. For any pair of patent documents and the same keywords mentioned therein, an edge connecting the two nodes is created in an undirected network. The weight of the edge is set according to the number of the same keywords or the frequency of occurrence of the same keywords in the patent document. All patent documents in the target patent portfolio are processed to form an undirected network including all patent nodes and corresponding edges.
4. The method for evaluating patent importance in a patent portfolio according to claim 1, characterized in that: The preset algorithm is at least a degree centrality algorithm, a betweenness centrality algorithm or a eigenvector centrality algorithm.
5. The method for evaluating patent importance in a patent portfolio according to claim 4, characterized in that: Using a degree centrality algorithm, based on the undirected network, the importance of each patent document in the target patent portfolio is determined, including: Based on the undirected network, the formula is adopted: CD(N i )=∑{j=1}^{n}x{ij}(i≠j) Determine the importance of each patent document in the target patent portfolio; Among them, CD(N i ) represents the degree centrality of node i in an undirected network, n represents the total number of nodes in the undirected network, x{ij} represents the relationship between node i and node j, if there is the same keyword between node i and node j, then x{ij}=1; otherwise, x{ij}=0.
6. The method for evaluating patent importance in a patent portfolio according to claim 4, characterized in that: Using the betweenness centrality algorithm, based on the undirected network, the importance of each patent document in the target patent portfolio is determined, including: Based on the undirected network, the formula is adopted: Determine the importance of each patent document in the target patent portfolio; Among them, C b (v i ) represents the betweenness centrality of node i in an undirected network, g ij (v i ) are all the shortest paths from node i to node j that pass through node v i The number of paths, and g ij is the total number of shortest paths from node i to node j.
7. The method for evaluating patent importance in a patent portfolio according to claim 4, characterized in that: Using the eigenvector centrality algorithm, based on the undirected network, the importance of each patent document in the target patent portfolio is determined, including: Based on the undirected network, the formula is adopted: Determine the importance of each patent document in the target patent portfolio; Among them, assuming that C e =(C e (v1),C e (v2),…,C e (v n )) T is the center vector of all nodes in the undirected network, then λC e =A T C e , C e is the adjacency matrix A T λ is the eigenvector of , λ is the eigenvalue, λ has multiple solutions, to ensure that the central value is greater than 0, take the maximum λ.
8. The method for evaluating patent importance in a patent portfolio according to claim 1, characterized in that: The step of using a preset algorithm to determine the importance of each patent document in the target patent portfolio based on the undirected network further includes: According to the importance of each patent document in the target patent portfolio, the display order of each patent document in the target patent portfolio is adjusted so that patent documents with higher importance are displayed first; Alternatively, according to the importance of each patent document in the target patent portfolio, identification information is assigned to patent documents of different importance and displayed; the identification information at least includes: color identification, graphic elements or icons; Alternatively, according to the importance of each patent document in the target patent combination, the size scaling ratio of different patent documents in the target patent combination on the display interface is changed and displayed; the larger the display ratio, the higher the importance of the patent document.
9. A device for evaluating patent importance in a patent portfolio, characterized in that the device include: A target patent portfolio acquisition module is used to acquire a target patent portfolio from a patent database; A subject word determination module, used to determine target information of each patent document in the target patent portfolio, wherein the target information at least includes a subject word, an application subject or a classification number of the patent document; An undirected network construction module, used to construct an undirected network with patents as nodes and the same target information of any two patent documents in the target patent portfolio as edges; The patent document importance determination module is used to determine the importance of each patent document in the target patent portfolio based on the undirected network using a preset algorithm.
10. A device for evaluating patent importance in a patent portfolio, characterized in that the device include: A memory, a processor, and a communication interface coupled to the processor; the memory stores a computer program executable by the processor; When the processor runs the computer program, it executes a method for evaluating the importance of patents in a patent combination as described in any one of claims 1 to 8.
Citation Information
Cited By
Patent maintenance decision method and system based on patent portfolio synergy analysis
CN122757524A