Target patent identification method and device
By clustering and thread matching methods of massive patent data, the accuracy and efficiency of identifying specific patents are solved, and automated identification and more accurate patent potential identification are achieved.
Patent Information
- Application Number
- CN202510053179.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately and quickly identify specific patents from massive patent data, and there are strong subjectivity and efficiency problems.
By determining candidate patents and their patent citation information, the patent groups are divided based on the application year, clustering is performed to form the first and second cluster clusters, and target patents are determined in combination with cluster cluster threads.
It achieves more accurate identification of potential target patents, automated processing of large amounts of patent data, avoids the inefficiency and deviation of traditional manual classification, and provides enterprises with scientific patent strategies, technology research and development and investment decision-making basis.
Smart Images

Figure CN119988636A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of data processing and information retrieval technology, and in particular to a method and device for identifying a target patent. Background Art
[0002] In modern patent analysis, unique patents, as a form of intellectual property with significant innovation and disruptive potential, are gaining increasing attention from both academia and industry. These patents typically represent patents that differ significantly from the majority of patents in a given technology field. Although relatively rare, these patents often have a profound impact on technological development, industrial transformation, and market competition. Unlike traditional patent analysis methods, unique patents do not necessarily conform to mainstream technological paths or industry trends, but their emergence may signal future technological shifts or potential market opportunities. Therefore, identifying unique patents is of great significance to businesses, research institutions, and government agencies.
[0003] Currently, the discovery and analysis of unique patents primarily relies on expert experience and domain knowledge, often employing qualitative analysis to evaluate patents. However, this approach is highly subjective, and the accuracy and objectivity of the analysis results are difficult to guarantee. This is especially true when faced with massive patent datasets, where the efficiency and reliability of manual analysis are limited. Therefore, how to accurately and quickly identify unique patents from within this vast amount of patent data has become a pressing technical challenge. Summary of the Invention
[0004] The present application aims to solve one of the technical problems in the related art at least to a certain extent.
[0005] To this end, one purpose of the present application is to propose a method for identifying a target patent, including: determining multiple candidate patents, and determining patent citation information for each candidate patent; dividing the candidate patents based on the application year of the candidate patents to obtain multiple patent groups; for any patent group, clustering the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establishing a cluster thread based on the first cluster; taking each candidate patent as a node, establishing an edge based on the patent citation information of the candidate patent, constructing a patent knowledge graph, and clustering the patent knowledge graph to obtain multiple second clusters; matching the first cluster and the second cluster, combining the cluster threads to determine the target patent.
[0006] The second purpose of this application is to propose an identification device for the target patent.
[0007] The third objective of this application is to provide an electronic device.
[0008] A fourth object of the present application is to provide a non-transitory computer-readable storage medium.
[0009] A fifth object of this application is to provide a computer program product.
[0010] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a method for identifying a target patent, including: determining multiple candidate patents, and determining the patent citation information of each candidate patent; dividing the candidate patents based on the application year of the candidate patents to obtain multiple patent groups; for any patent group, clustering the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establishing a cluster thread based on the first cluster; taking each candidate patent as a node, establishing an edge based on the patent citation information of the candidate patent, constructing a patent knowledge graph, and clustering the patent knowledge graph to obtain multiple second clusters; matching the first cluster and the second cluster, combining the cluster threads to determine the target patent.
[0011] According to one embodiment of the present application, candidate patents in a patent group are clustered based on patent citation information to obtain at least one first cluster corresponding to the patent group, including: determining a reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; calculating a first similarity between the reference patent sets pairwise, and clustering the candidate patents in the patent group based on the first similarity to obtain a first cluster corresponding to the patent group.
[0012] According to one embodiment of the present application, a clustering thread is established based on the first clustering cluster, including: combining the first clustering clusters in adjacent application years in pairs to obtain multiple first clustering cluster point pairs, wherein the first clustering cluster point pairs contain two first clustering clusters and the application years corresponding to the two first clustering clusters are adjacent; calculating the second similarity between the two first clustering clusters included in each first clustering cluster point pair; determining a target second similarity greater than a preset similarity threshold from the second similarity, and connecting the two first clustering clusters included in the first clustering cluster point pairs corresponding to the target second similarity to obtain a clustering thread.
[0013] According to one embodiment of the present application, candidate patents in a patent group are clustered based on patent citation information to obtain at least one first cluster corresponding to the patent group, including: determining a reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; setting clustering constraints, wherein the clustering constraints include: for any first cluster generated, the first cluster corresponds to a target reference patent set, and the intersection between the reference patent sets corresponding to any two candidate patents in the first cluster is the same as the target reference patent set; clustering the candidate patents in the patent group based on the clustering constraints to obtain the first cluster corresponding to the patent group.
[0014] According to one embodiment of the present application, a clustering thread is established based on the first clustering, including: combining first clusterings in adjacent application years in pairs to obtain multiple first clustering point pairs, wherein the first clustering point pairs contain two first clusterings and the application years corresponding to the two first clusterings are adjacent; obtaining target reference patent sets corresponding to the two first clusterings in the first clustering point pairs; calculating the number of intersection patents and the number of union patents between the target reference patent sets, and calculating the ratio of the number of intersection patents to the number of union patents; determining a target ratio greater than a preset ratio threshold from the ratios, and connecting the two first clusterings in the first clustering point pairs corresponding to the target ratio to obtain a clustering thread.
[0015] According to one embodiment of the present application, the first cluster and the second cluster are matched, and the target patent is determined in combination with the cluster thread, including: for any second cluster, dividing the second cluster according to the application year of the candidate patent in the second cluster to obtain multiple second cluster subsets corresponding to the second cluster; obtaining the specific potential value corresponding to the second cluster subset; for any second cluster, determining the target second cluster subset from the multiple second cluster subsets corresponding to the second cluster according to the specific potential value; determining the target application year corresponding to the target second cluster subset, and determining the target first cluster corresponding to the target application year from the first cluster; determining the target patent based on the target first cluster, the target second cluster subset and the cluster thread.
[0016] According to one embodiment of the present application, obtaining the specific potential value corresponding to the second cluster subset includes: obtaining the number of overlapping candidate patents between the first cluster and the second cluster subset corresponding to the same application year as the second cluster subset; using the second cluster subsets that are in the first N adjacent application years of the second cluster subset and belong to the same second cluster as the second cluster subset as the N associated subsets of the second cluster subset, and obtaining the total number of candidate patents contained in the N associated subsets; obtaining the difference between the number of overlapping candidate patents and the total number, and using the difference as the specific potential value corresponding to the second cluster subset.
[0017] According to one embodiment of the present application, based on the target first cluster, the target second cluster subset and the cluster thread, the target patent is determined, including: determining the overlapping patents between the target first cluster and the target second cluster subset; and screening out the overlapping patents in the starting cluster located in the cluster thread as the target patent.
[0018] According to one embodiment of the present application, a reference patent set corresponding to each candidate patent in the patent group is determined based on the patent citation information of each candidate patent in the patent group, including: determining an initial reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group; for any candidate patent, screening out valid initial reference patents from the initial reference patent set corresponding to the candidate patent to form a reference patent set corresponding to the candidate patent.
[0019] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a target patent identification device, including: a first determination module, used to determine multiple candidate patents, and determine the patent citation information of each candidate patent; a patent division module, used to divide the candidate patents based on the application year of the candidate patents to obtain multiple patent groups; a first clustering module, used to cluster the candidate patents in the patent group based on the patent citation information for any patent group, obtain at least one first cluster cluster corresponding to the patent group, and establish a cluster cluster thread based on the first cluster cluster; a second clustering module, used to take each candidate patent as a node, establish an edge based on the patent citation information of the candidate patent, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain multiple second cluster clusters; a second determination module, used to match the first cluster cluster and the second cluster cluster, and determine the target patent in combination with the cluster cluster thread.
[0020] According to one embodiment of the present application, the first clustering module is further used to: determine a reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; calculate the first similarity between the reference patent sets pairwise, and cluster the candidate patents in the patent group based on the first similarity to obtain a first cluster corresponding to the patent group.
[0021] According to one embodiment of the present application, the first clustering module is further used to: combine the first clusters in adjacent application years in pairs to obtain multiple first cluster point pairs, wherein the first cluster point pairs contain two first clusters and the application years corresponding to the two first clusters are adjacent; calculate the second similarity between the two first clusters included in each first cluster point pair; determine a target second similarity greater than a preset similarity threshold from the second similarity, and connect the two first clusters included in the first cluster point pairs corresponding to the target second similarity to obtain a cluster thread.
[0022] According to one embodiment of the present application, the first clustering module is further used to: determine the reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; set clustering constraints, wherein the clustering constraints include: for any first cluster cluster generated, the first cluster cluster corresponds to a target reference patent set, and the intersection between the reference patent sets corresponding to any two candidate patents in the first cluster cluster is the same as the target reference patent set; cluster the candidate patents in the patent group based on the clustering constraints to obtain the first cluster cluster corresponding to the patent group.
[0023] According to one embodiment of the present application, the first clustering module is further used to: combine the first clusters in adjacent application years in pairs to obtain multiple first cluster point pairs, wherein the first cluster point pairs contain two first clusters and the application years corresponding to the two first clusters are adjacent; obtain the target reference patent sets corresponding to the two first clusters in the first cluster point pairs; calculate the number of intersection patents and the number of union patents between the target reference patent sets, and calculate the ratio of the number of intersection patents to the number of union patents; determine a target ratio greater than a preset ratio threshold from the ratio, and connect the two first clusters in the first cluster point pairs corresponding to the target ratio to obtain a cluster thread.
[0024] According to one embodiment of the present application, the second determination module is further used to: for any second cluster cluster, divide the second cluster cluster according to the application year of the candidate patent in the second cluster cluster to obtain multiple second cluster cluster subsets corresponding to the second cluster cluster; obtain the specific potential value corresponding to the second cluster cluster subset; for any second cluster cluster, determine the target second cluster cluster subset from the multiple second cluster cluster subsets corresponding to the second cluster cluster according to the specific potential value; determine the target application year corresponding to the target second cluster cluster subset, and determine the target first cluster cluster corresponding to the target application year from the first cluster cluster; determine the target patent based on the target first cluster cluster, the target second cluster cluster subset and the cluster cluster thread.
[0025] According to one embodiment of the present application, the second determination module is also used to: obtain the number of overlapping candidate patents between the first cluster and the second cluster subset corresponding to the same application year as the second cluster subset; use the second cluster subsets that are in the first N adjacent application years of the second cluster subset and belong to the same second cluster as the second cluster subset as the N associated subsets of the second cluster subset, and obtain the total number of candidate patents contained in the N associated subsets; obtain the difference between the number of overlapping candidate patents and the total number, and use the difference as the specific potential value corresponding to the second cluster subset.
[0026] According to one embodiment of the present application, the second determination module is further used to: determine overlapping patents between the target first cluster and the target second cluster subset; and screen out overlapping patents in the starting cluster in the cluster thread as target patents.
[0027] According to one embodiment of the present application, the first clustering module is further used to: determine the initial reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group; for any candidate patent, screen out valid initial reference patents from the initial reference patent set corresponding to the candidate patent to form the reference patent set corresponding to the candidate patent.
[0028] To achieve the above-mentioned purpose, the third aspect embodiment of the present application proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to implement the target patent identification method as described in the first aspect embodiment of the present application.
[0029] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the target patent identification method as described in the first aspect embodiment of the present application.
[0030] To achieve the above-mentioned purpose, the fifth embodiment of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements the target patent identification method as described in the first embodiment of the present application.
[0031] This application achieves at least the following beneficial effects: by combining two clustering methods, this application can more accurately identify target patents with potential, automatically process large amounts of patent data, and avoid the inefficiency and bias of traditional manual classification; in addition, it can also provide a scientific basis for enterprises in patent strategies, technology research and development, and investment decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0033] Figure 1 It is a schematic diagram of an exemplary implementation of a method for identifying a target patent shown in an embodiment of the present application.
[0034] Figure 2 It is a schematic diagram of an exemplary implementation of a method for identifying a target patent shown in an embodiment of the present application.
[0035] Figure 3 This is a schematic diagram of a first cluster, a second cluster, and a cluster thread shown in an embodiment of the present application.
[0036] Figure 4 This is a schematic diagram illustrating an embodiment of the present application for obtaining a specific potential value corresponding to a second cluster subset.
[0037] Figure 5 It is a schematic diagram of an exemplary implementation of a method for identifying a target patent shown in an embodiment of the present application.
[0038] Figure 6 This is a schematic diagram of an identification device of a target patent shown in an embodiment of the present application.
[0039] Figure 7 This is a schematic diagram of an electronic device shown in one embodiment of the present application. DETAILED DESCRIPTION
[0040] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0041] Figure 1 is a schematic diagram of an exemplary embodiment of a method for identifying a target patent shown in this application, such as Figure 1 As shown, the method for identifying the target patent includes the following steps:
[0042] S101, determining multiple candidate patents, and determining patent citation information for each candidate patent.
[0043] For example, patents within a preset time period can be determined from the patent library as candidate patents. For example, all patents from 2014 to 2023 (generally, the patent application year is used as the basis, and the patent application year is also the year in which the patent application date is located) can be determined from the patent library as candidate patents.
[0044] Furthermore, after determining the candidate patents, it is necessary to determine the patent citation information of each candidate patent.
[0045] For example, patent citation information can be expressed as: candidate patent 1 cites the previous patent A.
[0046] For example, patent citation information can be expressed as: candidate patent 2 cites previous patents A, B, and C.
[0047] S102: Divide the candidate patents based on their application years to obtain multiple patent groups.
[0048] Since the authorization date of a patent may not reflect the iteration of technology in a timely manner due to different review years, in this application, the candidate patents are divided according to the application year of the candidate patents to obtain multiple patent groups.
[0049] For example, if we divide the candidate patents from 2014 to 2023 according to the application year, we can get 10 patent groups. They are represented as: C 2014 、C 2015 ...C 2023 The following examples are based on this.
[0050] Each patent group includes all candidate patents in the corresponding application year of the patent group, which can be generally expressed as:
[0051] C year ={patent1, patent2,..., patent n}
[0052] In the above formula, patent1, patent2...patent n For patent group C year n candidate patents within.
[0053] S103: For any patent group, cluster the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establish a cluster thread based on the first cluster.
[0054] For example, for patent group C 2014 , according to Patent Group C 2014 The patent citation information of each candidate patent contained in the patent group C 2014 Cluster the candidate patents in the patent group C 2014 Corresponding to at least one first cluster.
[0055] For example, for patent group C 2015 , according to Patent Group C 2015 The patent citation information of each candidate patent contained in the patent group C 2015 Cluster the candidate patents in the patent group C 2015 Corresponding to at least one first cluster.
[0056] And so on, until you get C 2014 、C 2015 ...C 2023 There are 10 patent groups corresponding to the first cluster.
[0057] Among them, the first cluster corresponding to the patent group in each application year can be generally expressed as:
[0058] CC-Cluster year ={cluster1,cluster2,...,cluster i}
[0059] In the above formula, CC-Cluster year Representative Patent Group C year The corresponding first cluster set contains cluster1, cluster2...cluster i There are i first clusters in total.
[0060] After obtaining the first cluster corresponding to each patent group, a cluster thread is established based on the first cluster. Optionally, based on the similarity between first clusters in adjacent application years, first clusters with similar relationships are connected to obtain a cluster thread.
[0061] S104: Take each candidate patent as a node, establish edges based on the patent citation information of the candidate patent, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain multiple second clusters.
[0062] For example, each candidate patent from 2014 to 2023 is taken as a node V, and an edge E is established based on the patent citation information of the candidate patent to construct a patent knowledge graph G = (V, E).
[0063] Among them, the patent knowledge graph G is an undirected graph.
[0064] It should be noted that, unlike the above-mentioned method of obtaining the first cluster, when obtaining the second cluster, there is no need to divide the candidate patents based on the application year. What is directly obtained here is the entire patent knowledge graph corresponding to 2014 to 2023, thereby clustering the entire patent knowledge graph to obtain multiple second clusters.
[0065] In some embodiments, the patent knowledge graph is clustered using the Leiden algorithm to obtain multiple second clusters. The general expression is:
[0066] DC-Cluster=leiden(G)
[0067] In the above formula, DC-cluster represents the second cluster, leiden represents the leiden algorithm, and G represents the patent knowledge graph.
[0068] S105: Match the first cluster and the second cluster, combine the cluster threads, and determine the target patent.
[0069] The target patent in this application can be understood as a specific patent with significant value. After the first cluster, the second cluster, and the cluster thread are determined, specific patent identification processing is performed on them to obtain the target patent.
[0070] The embodiment of the present application proposes a method for identifying a target patent, including: determining multiple candidate patents, and determining the patent citation information of each candidate patent; dividing the candidate patents based on the application year of the candidate patents to obtain multiple patent groups; for any patent group, clustering the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establishing a cluster thread based on the first cluster; taking each candidate patent as a node, establishing an edge based on the patent citation information of the candidate patent, constructing a patent knowledge graph, and clustering the patent knowledge graph to obtain multiple second clusters; matching the first cluster and the second cluster, combining the cluster threads to determine the target patent. This application combines two clustering methods to more accurately identify target patents with potential, automatically process large amounts of patent data, and avoid the inefficiency and bias of traditional manual classification; in addition, it can also provide a scientific basis for enterprises in patent strategy, technology research and development, and investment decisions.
[0071] Figure 2 is a schematic diagram of an exemplary embodiment of a method for identifying a target patent shown in this application, such as Figure 2 As shown, the method for identifying the target patent includes the following steps:
[0072] S201, determining multiple candidate patents, and determining patent citation information for each candidate patent.
[0073] S202: Divide the candidate patents based on their application years to obtain multiple patent groups.
[0074] Regarding the specific implementation of steps S201 to S202, reference may be made to the detailed introduction of the relevant parts in the above embodiment, which will not be elaborated here.
[0075] S203: For any patent group, determine a reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent.
[0076] In some embodiments, for any candidate patent in any patent group, all reference patents cited by the candidate patent as indicated by the patent citation information of the candidate patent are constructed to form a reference patent set corresponding to the candidate patent. 2014 There are 10,000 candidate patents. For any candidate patent among the 10,000 candidate patents, if the candidate patent cites the previous patents A, B, and C, then the reference patent set corresponding to the candidate patent is {Patent A, Patent B, Patent C}.
[0077] In some embodiments, considering that patents with too few citations will affect subsequent clustering results, in this application, the initial reference patent set corresponding to each candidate patent in the patent group is determined based on the patent citation information of each candidate patent in the patent group; for any candidate patent, valid initial reference patents are screened from the initial reference patent set corresponding to the candidate patent to form the reference patent set corresponding to the candidate patent.
[0078] Optionally, the criteria for determining whether the initial reference patent is valid can be set to satisfy one of the following conditions:
[0079] 1. If the initial reference patent was published less than 3 years ago, the initial reference patent that has at least the same publication years plus one citation after the initial reference patent is valid. For example, if the initial reference patent was published 1 year ago, it must have been cited at least 2 times to be valid; if the initial reference patent was published 2 years ago, it must have been cited at least 3 times to be valid; if the initial reference patent was published 3 years ago, it must have been cited at least 4 times to be valid.
[0080] 2. If the initial reference patent was published more than 3 years ago, it is valid if it is cited five times or more in a single year after the publication of the initial reference patent.
[0081] For example, if a candidate patent cites the previous patents A, B, and C, the initial reference patent set corresponding to the candidate patent is determined to be {Patent A, Patent B, Patent C}. After judgment, if Patent A does not meet any of the above criteria, Patent A will be removed from the above initial reference patent set. That is, the reference patent set that the candidate patent finally corresponds to is determined to be {Patent B, Patent C}, and does not include Patent A.
[0082] S204: Calculate first similarities between reference patent sets in pairs, and cluster candidate patents in the patent group according to the first similarities to obtain first clusters corresponding to the patent group.
[0083] For example, if patent group C 2014 There are 10,000 candidate patents in the patent group. Taking the ideal case, that is, each candidate patent corresponds to a reference patent set, that is, a total of 10,000 reference patent sets are obtained, then the first similarities between the reference patent sets are calculated pairwise (10,000×10,000 first similarities can be obtained, which can be expressed as a 10,000×10,000 first similarity matrix), and then a clustering algorithm is used to cluster the candidate patents in the patent group according to the first similarity to obtain the first clustering cluster corresponding to the patent group. The patents in each first clustering cluster have high similarity.
[0084] It is not difficult to understand that in practice, not every candidate patent will cite a previous patent, that is, not every candidate patent will correspond to a reference patent set. In this case, the reference patent set of the candidate patent can be regarded as 0. At the same time, the first similarity between the reference patent set of the candidate patent and the reference patent sets of other candidate patents is also regarded as 0.
[0085] S205 , combining first clusters in adjacent application years in pairs to obtain a plurality of first cluster point pairs, wherein the first cluster point pairs include two first clusters and the application years corresponding to the two first clusters are adjacent.
[0086] For example, if patent group C 2015 Corresponding to the 100 first clusters, patent group C 2016 Corresponding to the 50 first clusters, the patent group C 2015 The corresponding first cluster and patent group C 2016 The corresponding first clusters are combined in pairs to obtain multiple first cluster point pairs (here for patent group C 2015 and Patent Group C 2016 100×50 first cluster point pairs will be obtained, where each first cluster point pair contains 1 patent group C 2015 The corresponding first cluster and 1 patent group C 2016 The corresponding first cluster.
[0087] Similarly, for the first clusters in adjacent application years, all first cluster point pairs are obtained by combining them in pairs.
[0088] S206: Calculate and obtain a second similarity between the two first clusters included in each first cluster point pair.
[0089] After the plurality of first cluster point pairs are determined, for any first cluster point pair, the similarity between the two first clusters included in the first cluster point pair is calculated and obtained as the second similarity.
[0090] The second similarity may be similarity based on cluster centers, Jaccard similarity, or the like.
[0091] S207 , determining a target second similarity greater than a preset similarity threshold from the second similarities, and connecting two first clusters included in the first cluster point pair corresponding to the target second similarity to obtain a cluster thread.
[0092] After determining the second similarity corresponding to each first cluster point pair as described above, each second similarity is compared with a preset similarity threshold, a second similarity greater than the preset similarity threshold is determined as a target second similarity, and the two first clusters included in the first cluster point pair corresponding to the target second similarity are connected to obtain a cluster thread.
[0093] For example, if a first cluster point pair contains 1 patent group C 2015 The corresponding first cluster and 1 patent group C 2016 If the first cluster point pair corresponds to the first cluster point pair, and the second similarity corresponding to the first cluster point pair is greater than the preset similarity threshold, then the patent group C included in the first cluster point pair is 2015 The corresponding first cluster and patent group C 2016The corresponding first cluster is connected. And so on.
[0094] S208: Take each candidate patent as a node, establish edges based on the patent citation information of the candidate patent, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain multiple second clusters.
[0095] For example, each candidate patent from 2014 to 2023 is taken as a node V, and an edge E is established based on the patent citation information of the candidate patent to construct a patent knowledge graph G = (V, E).
[0096] Among them, the patent knowledge graph G is an undirected graph.
[0097] Figure 3 This is a schematic diagram of a first cluster, a second cluster, and a cluster thread shown in this application, such as Figure 3 As shown, the patents from 2014 to 2023 in the patent database are used as candidate patents. Figure 3 The circles on the left represent the first clusters of the corresponding year, and the connecting lines between the circles represent that the two first clusters are threaded together. Figure 3 Multiple second clusters are shown on the right.
[0098] S209: Match the first cluster and the second cluster, combine the cluster threads, and determine the target patent.
[0099] The matching of the first cluster and the second cluster in S209 and the determination of the target patent in combination with the cluster threads may specifically include the following steps:
[0100] S2091 , for any second cluster, dividing the second cluster according to the application year of the candidate patents in the second cluster to obtain a plurality of second cluster subsets corresponding to the second cluster.
[0101] As can be seen above, when clustering the second clusters, candidate patents were not divided according to application year, so the second clusters may include candidate patents from multiple application years. In this application, for any second cluster, the second cluster is divided according to the application year of the candidate patents in the second cluster to obtain multiple second cluster subsets corresponding to the second cluster.
[0102] Figure 4 This is a schematic diagram of obtaining the specific potential value corresponding to the second cluster subset shown in the present application, such as Figure 4As shown, if a second cluster includes candidate patents from 2017 to 2023, after dividing the second cluster according to the application year of the candidate patents in the second cluster, the second cluster subset corresponding to each year from 2017 to 2023 can be obtained.
[0103] S2092: Obtain the specific potential value corresponding to the second cluster subset.
[0104] Specifically, the number of overlapping candidate patents between the first cluster and the second cluster subset corresponding to the same application year as the second cluster subset is obtained; the second cluster subsets that are in the first N adjacent application years of the second cluster subset and belong to the same second cluster as the second cluster subset are used as the N associated subsets of the second cluster subset, and the total number of candidate patents contained in the N associated subsets is obtained; the difference between the number of overlapping candidate patents and the total number is obtained, and the difference is used as the specific potential value corresponding to the second cluster subset.
[0105] For example, Figure 4 Take the second cluster subset of the 2020 application year as an example. The second cluster subset includes 39 candidate patents. Get all the first clusters corresponding to the 2020 application year, and get the number of candidate patents that overlap between all the first clusters corresponding to the 2020 application year and the second cluster subset. If the number of candidate patents that overlap is 18 ( Figure 4 The two red boxes in the figure represent overlapping candidate patents); taking N as 3 as an example, the total number of candidate patents contained in the second cluster subsets that belong to the same second cluster as the second cluster subset and are located in the first three years of the application year 2020 is obtained; Figure 4 For example, the total number of candidate patents contained in the associated subsets of the first three years corresponding to the second cluster subset of the 2020 application year is 2+4+3=9, then Figure 4 , the specific potential value corresponding to the second cluster subset for the 2020 application year is 18-9=9.
[0106] The general formula for calculating the specific potential value corresponding to the second cluster subset can be expressed as:
[0107]
[0108] In the above formula, OP represents the specific potential value of a second cluster corresponding to the second cluster subset in year year; DC-Cluster year Represents the second cluster subset corresponding to a second cluster in year year, CC-cluster year Representative Patent Group C year The corresponding first cluster set.
[0109] S2093 : For any second cluster, determine a target second cluster subset from a plurality of second cluster subsets corresponding to the second cluster according to the specific potential value.
[0110] In some embodiments, for any second cluster, after obtaining the specific potential value of each second cluster subset corresponding to the second cluster, the specific potential value is compared with a preset specific potential value threshold, and the second cluster subset corresponding to the specific potential value greater than the preset specific potential value threshold is determined as the target second cluster subset corresponding to the second cluster.
[0111] In some embodiments, for any second cluster, after obtaining the specific potential value of each second cluster subset corresponding to the second cluster, the second cluster subset with the largest specific potential value corresponding to the second cluster is determined as the target second cluster subset corresponding to the second cluster.
[0112] S2094: Determine the target application year corresponding to the target second cluster subset, and determine the target first cluster corresponding to the target application year from the first cluster.
[0113] For example, continue with Figure 4 For example, if Figure 4 The second cluster subset for the application year 2020 is the target second cluster subset, then all first clusters for the application year 2020 are obtained as the target first clusters corresponding to the target second cluster subset.
[0114] S2095: Determine a target patent based on the target first cluster, the target second cluster subset, and the cluster thread.
[0115] Specifically, the overlapping patents between the target first cluster corresponding to the target second cluster subset and the target second cluster subset are determined; and the overlapping patents in the starting cluster in the cluster thread are screened out as target patents.
[0116] Continue with Figure 4 For example, if Figure 4 The second cluster subset of the 2020 application year is the target second cluster subset, then all the first clusters of the 2020 application year are obtained as the target first cluster, and the overlapping patents between all the first clusters of the 2020 application year and the target second cluster subset are obtained, and the overlapping patents in the starting cluster in the cluster thread are taken as the target patents (that is, the starting node in a certain thread).
[0117] The embodiment of the present application combines two clustering methods to more accurately identify target patents with potential, automatically process large amounts of patent data, and avoid the inefficiency and bias of traditional manual classification. In addition, it can also provide a scientific basis for enterprises in patent strategies, technology research and development, and investment decisions.
[0118] Figure 5 is a schematic diagram of an exemplary embodiment of a method for identifying a target patent shown in this application, such as Figure 5 As shown, the method for identifying the target patent includes the following steps:
[0119] S501, determining multiple candidate patents, and determining patent citation information for each candidate patent.
[0120] S502: Divide the candidate patents based on their application years to obtain multiple patent groups.
[0121] S503: For any patent group, determine a reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent.
[0122] Regarding the specific implementation of steps S501 to S503, please refer to the detailed introduction of the relevant parts in the above embodiment, which will not be repeated here. The specific implementation of the above S503 has been introduced in detail in S203.
[0123] S504, setting clustering constraints, wherein the clustering constraints include: for any generated first cluster, the first cluster corresponds to a target reference patent set, and the intersection between the reference patent sets corresponding to any two candidate patents in the first cluster is the same as the target reference patent set.
[0124] For example, if the target reference patent set corresponding to a first cluster is a set of m reference patents, then the intersection of the reference patent sets corresponding to any two candidate patents in the first cluster is the m reference patents. The general formula is as follows:
[0125]
[0126] In the above formula, cluster i represents the i-th first cluster, and the target reference patent set corresponding to the first cluster is a set formed by m reference patents. Represents the n candidate patents in the i-th first cluster.
[0127] S505: Cluster the candidate patents in the patent group based on the clustering constraint condition to obtain a first cluster corresponding to the patent group.
[0128] Based on the clustering constraints set above, the candidate patents in the patent group are clustered to obtain a first cluster corresponding to the patent group.
[0129] S506 , combining first clusters in adjacent application years in pairs to obtain a plurality of first cluster point pairs, wherein the first cluster point pairs include two first clusters and the application years corresponding to the two first clusters are adjacent.
[0130] For example, if patent group C 2015 Corresponding to the 100 first clusters, patent group C 2016 Corresponding to the 50 first clusters, the patent group C 2015 The corresponding first cluster and patent group C 2016 The corresponding first clusters are combined in pairs to obtain multiple first cluster point pairs (here for patent group C 2015 and Patent Group C 2016 100×50 first cluster point pairs will be obtained, where each first cluster point pair contains 1 patent group C 2015 The corresponding first cluster and 1 patent group C 2016 The corresponding first cluster.
[0131] Similarly, for the first clusters in adjacent application years, all first cluster point pairs are obtained by combining them in pairs.
[0132] S507, obtaining the target reference patent sets corresponding to the two first clusters in the first cluster point pair, calculating the number of intersection patents and the number of union patents between the target reference patent sets, and calculating the ratio of the number of intersection patents to the number of union patents.
[0133] From the above, it can be seen that each first cluster point pair contains two first clusters, and each first cluster corresponds to a target reference patent set. In this application, for any first cluster point pair, the target reference patent sets corresponding to the two first clusters in the first cluster point pair are obtained, and the number of intersection patents and the number of union patents between the target reference patent sets are calculated, and the ratio of the number of intersection patents to the number of union patents is calculated.
[0134] S508 : Determine a target ratio greater than a preset ratio threshold from the ratios, and connect two first clusters in the first cluster point pair corresponding to the target ratio to obtain a cluster thread.
[0135] Specifically, a ratio threshold is pre-set. A ratio greater than the preset threshold is determined from the ratios corresponding to each first cluster point pair as the target ratio. The two first clusters in the first cluster point pair corresponding to the target ratio are then connected to obtain cluster threads. High overlap indicates a high degree of continuity among the research questions of the candidate patents in these first clusters, which also means that specificity is more likely to occur.
[0136] For example, the general formula of a thread can be expressed as follows:
[0137]
[0138] In the above formula, thread represents a thread, which is formed by connecting k first clusters. The ratio of the first cluster point pairs formed by any two first clusters of adjacent years among the k first clusters is greater than the preset ratio threshold; threshold represents the preset ratio threshold.
[0139] In some embodiments, the preset ratio threshold may be any value between 20% and 60%.
[0140] Preferably, the preset ratio threshold is set to 30%.
[0141] S509: Take each candidate patent as a node, establish edges based on the patent citation information of the candidate patent, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain multiple second clusters.
[0142] S510: Match the first cluster and the second cluster, combine the cluster threads, and determine the target patent.
[0143] Regarding the specific implementation of steps S509 to S510, reference may be made to the detailed introduction of the relevant parts in the above embodiment, which will not be elaborated here.
[0144] This application combines two clustering methods to more accurately identify potential target patents, automatically process large amounts of patent data, and avoid the inefficiency and bias of traditional manual classification. In addition, it can also provide a scientific basis for enterprises in patent strategies, technology research and development, and investment decisions.
[0145] Figure 6 This is a schematic diagram of an identification device of a target patent shown in this application, such as Figure 6 As shown, the target patent identification device 600 includes a first determination module 601, a patent division module 602, a first clustering module 603, a second clustering module 604 and a second determination module 605, wherein:
[0146] The first determining module 601 is used to determine multiple candidate patents and determine patent citation information of each candidate patent.
[0147] The patent division module 602 is used to divide the candidate patents based on their application years to obtain multiple patent groups.
[0148] The first clustering module 603 is used to cluster the candidate patents in any patent group based on the patent citation information, obtain at least one first cluster corresponding to the patent group, and establish a cluster thread based on the first cluster.
[0149] The second clustering module 604 is used to take each candidate patent as a node, establish edges based on the patent citation information of the candidate patent, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain multiple second clusters.
[0150] The second determination module 605 is used to match the first cluster and the second cluster, and determine the target patent by combining the cluster threads.
[0151] By combining two clustering methods, this device can more accurately identify potential target patents and automatically process large amounts of patent data, avoiding the inefficiency and bias of traditional manual classification. In addition, it can also provide a scientific basis for companies in patent strategies, technology research and development, and investment decisions.
[0152] Furthermore, the first clustering module 603 is also used to: determine the reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; calculate the first similarity between the reference patent sets pairwise, and cluster the candidate patents in the patent group according to the first similarity to obtain the first cluster corresponding to the patent group.
[0153] Furthermore, the first clustering module 603 is also used to: combine the first clusters in adjacent application years in pairs to obtain multiple first cluster point pairs, wherein the first cluster point pairs contain two first clusters and the application years corresponding to the two first clusters are adjacent; calculate the second similarity between the two first clusters included in each first cluster point pair; determine a target second similarity greater than a preset similarity threshold from the second similarity, and connect the two first clusters included in the first cluster point pairs corresponding to the target second similarity to obtain a cluster thread.
[0154] Furthermore, the first clustering module 603 is also used to: determine the reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; set clustering constraints, wherein the clustering constraints include: for any first cluster cluster generated, the first cluster cluster corresponds to a target reference patent set, and the intersection between the reference patent sets corresponding to any two candidate patents in the first cluster cluster is the same as the target reference patent set; cluster the candidate patents in the patent group based on the clustering constraints to obtain the first cluster cluster corresponding to the patent group.
[0155] Furthermore, the first clustering module 603 is also used to: combine the first clusters in adjacent application years in pairs to obtain multiple first cluster point pairs, wherein the first cluster point pair contains two first clusters and the application years corresponding to the two first clusters are adjacent; obtain the target reference patent sets corresponding to the two first clusters in the first cluster point pair; calculate the number of intersection patents and the number of union patents between the target reference patent sets, and calculate the ratio of the number of intersection patents to the number of union patents; determine a target ratio greater than a preset ratio threshold from the ratio, and connect the two first clusters in the first cluster point pair corresponding to the target ratio to obtain a cluster thread.
[0156] Furthermore, the second determination module 605 is also used to: for any second cluster cluster, divide the second cluster cluster according to the application year of the candidate patents in the second cluster cluster to obtain multiple second cluster cluster subsets corresponding to the second cluster cluster; obtain the specific potential value corresponding to the second cluster cluster subset; for any second cluster cluster, determine the target second cluster cluster subset from the multiple second cluster cluster subsets corresponding to the second cluster cluster according to the specific potential value; determine the target application year corresponding to the target second cluster cluster subset, and determine the target first cluster cluster corresponding to the target application year from the first cluster cluster; determine the target patent based on the target first cluster cluster, the target second cluster cluster subset and the cluster cluster thread.
[0157] Furthermore, the second determination module 605 is also used to: obtain the number of overlapping candidate patents between the first cluster and the second cluster subset corresponding to the same application year as the second cluster subset; use the second cluster subsets that are in the first N adjacent application years of the second cluster subset and belong to the same second cluster as the second cluster subset as the N associated subsets of the second cluster subset, and obtain the total number of candidate patents contained in the N associated subsets; obtain the difference between the number of overlapping candidate patents and the total number, and use the difference as the specific potential value corresponding to the second cluster subset.
[0158] Furthermore, the second determination module 605 is further configured to: determine overlapping patents between the target first cluster and the target second cluster subset; and select overlapping patents in the starting cluster in the cluster thread as target patents.
[0159] Furthermore, the first clustering module 603 is also used to: determine the initial reference patent set corresponding to each candidate patent in the patent group based on the patent citation information of each candidate patent in the patent group; for any candidate patent, screen out valid initial reference patents from the initial reference patent set corresponding to the candidate patent to form the reference patent set corresponding to the candidate patent.
[0160] In order to implement the above embodiment, the present application also provides an electronic device 700, such as Figure 7 As shown, the electronic device 700 includes: a processor 701 and a memory 702 communicatively connected to the processor, the memory 702 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 701 to implement the target patent identification method as shown in the above embodiment.
[0161] In order to implement the above embodiment, the embodiment of the present application also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to implement the target patent identification method shown in the above embodiment.
[0162] In order to implement the above embodiments, the embodiments of the present application also propose a computer program product, including a computer program, which implements the target patent identification method shown in the above embodiments when executed by a processor.
[0163] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0164] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0165] In the description of this specification, reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application.
[0166] In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0167] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for identifying a target patent, characterized in that: include: determining a plurality of candidate patents, and determining patent citation information for each of the candidate patents; Dividing the candidate patents based on the application years of the candidate patents to obtain multiple patent groups; For any of the patent groups, clustering the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establishing a cluster thread according to the first cluster; Taking each of the candidate patents as a node, establishing edges based on the patent citation information of the candidate patents, constructing a patent knowledge graph, and clustering the patent knowledge graph to obtain a plurality of second clusters; The first cluster and the second cluster are matched, and the target patent is determined by combining the cluster threads.
2. The method according to claim 1, characterized in that The clustering of the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group includes: Determine, according to the patent citation information of each candidate patent in the patent group, a reference patent set corresponding to each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; The first similarities between the reference patent sets are calculated pairwise, and the candidate patents in the patent group are clustered according to the first similarities to obtain a first cluster corresponding to the patent group.
3. The method according to claim 2, characterized in that The step of establishing a cluster thread according to the first cluster comprises: Combining first clusters in adjacent application years in pairs to obtain a plurality of first cluster point pairs, wherein the first cluster point pairs include two first clusters and the application years corresponding to the two first clusters are adjacent; Calculate and obtain a second similarity between two first clusters included in each of the first cluster point pairs; A target second similarity greater than a preset similarity threshold is determined from the second similarities, and two first clusters included in the first cluster point pair corresponding to the target second similarity are connected to obtain a cluster thread.
4. The method according to claim 1, characterized in that: The clustering of the candidate patents in the patent group based on the patent citation information to obtain at least one first cluster corresponding to the patent group includes: Determine, according to the patent citation information of each candidate patent in the patent group, a reference patent set corresponding to each candidate patent in the patent group, wherein the reference patent set includes at least one reference patent; Setting clustering constraints, wherein the clustering constraints include: for any generated first cluster, the first cluster corresponds to a target reference patent set, and the intersection between the reference patent sets corresponding to any two candidate patents in the first cluster is the same as the target reference patent set; The candidate patents in the patent group are clustered based on the clustering constraint conditions to obtain a first cluster corresponding to the patent group.
5. The method according to claim 4, characterized in that The step of establishing a cluster thread according to the first cluster comprises: Combining first clusters in adjacent application years in pairs to obtain a plurality of first cluster point pairs, wherein the first cluster point pairs include two first clusters and the application years corresponding to the two first clusters are adjacent; Obtaining target reference patent sets corresponding to two first clusters in the first cluster point pair respectively; Calculate the number of intersection patents and the number of union patents between the target reference patent sets, and calculate the ratio of the number of intersection patents to the number of union patents; A target ratio greater than a preset ratio threshold is determined from the ratios, and two first clusters in the first cluster point pair corresponding to the target ratio are connected to obtain a cluster thread.
6. The method according to claim 3 or 5, characterized in that: The matching of the first cluster and the second cluster, and combining the cluster threads to determine the target patent, includes: For any of the second clusters, the second clusters are divided according to the application years of the candidate patents in the second clusters to obtain a plurality of second cluster subsets corresponding to the second clusters; Obtaining a specific potential value corresponding to the second cluster subset; For any of the second clusters, determining a target second cluster subset from a plurality of second cluster subsets corresponding to the second cluster according to the specific potential value; Determine a target application year corresponding to the target second cluster subset, and determine a target first cluster corresponding to the target application year from the first clusters; A target patent is determined based on the target first cluster, the target second cluster subset and the cluster thread.
7. The method according to claim 6, characterized in that The obtaining of the specific potential value corresponding to the second cluster subset includes: Obtaining the number of overlapping candidate patents between the first cluster and the second cluster subset corresponding to the same application year as the second cluster subset; The second cluster subsets that are in the first N adjacent application years of the second cluster subset and belong to the same second cluster as the second cluster subset are used as N associated subsets of the second cluster subset, and the total number and sum of the candidate patents included in the N associated subsets are obtained; The difference between the number of overlaps of the candidate patents and the total number is obtained, and the difference is used as the specific potential value corresponding to the second cluster subset.
8. The method according to claim 7, characterized in that The determining of a target patent based on the target first cluster, the target second cluster subset and the cluster thread comprises: Determining overlapping patents between the target first cluster and the target second cluster subset; The overlapping patents in the starting cluster in the cluster thread are screened out as target patents.
9. The method according to claim 2 or 4, characterized in that: Determining the reference patent set corresponding to each candidate patent in the patent group according to the patent citation information of each candidate patent in the patent group includes: Determine, based on the patent citation information of each candidate patent in the patent group, an initial reference patent set corresponding to each candidate patent in the patent group; For any of the candidate patents, valid initial reference patents are screened out from the initial reference patent set corresponding to the candidate patent to form the reference patent set corresponding to the candidate patent.
10. A target patent identification device, characterized in that: include: A first determination module is used to determine a plurality of candidate patents, and to determine patent citation information of each of the candidate patents; A patent division module, used to divide the candidate patents based on the application years of the candidate patents to obtain multiple patent groups; A first clustering module is used for clustering candidate patents in any of the patent groups based on the patent citation information to obtain at least one first cluster corresponding to the patent group, and establishing a cluster thread according to the first cluster; A second clustering module is used to take each of the candidate patents as a node, establish edges based on the patent citation information of the candidate patents, construct a patent knowledge graph, and cluster the patent knowledge graph to obtain a plurality of second clusters; The second determination module is used to match the first cluster and the second cluster, and determine the target patent in combination with the cluster threads.
11. An electronic device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.