Unstructured patent data classification method and device and storage medium
Through the improved K-Means clustering algorithm, combined with the variance maximization strategy and the error sum of squares minimization method, the problem of unstructured data classification in massive high-dimensional patent text data is solved, efficient and accurate patent cluster acquisition is achieved, and more accurate patent technology development status analysis is supported.
Patent Information
- Application Number
- CN202510076514.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
In massive high-dimensional patent text data, how to efficiently and accurately obtain patent development status from unstructured patent data, and then conduct patent evolution research more accurately.
Through the K-Means clustering algorithm improved based on the variance maximization strategy, combined with the error squared minimization and the contour coefficient method, the optimal K value and the best center of quality are determined, thereby achieving efficient classification of unstructured patent data.
This method can effectively avoid structured data such as IPC, retain the substantial content of patents in unstructured patent data, improve clustering effect, improve convergence speed and computing efficiency, and help more accurately explore the technological development status of patents.
Smart Images

Figure CN119988609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patent data classification, and in particular to a method, device and storage medium for classifying unstructured patent data. Background Art
[0002] Patent data is divided into structured data such as patent citation relationship, IPC number, patent owner, and unstructured data such as title, abstract, and claims. Among them, unstructured data contains in-depth information about technological evolution. Therefore, technology evolution analysis based on unstructured patent data can more effectively capture the essential information in patents. In the process of technology evolution analysis, after the patent text is efficiently represented by unstructured patent data feature extraction, how to find patent clusters from the internal information of the patent text is a key issue.
[0003] Cluster analysis is based on a similarity measurement method, which clusters samples with similar characteristics into a cluster, so that the characteristic differences of samples within a cluster are small, while the characteristic differences of samples between clusters are large. Cluster analysis divides samples into different clusters based on the information describing samples and their relationships found in the data. High-dimensional data requires efficient computing methods to ensure that samples within the same cluster are similar to each other and samples between different clusters are different. The greater the similarity within a cluster, the greater the gap between clusters, indicating that the clustering effect is better. Clustering algorithm is one of the keys that affects the clustering effect and can be applied to the screening and classification of patent clusters.
[0004] However, clustering algorithms divide samples into different clusters according to their similarity, and the algorithm categories are divided into hierarchical clustering, partition clustering, density clustering, grid clustering and model clustering according to the characteristics of the clusters, each of which has different characteristics. Faced with massive high-dimensional patent text data, how to efficiently and accurately obtain the substantive content of the patent from the unstructured data of the patent text, and then more accurately mine the technical development status of the patent from the content of the patent text, has become a problem that needs to be solved in this field. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a method, device and storage medium for classifying unstructured patent data. It is implemented based on the K-Means clustering algorithm improved by the variance maximization strategy, which can effectively avoid structured data such as IPC, and more accurately mine the technical development status of the patent from the patent text content, providing support for patent evolution research.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] According to a first aspect of the present invention, there is provided an unstructured patent data classification method based on an improved variance decision tree, comprising the following steps: obtaining all sample sets, wherein the all sample sets include unstructured patent data; based on the all sample sets, obtaining an optimal K value by combining error sum of square minimization with a silhouette coefficient method; based on the optimal K value, obtaining a best quality center by using a decision tree based on variance maximization; based on the optimal K value and the best quality center, obtaining a final patent cluster by using a K-Means clustering algorithm.
[0008] As an optimal technical solution, a decision tree based on variance maximization is used to obtain the best quality center, which specifically includes: constructing a decision tree and generating nodes; when the sample categories in the entire sample set are different, the node is a leaf node, and based on a preset variance maximization strategy, the optimal partitioning attribute and the corresponding mean of the entire sample set are determined; based on the mean, the entire sample set is divided into a first subset and a second subset, the first subset contains samples with attribute values less than the mean, and the second subset contains samples with attribute values greater than the mean; based on the preset variance maximization strategy, the optimal partitioning attribute is determined again for each subset, and recursive splitting is performed; when the number of leaf nodes reaches the optimal K value, the splitting is terminated.
[0009] As a preferred technical solution, during the recursive splitting process, when the preset recursive return condition is met, a recursive return is performed, and the recursive return condition is one of the following two situations: situation one, the current attribute set is empty, or all samples have the same values on all attributes and cannot be divided; situation two, the sample set contained in the current node is empty and cannot be divided.
[0010] As a preferred technical solution, the optimal partitioning attribute is defined as:
[0011]
[0012] In the formula, f * represents the optimal partitioning attribute in the sample set, Indicates that the sample is concentrated The largest attribute, σ 2 is the variance, μ is the mean, s is the feature number, and R is the number of features for each sample.
[0013] As a preferred technical solution, the optimal K value is obtained by combining the minimization of the sum of squared errors with the silhouette coefficient method, specifically including: determining the range of K values; for each K value, calculating and comparing the sum of squared errors of the distances from the points in the corresponding cluster to the center; for each K value, calculating and comparing the corresponding silhouette coefficient values; and comprehensively determining the optimal K value based on the K value corresponding to the maximum decrease in the sum of squared errors and the K value corresponding to the maximum silhouette coefficient.
[0014] As a preferred technical solution, the optimal K value is determined comprehensively, specifically including: when the K value corresponding to the maximum decrease in the sum of squared errors is consistent with the K value corresponding to the maximum silhouette coefficient, the consistent K value is directly used as the optimal K value; when the K value corresponding to the maximum decrease in the sum of squared errors is inconsistent with the K value corresponding to the maximum silhouette coefficient, the comprehensive score is calculated using the balanced weight method, and the K value with the largest comprehensive score is used as the optimal K value.
[0015] As an optimal technical solution, it specifically includes: when the K value corresponding to the maximum decrease in the sum of squared errors is consistent with the K value corresponding to the maximum silhouette coefficient, the consistent K value is directly used as the optimal K value; when the K value corresponding to the maximum decrease in the sum of squared errors is inconsistent with the K value corresponding to the maximum silhouette coefficient, for the two K values, repeated clustering is performed multiple times, the clustering results corresponding to the two K values are compared, and the K value with better overall effect is selected as the optimal K value.
[0016] As a preferred technical solution, the error sum of squares is defined as:
[0017]
[0018] In the formula, D represents the sample, C k represents the kth cluster, D i It is cluster C k The i-th sample in c k is the mean of the samples in the kth cluster, that is, the center of the cluster, and N is the number of samples.
[0019] According to a second aspect of the present invention, there is provided an unstructured patent data classification device based on an improved variance decision tree, comprising a memory, a processor, and a program stored in the memory, wherein the method described is implemented when the processor executes the program.
[0020] According to a third aspect of the present invention, there is provided a storage medium having a program stored thereon, wherein the program implements the method described above when executed.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] 1. The present invention proposes a K-means clustering algorithm based on an improved variance decision tree according to unlabeled tasks, and uses the combination of minimization of the sum of squared errors and the silhouette coefficient method to obtain the optimal K value, and guides the decision tree based on the variance maximization strategy to obtain the best quality center for unlabeled data partition. The improved K-means clustering algorithm is used for the classification of unstructured patent data to obtain patent clusters, and the initial K value and initial mass center can be reasonably determined to improve the convergence speed and clustering effect, thereby efficiently and accurately obtaining the substantive content of the patent from the unstructured data of the patent text, which is helpful to more accurately mine the technical development status of the patent from the content of the patent text;
[0023] 2. The present invention uses the variance maximization strategy to improve the decision tree. The determination of the optimal partition attribute is related to the ratio of the variance to the mean. The partition criterion has a preference for attributes with a large number of possible values. The ratio of the variance to the mean is used to avoid the preference effect caused by the number of attributes. It can effectively avoid structured data such as IPC, retain the substantive content of the patent contained in the unstructured patent data, and improve the final clustering effect.
[0024] 3. When determining the optimal K value, the present invention combines the minimization of the sum of squared errors with the silhouette coefficient method, and uses the balanced weight method or the multiple test comparison method to comprehensively determine the optimal K value, which can further improve the clustering effect and computational efficiency of the improved K-means clustering algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic flow chart of the method provided by the present invention;
[0026] Figure 2 This is a principle diagram of decision tree partitioning based on variance maximization in Example 1 of the present invention;
[0027] Figure 3 The present invention provides an implementation process of the method in embodiment 1. DETAILED DESCRIPTION
[0028] In the face of massive amounts of high-dimensional patent text data, clustering algorithms with lower time and space complexity are more applicable. K-means clustering is a classic partitioning clustering algorithm that measures similarity by the distance between samples and calculates the distance between sample points and cluster centroids, thereby dividing sample points into clusters close to the centroids. In actual clustering work, the final cluster is obtained through rounds of iterations. The steps are roughly based on the initial set number of clusters K and the corresponding k initial centroids. The first round of calculation of the distance between the sample and the centroid and division of the cluster, recalculation of the mean of the new cluster as the new centroid, calculation of the distance between the second round of sample points and the new centroid, and continuous iterative calculation until the centroid of the sample no longer changes, then clustering is completed. Its goal is to minimize the distance within the cluster and maximize the distance between clusters. Since the K-means clustering algorithm follows the principle that the farther the distance between two samples, the lower the similarity, the samples are divided according to certain data features, and finally different clusters are obtained. Under this greedy strategy, the approximate optimal method is found through iteration. Therefore, the K-means clustering algorithm is more suitable for high-dimensional and massive patent text analysis tasks. It can be used to avoid structured data such as IPC, obtain the substantive content of patents from unstructured patent data, and then more accurately mine the technical development status of patents from the content of patent texts. However, the clustering effect of K-means clustering is affected by the initial K value and the initial centroid. In addition, in machine learning, decision trees are mainly used to perform classification tasks. The tree-type decision model constructed according to data attributes is extended from the root node from top to bottom in sequence. The root node contains all samples, and the samples are divided into different leaf nodes as the sample attributes are judged.
[0029] In order to reasonably determine the initial K value and the initial centroid to improve the convergence speed and improve the clustering effect, and to efficiently and accurately obtain the substantive content of the patent from the unstructured data of the patent text, the present invention improves the traditional decision tree and proposes an unstructured patent data classification method based on the variance decision tree improvement. Specifically, according to the unlabeled task, a K-means clustering algorithm based on the variance decision tree improvement is proposed, and the optimal K value is obtained by using the minimum sum of squared errors (SSE) principle, guiding the mean variance decision tree to obtain the best centroid for unlabeled data division, thereby improving the convergence speed and clustering effect.
[0030] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0031] Embodiment 1:
[0032] like Figure 1As shown, the unstructured patent data classification method based on variance decision tree improvement provided by this embodiment includes: obtaining all sample sets including unstructured patent data; based on all sample sets, obtaining the optimal K value by combining error sum of square minimization with silhouette coefficient method; based on the optimal K value, obtaining the best quality center by using decision tree based on variance maximization; based on the optimal K value and the best quality center, obtaining the final patent cluster by using K-Means clustering algorithm. The method is described in detail below:
[0033] 1. Determination of the optimal K value
[0034] K-means clustering sets the optimal K value and the best quality core to improve the convergence speed and clustering effect. The first problem to be overcome is how to determine the optimal K value. The commonly used method to determine the K value of the number of clusters is to compare the average silhouette method (ASM) with SSE to determine the optimal K value.
[0035] ASM is the ratio of the distance between cluster samples (separation) to the distance between cluster samples (cohesion), and is usually used as an evaluation index of clustering effect. The value is [-1, 1]. The larger the value, the better the clustering effect. The cohesion a(D) is shown in formula (1), the separation b(D) is shown in formula (2), and the ASM is shown in formula (3).
[0036]
[0037] In the formula, a(D) represents the cohesion between samples in the cluster, D is the sample in the data set, C k Indicates the cluster where sample D is located, dist(D,D ′ ) represents cluster C k The distance between sample D and other samples in the cluster. The smaller a(D) is, the closer the distance between samples in the cluster is, and the better the clustering effect is.
[0038] In the formula, b(D) represents the separation between clusters, dist(D,D ′ ) reflects cluster C k The minimum distance between sample D and other cluster samples. The smaller the value, the higher the separation.
[0039] In the formula, ASM represents the silhouette coefficient value of the clustering algorithm, D i It is cluster C k The i-th sample in Represents sample D i The larger the ASM value, the more obvious the clustering effect.
[0040] Minimizing the sum of squared errors (SSE) requires calculating the sum of squared errors of the distances from points in the cluster to the center for different clusters K, and recording and comparing them. SSE is defined as:
[0041]
[0042] In the formula, c k is the mean of the samples in the kth cluster, i.e., the center of the cluster, and N is the number of samples. The smaller the SSE value, the better the clustering effect. As the K value increases, the position where the SSE value decreases the most corresponds to the optimal number of clusters.
[0043] After obtaining the SSE values and ASM values corresponding to each K value, the optimal K value is obtained by combining SSE minimization with the ASM method. The specific process is as follows:
[0044] 1) Determine the K value range;
[0045] 2) For each K value, calculate the sum of squared errors of the distances from the corresponding cluster points to the center and compare them to obtain the SSE with the largest decrease, i.e., maxΔSSE;
[0046] 3) For each K value, calculate the corresponding silhouette coefficient value and compare them to obtain the maximum silhouette coefficient;
[0047] 4) The optimal K value is determined comprehensively based on the K value corresponding to the maximum decrease in the sum of squared errors and the K value corresponding to the maximum silhouette coefficient.
[0048] When comprehensively determining the optimal K value, there are two situations:
[0049] Case 1: When the K value corresponding to the maximum decrease in the sum of squared errors is consistent with the K value corresponding to the maximum silhouette coefficient, the consistent K value is directly used as the optimal K value;
[0050] Scenario 2: When the K value corresponding to the maximum decrease in the sum of squared errors is inconsistent with the K value corresponding to the maximum silhouette coefficient, it can be determined by a variety of methods. In some embodiments, the balanced weight method can be used to calculate the comprehensive score, and the K value with the largest comprehensive score is used as the optimal K value. Specifically, different weights are assigned to SSE and ASM, and the comprehensive score is calculated according to the formula Score(K)=w1·Normalized SSE(K)+w2·Normalized ASM(K), and then the K value with the largest comprehensive score is selected as the optimal K value. In the formula, Normalized SSE means that the SSE value is normalized to the [0,1] interval, the smaller the better; Normalized ASM means that the ASM value is normalized to the [0,1] interval, the larger the better; weights w1 and w2 can be set according to demand. This comprehensive judgment method can take into account the compactness and separability of clustering, making the results more balanced. In other embodiments, multiple experiments can be compared, and multiple repeated clustering is performed for the two K values, and the clustering results corresponding to the two K values are compared, and the K value with better overall effect is selected as the optimal K value. This comprehensive judgment method is suitable for processing patent text data with large data volume.
[0051] 2. Determination of the Best Quality
[0052] The second problem that K-means clustering needs to overcome is how to determine the best quality center. In the method proposed in this embodiment, the decision tree based on variance maximization is obtained by improving the basic decision tree using the variance maximization strategy, so as to determine the best quality center according to the optimal K value. The principle is that the key to decision tree theory is how to make sample data better divided, and the division attributes at the high-quality nodes can make the samples in the branch nodes belong to the same category as much as possible. Since many random phenomena are close to Gaussian distribution to a great extent, the variance σ 2 Describes the discreteness of the random variable, σ 2 The larger the variance, the greater the irrelevance between variables. Therefore, this method proposes a decision tree algorithm based on the variance maximization strategy, which is suitable for the decision tree partitioning task of unlabeled text data. The variance maximization partitioning criterion has a preference for attributes with a large number of possible values. 2 The ratio to the mean μ avoids the preference effect caused by the number of attributes.
[0053] The optimal partitioning attribute under the variance maximization strategy is defined as:
[0054]
[0055] In the formula, f * represents the optimal partitioning attribute in the sample set, Indicates that the sample is concentrated The largest attribute, σ 2is the variance, μ is the mean, s is the feature number, and R is the number of features for each sample.
[0056] Optimal partition attribute f * The sample set can be divided into subsets according to the mean μ of the attribute and (i.e., the first subset and the second subset), where, Contains attribute f * Samples with values less than μ, Contains attribute f * Samples with values greater than μ. According to the variance maximization strategy, the optimal partitioning attributes are selected again for the sample subsets and split recursively. There are two conditions for recursive return:
[0057] (1) The current attribute set is empty, or all samples have the same values on all attributes and cannot be divided.
[0058] (2) The sample set contained in the current node is empty and cannot be divided.
[0059] like Figure 2 As shown in the figure, the decision tree based on variance maximization is used to obtain the best quality center. The specific implementation process is as follows:
[0060] 1) Generate Tree;
[0061] 2) Generate node node;
[0062] 3) When all samples in the sample set belong to the same category, the node is a leaf node;
[0063] 4) When the sample categories in all sample sets are different, the optimal partition attribute f of all sample sets is determined based on the preset variance maximization strategy * And the corresponding mean μ, and divide the first subset and the second subset, specifically:
[0064] calculate The maximum The attribute of * , divide the entire sample set into subsets by μ and
[0066] 5) For each subset and The optimal partitioning attribute is determined again using the variance maximization strategy, and recursive splitting is performed until any of the aforementioned recursive return conditions is met;
[0067] 6) When the number of leaf nodes reaches the optimal K value, the splitting is terminated.
[0068] 3. Unstructured patent data classification method based on variance decision tree improvement
[0069] Based on the combination of SSE and ASM to determine the optimal K value and the improved decision tree proposed in this embodiment to determine the best quality center, the unstructured patent data classification method based on the improved variance decision tree proposed in this embodiment can divide the samples into corresponding classes according to the variance maximization strategy. The termination condition of the decision tree splitting is the number of leaf nodes K. The K value used as the termination condition here is the optimal K value determined by combining SSE and ASM, that is, when the decision tree splits to obtain the maximum number of leaf nodes K, the splitting is terminated; then, the average of the samples in the K leaf nodes is the best quality center; finally, the optimal K value and the best quality center are input into K-means clustering to obtain K text clusters, that is, the final patent cluster. One implementation process of this method is as follows: Figure 3 As shown, its essence is the implementation of the K-means clustering algorithm improved based on the variance decision tree. The pseudo code of the algorithm is shown in Algorithm 1.
[0070]
[0071]
[0072] 4. Validation of the Method
[0073] The unstructured patent data classification method based on the improved variance decision tree proposed in this embodiment is essentially the implementation of the K-means clustering algorithm based on the improved variance decision tree: first, the optimal K value of clustering is determined in combination with ASM according to the position where the SSE value of the clustering objective function decreases the most, and then the samples are input into the decision tree, and the samples are divided into leaf nodes according to the attributes in turn according to the variance maximization strategy. The termination splitting condition of the decision tree is that the number of leaf nodes is the optimal K value, and the data in the leaf nodes that meet the splitting condition are averaged, which is the best quality center, and the obtained K and the best quality center are input into K-means clustering to obtain k patent clusters.
[0074] In the sample set, the unstructured data is subjected to feature extraction, and the text features are mapped to numerical vectors. The clustering algorithm proposed in this embodiment is still based on numerical objects. In order to fully verify the effectiveness of the improved clustering algorithm, the experimental data is selected as the data set in the UCI general database, which is the UCI standard experimental data with different six feature dimensions of Pen, Glass, Led Display Domain, Libras Movement, Heart-Cleveland and Aggregation for clustering comparison. The feature distribution of the six data sets is shown in Table 1.
[0075] Table 1 Dataset characteristics
[0076]
[0077] There are missing values in the data set, so preprocessing is performed before training the K-means clustering algorithm model based on variance decision tree. The missing values in the UCI data set are filled with 0. The first column is used as the category label, and the second column and above are the feature items of the data.
[0078] In order to verify that the convergence speed and clustering effect of the method proposed in this embodiment are significant, it is compared with the two most classic initial centroid setting methods of K-means++ and Random in the K-means clustering algorithm. The three algorithm ideas are shown in Table 2.
[0079] Table 2 Comparison of three clustering algorithm ideas
[0080]
[0081] After constructing the algorithm model corresponding to the method proposed in this embodiment (represented by DT-K in the table for convenience of recording), K-means++ centroid selection strategy K-means model and Random centroid selection strategy K-means model, the UCI standard data was used for comparative verification. During verification, the UCI standard data was input into the above three models respectively, and the evaluation results of different experimental data in the model are shown in Table 3.
[0082] There are two types of cluster analysis effectiveness evaluation: one is the effectiveness evaluation based on external criteria, that is, measuring the consistency between clustering results and data category labels; the other is the effectiveness evaluation based on internal criteria, that is, evaluating the clustering effect only from the distribution and morphology of the cluster itself. In Table 3, the external indicators include the comprehensive evaluation index V-measure (i.e., V-meas) that integrates homogeneity and integrity, the Adjusted Rand Index (ARI), the Adjusted Mutual Information (AMI), and Purity; the internal indicators include the silhouette coefficient (Average Silhouette Method, ASM).
[0083] Table 3 Comparison of clustering effects of three clustering algorithms
[0084]
[0085]
[0086] From Table 3, we can see that:
[0087] Compared with K-means++ and random initial centroid selection methods, the Pen dataset is clustered using the method proposed in this embodiment (DT-K), and the four external indicators are improved by 0.242, 0.234, 0.241, and 0.051 respectively, and the internal indicator ASM is improved by 0.051, indicating that the clustering result of the method DT-K proposed in this embodiment on the Pen dataset is consistent with the actual situation. Compared with the other two methods, the Glass dataset is clustered using the method proposed in this embodiment, and the internal indicators of ASM are improved by 0.092 and 0.087 respectively, proving that the method proposed in this embodiment makes the similarity within the cluster greater and the difference between clusters greater. Compared with the other two methods, the Led Display Domain and Libras Movement obtained higher scores in external indicators when clustering using the method proposed in this embodiment, indicating that the cluster results are highly consistent with the actual situation. In addition, the method proposed in this embodiment has the shortest running time in calculating UCI data with different six feature dimensions of Pen, Glass, Led Display Domain, Libras Movement, Heart-Cleveland and Aggregation, which proves that the method proposed in this embodiment uses the improved K-mwans algorithm to effectively improve the convergence speed.
[0088] In summary, compared with the K-means++ and random initial centroid selection methods, the method proposed in this embodiment obtains higher external and internal evaluations of V-measure, ARI, AMI, Purity, and ASM on six data sets selected from the UCI standard database, and consumes the shortest time. The above fully demonstrates that the use of the method proposed in this embodiment to calculate the optimal K value and the best centroid can effectively improve the clustering effect and increase the convergence speed of K-means.
[0089] Embodiment 2:
[0090] This embodiment uses unstructured patent text data to verify the effectiveness of the method proposed in the present invention. Specifically, the sample set uses unstructured patent data. After using a pre-designed feature extraction algorithm to map the unstructured patent data into a low-dimensional feature space and retain the complete text content, the classification method proposed in the previous embodiment is used to construct a decision tree based on the unlabeled characteristics of the patent text content, divide the unstructured patent data into different categories, and calculate the best quality center so that the clustering result is consistent with the actual situation. The specific verification process of this embodiment is basically consistent with the method steps in Example 1, and will not be repeated here.
[0091] The unstructured patent data used in this embodiment is obtained from the existing commercial patent database. Specifically, 2,000 patent data of different parts, categories, and groups A, B, C, and F are exported by IPC number. External evaluation indicators are used to verify the clustering effect. The "part" of the IPC number is manually labeled as "category 1", "category 2", "category 3", and "category 4" and added to the first column of the text-word feature matrix of the pre-constructed text distributed representation. The text content features of the patent are from the second column to the last column. During verification, the features of the unstructured patent data are extracted, and the original data features are mapped to the new feature space and then input into the algorithm model corresponding to the method proposed in this embodiment to obtain the final patent cluster. The clustering effect evaluation is shown in Table 4.
[0092] Table 4 Clustering results of unstructured patent data
[0093]
[0094] The patent text data after feature extraction is input into the algorithm model corresponding to the method proposed in this embodiment. First, the silhouette coefficient and SSE are used to calculate the optimal number of clusters, which is 2. Secondly, the variance and mean ratio of each attribute in the patent text data are calculated according to the variance maximization strategy. The mean square error is used to eliminate the preference effect of the number of attributes. The attribute with the largest ratio is used as the optimal partition attribute. The mean is used as the basis for binary division to build a decision tree. The patent text is divided in turn according to the irrelevance of the attributes. The split is terminated when the maximum leaf node is 2. The average value of the text data in the leaf node is the best quality center. Finally, the best quality center and K value are used to guide K-means clustering to obtain the final two patent clusters. Among the external indicators, the V-measure score is 0.374, the ARI score is 0.212, the AMI score is 0.374, and the Purity score is 0.452, indicating that the clusters obtained by the method proposed in this embodiment are consistent with the actual situation. The internal evaluation index ASM is 0.313, proving that the method proposed in this embodiment makes the similarity within the cluster high and the difference between clusters large. According to statistics, the two patent groups obtained by clustering contain 6281 and 1719 patents respectively, totaling 8000 patent text data. The two patent clusters can be further used for technology evolution analysis in practice.
[0095] Example 3
[0096] This embodiment provides an unstructured patent data classification device based on variance decision tree improvement, including a memory, a processor, and a program stored in the memory, and the processor implements the method in the aforementioned embodiment when executing the program. The device processor includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit to a random access memory (RAM). In RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. Multiple components in the device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks. The processing unit performs the various methods and processes described above, such as one or more steps in the aforementioned embodiments.
[0097] Further, the present embodiment also provides a storage medium on which a program is stored, and the aforementioned method is implemented when the program is executed. The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine or completely on a remote machine or server. In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk-read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for classifying unstructured patent data based on improved variance decision tree, characterized in that: The following steps are involved: Obtaining a complete sample set, wherein the complete sample set includes unstructured patent data; Based on the entire sample set, the optimal K value is obtained by combining the minimization of the sum of squared errors with the silhouette coefficient method; Based on the optimal K value, the best quality center is obtained by using a decision tree based on variance maximization; Based on the optimal K value and the best quality center, the final patent cluster is obtained using the K-Means clustering algorithm.
2. The unstructured patent data classification method based on variance decision tree improvement according to claim 1 is characterized in that: Use decision trees based on variance maximization to obtain the best quality, including: Build a decision tree and generate nodes; When the sample categories in all the sample sets are different, the node is a leaf node, and based on a preset variance maximization strategy, the optimal partitioning attribute and the corresponding mean of all the sample sets are determined; Based on the mean, the entire sample set is divided into a first subset and a second subset, the first subset includes samples whose attribute values are less than the mean, and the second subset includes samples whose attribute values are greater than the mean; Based on the preset variance maximization strategy, the optimal partitioning attributes are determined again for each subset, and recursive splitting is performed; When the number of leaf nodes reaches the optimal K value, the splitting is terminated.
3. The unstructured patent data classification method based on variance decision tree improvement according to claim 2 is characterized in that: During the recursive splitting process, when the preset recursive return condition is met, recursive return is performed, and the recursive return condition is one of the following two situations: Case 1: The current attribute set is empty, or all samples have the same values on all attributes and cannot be divided; Case 2: The sample set contained in the current node is empty and cannot be divided.
4. The unstructured patent data classification method based on variance decision tree improvement according to claim 2 is characterized in that: The optimal partitioning attribute is defined as: In the formula, f * represents the optimal partitioning attribute in the sample set, Indicates that the sample is concentrated The largest attribute, σ 2 is the variance, μ is the mean, s is the feature number, and R is the number of features for each sample.
5. The unstructured patent data classification method based on variance decision tree improvement according to claim 1 is characterized in that: The optimal K value is obtained by combining the minimization of the sum of squared errors with the silhouette coefficient method, including: Determine the K value range; For each K value, calculate the sum of squared errors of the distances from the corresponding points in the cluster to the center and compare them; For each K value, calculate the corresponding silhouette coefficient value and compare; The optimal K value is determined comprehensively based on the K value corresponding to the maximum decrease in the sum of squared errors and the K value corresponding to the maximum silhouette coefficient.
6. The unstructured patent data classification method based on variance decision tree improvement according to claim 5 is characterized in that: Comprehensively determine the optimal K value, including: When the K value corresponding to the maximum decrease in the sum of squared errors is consistent with the K value corresponding to the maximum silhouette coefficient, the consistent K value is directly used as the optimal K value; When the K value corresponding to the maximum decrease in the sum of squared errors is inconsistent with the K value corresponding to the maximum silhouette coefficient, the balanced weight method is used to calculate the comprehensive score, and the K value with the largest comprehensive score is taken as the optimal K value.
7. The unstructured patent data classification method based on variance decision tree improvement according to claim 5 is characterized in that: Comprehensively determine the optimal K value, including: When the K value corresponding to the maximum decrease in the sum of squared errors is consistent with the K value corresponding to the maximum silhouette coefficient, the consistent K value is directly used as the optimal K value; When the K value corresponding to the maximum decrease in the sum of squared errors is inconsistent with the K value corresponding to the maximum silhouette coefficient, repeated clustering is performed for the two K values, and the clustering results corresponding to the two K values are compared. The K value with better overall effect is selected as the optimal K value.
8. The unstructured patent data classification method based on variance decision tree improvement according to claim 5 is characterized in that: The error sum of squares is defined as: In the formula, D represents the sample, C k represents the kth cluster, D i It is cluster C k The i-th sample in c k is the mean of the samples in the kth cluster, that is, the center of the cluster, and N is the number of samples.
9. An unstructured patent data classification device based on variance decision tree improvement, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.