A patent document clustering method
A patent document, clustering method technology, applied in special data processing applications, instruments, electronic digital data processing, etc., can solve the problems of missing hidden information, insufficient patent analysis, etc., to achieve the effect of avoiding dimensional disasters
Patent Information
- Authority / Receiving Office
- CN · China
- Current Assignee / Owner
- Publication Date
- 2017-10-17
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention relates to a method for clustering patent document corpus, in particular to a method for clustering patent document. Background technique
[0002] In the current economic environment, patents play an increasingly important role in enhancing corporate value. By applying for a patent, the intellectual property rights of the enterprise can be protected, thereby protecting the core competitiveness of the enterprise. At present, scholars have conducted a lot of research on patent documents, such as labeling patent abstracts, extracting key technologies of patents, and performing cluster analysis on patents.
[0003] In recent years, in the field of data mining, research on text clustering has achieved many results. Many of these methods are based on expressing documents as vectors, and use clustering algorithms to cluster and analyze documents. Patent documents contain a large number of unstructured forms of information, so clustering can b...
Examples
Embodiment
[0037] S1. Corpus collection and preprocessing:
[0038] a1. Corpus collection:
[0039]Select the automotive field, and use crawler technology to crawl patent document information in each category according to the eight categories of patent IPC classification numbers A-H from the "State Intellectual Property Office Patent Database" to form a corpus. Patent document information includes patent titles, IPC classification numbers, and patent abstracts; the patent abstracts of all patent documents in the extracted corpus are stored as word vector training corpus; the patent abstracts of 1,000 patent documents in the extracted corpus are stored as attribute and attribute value model training corpus Set, attribute and attribute value model training corpus contains patent abstracts of eight categories A-H and extracts 125 patent abstracts for each category; extracts patent titles, patent abstracts and IPC classification numbers of 640 patent documents from the corpus and stores them...