A patent document clustering method

A patent document, clustering method technology, applied in special data processing applications, instruments, electronic digital data processing, etc., can solve the problems of missing hidden information, insufficient patent analysis, etc., to achieve the effect of avoiding dimensional disasters

CN104881401BActive Publication Date: 2017-10-17DALIAN UNIV OF TECH
4 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Current Assignee / Owner
Publication Date
2017-10-17

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

A patent literature clustering method comprises the following steps that S1, a corpus set is collected and preprocessed; S2, clustering analysis is carried out on feature word extraction of corpuses; S3, vector representation of data patients is analyzed based on clustering of term vectors; S4, clustering is carried out; S5, a clustering result is evaluated. The title and summary information of patent literature is comprehensively considered through the patent literature clustering method, the patient summary information is utilized from different angles, overall information of patent summary texts is considered, meanwhile, information of attributes and attribute values in patent summaries are considered, and connotative semantic information in the patent text summaries is fully mined. Information hidden in large-scale corpuses is fully utilized, the large-scale corpuses are utilized for characteristic training, words are expressed in a low-latitude vector form, the curse of dimensionality is avoided, and meanwhile information in the texts is extracted better. Different weights are set, data of three forms are fused through the titles, the summaries and the attribute values of the summaries, and a good patent clustering effect is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to a method for clustering patent document corpus, in particular to a method for clustering patent document. Background technique

[0002] In the current economic environment, patents play an increasingly important role in enhancing corporate value. By applying for a patent, the intellectual property rights of the enterprise can be protected, thereby protecting the core competitiveness of the enterprise. At present, scholars have conducted a lot of research on patent documents, such as labeling patent abstracts, extracting key technologies of patents, and performing cluster analysis on patents.

[0003] In recent years, in the field of data mining, research on text clustering has achieved many results. Many of these methods are based on expressing documents as vectors, and use clustering algorithms to cluster and analyze documents. Patent documents contain a large number of unstructured forms of information, so clustering can b...

Examples

Embodiment

[0037] S1. Corpus collection and preprocessing:

[0038] a1. Corpus collection:

[0039]Select the automotive field, and use crawler technology to crawl patent document information in each category according to the eight categories of patent IPC classification numbers A-H from the "State Intellectual Property Office Patent Database" to form a corpus. Patent document information includes patent titles, IPC classification numbers, and patent abstracts; the patent abstracts of all patent documents in the extracted corpus are stored as word vector training corpus; the patent abstracts of 1,000 patent documents in the extracted corpus are stored as attribute and attribute value model training corpus Set, attribute and attribute value model training corpus contains patent abstracts of eight categories A-H and extracts 125 patent abstracts for each category; extracts patent titles, patent abstracts and IPC classification numbers of 640 patent documents from the corpus and stores them...