An industrial Internet patent identification method based on multi-instance learning

By representing patent data as sentence packets and using multiple example learning methods, the problem of low patent recognition efficiency in industrial Internet in the prior art is solved, and efficient and accurate patent recognition and classification are achieved.

CN114330314BActive Publication Date: 2025-05-02HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111593675.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-05-02
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and classify industrial Internet patents, resulting in low manual review efficiency and low patent analysis efficiency.

Method used

Using a multi-example learning method, the patent data is represented as a sentence package, the sentence similarity is calculated using Jaccard coefficients, and classified prediction is performed in combination with the K nearest neighbor algorithm to achieve efficient identification of industrial Internet patents.

Benefits of technology

It significantly reduces the time of manual review, improves patent analysis efficiency, effectively identifies industrial Internet patents, and improves the accuracy of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330314B_ABST
    Figure CN114330314B_ABST
Patent Text Reader

Abstract

The present invention relates to an industrial Internet patent identification method based on multi-example learning. Natural language processing technology is used to divide the abstract information in the patent into sentences, and a text topic sentence extraction algorithm based on a sentence relationship graph is used to extract the topic sentence in the abstract, which can effectively reduce the computational overhead. At the same time, by combining the topic sentences extracted from the title and the abstract, the patent is converted into a sentence package, where each patent is regarded as a package and each sentence in the package is regarded as an example. Finally, the category of the new sample is predicted by adopting the K Nearest Neighbors (KNN) algorithm. The method of the present invention can effectively improve the industrial Internet patent identification effect, greatly reduce the cost of manual review, and has a very important significance for patent retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial Internet patent identification method based on multi-example learning, which belongs to the technical field of data mining. Background Art

[0002] As a new thing born from the integration of my country's new generation of network information technology and modern industry, the Industrial Internet is a key support for realizing all factors, the entire industrial chain, and the entire value chain in the manufacturing and production fields. The Industrial Internet is an important infrastructure for promoting the digitalization, networking, and intelligent development of my country's industrial economy. It is also the core carrier for Internet technology to cross from the consumer field to the production field and from the virtual economy to the real economy. Compared with the traditional Internet industry, the development of the Industrial Internet has more industrial characteristics such as high R&D investment and long payback cycle. This has caused most domestic industrial Internet software companies to be less active in R&D, and the technology investment in the short term is limited. R&D costs and other problems. Therefore, how to correctly guide the development of the Industrial Internet industry from the perspective of patent layout has become an imminent issue.

[0003] Carrying out patent navigation work can not only give full play to the role of the patent system in allocating industrial innovation resources, but also give play to the guiding role of patent information analysis in industrial innovation decision-making, thereby further improving the industry's innovation capabilities, avoiding intellectual property risks and improving the overall competitiveness of the industry. However, in order to smoothly carry out patent navigation work, data identification is particularly important. Industrial Internet patent navigation requires a large amount of data support, which requires us to identify relevant patent data by selecting the main classification number, title, abstract and other factors that affect the patent category as basic features, and then establish a dedicated database to store these patent data. Industrial Internet patent data recognition is essentially a text classification problem. Good recognition effect can greatly reduce manual review time and improve patent analysis efficiency.

[0004] At present, more and more scholars are investing in the research of Chinese text classification. After the research of many scholars, various Chinese text classification methods based on machine learning have emerged one after another. Many deep learning models such as convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), etc. have also been applied to text classification. However, most of the current research on Chinese patent text classification focuses on the classification of patent classification numbers. In a certain sub-field, especially in the field of industrial Internet, no scholar has proposed a relevant and feasible classification scheme for patent identification. In recent years, the industrial Internet industry has been developing in full swing, and a large number of innovative patent technologies have emerged. Industrial Internet patent identification has a large research space and application prospects. It is particularly important to use effective algorithms to identify industrial Internet patents. Summary of the invention

[0005] The purpose of the present invention is to address the above-mentioned problems and provide an industrial Internet patent identification method based on multi-instance learning, which can greatly reduce the efficiency of manual review and effectively identify industrial Internet patent data.

[0006] To achieve the above object, the technical solution of the present invention is:

[0007] An industrial Internet patent identification method based on multi-instance learning includes the following steps:

[0008] Step 1: Obtain patent data P from the dataset = (P1, P2, ..., P n ), n is the number of patents, and each sample (i.e., patent) is represented by P i =(id,pnun,title,abstract), where id represents the patent number, pnun represents the patent application number, title represents the patent title, and abstract represents the patent abstract;

[0009] Step 2: Data filtering: There are a large number of duplicate patents in the data set. These patents are caused by the applicants' first applications that were not approved due to patent quality issues and then re-applied after modification. Patent data is deduplicated by using the patent application number pnum. That is, if multiple patents have the same patent application number, only one patent related to the patent application number is retained.

[0010] Step 3: Sentence segmentation: For the abstract text content in the patent, divide it into sentences according to ".", ";", "!", "?", etc., and represent the abstract as a set of sentences abstract = (s1, s2, ..., s m ), where m represents the number of sentences contained in abstract;

[0011] Step 4: Data preprocessing: First, use the Chinese online word segmentation tool LTP (Language Technology Platform) to segment the text content in the patent; then, delete the noise information such as numbers and punctuation contained in the text content; finally, use the Chinese stop word list to remove the stop words contained in the text content; after preprocessing, each sample is represented as P i =<id,pnum,preTitle,preAbstract> , where preTitle and preAbstract represent the preprocessed title information and abstract information respectively;

[0012] Step 5: Sentence similarity calculation: As a short text, a sentence contains limited words and is not suitable for similarity calculation methods for long texts, such as the space vector model. The present invention uses the Jaccard coefficient to calculate the similarity, which represents the sample as a bag-of-words model, and then measures the similarity based on the ratio of the number of identical words contained in the two samples to the number of all different words.

[0013] Step 6: Abstract topic sentence extraction: In a text, there may be many sentences with repeated meanings in order to describe a topic. The sentence that best represents the content of the text should be selected, i.e., the topic sentence. The present invention uses a text topic sentence extraction algorithm based on a sentence relationship graph to extract the topic sentence. j (j=1,2,…,m), if there are many sentences with a large similarity, then the sentence is more likely to be selected as the topic sentence; finally, the set of topic sentences topSen={s1,s2,…s m'}, m' is the number of topic sentences finally selected;

[0014] Step 7: Bag of Sentences Representation: For each patent P i , borrowing the concept of bag in multi-instance learning theory, using sentence bag to represent patents; treating title as a separate topic sentence, merging it with the topic sentence in topSen, and combining patent P i Expressed as Thus, patent P i As a package, P i Each sentence in is considered as an example, l i For package P i The number of examples in ;

[0015] Step 8: Inter-package similarity calculation: For any two sentence packages P a and P b , l a and l b Respectively represent the package P a and P b The number of examples in P a and P b The similarity is:

[0016]

[0017] and

[0018]

[0019]

[0020] Among them, Sim(P a ,P b ) indicates Pa and P b similarity;

[0021] Step 9: Training set division: Divide the data set into a positive sample set and a negative sample set, with industrial Internet patents as positive samples and non-industrial Internet patents as negative samples; adopt the 10-fold cross-validation method, that is, divide the positive and negative samples into 10 folds, and merge the positive and negative samples of each fold. The merged 9 folds are used as the training set and 1 fold is used as the test set;

[0022] Step 10: Classification prediction: For the new patent sample P u First, the title and abstract are preprocessed, and then the topic sentence is extracted from the preprocessed abstract. The title is regarded as a separate topic sentence, and the patent P u Represented as a sentence bag form; K Nearest Neighbors (KNN) algorithm is used to predict P u Category.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] The present invention relates to an industrial Internet patent identification method based on multi-instance learning, and constructs an industrial Internet patent classification method based on multi-instance learning, which can effectively identify industrial Internet patents and improve the accuracy of classification. In the present invention, the patent is represented in the form of a sentence package, and each sentence is regarded as an example, which can effectively avoid the ambiguity problem of each example category in the positive package. The present invention uses a text topic sentence extraction algorithm based on a sentence relationship graph to extract the topic sentence, which effectively overcomes the high computational cost problem brought by long text. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0026] Figure 1 This is a flow chart of the industrial Internet patent identification method based on multi-instance learning of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] Data source acquisition: The patent data involved in this experiment was obtained through the Innojoy patent search engine, and at the same time, it was obtained through the Thomson Innovation patent search and analysis system, Innography patent search and analysis system, SooPAT patent search and analysis system and other platforms developed by Thomson Reuters. The search period was from January 1, 2000 to September 30, 2020. Through keyword search, 69,658 patent data were finally obtained. For the collected data, 3 graduate students were invited to annotate them. For each patent sample, each graduate student annotated it, that is, to determine whether it is an industrial Internet patent. If the three annotation results of a patent are consistent, the annotation result of the patent is considered acceptable; otherwise, the patent needs to be submitted to experts for decision-making. Experts review all patents with inconsistent annotation results and determine their final labels, that is, whether they are industrial Internet patents. Finally, 24,880 pieces of industrial Internet patent data were obtained.

[0029] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the following Figure 1 , the industrial Internet patent identification method based on multi-instance learning provided by the patent of this invention is described in detail, including the following steps:

[0030] Step 1: Obtain patent data P from the dataset = (P1, P2, ..., P n ), n is the number of patents, and each sample (i.e., patent) is represented by P i =(id,pnun,title,abstract), where id represents the patent number, pnun represents the patent application number, title represents the patent title, and abstract represents the patent abstract;

[0031] Step 2: Data filtering: De-duplicate patent data by using the patent application number pnum. That is, if multiple patents have the same patent application number, only one patent related to the patent application number is retained;

[0032] Step 3: Sentence segmentation: For the abstract text content in the patent, divide it into sentences according to ".", ";", "!", "?", etc., and represent the abstract as a set of sentences abstract = (s1, s2, ..., s m), where m represents the number of sentences contained in abstract;

[0033] Step 4: Data preprocessing: First, use the Chinese online word segmentation tool LTP (Language Technology Platform) to segment the text content in the patent; then, delete the noise information such as numbers and punctuation contained in the text content; finally, use the Chinese stop word list to remove the stop words contained in the text content; after preprocessing, each sample is represented as P i =<id,pnum,preTitle,preAbstract> , where preTitle and preAbstract represent the preprocessed title information and abstract information respectively;

[0034] Step 5: Sentence similarity calculation: After preprocessing, each sentence in the abstract is actually a set of words; for any two sentences s in the abstract a and b , the Jaccard coefficient is used to calculate the sentence similarity, the formula is as follows:

[0035]

[0036] Where || represents the number of words contained in the set;

[0037] Step 6: Abstract topic sentence extraction: for patent P i In the abstract, the text topic sentence extraction algorithm based on the sentence relationship graph is used to extract the topic sentence. The detailed steps are as follows:

[0038] 6-1. For abstract=(s1,s2,…,s m ), calculate the similarity Sim(s) between all sentences based on the Jaccard coefficient j ,s k ), where j≠k and j,k=1,2,…,m, construct the similarity matrix X m×m ;

[0039] 6-2. Set the similarity threshold δ1 (δ1 = 0.3) and construct the matrix Y m×m , where each element Y jk (Y jk ∈Y m×m ) is:

[0040]

[0041] 6-3. Construct a row vector Z 1×m , corresponding to the component Z j(j=1,2,…,m) represents sentence s j The importance of Y m×m Each row of corresponds to; Z j The value is Y m×m The number of elements in the jth row whose values ​​are greater than 0. The larger the value, the wider the content covered by the corresponding sentence.

[0042] 6-4. Sort the sentences by importance from large to small, select the threshold δ2 (δ2 = 0.5) and the compression ratio R (R = 0.2), initialize the set topSen = {}, and process each sentence s in turn according to the importance of the sentence j :

[0043] 6-5. If the sentence to be processed is s j If it is not marked as "processed", j Put it into topSen and mark it as "processed", and scan the matrix Y at the same time m×m In row j, if Y jk ≥δ2, then sentence s k Also marked as "processed"; if the sentence to be processed is s j If it has been marked as "processed", no action will be taken;

[0044] 6-6. Repeat step 6-5 until there are no sentences to be processed or the number of sentences contained in topSen has reached R×m;

[0045] 6-7. Finally, we get the abstract topic sentence set topSen = {s1, s2, ...s m'}, m' is the number of topic sentences finally selected;

[0046] Step 7: Bag of Sentences Representation: For each patent P i , borrowing the concept of bag in multi-instance learning theory, using sentence bag to represent patents; treating title as a separate topic sentence, merging it with the topic sentence in topSen, and combining patent P i Expressed as Thus, patent P i As a package, P i Each sentence in is considered as an example, l i For package P i The number of examples in ;

[0047] Step 8: Inter-package similarity calculation: For any two sentence packages P a and P b , l a and l b Respectively represent the package P a and P b The number of examples in Pa and P b The similarity is:

[0048]

[0049] and

[0050]

[0051]

[0052] Among them, Sim(P a ,P b ) indicates P a and P b similarity;

[0053] Step 9: Training set division: Divide the data set into a positive sample set and a negative sample set, with industrial Internet patents as positive samples and non-industrial Internet patents as negative samples; adopt the 10-fold cross-validation method, that is, divide the positive and negative samples into 10 folds, and merge the positive and negative samples of each fold. The merged 9 folds are used as the training set and 1 fold is used as the test set;

[0054] Step 10: Classification prediction: For the new patent sample P u First, preprocess the title and abstract, then extract the topic sentence from the preprocessed abstract, and treat the title as a separate topic sentence; use the K Nearest Neighbors (KNN) algorithm to predict the category of the new sample. The steps are as follows:

[0055] 10-1. Represent the patent to be classified as a sentence package l u For package P u The number of examples in ;

[0056] 10-2. Calculate the patent P to be classified u The similarity between each sample in the training set;

[0057] 10-3. From the training set, the patent P to be classified u Select K samples with the largest similarity to form a set CadSet = {P1, P2, ..., P K};

[0058] 10-4. Count the patents to be classified u Relative to each category c d The weight is calculated as follows:

[0059]

[0060] Where w(P u,c d ) indicates P u Belongs to class c d The weight of y(P v ,c d ) is the category attribute function, that is, if P v Belongs to class c d , then its value is 1, otherwise it is 0;

[0061] 10-5. Finally, the class with the largest class weight is the class to which the sample to be classified belongs. The above describes the implementation of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described implementation. For those skilled in the art, various changes, modifications, substitutions and variations of these implementations without departing from the principles and spirit of the present invention still fall within the scope of protection of the present invention.

Claims

1. An industrial Internet patent identification method based on multi-instance learning, characterized by: The following steps are involved: Step 1: Obtain patent data P from the dataset = (P1, P2, ..., P n ), n is the number of patents, and each patent sample is represented as P i =(id,pnun,title,abstract), Where id represents the patent number, pnun represents the patent application number, title represents the patent title, and abstract represents the patent abstract; Step 2: Data filtering: De-duplicate patent data by patent application number pnum, and only keep patents related to one patent application number; Step 3: Sentence segmentation: For the abstract text content in the patent, divide it into sentences according to ".", ";", "!", "?", etc., and represent the abstract as a set of sentences abstract = (s1, s2, ..., s m ), where m represents the number of sentences contained in abstract; Step 4: Data preprocessing: Use the online word segmentation tool LTP to segment the text content in the patent; then delete the noise information contained in the text content; then use the stop word list to remove the stop words contained in the text content; after preprocessing, each sample is represented as P i =<id,pnum,preTitle,preAbstract> , where preTitle and preAbstract represent the preprocessed title information and abstract information respectively; Step 5: Sentence similarity calculation: Use the Jaccard coefficient to calculate the similarity, represent the sample as a bag-of-words model, and measure the similarity based on the ratio of the number of identical words in the two samples to the number of all different words; Step 6: Abstract topic sentence extraction: Use the text topic sentence extraction algorithm based on sentence relationship graph to extract the topic sentence. j (j=1,2,…,m), get the set of topic sentences topSen= {s1,s2,…s m' }, m' is the number of topic sentences finally selected; Step 7: Bag of Sentences Representation: For each patent P i , using bags of sentences to represent patents; Treat title as a separate topic sentence and merge it with the topic sentence in topSen. i Expressed as Thus, patent P i As a package, P i Each sentence in is considered as an example, i For package P i The number of examples in ; Step 8: Inter-package similarity calculation: For any two sentence packages P a and P b , l a and l b Respectively represent the package P a and P b The number of examples in P a and P b The similarity is: and Among them, Sim(P a ,P b ) indicates P a and P b similarity; Step 9: Training set division: Divide the data set into positive sample sets and negative sample sets, where industrial Internet patents are used as positive samples and non-industrial Internet patents are used as negative samples: Step 10: Classification prediction: For the new patent sample P u , preprocess the title and abstract, then extract the topic sentence from the preprocessed abstract, treat the title as a separate topic sentence, and extract the topic sentence from the patent P u Represented as a bag of sentences; Use K nearest neighbor algorithm to predict P u Category of; The step six comprises the following steps: Step (6-1): For abstract=(s1,s2,…,s m ), calculate the similarity Sim(s) between all sentences based on the Jaccard coefficient j ,s k ), where j≠k and j,k=1,2,…,m, construct the similarity matrix X m×m ; Step (6-2): Set the similarity threshold δ1 (δ1 = 0.3) and construct the matrix Y m×m , where each element Y jk (Y jk ∈Y m×m ) is: Step (6-3): Construct a row vector Z 1×m , corresponding to the component Z j (j=1,2,…,m) represents sentence s j The importance of Y m×m Each row of corresponds to; Z j The value is Y m×m The number of elements in the j-th row whose value is greater than 0; Step (6-4): Sort the sentences in descending order of importance, select the threshold δ2 (δ2 = 0.5) and the compression ratio R (R = 0.2), initialize the set topSen = {}, and process each sentence s in turn according to the importance of the sentence j ; Step (6-5): If the sentence to be processed is s j If it is not marked as "processed", j Put it into topSen and mark it as "processed", and scan the matrix Y at the same time m×m In row j, if Y jk ≥δ2, then sentence s k Also marked as "processed"; if the sentence to be processed is s j If it has been marked as "processed", no action will be taken; Step (6-6): Repeat step 6-5 until there are no sentences to be processed or the number of sentences contained in topSen has reached R×m; Step (6-7): Get the abstract topic sentence set topSen = {s1, s2, ...s m' }, m' is the number of topic sentences finally selected.

2. According to the method for industrial Internet patent identification based on multi-instance learning according to claim 1, it is characterized by: The noise information in step 4 includes numbers and punctuation marks.

3. The method for industrial Internet patent identification based on multi-instance learning according to claim 1 or 2 is characterized in that: In step 5, for any two sentences s in abstract a and b , the Jaccard coefficient is used to calculate the sentence similarity, the formula is as follows: Here || represents the number of words in the set.

4. According to the method for industrial Internet patent identification based on multi-instance learning according to claim 1, it is characterized by: The step ten comprises the following steps: Step (10-1): Represent the patent to be classified as a sentence bag l u For package P u The number of examples in ; Step (10-2): Calculate the patents to be classified P u The similarity between each sample in the training set; Step (10-3): Select the patent P to be classified from the training set u Select K samples with the largest similarity to form a set CadSet = {P1, P2, ..., P K }; Step (10-4): Count the patents to be classified u Relative to each category c d The weight is calculated as follows: Where w(P u ,c d ) indicates P u Belongs to class c d The weight of y(P v ,c d ) is the category attribute function, that is, if P v Belongs to class c d , then its value is 1, otherwise it is 0; Step (10-5): The class with the largest class weight is the class to which the sample to be classified belongs.

Citation Information

Patent Citations

  • Abstract generation method based on single long text

    CN111858912A

  • Semi-automatic construction method of Chinese patent corpus based on TRIZ

    CN112487192A