Unstructured patent text representation method and device and storage medium
By introducing word distribution state and word intensity in unstructured patent text representation, the weighting method of word vectors is improved, the problems of sparse and dimensional disasters in data processing of unstructured patent text are solved, the distinction and calculation efficiency of text are improved, and the accuracy of technical evolution analysis is improved.
Patent Information
- Application Number
- CN202510076516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has sparse and dimensional disaster problems when processing unstructured patent text data, resulting in low distinction and low computational efficiency, which in turn affects the accuracy and cost of technical evolution analysis.
The unstructured patent text representation method based on word distribution state improvement is adopted. By obtaining the unstructured patent text data set for pre-processing, the word vector is trained using the Word2vec method, the Gini impurity metric word distribution state is calculated, and the word intensity is calculated based on the word distribution state, and the word vector is finally weighted to obtain the distributed representation of unstructured patent text.
It effectively solves the problems of sparse and dimensional disasters, makes the text more distinctive, improves the calculation efficiency of text data, and thus improves the accuracy and efficiency of technical evolution analysis.
Smart Images

Figure CN119990121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of feature extraction technology, and in particular to an unstructured patent text representation method, device and storage medium. Background Art
[0002] Patent data is divided into structured data such as patent citation relationship, IPC number, patent owner, and unstructured data such as title, abstract, and claims. Among them, unstructured data contains in-depth information about technological evolution. Therefore, conducting technological evolution analysis based on unstructured patent data can more effectively capture the essential information in patents.
[0003] In the process of technology evolution analysis, how to effectively extract text features from massive patents is one of the key issues. At present, machine learning has been widely used in the field of data mining. However, most of the existing technologies are based on structured data for data mining, relying on expert experience knowledge or relying on statistical learning. The same machine learning method is directly applied to unstructured data for feature extraction, which is prone to sparsity and dimensionality disaster problems, which reduces the text's discriminability and the computational efficiency of feature extraction, resulting in low accuracy and high cost in the final technology evolution analysis. Therefore, how to make text more discriminative and improve the computational efficiency of text data has become a problem that needs to be solved in this field. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a method, device and storage medium for representing unstructured patent text, which can quickly and efficiently process the massive and high-dimensional information in unstructured patent data, solve the sparsity and dimensionality disaster problems existing in the process of processing text data, make the text more distinctive, and at the same time improve the computational efficiency of text data, thereby helping to achieve accurate and efficient technology evolution analysis.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to a first aspect of the present invention, there is provided an unstructured patent text representation method based on word distribution state improvement, comprising the following steps: obtaining an unstructured patent text dataset and preprocessing it; based on the preprocessed dataset, training word vectors using the Word2vec method; for each word, calculating the word distribution state of the Gini impurity measure, and calculating the corresponding word strength based on the word distribution state; using the word strength to weight the corresponding word vector to obtain the final unstructured patent text distributed representation.
[0007] As a preferred technical solution, the word strength is expressed as:
[0008]
[0009] In the formula, HgIDF(t j ) represents the word strength improved based on the word distribution state, H(t j ) represents the word t in the corresponding text j The information entropy of g(t j ) represents word t j The impurity of N is the total number of texts in the dataset, and n is the number of words containing t j The number of texts.
[0010] As a preferred technical solution, the word distribution state is expressed as:
[0011]
[0012] In the formula, gIDF(t j ) represents word t j The distribution state in the text, g(t j ) represents word t j The impurity of N is the total number of texts, and n is the number of words containing t j The number of texts.
[0013] As a preferred technical solution, the method also includes: calculating word sense similarity and dynamically adjusting the word sense similarity according to the meaning of the word in the context; using the word strength and the word sense similarity to weight the corresponding word vector to obtain the final distributed representation of the unstructured patent text.
[0014] As a preferred technical solution, the Word2vec method uses a preset Skip-gram model to train and obtain word vectors.
[0015] As a preferred technical solution, the Word2vec method uses a preset CBOW model to train and obtain word vectors.
[0016] As a preferred technical solution, the specific process of the preprocessing includes: after removing stop words and punctuation marks, performing word segmentation on the text and building a word library.
[0017] As a preferred technical solution, the final distributed representation of the unstructured patent text is input into five traditional classifiers, namely KNN, LR, Bayes, RF and DT, for classification and discrimination to obtain a discrimination result.
[0018] According to a second aspect of the present invention, there is provided an unstructured patent text representation device based on word distribution state improvement, comprising a memory, a processor, and a program stored in the memory, wherein the processor implements the method described when executing the program.
[0019] According to a third aspect of the present invention, there is provided a storage medium having a program stored thereon, wherein the program implements the method as described above when executed.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] 1. Based on the distributed representation of text, the present invention fully considers the text word strength factor and proposes a feature extraction algorithm based on the improved word distribution state to characterize the word strength of patent text, which can eliminate redundant information, retain high-value information and improve computing efficiency, thereby mapping text features to a new feature space, and obtaining the final distributed representation of unstructured patent text, providing accurate data support for technology evolution analysis;
[0022] 2. The present invention introduces Gini impurity to measure the word distribution state. The proposed Gini impurity does not consider the category label on the basis of the Gini index algorithm, calculates the word impurity from the distribution state of the word in the text, measures the contribution of a word in the text, and expresses the word appropriately in the text representation link to obtain a relatively comprehensive text representation and thus improve the accuracy of text feature extraction. Based on this, the traditional TF-IDF algorithm is improved in combination with information entropy, starting from the word distribution and introducing global considerations. The word strength is reflected from the perspectives of word impurity, chaos, etc. to measure the word representation ability in the text, and the calculated word strength is used to weight the corresponding word vector to improve the distributed representation method, which can solve the sparsity and dimensionality disaster problems existing in the process of processing text data, making the text more distinctive, while improving the calculation efficiency of text data, thereby providing data support for the study of clustering analysis evolution path of unstructured patent texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the process of the present invention;
[0024] Figure 2 A schematic diagram of the implementation flow of the method provided for Example 1 of the present invention. DETAILED DESCRIPTION
[0025] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0026] Embodiment 1:
[0027] Text distributed representation generally uses a representation algorithm model to map text data into a low-dimensional dense vector, avoiding problems such as low computational efficiency caused by sparsity and high dimensionality, while fully retaining text information. Therefore, this embodiment provides an unstructured patent text representation method based on improved word distribution state, which fully considers the text word strength factor on the basis of text distributed representation, proposes a word vector feature extraction algorithm based on improved word distribution state, and characterizes the word strength of patent text. The algorithm has the advantages of eliminating redundant information, retaining high-value information, and improving computational efficiency, thereby mapping text features to a new feature space.
[0028] like Figure 1 As shown, the above method specifically includes: obtaining an unstructured patent text dataset and preprocessing it; based on the preprocessed dataset, training word vectors using the Word2vec method; for each word, calculating the word distribution state of the Gini impurity measure, and calculating the corresponding word strength based on the word distribution state; using the word strength to weight the corresponding word vector to obtain the final distributed representation of the unstructured patent text.
[0029] The word vector feature extraction algorithm based on word distribution improvement provided in this embodiment introduces Gini impurity and information entropy to improve the traditional TF-IDF algorithm. The basic principle of the traditional TF-IDF algorithm is to multiply the high-frequency words (TF) in the text by the low text frequency (IDF) of the word in the entire text collection to obtain the TF-IDF of the word, which reflects whether the word has the ability to distinguish. The algorithm is expressed as:
[0030]
[0031] In the formula, TF-IDF represents the product of the frequency of a word in a text and the inverse text frequency, N represents the total number of texts in the text set, and n represents the number of texts containing the word. From the perspective of text, the higher the frequency of the word in a text, the greater the TF, and if the fewer texts contain the word, the greater the IDF, it means that the word has good distinguishing ability.
[0032] However, traditional algorithms cannot reflect high-value information in the text. The text representation form constructed by the bag-of-words model based solely on word frequency and inverse text is often too simple and cannot accurately reflect the distribution of words in the text. The text feature strength expressed by the distribution of different words in the text is also inconsistent. Therefore, this embodiment further explores the distribution of words in the text from statistical theory and introduces a new word strength measurement method to obtain a more accurate feature extraction algorithm.
[0033] The Gini impurity principle is to calculate the degree of inconsistency between two randomly extracted terms t and t′ in a text, where t represents the target term extracted from the text and t′ represents the non-target term in the two extracted terms. i}, i = 1, 2, ..., N, construct the word library T = {t j}, j = 1, 2, ..., M. The word distribution is defined as formula (2), and the Gini impurity is defined as shown in formula (3).
[0034]
[0035] In the formula, Representation word t j The number of times a word appears in a text. Representation word t j Distribution in this text.
[0036]
[0037] In the formula, g(t j ) represents word t j The impurity, g(t j )The larger the value, the more likely the word t j The higher the distribution impurity in the text, the more likely the word t j Strong, word t j It makes a great contribution in distinguishing the relationships between texts.
[0038] The traditional TF-IDF algorithm only considers the form of word representation in text from the word frequency. The feature extraction algorithm provided in this embodiment makes up for the single disadvantage of the traditional algorithm by introducing the Gini value, and uses the word impurity to measure the distribution state of the word. Combining the basic principle of TF-IDF, that is, formula (1) and the Gini value formula, that is, formula (3), considers introducing Gini impurity to improve TF-IDF weight calculation, that is, starting from local text features and considering global text features. Specifically, after introducing Gini impurity to improve TF-IDF weight calculation, the word t j The distribution state in the text is expressed as:
[0039]
[0040] In the calculation process, the word t j The higher the distribution impurity, the higher the gIDF(t j ) is larger, the word t j The greater the intensity; on the contrary, if the word t j The smaller the impurity, the corresponding gIDF(t j), the smaller the word strength is. In formula (4), the denominator is normalized based on the TFC algorithm to eliminate the influence of text length on text representation. The length of each text is different, and the amount of information in the text is also inconsistent, so as to avoid the text length diluting the importance of the word.
[0041] Word strength is expressed as:
[0042]
[0043] In the formula, HgIDF(t j ) represents the word t improved based on the word distribution state j The strength, H(t j ) represents the word t in the corresponding text j The information entropy is used to measure the uncertainty of information. For the word t j The probability of occurrence.
[0044] In formula (5), both the distribution of words in a single text and the global text are considered. j ) is larger, the more chaotic the word distribution is. j ) represents the word t j The higher the impurity, the higher the HgIDF(t j ) The larger the word t j The greater the strength, the more distinguishing ability it has in the text. Among the feature words of many text representations, this word has great strength and high contribution. This word is selected as the feature word first, and the weight of this word in the text distributed representation method is naturally greater.
[0045] The feature extraction algorithm provided in this embodiment can avoid the singleness of the word frequency weight measurement by introducing the word distribution state to measure the word strength. Whether it is the Gini value or entropy, it only considers the word distribution, and has nothing to do with the meaning expression of the word itself. In the text representation link, in order to obtain text information more comprehensively, as many words as possible are usually retained, but in actual tasks, different words express the same or similar meanings. Relying only on the word distribution state to measure the contribution of words to represent the text is easy to cause dimensional disasters. Therefore, the text representation method of this embodiment uses a neural network to learn word vectors on the basis of proposing an improved word vector feature extraction algorithm based on the word distribution state. Word2vec is one of the more typical deep learning models used in the field of natural language processing, mainly including Skip-gram and CBOW models. The word vector representation method of the CBOW model does not consider grammar and word order, while the word vector constructed by Skip-gram strengthens the semantic similarity. Therefore, this embodiment uses the Skip-gram model to construct the word vector of the patent text information.
[0046] In the Skip-gram model, a probability distribution is defined, that is, given a central word, the probability of other words appearing in the context of this central word. The vector representation of the word is usually selected to maximize the probability distribution value, and for a certain word, there is only one probability distribution, which is the output of the upper and lower words around the central word. The loss function J(Θ) during the iteration process of the model can be expressed as:
[0047]
[0048] In the formula, M represents the total number of words in the vocabulary, j represents the number of the central word in the vocabulary, o represents the window size, [-m, m] is the window size range, P(t (j+o) |t (j) ) represents the probability of the context words appearing around the central word within the window size of o. The smaller the loss function J(θ), the better the training effect. During the training process, the back propagation algorithm is used to adjust the parameter matrix according to the loss function, and finally the loss function is minimized.
[0049] When the Skip-gram model is implemented, the word vector of each word is weighted by the position weight, retaining the contextual semantic relationship information. Therefore, using this model for word vector learning can reduce the word size while improving semantic information.
[0050] The text representation method provided in this embodiment uses the word strength measurement method of the word distribution state to weight the word vector learned by the Skip-gram model of Word2vec, maps the text set features to a new feature space, enhances the word relationship information and introduces the text distribution state to accurately describe the text. The weighted algorithm principle is shown in equations (7) and (8).
[0051]
[0052] In formula (7), gIDF(t j )_weight_Word2vec(t j ) is gIDF(t j )The weighted word t j The mapped Word2vec vector, the molecule represents word t j gIDF(t j ) weighted vector, the denominator is the sum of the weights of all words;
[0053] In formula (8), HgIDF(t j )_weight_Word2vec(t j ) is HgIDF(t j )The weighted word t jThe mapped Word2vec vector, the molecule represents word t j HgIDF(t j ) weighted vector, the denominator is the sum of the weights of all words.
[0054] According to the above argumentation process, the pseudo code of the word vector feature extraction algorithm based on the improved word distribution state is obtained, as shown in Algorithm 1.
[0055]
[0056] Next, the effectiveness of the method provided in this embodiment is verified through experiments.
[0057] The experimental data uses a general data set in the field of natural language processing (NLP). The 10 categories in a news corpus are sports, entertainment, home, real estate, education, fashion, current affairs, games, technology, and finance. There are 500 texts in each category, and a total of 5000 texts are used for experiments. Using five classifiers, K-Nearest Neighbor (KNN), Logistic Regression (LR), Naive Bayes (Bayes), Random Forest (RF), and Decision Tree (DT), the training set and test set of the data are divided and repeated experiments are performed using 100 ten-fold crossovers to verify the effectiveness of the method proposed in this embodiment.
[0058] Figure 2 The specific experimental process is shown:
[0059] First, obtain the sample p = {D1, ..., D i ,…,D n}, i = 1, 2, ..., N, the news text with category labels is preprocessed, stop words and punctuation are eliminated, and the text is segmented using the Jieba word segmentation library to obtain D i ={t1,t2,…,t j}, j = 1, 2, ..., M, construct the vocabulary;
[0060] Secondly, after building the vocabulary, Word2vec is used to train the word vector (V j =(f1,f2,…,f s )), the word distribution state gIDF(t j ) and word strength HgIDF(t j ), and weight the word vectors by equations (7) and (8) respectively, and we can get the text word vector representation:
[0061]
[0062] Among them, HgIDF(t j ) can also be replaced by gIDF(t j ).
[0063] Next, the text distribution representation after the word vector is improved by the word distribution state is obtained by the word strength / word distribution state and word meaning similarity:
[0064]
[0065] Finally, the obtained text vector is input into five traditional classifiers: KNN, LR, Bayes, RF, and DT for classification and discrimination. Specifically:
[0066] In the experiment, the classifier parameter settings are shown in Table 1.
[0067] Table 1 Classifier parameter settings
[0068]
[0069] The effectiveness of the method is verified by comparing the classification accuracy of the method proposed in this embodiment with that of the traditional algorithm in five traditional classifiers. The calculation formula of the accuracy is as follows:
[0070]
[0071] In the formula, TP (True Positive) means predicting the positive class as the positive class; FN (False Negative) means predicting the positive class as the negative class; FP (False Positive) means predicting the negative class as the positive class; TN (True Negative) means predicting the negative class as the negative class.
[0072] Five traditional classifiers, KNN, Logistic Regression, Bayes, Random Forest, and Decision Tree, were used to verify the classification effects of gIDF_weight_Word2vec, HgIDF_weight_Word2vec, and TF-IDF. The accuracy rates obtained by 100 times of ten-fold cross validation are shown in Table 2.
[0073] Table 2 Comparison of classification accuracy between this method and traditional algorithms
[0074]
[0075] It can be seen from Table 2 that the results (HgIDF_weight_w2v, gIDF_weight_w2v) obtained by the method provided in this embodiment are compared with the traditional TF-IDF, and a higher accuracy is obtained in the traditional classifiers of KNN, LR, Bayes, RF, and DT. Specifically, in the KNN classifier, the HgIDF_weight_w2v feature extraction algorithm improves the classification accuracy by 1.61% over the gIDF_weight_w2v and by 65.47% over the TF-IDF. In the LR classifier, the HgIDF_weight_w2v feature extraction algorithm improves the classification accuracy by 0.96% over the gIDF_weight_w2v and by 61.13% over the TF-IDF. In the Bayes classifier, the HgIDF_weight_w2v feature extraction algorithm improves the classification accuracy by 2.04% over the gIDF_weight_w2v and by 68.35% over the TF-IDF. In the Random Forest classifier, the HgIDF_weight_w2v feature extraction algorithm improves the classification accuracy by 0.74% over gIDF_weight_w2v and 53.74% over TF-IDF. In the DT classifier, the HgIDF_weight_w2v feature extraction algorithm improves the classification accuracy by 3.22% over gIDF_weight_w2v and 58.04% over TF-IDF. The experimental results show that the method provided in this embodiment has a higher classification accuracy than the traditional TF-IDF feature extraction algorithm, proving that the method provided in this embodiment has more efficient text representation characteristics.
[0076] Embodiment 2:
[0077] This embodiment uses unstructured patent text data to verify the effectiveness of the method proposed in the present invention. The specific implementation process is basically the same as the steps in Example 1, and will not be repeated here. This embodiment continues to use classifiers such as KNN and Bayes to verify the effectiveness of the method proposed in the present invention. The experimental data uses 1,895 pieces of unstructured patent data related to the scientific and technological research of new power systems derived from a patent database.
[0078] In the unstructured patent data, the title and abstract contain the patent summary content, and the independent claim summarizes the technical content with the concept of legal authority in more detail. Therefore, the unstructured patent data adopted in this embodiment mainly consists of three parts: title, abstract, and independent claim. In order to show the fit of the method proposed in the present invention in the patent field, the same classifier is used based on the classification accuracy and other evaluation criteria, and the IPC number "section" is calculated as the category label. The accuracy experimental results of the method proposed in the present invention are shown in Table 3.
[0079] Table 3 Accuracy
[0080]
[0081] As shown in Table 3, the accuracy of the proposed method in the classification experiment of unstructured patent data using six classifiers is above 90%, and its accuracy is as high as 98% in the LR classifier, and the variance is controlled at 0.03%, indicating stability. The experimental verification fully demonstrates the effectiveness of the proposed method in unstructured patent data, which is a more ideal text representation method.
[0082] Embodiment 3:
[0083] This embodiment provides an unstructured patent text representation device based on word distribution state improvement, including a memory, a processor, and a program stored in the memory, and the processor implements the method in the aforementioned embodiment when executing the program. The processor of the device includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit to a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. Multiple components in the device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks. The processing unit performs the various methods and processes described above, such as one or more steps in the aforementioned embodiment.
[0084] Further, the present embodiment also provides a storage medium on which a program is stored, and the aforementioned method is implemented when the program is executed. The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine or completely on a remote machine or server. In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk-read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0085] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for representing unstructured patent text based on improved word distribution, characterized in that: The following steps are involved: Obtain unstructured patent text datasets and perform preprocessing; Based on the preprocessed data set, the word vector is trained using the Word2vec method; For each word, calculate the word distribution state of the Gini impurity measure, and calculate the corresponding word strength based on the word distribution state; The word strength is used to weight the corresponding word vectors to obtain the final distributed representation of the unstructured patent text.
2. The unstructured patent text characterization method based on word distribution state improvement according to claim 1 is characterized in that: The word strength is expressed as: In the formula, HgIDF(t j ) represents the word strength improved based on the word distribution state, H(t j ) represents the word t in the corresponding text j The information entropy of g(t j ) represents word t j The impurity of N is the total number of texts in the dataset, and n is the number of words containing t j The number of texts.
3. The unstructured patent text characterization method based on word distribution state improvement according to claim 1 is characterized in that: The word distribution state is expressed as: In the formula, gIDF(t j ) represents word t j The distribution state in the text, g(t j ) represents word t j The impurity of N is the total number of texts, and n is the number of words containing t j The number of texts.
4. The unstructured patent text characterization method based on word distribution state improvement according to claim 1 is characterized in that: The method further comprises: Calculating word sense similarity and dynamically adjusting the word sense similarity according to the meaning of the word in the context; The corresponding word vectors are weighted using the word strength and the word meaning similarity to obtain a final distributed representation of the unstructured patent text.
5. The unstructured patent text characterization method based on word distribution state improvement according to claim 1 is characterized in that: The Word2vec method uses a preset Skip-gram model to train word vectors.
6. The unstructured patent text representation method based on word distribution state improvement according to claim 1 is characterized in that: The Word2vec method uses a preset CBOW model to train word vectors.
7. The unstructured patent text characterization method based on word distribution state improvement according to claim 1 is characterized in that: The specific process of the preprocessing includes: after removing stop words and punctuation marks, performing word segmentation on the text and building a word library.
8. The unstructured patent text representation method based on word distribution state improvement according to claim 1 is characterized in that: The final distributed representation of the unstructured patent text is input into five traditional classifiers, namely KNN, LR, Bayes, RF and DT, for classification and discrimination to obtain discrimination results.
9. An unstructured patent text representation device based on word distribution state improvement, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.