An academic keyword batch recognition system
By constructing an academic vocabulary list and a deep semantic model, combining statistical features and semantic similarity calculations, the problems of low batch processing efficiency and noisy words in the existing technology are solved, and efficient and accurate extraction of academic keywords is achieved.
Patent Information
- Application Number
- CN202211119575.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-15
AI Technical Summary
The existing keyword extraction system cannot be batch processed, has low efficiency and a large number of noisy words, which cannot meet the needs of fast and accurate extraction of large-scale academic paper data.
Academic vocabulary list is constructed, combined with N-Gram word frequency, point mutual information, left and right entropy and time influence factors to calculate word formation probability, and uses TF-IDF and deep semantic models to sort keywords, and realize batch recognition through word segmentation module, keyword rough batch processing module and fine batch processing module.
It improves the accuracy and efficiency of keyword extraction, filters out candidate keywords that are not related to semantics, avoids the emergence of unlogged words, makes full use of the characteristics of academic documents, and improves the accuracy and generalization ability of academic keyword extraction.
Smart Images

Figure CN115392244B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to an academic keyword batch recognition system. Background Art
[0002] Today, with the highly developed Internet, the acquisition of data has become very convenient. However, for the huge amount of paper data resources, when users search for articles, they often need to view papers among a large number of retrieval results, consuming time and energy. And batch and accurately extracting keywords of a large number of papers during retrieval is beneficial for users to quickly understand the relevance of papers and improve the user's retrieval experience. Existing keyword extraction algorithms include TF-IDF, TextRank[1], Yake[2], AutoPhrase[3], KeyBert, etc. But existing algorithms often have a large number of noise words when extracting keywords, misidentifying non-keywords as keywords, and the calculation efficiency is very low, which is not suitable for batch keyword extraction. Therefore, for the academic fields such as a huge number of paper literatures, there is an urgent need for a keyword batch recognition system with high efficiency and strong accuracy to achieve batch, fast, and accurate extraction of academic text keywords. Summary of the Invention
[0003] In view of the above analysis, the present invention aims to provide an academic keyword batch recognition system; it solves the problems that the existing keyword extraction system cannot perform batch processing, has low efficiency, and has noise words.
[0004] The object of the present invention is mainly achieved by the following technical solutions:
[0005] The present invention discloses an academic keyword batch recognition system, including a word segmentation module, a keyword rough sorting batch processing module, and a keyword fine sorting batch processing module; wherein,
[0006] The word segmentation module includes a word list construction unit and a keyword recall unit; the word list construction unit is used to perform word frequency statistics on the titles and abstracts of all papers to be recognized and calculate the word formation probability, and construct an academic word list according to the word formation probability; the keyword recall unit is used to segment the titles and abstracts of all papers to be recognized based on the academic word list to obtain the recalled keywords of each paper;
[0007] The keyword rough sorting batch processing module is used to perform batch sorting processing on the recalled keywords of all papers to be recognized according to the statistical scores of the recalled keywords to obtain the candidate keywords corresponding to each paper;
[0008] The keyword fine-ranking batch processing module is used to batch rank the semantic similarities between the candidate keywords corresponding to all papers to be recognized, and the corresponding titles and abstracts, and obtain the academic keywords corresponding to each paper based on the semantic similarities; the semantic similarities are calculated by a pre-trained deep semantic model.
[0009] Further, constructing the academic vocabulary includes:
[0010] Constructing a paper corpus, where the paper corpus includes the titles and corresponding abstracts of all papers to be recognized;
[0011] Performing word frequency statistics on the titles and abstracts of the papers in the paper corpus;
[0012] Calculating the word formation probability of each word based on N-Gram word frequency, point mutual information, left and right entropy, and time influence factor, and selecting the words with word formation probability greater than the probability threshold to construct the academic vocabulary.
[0013] Further, the time influence factor is calculated based on the average time span between the publication time of the papers containing the word and the first appearance time of the word, and the calculation formula is:
[0014]
[0015] where n represents the number of papers containing the word x, t i represents the publication year of the i-th paper containing the word x, and t v represents the publication time of the paper in which the word x first appears in the paper corpus.
[0016] Further, the value of the point mutual information PMI(x,y) is calculated by the following formula:
[0017]
[0018] where x and y are respectively words or characters obtained by the N-Gram algorithm, p(x,y) is the probability that the combination of x and y appears in the text, p(x) is the probability that x appears in the text, and p(y) is the probability that y appears in the text.
[0019] Further, the word formation probability is calculated by the following formula:
[0020]
[0021] where |D| represents the total number of papers in the paper corpus, |{d∈D:x∈d}| represents the number of papers in the paper corpus containing the word x, freq(x) represents the probability that the word x appears in the paper corpus, PMI(x) represents the point mutual information of the word x, and H l (x) represents the left neighbor word information entropy, and Hr(x) Represents the information entropy of the right adjacent word.
[0022] Furthermore, input the titles of the papers in the paper corpus, the abstracts corresponding to the titles, and n randomly selected abstracts from the paper corpus into the dual-tower structure model of DSSM, calculate the similarity between the title and the abstract of the paper, and through iterative update of the loss function, maximize the semantic similarity between the title and the abstract corresponding to the title to train a deep semantic model; n is an integer greater than 1.
[0023] Furthermore, the deep semantic model trained by using the dual-tower structure model of DSSM includes an input layer, a representation layer, and a matching layer;
[0024] The input layer uses the N-Gram model to respectively reduce the dimensions of the input title and abstract to obtain low-dimensional semantic vectors after dimensionality reduction and compression;
[0025] The representation layer includes three fully connected layers, and each layer is activated using a non-linear activation function to perform feature integration on the low-dimensional semantic vectors to obtain representation layer hidden vectors with a fixed dimension;
[0026] The matching layer calculates the semantic similarity between the title and the abstract based on the representation layer hidden vectors.
[0027] Furthermore, perform keyword rough sorting on the paper to be recognized, and the obtained candidate keywords include:
[0028] Segment the title and abstract of the paper to be recognized based on the academic vocabulary list; calculate the comprehensive score of each word according to the word length, word position, and TF-IDF score of each word obtained after segmentation; obtain candidate keywords based on the comprehensive score;
[0029] Among them, the comprehensive score of each word is calculated through the following formula:
[0030] Score(x i ) = Length(x i ) · Position(x i ) · tfidf(x i )
[0031] Among them, i represents the index value of the word, Length(x i ) represents the length of the word x i , Position(x i ) represents the position score of x i , and tfidf(x i ) represents the TF-IDF score.
[0032] Furthermore, the TF-IDF weight statistical score is calculated through the following formula:
[0033]
[0034]
[0035] tfidf(t, d, D) = tf(t, d)·idf(t, D);
[0036] where t is the word obtained through N - Gram processing, d is the paper to be processed where the word t is located, t' is any word included in the paper d, f t,d is the frequency of the word t appearing in the paper d, f t′,d is the frequency of any word included in the paper d appearing in the paper d, D is the dataset of papers to be recognized in the paper corpus, |{d ∈ D: t ∈ d}| represents the number of documents in the paper corpus that contain the word t, tf(t, d) represents the term frequency, idf(t, D) represents the inverse document frequency, and tfidf(t, d, D) represents the TF - IDF score.
[0037] Furthermore, a position score is calculated according to the position of the word in the title and abstract, and the calculation formula of the position score is:
[0038]
[0039] where i represents the index value of the word.
[0040] Advantages of this technical solution:
[0041] 1. The academic keyword batch recognition system of the present invention performs word segmentation using the academic word list constructed from the papers to be batch - processed. By combining the keyword extraction method of statistics and semantic calculation, it filters out the semantically irrelevant candidate keywords sorted based on statistical features and outputs the final keywords, greatly improving the accuracy of keyword extraction.
[0042] 2. When constructing the academic word list, the influences of N - Gram word frequency, point - wise mutual information (PMI), time influence factor, and left - right entropy (Entropy) are taken into account, improving the accuracy of word segmentation in the text pre - processing stage.
[0043] 3. In the keyword rough ranking stage of the present invention, considering the word length, the position of the word, and the TF - IDF score of the keyword, the candidate keywords are comprehensively ranked, and the words higher than the threshold are returned as candidate keywords. A large number of non - keywords are filtered in the keyword rough ranking stage, improving the efficiency and precision of keyword extraction.
[0044] 4. The present invention directly constructs an academic vocabulary using the papers to be recognized that need to be processed in batches, and trains a deep semantic model through the titles and corresponding abstracts of the same batch of papers to be recognized, avoiding the appearance of out-of-vocabulary words, making full use of the characteristics of academic documents, and improving the accuracy of academic keyword extraction.
[0045] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0046] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0047] Figure 1 It is a block diagram of an academic keyword batch recognition system according to an embodiment of the present invention;
[0048] Figure 2 It is a schematic structural diagram of a semantic model according to an embodiment of the present invention; Detailed Embodiments
[0049] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.
[0050] An academic keyword batch recognition system in this embodiment, as Figure 1 shown, includes a word segmentation module, a keyword rough sorting batch processing module, and a keyword fine sorting batch processing module;
[0051] The word segmentation module includes a vocabulary building unit and a keyword recall unit; the vocabulary building unit is used to perform word frequency statistics on the titles and abstracts of all papers to be recognized and calculate the word formation probability, and construct an academic vocabulary based on the word formation probability; the keyword recall unit is used to perform word segmentation on the titles and abstracts of all papers to be recognized based on the academic vocabulary to obtain the recalled keywords of each paper;
[0052] The keyword rough sorting batch processing module is used to perform batch sorting processing on the recalled keywords of all papers to be recognized according to the statistical scores of the recalled keywords to obtain the candidate keywords corresponding to each paper;
[0053] The keyword fine sorting batch processing module is used to perform batch sorting processing on the semantic similarity between the candidate keywords corresponding to all papers to be recognized and the corresponding titles and abstracts, and obtain the academic keywords corresponding to each paper based on the semantic similarity; the semantic similarity is calculated by a pre-trained deep semantic model.
[0054] Specifically, when constructing the academic vocabulary, the vocabulary construction unit of this embodiment uses all the papers to be recognized to construct a paper corpus, which includes the titles and abstracts of all the papers to be recognized. By uniformly performing word frequency statistics on the titles and abstracts of all the papers to be recognized in the paper corpus and calculating the word formation probability, words with a word formation probability greater than the threshold are selected to construct. First, N-Gram word frequency statistics are performed on the titles and abstracts of all the papers to be recognized using the N-Gram algorithm; based on the N-Gram word frequency, point mutual information, left and right entropy, and time influence factor, the word formation probability of each word obtained after segmentation by the N-Gram algorithm is calculated, and words with a word formation probability greater than the probability threshold are selected to construct the academic vocabulary.
[0055] Preferably, this embodiment constructs the academic vocabulary through the titles and abstracts of 500,000 papers that need to be batch recognized. First, the N-Gram algorithm is used to perform word frequency statistics on the titles and abstracts of 500,000 papers. The words obtained after segmentation by the N-Gram algorithm will include inaccurate words or stop words. Stop words refer to words that appear frequently in the text but are not relevant to the text content. Therefore, it is necessary to calculate the word formation probability of each word based on the N-Gram word frequency, point mutual information, left and right entropy, and time influence factor to exclude the influence of inaccurate words and stop words and improve the quality of subsequent word segmentation based on the academic vocabulary. In addition, a small number of English words will be obtained during word frequency statistics, and the obtained English words can also be normalized, including deleting duplicate spaces and punctuation marks, unifying case, abbreviation / synonym replacement, spelling correction, word form reduction, etc.
[0056] Among them, point mutual information is a measure of the mutual dependence between two words or characters. The value of the point mutual information of word x and word y, PMI(x,y), is calculated by the following formula:
[0057]
[0058] Among them, x and y are respectively words or characters obtained by the N-Gram algorithm, p(x,y) is the probability that the combination of x and y appears in the paper corpus, p(x) is the probability that x appears in the paper corpus, and p(y) is the probability that y appears in the paper corpus.
[0059] For example, the probabilities of the two words "machine" and "learning" appearing in the paper corpus are 0.000125 and 0.0001871 respectively. Theoretically, if "machine" and "learning" are completely unrelated, the probability that they happen to be put together should be 0.000125×0.0001871, approximately 2.339×10 -8 . In fact, the probability that "machine learning" appears in this paper corpus is 7.263×10 -6。The probability is much higher than the predicted probability. Therefore, the logarithm of the ratio of the true occurrence probability of a word to the predicted probability is called point mutual information. The higher this value is, the higher the probability that the word group forms a word independently.
[0060] Information entropy describes the degree of chaos of information and is also called the degree of uncertainty. The calculation formula is as follows:
[0061]
[0062] Among them, H x (X) represents the entropy of the neighboring words of word x, p(w) represents the probability of the neighboring word w of word x appearing, and X represents the set of all neighboring words of word x.
[0063] The left and right entropy represents the entropy of the left neighboring words and the entropy of the right neighboring words of a word. When calculating the probability of forming a word, the entropy calculation formula is introduced to calculate the entropy of the left neighboring words and the entropy of the right neighboring words respectively; when calculating the entropy of the left neighboring words, X represents the set of all left neighboring words of word x; when calculating the entropy of the right neighboring words, X represents the set of all right neighboring words of word x. Among them, the greater the entropy of the left neighboring words and the entropy of the right neighboring words of a word, the greater the probability of forming a word.
[0064] For example, in the sentence "When eating grapes, don't spit out the grape skins. When not eating grapes, spit out the grape skins", the left neighboring words of "grapes" include {eat, spit, eat, spit}, and the right neighboring words include {don't, skins, instead, skins}.
[0065] Calculated by the formula of information entropy, the left entropy of "grapes" is:
[0066] –(1 / 2)·log(1 / 2)–(1 / 2)·log(1 / 2)≈0.69;
[0067] Its right entropy is:
[0068] –(1 / 2)·log(1 / 2)–(1 / 4)·log(1 / 4)–(1 / 4)·log(1 / 4)≈1.04;
[0069] The left and right entropy values indicate the richness of the information of the left and right neighboring words of a word. A string with a high probability of forming a word should have rich information of the left and right neighboring words.
[0070] Generally, in addition to being able to freely combine with other words and appear frequently, a word also needs to be widely mentioned in a large number of papers within a certain period of time. Therefore, time is an important indicator to measure whether a string forms a word. By calculating the average time span between the publication time of the papers containing the word and the first appearance time of the word as the time influence factor, the calculation formula of the time influence factor is shown as follows:
[0071]
[0072] Among them, n represents the number of papers containing the word x, and t i represents the publication year of the i-th paper containing the word x, and t v represents the time when the paper in which the word x first appears in the paper corpus is published.
[0073] In order to reduce the influence of stop words, the present invention uses the inverse document frequency to weight and calculate the word formation probability. The more documents contain a word, the lower the importance of this word. After the inverse document frequency weighting calculation, the influence of such words is excluded.
[0074] The word formation probability calculation formula is as follows:
[0075]
[0076] Among them, |D| represents the total number of papers, and |{d ∈ D: x ∈ d}| represents the number of papers containing the word x in the paper corpus. represents the inverse document frequency; freq(x) represents the N-Gram word frequency of the word x, that is, the frequency of the word x appearing in the paper corpus, and PMI(x) represents the point mutual information of the word x. H xl (X l ) represents the left neighbor word information entropy, and H xr (X r ) represents the right neighbor word information entropy. X l represents the set of all left neighbor words of the word x, and Xr represents the set of all right neighbor words of the word x.
[0077] According to the word formation probabilities of all words obtained by word frequency statistics through the N-Gram algorithm for the titles and abstracts of all papers to be recognized, words with word formation probabilities greater than the threshold are selected to construct an academic word list. Preferably, in this embodiment, the word formation probability threshold is set to 0.5.
[0078] Preferably, the keyword recall unit performs word segmentation on the titles and abstracts of all papers to be recognized based on the constructed academic word list to obtain the recall keywords of each paper;
[0079] In this embodiment, the Jieba word segmentation tool is used to perform batch word segmentation processing on the titles and abstracts of all papers to be recognized. For example, for the input text "Research on Human Behavior Recognition in Complex Scenarios Based on Deep Learning", after being processed by the word segmentation module, it becomes: "Based on", "Deep Learning", "of", "Complex Scenarios", "under", "Human Body", "Behavior Recognition", "Research".
[0080] It should be noted that when constructing an academic vocabulary, if academic papers in as many fields as possible are selected for construction, the academic keyword batch recognition method of the present invention also has a certain generalization ability and can be directly applied to academic keyword recognition. In this embodiment, the vocabulary is constructed using the papers to be recognized, which improves the accuracy of keyword batch recognition to a certain extent.
[0081] Further, the keyword rough ranking batch processing module performs batch ranking processing on the recalled keywords of all papers to be recognized according to the statistical scores of the recalled keywords, and obtains candidate keywords corresponding to each paper;
[0082] The statistical score of the recalled keyword is calculated according to the word length, word position and TF-IDF score of each word;
[0083] Specifically, statistical features such as word length, word position, and TF-IDF weight of the word can be used for weighting, and the words obtained after the division are roughly ranked as keywords according to the weighted statistical scores. According to the rough ranking results of the keywords, candidate keywords are obtained. Among them, the IDF weight of the TF-IDF needs to be statistically calculated for all papers to be recognized, and then for each text to be recognized, the word frequency of each word in the text to be recognized is statistically calculated and multiplied by the IDF weight of the word to obtain the final TF-IDF score. The TF-IDF calculation formula is as follows:
[0084]
[0085]
[0086] tfidf(t, d, D) = tf(t, d) · idf(t, D)
[0087] where t is the word obtained through N-Gram processing, d is the paper to be processed where the word t is located, t′ is any word included in the paper d, f t,d is the frequency of the word t appearing in the paper d, f t′,d is the frequency of any word included in the paper d appearing in the paper d, D is the set of all papers to be recognized in the paper corpus, |{d ∈ D: t ∈ d}| represents the number of documents in the paper corpus that contain the word t, tf(t, d) represents the word frequency, idf(t, D) represents the inverse document frequency, and tfidf(t, d, D) represents the TF-IDF score; among them, the more documents contain the word t, the lower the idf value; the tfidf value is the product of the tf value and the idf value, and the higher the word frequency and the fewer the number of documents containing the word, the higher the score.
[0088] The position score is calculated according to the position of the word in the title and abstract, and the calculation formula of the position score is:
[0089]
[0090] The comprehensive score of each word is calculated by the following formula:
[0091] Score(x i ) = Length(x i ) · Position(x i ) · tfidf(x i )
[0092] where i represents the index value of the word, Length(x i ) represents the length of the word x i , and Position(x i ) represents the position score of x i . If the candidate word is in the title, the position weight is a constant 2. If the candidate word is in the abstract, the earlier the position of the word, the relatively higher the score.
[0093] For example, for the input paper title "Research on Human Behavior Recognition in Complex Scenarios Based on Deep Learning", after word segmentation, the comprehensive scores calculated according to the word length, word position, and TF-IDF scores are shown in Table 1:
[0094] Table 1 Example of Comprehensive Scores of Keywords
[0095]
[0096]
[0097] Sort according to the comprehensive scores, and return the words higher than the threshold as candidate keywords.
[0098] In this embodiment, the threshold is set to 1.2. Then, the candidate words "based on", "of", and "under" are filtered out in the initial keyword ranking stage. The remaining candidate keywords include "deep learning", "complex scenario", "human body", "behavior recognition", and "research".
[0099] Furthermore, through the keyword fine-ranking batch processing module, batch ranking processing is performed on the semantic similarity between the candidate keywords corresponding to all papers to be recognized and the corresponding titles and abstracts, and academic keywords corresponding to each paper are obtained based on the semantic similarity;
[0100] where the semantic similarity is calculated by a pre-trained deep semantic model;
[0101] First, use the titles and abstracts of all papers to be recognized to train a DSSM two-tower structure model to obtain a deep semantic model. Then, input all candidate keywords and the titles and abstracts of the papers to be recognized into the deep semantic model, calculate the semantic similarity between the candidate keywords and the titles and abstracts of the papers, sort the semantic similarities, and select the keywords with semantic similarities greater than the threshold to obtain academic keywords.
[0102] Preferably, the training of the deep semantic model includes:
[0103] Input the titles of the papers to be recognized, the abstracts corresponding to the titles, and n randomly selected abstracts from the papers to be recognized into the two-tower structure model of DSSM. Among them, the abstract corresponding to the title is the positive sample, and the randomly selected abstracts are negative samples. Calculate the similarity between the input titles and abstracts, and through iterative update of the loss function, maximize the semantic similarity between the input titles and the positive samples to obtain the trained deep semantic model, where n is an integer greater than 1.
[0104] The deep semantic model trained using the two-tower structure model of DSSM includes an input layer, a representation layer, and a matching layer;
[0105] The input layer uses the N-Gram model to reduce the dimensions of the input titles and abstracts respectively to obtain low-dimensional semantic vectors after dimensionality reduction and compression;
[0106] The representation layer includes three fully connected layers, each layer is activated using a non-linear activation function, and the low-dimensional semantic vectors are feature-integrated to obtain representation layer hidden vectors with a fixed dimension;
[0107] The matching layer calculates the semantic similarity between the titles and abstracts based on the representation layer hidden vectors.
[0108] It should be noted that the deep semantic model trained in this embodiment can perform keyword fine-ranking by calculating the semantic similarity between candidate keywords and the titles and abstracts of papers. Although the keyword rough-ranking in the foregoing steps can filter out some unimportant words from a statistical perspective, it will still misjudge words that are semantically irrelevant as keywords. Therefore, an unsupervised deep semantic model is used to perform fine-ranking on candidate keywords, calculate the semantics of the titles and abstracts and candidate keywords, so that the semantic distance between irrelevant candidate keywords and the titles and abstracts is large enough, and filtering can be performed by setting a threshold.
[0109] First, use the deep semantic model to encode the titles and abstracts and candidate keywords to obtain corresponding vector representations, and then use the cosine similarity to calculate the distance between the candidate keywords and the title and abstract vectors. The cosine similarity calculation formula is as follows:
[0110]
[0111] Where A and B represent the vectors of candidate keywords and title summary respectively.
[0112] Traditional semantic model training methods usually require similar sentence pairs for supervised learning, but similar sentence pairs require high manual annotation costs. Considering the structural characteristics of the paper, the present invention proposes a method for using the title and abstract as similar sentence pairs for semantic model training. The semantics of the title and abstract should be approximately equal. The abstract is a further description of the title, so the distance between them in the semantic space should be very small. In order to better model the deep semantic model, this embodiment uses the dual-tower structure model of DSSM for training, such as Figure 2 As shown in the figure, semantic similarity is calculated through the DSSM model. The model uses the title and abstract of the paper to be identified as input, uses a deep neural network model to express the title and abstract as low-dimensional semantic vectors, and calculates the distance between the two semantic vectors through cosine distance, and finally outputs the semantic similarity of the title and abstract. This model can be used to predict the semantic similarity of two sentences, and can also obtain the low-dimensional semantic vector expression of a sentence, thereby realizing the semantic similarity calculation of keywords and paper titles / abstracts.
[0113] Specifically, the deep semantic model trained by the dual-tower structure model of DSSM includes three layers: input layer, representation layer, and matching layer.
[0114] The input layer uses the N-Gram model to reduce the dimensionality of the input word, thereby achieving vector compression. When processing English papers, the tri-gram model is used for compression, that is, it is segmented according to every 3 characters. For example, the input word "algorithm" will be segmented into "#al", "alg", "lgo", "gor", "ori", "rit", "ith", "thm", "hm#". The advantage of this is that it can compress the space occupied by the word vector. The one-hot vector space of 500,000 words can be compressed into a 30,000-dimensional vector space through tri-gram; secondly, it enhances the generalization ability. The Chinese paper uses the uni-gram model, that is, each character is used as the smallest unit. For example, the input word "machine learning" will be segmented into "machine", "machine", "learning", and "learning". Using word vectors as input, the vector space is about 15,000 dimensions, where the dimension is determined by the number of commonly used Chinese characters.
[0115] The representation layer contains three fully connected layers, each of which is activated using a nonlinear activation function.
[0116] The matching layer uses cosine distance to calculate the similarity between positive and negative samples, and uses the negative log-likelihood loss function to optimize the neural network. The model is trained using the title and abstract of the paper as input data; among them, the positive sample is the abstract corresponding to the title, and the negative sample is the abstract randomly sampled from the papers to be recognized, and the randomly sampled abstract does not include the abstract corresponding to the title.
[0117] Both the middle network layer and the output layer adopt fully connected neural networks. Let W i represent the weight matrix of the i-th layer, and b i represent the bias term of the i-th layer. The hidden layer vector l i encoded by the i-th middle network layer and the output vector y encoded by the output layer can be respectively expressed as:
[0118] l i = f(W i l i-1 + b i ), i = 2,..., N - 1
[0119] y = f(W N l N-1 + b N )
[0120] where f represents the hyperbolic tangent activation function, and the hyperbolic tangent function is defined as follows:
[0121]
[0122] After encoding by the middle network layer and the output layer, a 128-dimensional semantic vector is obtained. The semantic similarity between the title and the abstract can be represented by the cosine similarity of these two semantic vectors:
[0123]
[0124] where y Q and y D respectively represent the vector representations of paper Q and paper D.
[0125] The semantic similarity between the title and the positive sample abstract can be transformed into a posterior probability through the softmax function:
[0126]
[0127] where γ is the smoothing factor of the softmax function, D+ is the abstract corresponding to the title Q, D' includes the abstract corresponding to the title Q and the randomly sampled abstracts, the R function represents the cosine distance function, and D is the entire sample space under the title.
[0128] During the training phase, through maximum likelihood estimation, we minimize the loss function so that after the normalization calculation by the softmax function, the similarity between the title and the positive sample abstract is maximized:
[0129]
[0130] Using the title and the abstract as similar pairs to train a deep semantic model, and then this model can be used to semantically encode the candidate keywords, and calculate the semantic similarity between the candidate keywords and the paper title abstract through the cosine distance.
[0131] For example, for the input paper title "Research on Human Behavior Recognition in Complex Scenarios Based on Deep Learning", and the candidate keywords "deep learning", "complex scenarios", "human body", "behavior recognition", "research", use the semantic models trained by the DSSM structure to encode them respectively, and then use the cosine distance to calculate the semantic similarity between the candidate keywords and the title respectively and sort them. The refined keyword results are shown in Table 2:
[0132] Table 2 Example of refined keywords
[0133] Candidate word Semantic similarity Behavior recognition 0.932 Deep learning 0.875 Complex scene 0.824 Human body 0.541 Research 0.323
[0134] Assuming that the semantic similarity threshold is 0.6, then the final output "behavior recognition", "deep learning", "complex scenarios" are used as the final keywords.
[0135] Experimental results of this embodiment:
[0136] (1) Efficiency
[0137] On 400,000 paper data, extract keywords from the titles and abstracts of the papers to verify the efficiency of the academic keyword extraction method of the present invention. The experimental results show that this method is much higher than the semantic-based keyword extraction algorithm in terms of efficiency. Although there is an additional semantic calculation step compared to simple statistical methods such as TF-IDF, the proposed keyword batch recognition method of the present invention does not have a significant decrease in speed, and the speed of batch recognition of keywords is about 100 times that of KeyBert.
[0138] (2) Generalization ability and accuracy
[0139] The keyword batch recognition system of the present invention not only has good academic keyword recognition effect on the papers to be recognized for constructing the academic vocabulary, but also has good generalization ability through the academic keyword recognition system constructed by the present invention. On the publicly available paper data, 500 Chinese papers were randomly selected for the comparative evaluation of the keyword extraction results. The accuracy rate of the method proposed by the present invention is 0.83, higher than 0.65 of TF-IDF and 0.78 of KeyBert. Therefore, the method proposed in this patent has high generalization ability and accuracy while ensuring efficiency.
[0140] In summary, the present invention discloses an academic keyword batch recognition system. This method combines statistical methods and semantic matching algorithms based on deep learning. It performs batch word segmentation on the text to be recognized through the constructed academic vocabulary, then sorts the candidate keywords using statistical features such as TF-IDF, and finally re-sorts the candidate keywords using an unsupervised semantic model trained with the dual-tower structure of DSSM to output the final keywords and their weights. The present invention directly constructs the academic vocabulary using the papers to be recognized that need to be processed in batches, and trains the deep semantic model through the titles and corresponding abstracts of the same batch of papers to be recognized, avoiding the appearance of out-of-vocabulary words, making full use of the characteristics of academic documents, and improving the accuracy of academic keyword extraction.
[0141] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.
[0142] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. An academic keyword batch recognition system, characterized in that, It includes a word segmentation module, a keyword rough sorting batch processing module, and a keyword fine sorting batch processing module; The word segmentation module includes a word list construction unit and a keyword recall unit; the word list construction unit is used to perform word frequency statistics on the titles and abstracts of all papers to be recognized and calculate the word formation probability, and construct an academic word list according to the word formation probability; The keyword recall unit is used to segment the titles and abstracts of all papers to be recognized based on the academic word list to obtain the recalled keywords of each paper; Constructing the academic word list includes: constructing a paper corpus, where the paper corpus includes the titles and corresponding abstracts of all papers to be recognized; performing word frequency statistics on the titles and abstracts of the papers in the paper corpus; calculating the word formation probability of each word based on N-Gram word frequency, point mutual information, left and right entropy, and time influence factor, and selecting words with a word formation probability greater than the probability threshold to construct an academic word list; the time influence factor is calculated based on the average time span between the publication time of the paper containing the word and the first appearance time of the word, and the calculation formula is: Among them, n represents the number of papers containing the word x, and t i represents the publication year of the i-th paper containing the word x, and t v represents the time of publication of the paper in which the word x first appears in the paper corpus; Calculate the word formation probability through the following formula: Among them, |D| represents the total number of papers, and |{d ∈ D: x ∈ d}| represents the number of papers in the paper corpus that contain the word x. represents the inverse document frequency; freq(x) represents the N-Gram word frequency of the word x, that is, the frequency of the word x appearing in the paper corpus, and PMI(x) represents the point mutual information of the word x, H xl (X l ) represents the left neighbor word information entropy, H xr (X r ) represents the right neighbor word information entropy, X l represents the set of all left neighbor words of the word x, and Xr represents the set of all right neighbor words of the word x. The keyword rough sorting batch processing module is used to perform batch sorting processing on the recalled keywords of all papers to be recognized according to the statistical scores of the recalled keywords to obtain the candidate keywords corresponding to each paper; The keyword fine sorting batch processing module is used to perform batch sorting processing on the semantic similarity between the candidate keywords corresponding to all papers to be recognized and the corresponding titles and abstracts, and obtain the academic keywords corresponding to each paper based on the semantic similarity; the semantic similarity is calculated through a pre-trained deep semantic model.
2. The academic keyword batch recognition system according to claim 1, characterized in that Calculate the value of the point mutual information PMI(x,y) through the following formula: Among them, x and y are respectively words or characters obtained through the N-Gram algorithm, p(x,y) is the probability that the combined phrase of x and y appears in the paper corpus, p(x) is the probability that x appears in the paper corpus, and p(y) is the probability that y appears in the paper corpus.
3. The academic keyword batch recognition system according to claim 1, characterized in that Input the titles of the papers in the paper corpus, the abstracts corresponding to the titles, and n abstracts randomly selected from the paper corpus into the dual tower structure model of DSSM, calculate the similarity between the title and the abstract of the paper, and through iterative update of the loss function, maximize the semantic similarity between the title and the abstract corresponding to the title to train a deep semantic model; n is an integer greater than 1.
4. The academic keyword batch recognition system according to claim 3, characterized in that The deep semantic model trained using the dual tower structure model of DSSM includes an input layer, a representation layer, and a matching layer; The input layer uses the N-Gram model to reduce the dimensions of the input title and abstract respectively to obtain low-dimensional semantic vectors after dimensionality reduction and compression; The representation layer includes three fully connected layers, each layer is activated using a non-linear activation function, and the low-dimensional semantic vectors are feature integrated to obtain a representation layer hidden vector with a fixed dimension; The matching layer calculates the semantic similarity between the title and the abstract based on the representation layer hidden vector.
5. The academic keyword batch recognition system according to claim 1, characterized in that Performing keyword rough sorting on the papers to be recognized to obtain candidate keywords includes: Segment the title and abstract of the paper to be recognized based on the academic vocabulary; calculate the comprehensive score of each word according to the word length, word position, and TF-IDF score of each word obtained after segmentation; obtain candidate keywords based on the comprehensive score; Among them, the comprehensive score of each word is calculated through the following formula: Score(x i ) = Length(x i ) · Position(x i ) · tfidf(x i ) where i represents the index value of the word, Length(x i ) represents the length of the word x i , Position(x i ) represents the position score of x i , and tfidf(x i ) represents the TF-IDF score.
6. The academic keyword batch recognition system according to claim 4, characterized in that The TF-IDF weight statistical score is calculated through the following formula: tfidf(t, d, D) = tf(t, d) · idf(t, D); Among them, t is the word obtained through N-Gram processing, d is the paper to be processed where the word t is located, t′ is any word contained in the paper d, and f t,d is the frequency of the word t appearing in the paper d, and f t',d is the frequency of any word contained in the paper d appearing in the paper d. D is the dataset of papers to be recognized in the paper corpus. |{d ∈ D: t ∈ d}| represents the number of documents in the paper corpus that contain the word t. tf(t, d) represents the term frequency, idf(t, D) represents the inverse document frequency, and tfidf(t, d, D) represents the TF-IDF score.
7. The academic keyword batch recognition system according to claim 5, wherein Calculate the position score according to the position of the word in the title and abstract, and the calculation formula of the position score is: Among them, i represents the index value of the word.
Citation Information
Patent Citations
Intelligent patent similarity searching method and device based on semantic retrieval
CN113821646A