Program Language Recognition System and Method for Integrating Associated Information
By integrating associated information in the program language recognition system, combining part-of-speech features, semantic features, mutual information and dependent syntactic relationships, and using GCN and CRF for feature representation and decoding, the problem of low accuracy and efficiency of existing program language recognition methods is solved, and more efficient program language recognition is achieved.
Patent Information
- Application Number
- CN202210037262.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-13
AI Technical Summary
The existing programming language recognition methods have problems with low recognition accuracy and efficiency, and have high requirements for feature selection, resulting in poor generalization ability.
The programmatic recognition system that fuses association information is adopted. The basic feature extraction module uses the embedding layer in Torch and the GloVe word vector technology to generate part-of-speech features and semantic features, and combines the mutual information and dependent syntactic relationships between words for late fusion, input it into GCN for feature representation, and finally decoded through the CRF layer.
It improves the accuracy and efficiency of program language recognition, enhances the generalization ability of the model, and can more accurately identify program languages.
Smart Images

Figure CN114330338B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the recognition of formulaic language, and specifically to a formulaic language recognition system and method that integrates associated information. Background Art
[0002] Formulaic language is a multi-word combination with specific functions and semantics, and is generally recognized, stored, and extracted as a whole form. Research shows that most expressions in human language are essentially composed of formulaic language. The recognition of formulaic language, also known as "multi-word expression recognition", is a basic task in natural language processing, with a very wide range of applications, and has important theoretical and practical significance for computer-assisted language teaching, machine translation, etc.
[0003] In recent years, the research on formulaic language at home and abroad is in a booming stage. Scholars have obtained a large number of research results on formulaic language by means of corpus technology and computer application programs such as AntConc and AntGram. However, there are problems such as incomplete recognition criteria, low recognition accuracy and efficiency. Therefore, how to recognize formulaic language efficiently and accurately has become increasingly important. At present, formulaic language recognition methods mainly include statistical-based methods, rule-based methods, and machine learning methods. Statistical and rule-based recognition methods rely on pre-set standards and have poor portability. When facing complex texts, they cannot effectively recognize various types of formulaic language. With the rise of machine learning in the field of natural language processing, some scholars have tried to use classifiers such as random forests and support vector machines to recognize formulaic language through classification technology. However, this method has high requirements for feature selection and requires selecting a feature set that can effectively reflect the characteristics of formulaic language, resulting in poor generalization ability. Summary of the Invention
[0004] The main purpose of the present invention is to provide a formulaic language recognition system and method that integrates associated information.
[0005] According to one aspect of the present invention, there is provided a formulaic language recognition system that integrates associated information, including:
[0006] A basic feature extraction module, configured to use the embedding layer in Torch to generate word embedding vectors as part-of-speech features, and feature vectors trained by the GloVe word vector technology as semantic features, and the part-of-speech features and semantic features after late fusion as the basic features of the model;
[0007] An associated information extraction module, configured to use the mutual information between words and the dependency syntactic relationship of sentences as the associated information for recognizing formulaic language;
[0008] A label representation module, configured to represent labels.
[0009] According to another aspect of the present invention, there is provided a method for identifying formulaic language by integrating associated information, including:
[0010] A basic feature extraction method;
[0011] An associated information extraction method;
[0012] A label representation method.
[0013] Furthermore, the basic feature extraction method includes:
[0014] Feature selection;
[0015] Feature representation based on Bi-LSTM;
[0016] Late fusion of part-of-speech features and semantic features.
[0017] Still further, the feature selection includes generating word embedding vectors as part-of-speech features using the embedding layer in Torch and using feature vectors trained by GloVe to represent the semantic features of formulaic language:
[0018] Construct a co-occurrence matrix X according to the corpus, and each element X in the matrix ij represents the number of times word i and context word j co-occur within a context window of a specific size;
[0019] Construct an approximate relationship between the word vector and the co-occurrence matrix, as shown in Equation 1:
[0020]
[0021] Among them, wi and wj in the above formula are the word vectors we finally need to solve; and bi and bj are the bias terms of the two word vectors.
[0022] Construct a loss function, as shown in the formula:
[0023]
[0024] Among them, is the weight function. Its calculation formula is as shown in Equation 3:
[0025]
[0026] Among them, x represents the co-occurrence times, and x max represents the maximum co-occurrence times.
[0027] Still further, the feature representation based on Bi-LSTM includes:
[0028] Let the sentence , input it into the Bi-LSTM network, and the sentence representation of the hidden layer can be obtained . Each unit calculates based on the previous hidden vector and the current input vector to obtain the current hidden vector , and its operation is defined as follows:
[0029]
[0030]
[0031]
[0032]
[0033]
[0034] In the formula: it, ft, ct, ot, ht are the states of the memory gate, hidden layer, forget gate, cell nucleus, and output gate when inputting the t-th text respectively; W is the parameter of the model; b is the bias vector; is the Sigmoid function; tanh is the hyperbolic tangent function.
[0035] Furthermore, the late fusion of the part-of-speech feature and the semantic feature includes:
[0036] First, input the part-of-speech feature and the semantic feature into the Bi-LSTM respectively, and then splice the results of the two models to form a basic feature vector.
[0037] Furthermore, the method for extracting the association information includes:
[0038] Association information based on mutual information:
[0039] The definition of the mutual information (MI) of two discrete random variables X and Y is:
[0040]
[0041] where p(x, y) is the joint probability distribution function of X and Y, and p(x) and p(y) are the marginal probability distribution functions of X and Y respectively. If you want to measure the degree of association between any two words x, y in a certain dataset, you can calculate it like this: , where p(x), p(y) are the probabilities of x, y appearing independently in the dataset, which can be obtained by directly counting the word frequencies and then dividing by the total number of words; p(x, y) is the probability of x, y appearing in the dataset at the same time, which can be obtained by directly counting the number of times they appear at the same time and then dividing by the number of all unordered pairs;
[0042] Dependency syntactic analysis-based association information:
[0043] Dependency syntax reveals the dependency and collocation relationships between words in a sentence. One dependency relationship connects two words, one being the core word and the other being the modifier. Such relationships are interrelated with the semantic relationships of the sentence;
[0044] Feature representation based on graph convolutional neural network:
[0045] The relationships between words are represented by a graph through MI and dependency syntactic analysis. Therefore, a graph convolutional neural network is used to process the association information.
[0046] Given a graph G=(V, E), where V is the vertex set containing N nodes and E is the edge set including self-loop edges (i.e., each vertex is connected to itself), the feature information of the graph G(V, E) can be represented by the Laplacian matrix (L), as shown in Equation 11.
[0047]
[0048] Or use the symmetrically normalized Laplacian matrix:
[0049]
[0050] In the formula: A is the adjacency matrix of the graph; IN is the N-order identity matrix; D = diag(d) is the degree matrix of the vertices.
[0051]
[0052] Based on the Fourier transform of the graph, the graph convolution formula can be expressed as:
[0053]
[0054] In the formula: x is the basic feature vector of the nodes; g is the convolution kernel; U is the eigenvector matrix of the Laplacian matrix L.
[0055] The Chebyshev polynomial is used to simplify the graph convolution formula. Finally, the graph convolution layer propagation formula can be expressed as:
[0056]
[0057] In the formula: , ; is the activation function; W is the weight matrix to be trained.
[0058] Furthermore, the label representation method includes:
[0059] In CRF, for each sentence X = {x1, x2, …, xn}, there is a set of candidate tag sequences YX, and the final labeled sequence is determined by calculating the scores of each tag sequence Y = {y1, y2, …, yn} in the set. The process of calculating the scores is as shown in the formula:
[0060]
[0061] Among them, P is a score matrix, k is the number of all tags, and Pi,j represents the score of the i-th character in the sentence corresponding to the j-th tag; A is a transition matrix containing the start and end tags of the sentence, and Ai,j represents the transition score from tag i to tag j;
[0062] The scores of each tag sequence are normalized to obtain probabilities, and the tag sequence with the highest probability is the final sequence of the sentence. The normalization process is as shown in the formula.
[0063] .
[0064] Advantages of the present invention:
[0065] In order to represent the features of the text, the present invention proposes a late fusion model based on part-of-speech features and semantic features. The embedding layer in Torch is used to generate word embedding vectors as part-of-speech features, and the feature vectors trained by the GloVe word vector technology are used as semantic features, which can fully represent the features such as the high occurrence frequency of formulaic language and the fixed structure.
[0066] In order to further utilize the information between words, the mutual information between words is calculated and the dependency syntactic analysis of the sentence is performed. These two related information and basic features are input into the GCN for feature representation. Using the graph convolutional neural network to model the graph can capture the high-order neighbor information between words.
[0067] Regarding the recognition of formulaic language as a sequence labeling problem, the fused feature vectors are input into the CRF layer for decoding to obtain the label category of each character and obtain the formulaic language.
[0068] The present invention proposes a deep learning model to recognize formulaic language, represents the feature vectors through word embedding technology, fuses the related information that can represent the features of formulaic language, uses the graph convolutional neural network (GCN) to obtain deeper semantic features, and finally, considering the dependency relationship between tags, uses the conditional random field model for label decoding to achieve the purpose of recognizing formulaic language.
[0069] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. Description of the Drawings
[0070] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0071] Figure 1 It is a diagram of the GCN programming language recognition model that fuses associated information of the present invention;
[0072] Figure 2 It is a structural diagram of the Bi-LSTM model of the present invention;
[0073] Figure 3 It is a structural diagram of the late fusion model of the part-of-speech features and semantic features of the present invention;
[0074] Figure 4 It is a dependency syntactic relationship diagram of the sentence "evaluations play an invaluable role in X." of the present invention;
[0075] Figure 5 It is the adjacency matrix A constructed based on dependency syntactic analysis of the present invention;
[0076] Figure 6 It is a structural diagram of the graph convolutional neural network of the present invention;
[0077] Figure 7 It is a diagram of the ten-fold cross-validation results of the model of the present invention;
[0078] Figure 8 It is a diagram of the influence of different network layers of the present invention on the graph convolutional neural network based on dependency syntactic analysis;
[0079] Figure 9 It is a diagram of the influence of different network layers of the present invention on the graph convolutional neural network based on mutual information. Detailed Embodiments
[0080] In order to make the purposes, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0081] Reference Figure 1 and Figure 2 , a programming language recognition system that fuses associated information, includes:
[0082] The basic feature extraction module is used to generate word embedding vectors as part-of-speech features using the embedding layer in Torch, and feature vectors trained by the GloVe word vector technology as semantic features. The part-of-speech features and semantic features after late fusion are used as the basic features of the model.
[0083] The association information extraction module is used to adopt the mutual information between words and the dependency syntactic relationship of sentences as the association information for identifying formulaic language.
[0084] The label representation module is used to represent labels.
[0085] The formulaic language recognition method that fuses association information includes:
[0086] The basic feature extraction method;
[0087] The association information extraction method;
[0088] The label representation method.
[0089] The basic feature extraction method
[0090] In the process of natural language processing, the computer cannot directly use text data. The text data needs to be represented as feature vectors, and then the feature vectors are used as the input of the model. In the present invention, word embedding vectors are generated using the embedding layer in Torch as part-of-speech features, and feature vectors trained by the GloVe word vector technology as semantic features. The part-of-speech features and semantic features after late fusion are used as the basic features of the model.
[0091] Feature selection
[0092] The biggest difference between formulaic language and general multi-word expressions is that the structure of formulaic language is mostly fixed, often in the form of "verb + noun" or "subject + predicate + object", etc. Therefore, part-of-speech features are used as one of the features for identifying formulaic language. First, the text is analyzed for part-of-speech using the Stanford Part-of-Speech Tagger. An example of the part-of-speech analysis results is shown in Table 1. It can be found from the table that multi-word units with fixed sentence patterns are more likely to be formulaic language. Then, for each result after part-of-speech tagging, a unique code is assigned, thus converting the text data into vectors, and finally input into the embedding layer for training to generate word embedding vectors as part-of-speech features.
[0093] Table 1 Example of part-of-speech analysis results
[0094]
[0095] In addition, for the processing of text, words, phrases, and expressions mainly reflect the lexical information of the text rather than its semantic information, and thus cannot accurately express the content of the text. For formulaic language, which is a multi-word unit with a relatively high frequency of occurrence and has a relatively complete structure, meaning, and function, semantic features are important features of formulaic language. The present invention uses feature vectors trained by GloVe to represent the semantic features of formulaic language.
[0096] The full name of GloVe is Global Vectors for Word Representation. It is a word representation tool based on global word frequency statistics (count-based & overall statistics). It can represent a word as a vector composed of real numbers, and these vectors capture some semantic characteristics between words. Its implementation is divided into the following three steps:
[0097] (1) Construct a co-occurrence matrix X based on the corpus. Each element Xij in the matrix represents the number of times that word i and context word j co-occur within a context window of a specific size. Generally, the minimum unit of this number is 1, but GloVe does not think so. It proposes a decreasing weighting function: decay = 1 / d to calculate the weight according to the distance d between two words in the context window, that is, the weights of two words with a greater distance account for a smaller proportion of the total count.
[0098] (2) Construct an approximate relationship between the word vector and the co-occurrence matrix, as shown in Equation 1:
[0099]
[0100] Among them, wi and wj in the above formula are the word vectors to be finally solved; while bi and bj are the bias terms of the two word vectors.
[0101] (3) Construct a loss function, as shown in Equation 2:
[0102]
[0103] Among them, is the weight function. Its calculation formula is shown in Equation 3:
[0104]
[0105] Among them, x represents the co-occurrence times, and xmax represents the maximum co-occurrence times.
[0106] Feature Representation Method Based on Bi-LSTM
[0107] Since LSTM is good at capturing the long-distance long-term dependence relationship of the context information of sentences, it can better avoid the problems of gradient disappearance and gradient explosion, and has higher computational efficiency. However, this model cannot capture the bidirectional information of sentences. For the task of formulaic language recognition, if the forward information and backward information of sentences are added, the model can learn more semantic information when processing text. Therefore, Bi-LSTM is used to learn the hidden layer representation of the input sequence, expecting to obtain sentence features that can contain deeper semantic and syntactic information.
[0108] Let the sentence be input into the Bi-LSTM network, and the representation of the hidden layer of the sentence can be obtained . Each unit calculates according to the previous hidden vector and the current input vector to obtain the current hidden vector , and its operation is defined as follows:
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] In the formula: it, ft, ct, ot, ht are the states of the memory gate, hidden layer, forget gate, cell nucleus, and output gate at the time of inputting the t-th text respectively; W is the parameter of the model; b is the bias vector; is the Sigmoid function; tanh is the hyperbolic tangent function.
[0115] The Bi-LSTM model is composed of a forward LSTM and a backward LSTM model. Each layer of the LSTM network outputs a hidden state information respectively, and the parameters of the model are updated by backpropagation. The structure of the Bi-LSTM model is as Figure 2 shown:
[0116] Figure 2 In it, xt represents the input of the network at time t, and the LSTM in the box is the standard LSTM model, is the output of the forward LSTM at time t, is the output of the backward LSTM at time t, and ⊕ represents the concatenation operation. That is to say, the output representation of the BiLSTM at time t is defined as , that is, the output at time t is directly formed by concatenating the forward output and the backward output.
[0117] Late fusion of part-of-speech features and semantic features
[0118] Feature fusion includes two methods: early fusion and late fusion. Early fusion is to first fuse features of multiple layers and then train a model on the fused features (training is only unified after complete fusion). Compared with early fusion, late fusion first trains models with individual features separately and then fuses the results of multiple model trainings. The advantage of the late fusion method is that the results of the models can be flexibly selected, improving the fault tolerance of the system; the computational amount of fused information is reduced, and the real-time performance of the system is improved. The present invention adopts the late fusion method. First, the part-of-speech features and semantic features are respectively input into the Bi-LSTM, and then the results of the two models are concatenated to form a basic feature vector. The structure diagram is as Figure 3 shown:
[0119] Association information extraction method
[0120] The basic feature extraction module is trained in a large-scale text using a word embedding model and can obtain word vectors rich in text semantic features and part-of-speech features, but the syntactic structure information in the text is ignored. As the basis of language understanding, the syntactic structure can effectively represent the grammatical structure of the text and reveal the relationship between the components in the text. For formulaic language, it is a multi-word unit with a relatively high frequency of occurrence, and several words with high relevance may form formulaic language. Therefore, it is very important to select features that can represent the relationship between words for identifying formulaic language. Based on this, the present invention uses the mutual information between words and the dependency syntactic relationship of sentences as the association information for identifying formulaic language.
[0121] Association information based on mutual information
[0122] Mutual information is a measure of the correlation between two random variables, that is, the amount of information about another random variable contained in one random variable. The definition of the mutual information (MI) of two discrete random variables X and Y is:
[0123]
[0124] where p(x, y) is the joint probability distribution function of X and Y, and p(x) and p(y) are the marginal probability distribution functions of X and Y respectively. If you want to measure the degree of association between any two words x, y in a certain dataset, you can calculate it like this: , where p(x) and p(y) are the probabilities of x and y appearing independently in the dataset, which can be obtained by directly counting the word frequencies and then dividing by the total number of words; p(x,y) is the probability of x and y appearing simultaneously in the dataset, which can be obtained by directly counting the number of times they appear simultaneously and then dividing by the number of all unordered pairs. The mutual information is used to calculate the relationship between bigrams. The higher the mutual information, the higher the correlation between x and y, and the greater the possibility of forming a formulaic expression.
[0125] Association Information Based on Dependency Parsing
[0126] Dependency parsing reveals the dependency relationship and collocation relationship between words in a sentence. One dependency relationship connects two words, one is the head word and the other is the modifier. Such a relationship is interrelated with the semantic relationship of the sentence. The dependency relationships between words in a sentence include subject-predicate relationship (SBV), verb-object relationship (VOB), indirect object relationship (IOB), etc. The dependency parsing relationship of the sentence "evaluations play an invaluable role in X." is as Figure 4 shown. Among them, "play an invaluable role in" is a formulaic expression. It can be seen from the figure that there are intricate dependency relationships among these five words. Therefore, the dependency parsing between words can represent the dependency relationship between two words. The closer the relationship, the greater the possibility of forming a formulaic expression.
[0127] Feature Representation Based on Graph Convolutional Neural Network
[0128] In the identification of formulaic expressions using dependency parsing relationships, existing research mostly uses the dependency parsing relationships in the text to construct rules, extract features, etc., and then inputs them into a classifier for classification to identify formulaic expressions. Although such methods can achieve certain results, the non-linear semantic relationships between components in a sentence have not been learned and utilized. Spatially, the relationships between words can be represented by a graph through MI and dependency parsing. Therefore, a graph convolutional neural network is used to process the association information.
[0129] When combining GCN for natural language processing tasks, the dependency parsing structure, TF-IDF, mutual information, and sequence relationships of the text are usually used as one of the inputs of GCN. On the one hand, because these features can themselves be represented by a graph, and on the other hand, because they can enhance the information of the text. In constructing the graph convolutional neural network model of the present invention, the mutual information value and dependency parsing relationship between words are used to determine the word connection relationship. For the graph convolutional neural network with the mutual information value as the input, first, the mutual information value between words is calculated using the corpus as the dataset. Taking words as nodes and the mutual information value between nodes as the representation of the edge, its adjacency matrix A The value of the element aij in RN*N represents the mutual information value between the ith node and the jth node in the graph. For a graph convolutional neural network that takes syntactic dependency relations as input, first perform dependency syntactic analysis on the sentence. Use words as nodes and the dependency relations between words as the representation of edges, and its adjacency matrix A The value of the element aij in RN*N represents the dependency relation between the ith node and the jth node in the graph. If there is a dependency relation between two nodes, the value of aij is 1, otherwise it is 0. For example, in the example of dependency syntactic analysis "evaluations play an invaluable rolein X.", the adjacency matrix A constructed based on dependency syntactic analysis is as Figure 5 shown.
[0130] To directly perform deep learning modeling on graph data, the specific method uses a variant model of a convolutional neural network proposed - a graph convolutional neural network, and the structure is as Figure 6 shown. Specifically, given a graph G=(V, E), V is the vertex set containing N nodes, and E is the edge set including self-loop edges (that is, each vertex is connected to itself). The feature information of the graph G(V, E) can be represented by the Laplacian matrix (L), as shown in the formula.
[0131]
[0132] Or use the symmetrically normalized Laplacian matrix:
[0133]
[0134] In the formula: A is the adjacency matrix of the graph; IN is the N-order identity matrix; D = diag(d) is the degree matrix of the vertices.
[0135]
[0136] Based on the Fourier transform of the graph, the graph convolution formula can be expressed as:
[0137]
[0138] In the formula: x is the basic feature vector of the node; g is the convolution kernel; U is the eigenvector matrix of the Laplacian matrix L.
[0139] To reduce the computational complexity, in 2017, scholars used Chebyshev polynomials to simplify the graph convolution formula. Finally, the graph convolution layer propagation formula can be expressed as:
[0140]
[0141] In the formula: , ; is the activation function; W is the weight matrix to be trained.
[0142] Label representation module
[0143] The recognition of formulaic language is essentially a multi-classification problem. Therefore, the Softmax classifier is a commonly used method in the decoding stage. However, since this method is only a simple classification without considering the dependency relationship between labels, the conditional random field model (CRF) is used in the present invention.
[0144] CRF is a conditional probability distribution model of another set of output sequences given a set of input sequences, and has been widely used in natural language processing. In CRF, each sentence X = {x1, x2,..., xn} has a set of candidate label sequences YX, and the final annotation sequence is determined by calculating the scores of each label sequence Y = {y1, y2,..., yn} in the set. The process of calculating the scores is shown in Equation 17.
[0145]
[0146] Among them, P is a score matrix, k is the number of all labels, and Pi,j represents the score of the i-th character in the sentence corresponding to the j-th label; A is a transition matrix including the start and end labels of the sentence, and Ai,j represents the transition score from label i to label j.
[0147] Finally, the scores of each label sequence are normalized to obtain probabilities, and the label sequence with the highest probability is the final sequence of the sentence. The normalization process is as shown in the equation.
[0148]
[0149] The recognition method of the present invention is mainly divided into three parts: the basic feature extraction module, the associated information extraction module, and the label representation module. Its overall structure is as Figure 1 shown. First, the semantic features and part-of-speech features of the text are extracted and fused by the late fusion method. The fused result is used as the basic feature of the model. Then, the mutual information between words is calculated and the dependency syntactic analysis of the sentence is performed. The generated adjacency matrix and basic features are input into the GCN for feature representation. Finally, the feature vector is input into the CRF layer for decoding to obtain the label category of each character and obtain the formulaic language.
[0150] Experiments and analysis
[0151] Experimental environment
[0152] This invention was experimented on the Win64 operating system; the processor was an i5-7500U CPU @ 3.40GHz; the memory size was 16 GB. All neural network models were built using the deep learning framework PyTorch 1.2.0 for training and testing; the Python 3.6 programming language was used for coding.
[0153] Experimental data and annotation strategy
[0154] Thirty papers in the field of computer science were downloaded from Web of Science, and the text was preprocessed to remove references, pictures, formulas, etc., and then segmented into sentences, resulting in a total of 6,556 sentences, which were used as the dataset. The formulaic language in the formulaic language corpus of the University of Manchester was used as the standard, and three professors from the School of Foreign Languages were asked to extract the formulaic language in the sentences. There were 8,252 formulaic languages in total, and 4,136 formulaic languages remained after deduplication. Then, the sentences were tagged, and the annotation strategy adopted the "BIO annotation" method. "B" represents the starting position of the formulaic language, "I" represents the middle position of the formulaic language, and "O" represents the part that does not belong to the formulaic language.
[0155] Evaluation metrics
[0156] This invention uses the PRF metric to evaluate the experimental results of formulaic language recognition. P represents the accuracy rate of recognizing formulaic language (Precision); R refers to the proportion of the number of correctly recognized formulaic languages in the total number of formulaic languages in the corpus, which is called the recall rate (Recall); the F value is a comprehensive measure of the P value and the R value and is used as a comprehensive indicator for evaluating the formulaic language recognition effect. The three formulas correspond to formulas (3-9), (3-10), and (3-11) respectively:
[0157]
[0158]
[0159]
[0160] Among them, Nm represents the number of correctly recognized formulaic languages, Ntotal represents the total number of recognized formulaic languages, and Ncorrect represents the total number of formulaic languages manually annotated.
[0161] Parameter settings
[0162] Use the 300-dimensional word vectors pre-trained by Glove as semantic input features. For the part-of-speech features, the dimension of the word embedding vectors generated by the embedding layer is set to 300. Mini-batch stochastic gradient descent is adopted, with the batch size set to 16, the learning rate set to 0.001, the learning rate decay set to 0.9, and the optimization algorithm selected as the Adam algorithm. For all LSTM networks, there are 128 neurons in a single layer, so a two-layer LSTM has 256 neurons, and it is trained for 50 epochs. A two-layer GCN network structure is selected, and the output of the GCN layer is set to 64.
[0163] Experimental Settings and Analysis
[0164] Ablation Experiments on the Formulaic Language Recognition Model
[0165] To better verify the effectiveness of the GCN formulaic language recognition model that fuses correlation information, 7 comparative experiments were conducted by setting ablation experiments to determine which feature is more important for formulaic language recognition. The specific methods are introduced as follows:
[0166] (1) Before_Bi-LSTM: The part-of-speech features generated by PyTorch word embeddings and the semantic features generated by GloVe word embeddings are fused through early fusion. The fused feature vectors are input into Bi-LSTM to extract context semantic relationships, and finally input into CRF to complete the recognition of formulaic language.
[0167] (2) Before_CNN: The part-of-speech features generated by PyTorch word embeddings and the semantic features generated by GloVe word embeddings are fused through early fusion. The fused feature vectors are input into CNN, with the size of the convolutional kernel being 3*3, and two layers of CNN are set in total. Finally, it is input into CRF to complete the recognition of formulaic language.
[0168] (3) After_Bi-LSTM: The part-of-speech features generated by PyTorch word embeddings and the semantic features generated by GloVe word embeddings are fused through late fusion, that is, the two feature vectors are respectively input into Bi-LSTM, and then the processed vectors are fused. Finally, it is input into CRF to complete the recognition of formulaic language.
[0169] (4) After_Bi-LSTM_CNN: Based on After_Bi-LSTM, a layer of CNN with a convolutional kernel size of 3*3 is added between Bi-LSTM and CRF, and two layers of CNN are set in total.
[0170] (5) Bi-LSTM_SD_GCN: The part-of-speech features generated by PyTorch word embeddings and the semantic features generated by GloVe word embeddings are fused through late fusion. The fused feature vectors serve as basic features, and the matrix generated by syntactic dependency relations serves as associated features. They are jointly input into the GCN for feature representation and finally input into the CRF to complete the identification of formulaic language.
[0171] (6) Bi-LSTM_MI_GCN: The difference from Bi-LSTM_SD_GCN is that the matrix generated by dependency parsing is replaced with the matrix generated by the MI between words as the associated feature.
[0172] (7) Bi-LSTM_PMI_SD_GCN: This is the model proposed by the present invention. The part-of-speech features and semantic features through late fusion serve as basic features, and the matrices generated based on mutual information and dependency parsing serve as associated information. The basic features and associated information are input into the GCN, and finally feature decoding is performed through the CRF.
[0173] The seven methods are experimented on the dataset, and the experimental results are shown in Table 2.
[0174]
[0175] Analysis of experimental results:
[0176] (1) The difference between Experiment 1 and Experiment 2 lies in the comparison between Bi-LSTM and CNN. The experimental results show that Bi-LSTM has much better feature extraction effect than CNN. Since the key to formulaic language recognition is to analyze the relationships between words in a sentence, which is a typical sequence labeling problem. Bi-LSTM can capture the long-distance long-term dependence relationships of the context information of a sentence and can capture the bidirectional information of a sentence. However, CNN cannot capture long-distance dependence information well. Therefore, it is better to use Bi-LSTM in the formulaic language recognition task. It should be noted that the recall rate in the results of using CNN to extract features is relatively high, indicating that it can identify more formulaic languages, but it also identifies many non-formulaic languages, so the precision is not very high.
[0177] (2) The difference between Experiment 1 and Experiment 3 lies in the different ways of feature fusion. Experiment 1 adopts the early fusion method, while Experiment 3 adopts the late fusion method. The experimental results show that the precision rate of the late fusion method is higher, while the recall rate of the early fusion method is higher. The main reason for this is that when the early fusion method identifies formulaic language, it can identify more results, but at the same time, it also identifies many non-formulaic languages, so its precision rate is relatively low; while the features of the late fusion method are more accurate. Although the number of results identified by the late fusion method is less than that of the early fusion method, it can accurately identify formulaic language. From the F1 scores of the two methods, it can be seen that the effect of the late fusion method is better than that of the early fusion method.
[0178] (3) Experiment 4 adds a CNN layer on the basis of Experiment 3. However, the experimental results show that the results after adding CNN are worse than the previous ones. The main reason for this is that CNN is used to capture local correlations and extract local features. At the same time, the stride is fixed in each layer of CNN, so naturally this layer can only model information within a limited distance. However, Bi-LSTM has already obtained the long-distance dependencies of the context. Therefore, adding CNN after Bi-LSTM can perform deep-level abstraction on the features. But some text features require a wider receptive field to enable the model to combine more features. So after adding CNN, some correct formulaic languages are filtered out, resulting in a significant decrease in the precision rate and recall rate.
[0179] (4) Experiment 5 adds GCN feature extraction based on dependency parsing on the basis of Experiment 3. The experimental results show that after adding dependency syntactic features, the recall rate remains unchanged, but the precision rate becomes lower. Because syntactic dependency relationships mainly focus on the dependency relationships between two words in a sentence, it is easy to cause the extracted word strings not to belong to formulaic language, so the precision rate decreases.
[0180] Experiment 6 adds GCN feature extraction based on MI on the basis of Experiment 3. The experimental results show that after adding MI features, the recall rate increases significantly, indicating that the number of identified formulaic languages is more. In addition, from the F1 score, the F1 score is higher after adding MI features. Because MI focuses on the tight combination degree between two words, this feature can accurately represent the features of formulaic language and is very important for identifying formulaic language.
[0181] (5) Experiment 7 (the model of the present invention) inputs dependency syntax features and MI features into GCN for feature extraction. That is, compared with Experiment 6, which has the best experimental results, dependency syntax features are added. The experimental results show that although the dependency syntax features alone (Experiment 5) do not show a good effect, after combining the two, the precision and recall rates are significantly increased. The reason is that the dependency syntax analysis features and mutual information features each have advantages and disadvantages when identifying programming language. Combining the two features complements each other to achieve efficient extraction of programming language. It also shows that dependency syntax and mutual information are important features for measuring multi-word expressions.
[0182] In addition, a 10-fold cross validation was used to evaluate the reliability of the model of the present invention. The data set was divided into 10 parts, and 9 of them were used as training data and 1 as test data in turn for the experiment. The experimental results are shown in Figure 2. Figure 7 As shown, the stability of the model of the present invention can be seen from the figure.
[0183] In summary, through ablation experiments, the results verified the effect of the GCN programming language recognition model that integrates relevant information, that is, by using late-fused part-of-speech features and semantic features as basic features, and syntactic dependencies and mutual information as relevant information, the combined model can alleviate the errors of individual models and enhance their advantages.
[0184] Comparative experiments of different models
[0185] In order to verify the effectiveness of the model proposed in this invention, the CNN_Bi-LSTM_CRF model and the Bi-LSTM_CRF model were selected for comparison.
[0186] (1) CNN_Bi-LSTM_CRF: The present invention performs the task of named entity recognition. Since programming language recognition and named entity recognition tasks are similar, and this model performs well in the field of named entity recognition, this model is used as a comparative experiment. Word2vec is used to train word vectors, and the word vectors of the text data obtained after Word2vec training are spliced to generate a word vector matrix, which is then used as the input of the CNN convolutional layer. The CNN module extracts the spatial feature information of the text through convolution and vector matrix aggregation. Afterwards, the result is input into Bi-LSTM for forward and backward training. Finally, the vector with sentence feature information is put into the conditional random field for decoding and prediction to obtain the final sequence.
[0187] (2) Bi-LSTM_CRF: This paper describes the Deep-BGT system that participates in the PARSEME shared task, which is about automatically recognizing spoken multi-word expressions (VMWE). The author uses a bidirectional long short-term memory model with a conditional random field layer on top. The input layer includes word vectors generated by the fastText word embedding technology, POS, and dependencies. Each input vector is represented as a concatenation of these three features, similar to the early fusion technique. Because programming language is also a type of multi-word expression, the combination of Bi-LSTM and CRF is the mainstream method in the field of multi-word expression recognition, so this model is used as a comparative experiment.
[0188] The experimental results of the above two models and the model proposed in the present invention on the programming language recognition task are shown in Table 3.
[0189]
[0190] Experimental results analysis:
[0191] (1) Since the CNN_Bi-LSTM_CRF model is the main method in the field of named entity recognition, the experimental results show that this model does not perform well in the task of program language recognition. Therefore, in different tasks, although the two tasks are very similar, we should also start from the essence of the object in the task, explore the features that can represent the object under study, and design a dedicated model to treat the problem. At the same time, through the CNN_Bi-LSTM_CRF model, as well as Experiments 2 and 4 in the previous section, it can be found that the effect will not improve after adding CNN to the model, so CNN is not suitable for the task of program language recognition.
[0192] (2) The Bi-LSTM-CRF model is used to identify multi-word expressions. Compared with Experiment 1 in the previous section, the difference lies in the input features. The input features of Experiment 1 are part-of-speech features and GloVe word embedding features, while the input features of the Bi-LSTM-CRF model are part-of-speech features, fastText word embedding features, and syntactic dependencies. The experimental results show that the F1 score of the Bi-LSTM-CRF model is higher, but the recall rate of Experiment 1 is higher, indicating that syntactic dependencies are conducive to the recognition of programming language.
[0193] Meanwhile, in the comparison with the model of the present invention, the main difference is that the Bi-LSTM-CRF model only constructs the dependency syntactic relationship into a simple feature vector, concatenates it with other features and then trains, while the model of the present invention constructs a graph structure through the syntactic dependency tree, and then extracts features through GCN. The advantage of GCN is that it can be used to aggregate the information of all edges and nodes, thereby eliminating the boundary ambiguity between words, and any two non-adjacent nodes in the graph are second-order neighbors of each other, and non-local information of each other can be received through two node updates. The features aggregated by this method can more accurately represent formulaic language, so the effect of identifying formulaic language is better.
[0194] Comparative experiments with different numbers of network layers
[0195] Since the model of the present invention involves two graph convolutional neural networks, one is a graph convolutional neural network based on dependency syntactic analysis, and the other is a graph convolutional neural network based on mutual information. Therefore, when selecting the number of layers of the graph convolutional neural network, experiments are used to conduct two sets of comparative experiments, and the optimal number of network layers is selected through the results of the comparative experiments.
[0196] (1) Experiment 5 is the graph convolutional neural network structure based on dependency syntactic analysis. Experiments are respectively conducted by setting 1, 2, 3, 4, and 5 layers of graph convolutional neural networks. The experimental results are as Figure 8 shown. It can be seen from the figure that for dependency syntactic analysis, the effect of using 3 layers of graph convolution is the best.
[0197] (2) Experiment 6 is the graph convolutional neural network structure based on mutual information. Experiments are respectively conducted by setting 1, 2, 3, 4, and 5 layers of graph convolutional neural networks. The experimental results are as Figure 9 shown. It can be seen from the figure that for mutual information, the effect of using 2 layers of graph convolution is the best.
[0198] Through the above analysis, it can be known that for the dependency syntactic analysis feature, the effect of using 3 layers of graph convolution is the best, and for the mutual information feature, the effect of using 2 layers of graph convolution is the best. Therefore, in the GCN formulaic language recognition model that fuses correlation information in the present invention, the numbers of layers of the two graph convolutional neural networks are respectively set to two layers and three layers.
[0199] Conclusion
[0200] The present invention proposes a GCN formulaic language recognition model that fuses associated information. The part-of-speech features and semantic features fused through a late fusion method are used as basic features, and then the associated information is input into the GCN for feature representation. This combined representation can capture the syntactic and semantic structures of the multi-semantic network and can also perform more in-depth downstream semantic analysis. Finally, the fused feature vectors are input into the CRF layer for decoding to obtain the label category of each character and obtain the formulaic language. Multiple sets of comparative experiments on the scientific literature dataset show that, compared with the existing models, the model proposed by the present invention can improve the effect of formulaic language recognition, verifying the effectiveness of the model. In addition, it should be noted that the formulaic language recognition model proposed by the present invention can obtain powerful recognition performance only with a relatively small proportion of labeled text.
[0201] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying formulaic language by fusing associated information, characterized in that, it includes: A basic feature extraction method; An associated information extraction method; A label representation method; The basic feature extraction method includes: Feature selection; Feature representation based on Bi-LSTM; Late fusion of part-of-speech features and semantic features; The feature selection includes using the embedding layer in Torch to generate word embedding vectors as part-of-speech features and using feature vectors trained by GloVe to represent the semantic features of formulaic language: Construct a co-occurrence matrix X based on the corpus, where each element X ij in the matrix represents the number of times word i and context word j co-occur within a context window of a specific size; Construct an approximate relationship between the word vector and the co-occurrence matrix, and the relationship is shown in the following formula: where, wi and wj in the above formula are the word vectors we finally need to solve; while bi and bj are the bias terms of the two word vectors; Construct a loss function, as shown in the formula: Among them, f(X ij ) is a weight function, and its calculation formula is as follows: where x represents the co-occurrence times, and x max represents the maximum co-occurrence times; The feature representation based on Bi-LSTM includes: Let the sentence be s i = [x 1 , x 2 ,..., x t ,..., x n . Input it into the Bi-LSTM network to obtain the representation of the hidden layer of sentence s i {h 1 , h 2 ,..., h t ,... h n}; Each cell calculates the current hidden vector h t-1 based on the previous hidden vector h t and the current input vector x t . Its operation is defined as follows: i t = σ(W xi x t + W hi h t-1 + W ci c t-1 + b i ) f t = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f ) c t = f t c t-1 + i t tanh(W xc x t + W hc h t-1 + b c ) o t = σ(W x0 x t + W ho h t-1 + W co c t + b o ) h t = o t tanh(c t ) where: i t , f t , c t , o t , h t are the states of the memory gate, hidden layer, forget gate, cell nucleus, and output gate when inputting the t-th text, respectively; W are the parameters of the model; b is the bias vector; σ is the Sigmoid function; tanh is the hyperbolic tangent function; The late fusion of the part-of-speech features and semantic features includes: First, input the part-of-speech features and semantic features into Bi-LSTM respectively, and then splice the results of the two models to form a basic feature vector; The associated information extraction method includes: (1) Associated information based on mutual information: The definition of the mutual information of two discrete random variables X and Y is: where \(p(x,y)\) is the joint probability distribution function of \(X\) and \(Y\), and \(p(x)\) and \(p(y)\) are the marginal probability distribution functions of \(X\) and \(Y\) respectively; if we want to measure the degree of association between any two words \(x,y\) in a certain dataset, we calculate as follows: Among them, \(p(x)\) and \(p(y)\) are the probabilities of \(x\) and \(y\) appearing independently in the dataset, which can be obtained by directly counting the word frequencies and then dividing by the total number of words; \(p(x,y)\) is the probability that \(x\) and \(y\) appear in the dataset simultaneously, which can be obtained by directly counting the number of times they appear simultaneously and then dividing by the number of all unordered pairs; (2) Associated information based on dependency syntactic analysis: Dependency syntax reveals the dependency relationship and collocation relationship between words in a sentence. One dependency relationship connects two words, one is the core word and the other is the modifier. Such a relationship is interrelated with the semantic relationship of the sentence; (3) Feature representation based on graph convolutional neural network: The relationship between words is represented by a graph through MI and dependency syntactic analysis, so a graph convolutional neural network is used to process the associated information; Given a graph G=(V, E), V is the vertex set containing N nodes, and E is the edge set including self-loop edges. The feature information of the graph G(V, E) can be represented by the Laplacian matrix L, as shown in the following formula: L = D - A Or use the symmetrically normalized Laplacian matrix: L sys = I N - D -1 / 2 AD -1 / 2 where: A is the adjacency matrix of the graph; I N is the N-order identity matrix; D = diag(d) is the degree matrix of the vertices; Based on the Fourier transform of the graph, the graph convolution formula is expressed as: g * x = U[(U T g)·(U T x)] In the formula: x is the basic feature vector of the node; g is the convolution kernel; U is the eigenvector matrix of the Laplacian matrix L; Use Chebyshev polynomials to simplify the graph convolution formula, and finally the graph convolution layer propagation formula can be expressed as: In the formula: σ is the activation function; W is the weight matrix to be trained; The label representation method includes: In CRF, each sentence X={x1, x2, …, xn} has a set of candidate label sequences. The final annotation sequence is determined by calculating the scores of each label sequence Y={y1, y2, …, yn} in the set. The process of calculating the scores is shown in the formula: where \(P\in\mathbb{R}\) n×k is a scoring matrix, and \(k\) is the number of all labels; \(A\in\mathbb{R}\) (k+2)×(k+2) is a transition matrix that includes sentence start and end labels. Normalize the scores of each label sequence to obtain probabilities. The label sequence with the highest probability is the final sequence of the sentence. The normalization process is shown in the formula:
2. A system applying the method for identifying formulaic language by fusing associated information according to claim 1, characterized in that, it includes: The basic feature extraction module is used to generate word embedding vectors as part-of-speech features using the embedding layer in Torch, and feature vectors trained by the GloVe word vector technology as semantic features. The part-of-speech features and semantic features after late fusion are used as the basic features of the model; The associated information extraction module is used to adopt the mutual information between words and the dependency syntactic relationship of sentences as the associated information for identifying formulaic language; The label representation module is used to represent labels.
Citation Information
Patent Citations
GCN-based Chinese complex sentence implicit relationship analysis method and device
CN113378547A