A multi-task citation intent classification system, construction method and application

By building a multi-task citation intention classification system, using soft parameter sharing constraints and SciBERT pre-trained language model, combining heterogeneous feature sets and attention mechanisms, the problems of insufficient sample size and inaccurate classification in the existing technology are solved, and efficient citation intention classification is achieved.

CN116415171BActive Publication Date: 2025-05-13DALIAN UNIVERSITY OF FOREIGN LANGUAGES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310392056.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-05-13
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

The existing citation intention analysis system requires a large number of samples to be trained and the classification is inaccurate, making it difficult to deal with complex citation intention analysis of big data.

Method used

A multi-task citation intention classification system is proposed. By constructing a multi-task classification framework with soft parameter sharing constraints, combining SciBERT pre-trained language model and heterogeneous feature set, multi-headed attention mechanism and self-attention mechanism are used to calculate task weights, and multi-task joint learning is performed using sparse classification cross entropy as a loss function.

Benefits of technology

Improve the performance of citation intent classification, improve accuracy and recall on small data sets, and does not require a large number of sample training, and can perform citation intent classification more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116415171B_ABST
    Figure CN116415171B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for constructing a multi-task citation intention classification system, including: obtaining a task corpus and multiple auxiliary task corpora, inputting the main task corpus into the SciBERT pre-trained language model to obtain the SciBERT pre-trained language model representation of the main task corpus; inputting the multiple auxiliary task corpora into the SciBERT pre-trained language model respectively to obtain the SciBERT pre-trained language model representations of the multiple auxiliary task corpora; extracting the heterogeneous feature set of the main task corpus, and fusing the extracted heterogeneous feature set with the SciBERT pre-trained language model representation of the main task corpus to obtain Input intent ; using the SciBERT pre-trained language model representation of the auxiliary task corpus obtained in step S1 as Input auxiliary ; learning the hidden features in Inputint en through soft parameter sharing constraints and BiLSTM; learning the hidden features in Input auxiliary through BiLSTM; calculating the main task hidden feature weights using the multi-head attention mechanism and calculating the auxiliary task hidden feature weights using the self-attention mechanism; jointly learning multiple tasks to obtain a citation intention classification system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of citation intent classification, and in particular to a multi-task citation intent classification system, a construction method and an application.

[0002] Background

[0003] Citation intent refers to the inner psychological activity of authors when citing literature, reflecting the reasons and purposes of citing literature. Since citations play an important role in the evaluation of scientific literature and researchers, citation analysis is crucial to understanding academic literature. The purpose of citation intent classification research is to study how to classify the author's citation purpose into a specific category, which helps to conduct in-depth and systematic analysis of citation intent. In addition, in-depth research on citation intent classification is helpful for other bibliometric analysis tasks, such as citation importance analysis, academic impact assessment, and potential scholar relationship discovery. Traditional citation intent classification research mainly relies on manual qualitative analysis of small sample data, but manual qualitative methods are difficult to cope with the rapid growth of publicly available full-text literature. In addition, when considering generalization and scalability, manual analysis cannot handle complex citation intent analysis of big data. With the development of natural language processing technology, it has become possible to automatically extract semantic information from large document texts for citation intent analysis. Early research on automatic classification of citation intent was mainly based on feature engineering. Traditional citation intent features include content, location, sentence, syntax, structural features, etc. However, feature engineering-based methods analyze the citation context based on a set of predefined features, and it is difficult to extract enough information to represent the semantics in the citation context.

[0004] In the recent progress of citation analysis, deep learning is a direction worthy of further attention. Deep learning technology uses word vectors to represent citation text and uses pre-trained language models to capture semantic information in context. Pre-trained word vector technology saves the cost of feature engineering and performs better than machine learning citation intent analysis methods. However, due to insufficient training samples for citation intent analysis tasks, some deep learning models may not be able to effectively learn relevant features. Therefore, it is necessary to develop a system and method to improve the performance of citation intent analysis tasks on small data sets by exploring the information related to intent in the annotation information of citation analysis corpus related tasks. Summary of the invention

[0005] In order to solve the problems that the existing citation intent analysis system requires a large number of sample training for learning and training, and the citation intent classification is inaccurate, the present invention proposes a multi-task citation intent classification system, construction method and application. Considering the correlation between citation intent, citation position and citation value classification tasks, a multi-task citation classification framework with soft parameter sharing constraints is constructed, and independent models are constructed for multiple tasks to improve the performance of citation intent classification.

[0006] The present invention provides a method for constructing a multi-task citation intent classification system, comprising the following steps:

[0007] S1. Obtain a main task corpus of a main task and auxiliary task corpora of multiple auxiliary tasks, input the main task corpus into the SciBERT pre-trained language model, and obtain the SciBERT pre-trained language model representation of the main task corpus; input the multiple auxiliary task corpora into the SciBERT pre-trained language model respectively, and obtain the SciBERT pre-trained language model representation of the multiple auxiliary task corpora;

[0008] S2. Extract the heterogeneous feature set of the main task corpus, and fuse the extracted heterogeneous feature set with the SciBERT pre-trained language model representation of the main task corpus obtained in step S1 to obtain Input intent ; Take the SciBERT pre-trained language model representation of the auxiliary task corpus obtained in step S1 as input auxiliary ;

[0009] S3. Input obtained by soft parameter sharing constraints and BiLSTM learning step S2 intent The implicit features in ; Input is obtained through BiLSTM learning step S2 auxiliary Implicit features in

[0010] S4. After the learning of the implicit features of the main task and the implicit features of the auxiliary task in step S3 is completed, the main task uses a multi-head attention mechanism to calculate the implicit feature weight of the main task, and the auxiliary task uses a self-attention mechanism to calculate the implicit feature weight of the auxiliary task;

[0011] S5. Use multi-task joint learning, use sparse classification cross entropy as the loss function, regularize the main task and the auxiliary task, obtain the weight adjustment factors of the main task and the auxiliary task, and obtain a citation intent classification system.

[0012] Furthermore, in step S1, the main task is a citation intention classification task, and the main task corpus is a citation intention classification task corpus; the auxiliary tasks include a citation position classification task and a citation value classification task, the citation position classification task is used to classify the title of the paragraph where the citation is located, and the citation value classification task is used to classify whether a sentence includes a citation, and the auxiliary task corpus includes a citation position classification task corpus and a citation value classification task corpus, and the citation position classification task corpus and the citation value classification task corpus do not contain citation intention classification annotations.

[0013] Furthermore, in step S2, the Input intentAs shown in formula (1),

[0014] Input intent =(SciBERT(S),feature ij (S)) (1)

[0015] Among them, Input intent is the input of the main task citation intent classification task, SciBERT(S) is the SciBERT pre-trained language model representation of the main task corpus, i is the number of sentences, j is the number of words in the sentence, and feature ij (S) is the heterogeneous feature set of the main task corpus;

[0016] feature ij As shown in formula (2),

[0017] feature ij =cat([Onehot j (pos j ,pos_list),pattern j ,tfidf ij ,senti j ]) (2)

[0018] Among them, Onehot j (pos j ,pos_list) is the part-of-speech tagging feature represented by a one-hot vector, pattern j is the syntactic structure feature, tfidf ij represents the TF-IDF value of word j in sentence i, senti j It is the weighted embedding representation of the domain sentiment word vector multiplied by the TF-IDF vector.

[0019] pattern j It includes the following six syntactic structures:

[0020] i) reference + verb + verb [past / present / third person / past / past participle];

[0021] ii) verb [past / gerund / third person] + verb [gerund / past participle];

[0022] iii) Verb [all forms] + (Adverb [Comparative / Superlative]) + Verb + (Adverb [Comparative / Superlative]) + Past Participle;

[0023] iv) modal verb + (adverb [comparative / superlative]) + verb + (adverb [comparative / superlative]) + past participle;

[0024] v) (adverb [comparative / superlative]) + personal pronoun + (adverb [comparative / superlative]) + verb [all forms];

[0025] vi) gerund + (proper noun + juxtaposition conjunction + proper noun),

[0026] tfidf ij The calculation of is shown in (3),

[0027]

[0028] where d is an instance in the citation corpus, and f j,d is the frequency of word j in instance d, N is the number of instances in the citation corpus, and n j is the number of instances containing word j.

[0029] Furthermore, in step S2, the input of the citation position classification task of the auxiliary task is as shown in formula (4), Input section =SciBERT(S section ')(4)

[0030] Among them, Input section As the input for the citation position classification task, SciBERT (S section ') is the SciBERT pre-trained language model representation for the citation position classification task;

[0031] The input of the auxiliary task of citation position classification is shown in formula (5):

[0032] Input worth =SciBERT(S worth ')(5)

[0033] Among them, Input worth As the input of the citation value classification task, SciBERT (S worth ') is the SciBERT pre-trained language model representation for the citation value classification task.

[0034] Furthermore, in step S4, the main task divides the input into multiple subspaces during the attention weight calculation process, calculates the subspace weights through the weight parameter matrix of multiple subspace query vectors, subspace key information and subspace word embedding vectors, and shares the training parameters in these multiple subspaces to obtain the main task weight; the auxiliary task maps the input to the auxiliary task query vector, auxiliary task key information and auxiliary task word embedding vector during the attention weight calculation process, generates an auxiliary task attention weight map by the dot product of the auxiliary task query vector and the auxiliary task key information, and then generates an attention weighted feature by the dot product of the auxiliary task attention weight map and the auxiliary task word embedding vector, adjusts the dot product by the adjustment factor of the smoothed Softmax function, and obtains the auxiliary task weight.

[0035] Furthermore, in step S5, the citation value classification task uses the Sigmoid function as the activation function of multi-task joint learning, and the citation intent classification task and the citation position classification task use the Softmax function as the activation function of multi-task joint learning.

[0036] Furthermore, in step S5, sparse classification cross entropy is used as the loss function and regularization constraints are performed, as shown in formula (6),

[0037]

[0038] Among them, p i is the probability that the instance belongs to category i, is the probability that the instance does not belong to category i, w i is the parameter weight matrix of the fully connected multi-task learning layer,

[0039] The joint loss function L(w) is shown in formula (7):

[0040]

[0041] Among them, J j (w) is the loss function of task j, λ i It is the weight adjustment factor of each task.

[0042] The present invention also provides a multi-task citation intention classification system, which is obtained by adopting the above-mentioned multi-task citation intention classification system construction method, and includes an input module, a representation module, a context feature learning module, an attention module, a multi-task joint learning module and a citation intention output module;

[0043] The input module is used to input a data set and transmit the data set to the representation module, the input module includes a main task input module and an auxiliary task input module, the main task input module is used to input a main task data set, and the auxiliary task input module is used to input an auxiliary task data set;

[0044] The representation module is used to represent the input data set to obtain input data and transmit the input data to the context feature learning module. The representation module includes a main task representation module and an auxiliary task representation module. The main task representation module obtains the main task input data by fusing the heterogeneous feature set with the SciBERT pre-trained language model, and the auxiliary task representation module obtains the auxiliary task input data by using the SciBERT pre-trained language model.

[0045] The context feature learning module is used to learn the input vector to obtain implicit features and transmit the implicit features to the attention module. The context feature learning module includes a main task context feature learning module and an auxiliary task context feature learning module. The main task context feature learning module learns the main task implicit features through soft parameter sharing constraints and BiLSTM, and the auxiliary task context feature learning module learns the auxiliary task implicit features through BiLSTM.

[0046] The attention module is used to calculate the task weight according to the implicit feature and transmit the calculated task weight to the multi-task joint learning module, the attention module includes a main task attention module and an auxiliary task attention module, the main task attention module calculates the main task weight through a multi-head attention mechanism, and the auxiliary task attention module calculates the auxiliary task weight through a self-attention mechanism;

[0047] The multi-task joint learning module is used to use sparse classification cross entropy as a loss function, regularize the main task and the auxiliary task, obtain the weight adjustment factors of the main task and the auxiliary task, and transmit the weight adjustment factors to the citation intention output module;

[0048] The citation intent output module outputs a citation intent classification based on the weight adjustment factor.

[0049] The present invention also provides an application of a multi-task citation intent classification system in citation intent classification, wherein a data set to be classified for citation intent is input into the multi-task citation intent classification system to obtain citation intent classification data.

[0050] The present invention proposes a multi-task citation intent classification system, construction method and application, extracts heterogeneous feature sets and fuses them with pre-trained language models, the heterogeneous features provide useful information for citation intent prediction, based on soft parameter sharing constraints and BiLSTM learning implicit features, uses a multi-head attention mechanism to calculate the main task weight, and uses a self-attention mechanism to calculate the auxiliary task weight, which can extract implicit information in the feature space that is beneficial to citation intent classification. In addition, by setting auxiliary tasks and extracting heterogeneous features, the citation intent recognition performance for categories with few instances is improved, and the system and method of the present invention can classify citation intent more accurately. In addition, without learning through a large amount of data, only by learning a small data set, the precision and recall of the citation intent classification task can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic flow chart of a method for constructing a multi-task citation intent classification system according to embodiment 1 of the present invention;

[0052] Figure 2 A schematic diagram of the structure of a multi-task citation intent classification system according to embodiment 1 of the present invention;

[0053] Figure 3 Confusion matrix diagram of the multi-task citation intent classification system of Example 2 of the present invention;

[0054] Figure 4 Confusion matrix diagram of a multi-task citation intent classification system removing heterogeneous feature sets in Example 2 of the present invention;

[0055] Figure 5 Confusion matrix diagram of multi-task citation intent classification system with two auxiliary tasks removed in Example 2 of the present invention;

[0056] Figure 6 A visualization diagram of attention weights of a multi-task citation intent classification system in Example 2 of the present invention. DETAILED DESCRIPTION

[0057] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0058] Example 1

[0059] A method for constructing a multi-task citation intent classification system, such as Figure 1 As shown, the following steps are included:

[0060] S1. Obtain a main task corpus of a main task and auxiliary task corpora of multiple auxiliary tasks, input the main task corpus into the SciBERT pre-trained language model, and obtain the SciBERT pre-trained language model representation of the main task corpus; input the multiple auxiliary task corpora into the SciBERT pre-trained language model respectively, and obtain the SciBERT pre-trained language model representation of the multiple auxiliary task corpora;

[0061] In step S1, the main task is a citation intent classification task, and the main task corpus is a citation intent classification task corpus. As a preferred solution of this embodiment, the auxiliary tasks include a citation position classification task and a citation value classification task. The citation position classification task is used to classify the title of the paragraph where the citation is located, and the citation value classification task is used to classify whether a sentence includes a citation. The auxiliary task corpus includes a citation position classification task corpus and a citation value classification task corpus. The citation position classification task corpus and the citation value classification task corpus do not contain citation intent classification annotations.

[0062] Citation intention classification task: Given a quoted sentence, determine its citation intention. The citation intention is annotated into one of the following six categories: background, extension, use, motivation, comparison or future work / future, etc. Citation position classification task: Scientific and technological literature generally follows a certain structure, such as introducing the problem first, describing the method, discussing the results of the discovery, and finally summarizing the entire paper. Since the citation intention is related to the paragraph where the citation is located, this task is to classify the title of the paragraph where the citation is located. It is a multi-classification task, and the citation is divided into one of the five categories: introduction, related work, method, experiments, and conclusion. Citation value classification task: In scientific and technological literature, there are great differences in language features between sentences containing citations and general sentences. This task is to determine whether a sentence contains citations. It is a binary classification task. For example, a positive example (True) is a sentence containing citations, and a negative example (False) is a sentence that does not contain citations. For scientific and technological literature, there is a certain connection between the structure of the article and the citation intent. Therefore, in this embodiment, two auxiliary tasks related to the structure of the paper are proposed to help improve the performance of the main task of the citation intent classification task system.

[0063] S2. Extract the heterogeneous feature set of the main task corpus, and fuse the extracted heterogeneous feature set with the SciBERT pre-trained language model representation of the main task corpus obtained in step S1 to obtain Input intentThe SciBERT pre-trained language model representation of the main task corpus includes general semantics and scientific domain information from SciBERT, and the heterogeneous feature set includes heterogeneous citation information from grammatical, statistical, and sentiment features; the SciBERT pre-trained language model representation of the auxiliary task corpus obtained in step S1 is used as input auxiliary ,Similarly, the SciBERT pre-trained language model representation of the auxiliary task ,corpus includes general semantic and scientific domain information from SciBERT;

[0064] As a preferred solution of this embodiment, in step S2, Input intent As shown in formula (1),

[0065] Input intent =(SciBERT(S),feature ij (S)) (1)

[0066] Among them, Input intent is the input of the main task citation intent classification task, SciBERT(S) is the SciBERT pre-trained language model representation of the main task corpus, i is the number of sentences, j is the number of words in the sentence, and feature ij (S) is the heterogeneous feature set of the main task corpus;

[0067] feature ij As shown in formula (2),

[0068] feature ij =cat([Onehot j (pos j ,pos_list),pattern j ,tfidf ij ,senti j ]) (2)

[0069] Among them, Onehot j (pos j ,pos_list) is the part-of-speech tagging feature represented by a one-hot vector, pattern j is the syntactic structure feature, tfidf ij represents the TF-IDF value of word j in sentence i, senti j It is the weighted embedding representation of the domain sentiment word vector multiplied by the TF-IDF vector.

[0070] pattern j It includes the following six syntactic structures:

[0071] i) reference + verb + verb [past / present / third person / past / past participle];

[0072] ii) verb [past / gerund / third person] + verb [gerund / past participle];

[0073] iii) Verb [all forms] + (Adverb [Comparative / Superlative]) + Verb + (Adverb [Comparative / Superlative]) + Past Participle;

[0074] iv) modal verb + (adverb [comparative / superlative]) + verb + (adverb [comparative / superlative]) + past participle;

[0075] v) (adverb [comparative / superlative]) + personal pronoun + (adverb [comparative / superlative]) + verb [all forms];

[0076] vi) gerund + (proper noun + juxtaposition conjunction + proper noun),

[0077] tfidf ij The calculation of is shown in (3),

[0078]

[0079] where d is an instance in the citation corpus, and f j,d is the frequency of word j in instance d, N is the number of instances in the citation corpus, and n j is the number of instances containing word j.

[0080] In step S2, the input of the auxiliary task of citation position classification is shown in formula (4):

[0081] Input section =SciBERT(S section ') (4)

[0082] Among them, Input section As the input for the citation position classification task, SciBERT (S section ') is the SciBERT pre-trained language model representation for the citation position classification task;

[0083] The input vector of the auxiliary task citation position classification task is shown in formula (5):

[0084] Input worth =SciBERT(S worth ') (5)

[0085] Among them, Input worth As the input of the citation value classification task, SciBERT (Sworth ') is the SciBERT pre-trained language model representation for the citation value classification task.

[0086] S3. Input obtained by soft parameter sharing constraints and BiLSTM learning step S2 intent The implicit features in ; Input is obtained through BiLSTM learning step S2 auxiliary Implicit features in

[0087] S4. After the implicit features of the main task and the implicit features of the auxiliary task are learned in step S3, the main task uses a multi-head attention mechanism to calculate the weight of the implicit features of the main task. In addition to improving the training efficiency, the multi-head attention mechanism divides the input vector into multiple subspaces during the attention weight calculation process and shares the training parameters in these multiple subspaces, further optimizing the overall training performance. The auxiliary task uses a self-attention mechanism to calculate the weight of the implicit features of the auxiliary task;

[0088] As a preferred embodiment of this embodiment, in step S4, the main task divides the input into multiple subspaces during the attention weight calculation process, calculates the subspace weights through the weight parameter matrix of multiple subspace query vectors, subspace key information and subspace word embedding vectors, and shares the training parameters in these multiple subspaces to obtain the main task implicit feature weights; the auxiliary task maps the input to the auxiliary task query vector, auxiliary task key information and auxiliary task word embedding vector during the attention weight calculation process, generates an auxiliary task attention weight map by the dot product of the auxiliary task query vector and the auxiliary task key information, and then generates attention weighted features by the dot product of the auxiliary task attention weight map and the auxiliary task word embedding vector, adjusts the dot product by the adjustment factor of the smoothed Softmax function, and obtains the auxiliary task implicit feature weights.

[0089] S5. Use multi-task joint learning, use sparse classification cross entropy as the loss function, regularize the main task and the auxiliary task, obtain the weight adjustment factors of the main task and the auxiliary task, and obtain a citation intent classification system.

[0090] As a preferred embodiment of this embodiment, in step S5, the citation value classification task uses the Sigmoid function as the activation function of multi-task joint learning, and the citation intent classification task and the citation position classification task use the Softmax function as the activation function of multi-task joint learning.

[0091] In step S5, sparse classification cross entropy is used as the loss function and regularization constraints are performed, as shown in formula (6),

[0092]

[0093] Among them, pi is the probability that the instance belongs to category i, is the probability that the instance does not belong to category i, w i is the parameter weight matrix of the fully connected multi-task learning layer,

[0094] The joint loss function L(w) is shown in formula (7):

[0095]

[0096] Among them, J j (w) is the loss function of task j, λ i It is the weight adjustment factor of each task.

[0097] The present invention also provides a multi-task citation intention classification system, which is obtained by using the multi-task citation intention classification system construction method, such as Figure 2 As shown, it includes an input module, a representation module, a context feature learning module, an attention module, a multi-task joint learning module and a citation intention output module;

[0098] The input module is used to input a data set and transmit the data set to the representation module. The input module includes a main task input module and an auxiliary task input module. The main task input module is used to input a main task data set, and the auxiliary task input module is used to input an auxiliary task data set.

[0099] The representation module is used to represent the input data set to obtain input data and transmit the input data to the context feature learning module. The representation module includes a main task representation module and an auxiliary task representation module. The main task representation module obtains the main task input data by fusing the heterogeneous feature set with the SciBERT pre-trained language model, and the auxiliary task representation module obtains the auxiliary task input data through the SciBERT pre-trained language model.

[0100] The context feature learning module is used to learn the input vector to obtain implicit features and transmit the implicit features to the attention module. The context feature learning module includes a main task context feature learning module and an auxiliary task context feature learning module. The main task context feature learning module learns the main task implicit features through soft parameter sharing constraints and BiLSTM, and the auxiliary task context feature learning module learns the auxiliary task implicit features through BiLSTM.

[0101] The attention module is used to calculate the task weight according to the implicit features and transmit the calculated task weight to the multi-task joint learning module. The attention module includes the main task attention module and the auxiliary task attention module. The main task attention module calculates the main task weight through the multi-head attention mechanism, and the auxiliary task attention module calculates the auxiliary task weight through the self-attention mechanism.

[0102] The multi-task joint learning module is used to use sparse classification cross entropy as the loss function, regularize the main task and the auxiliary task, obtain the weight adjustment factors of the main task and the auxiliary task, and transmit the weight adjustment factors to the citation intention output module;

[0103] The citation intent output module outputs the citation intent classification based on the weight adjustment factor.

[0104] The present invention also provides an application of a multi-task citation intent classification system in citation intent classification, inputting a data set to be classified with citation intent into the multi-task citation intent classification system to obtain citation intent classification data.

[0105] Example 2

[0106] This embodiment uses the public citation analysis dataset ACL-ARC as a benchmark dataset to compare the performance of the system proposed in the present invention. In the ACL-ARC dataset, three tasks were manually annotated by domain experts in the field of natural language processing, namely, citation intent classification task, citation value classification task and citation location classification task. The corpus annotation data of each task in the corpus is different. Among them, the citation intent classification task corpus sub-dataset includes 1941 instances, which are divided into 6 categories according to the ACT annotation system, including: background, extension, use, motivation, comparison (Compare / Contrast) and future work (Future). The citation value classification task corpus sub-dataset contains 50,000 instances, which are divided into two categories: positive examples (True) are sentences containing citations, and negative examples (False) are sentences that do not contain citations. Among them, only 14% belong to the "True" positive example, that is, the "valuable" category, which belongs to an unbalanced dataset. The citation position sub-dataset has 47,757 instances, which are divided into five categories: Introduction, Related Work, Method, Experiments, and Conclusion, of which 45% belong to the “Introduction” abstract category among the five categories. Table 1 shows the data distribution of the ACL-ARC dataset in detail.

[0107] Table 1 Distribution of ACL-ARC dataset

[0108]

[0109] In the experiment, the ACL-ARC dataset was divided into three groups, of which 85% of the data was used for training, and the remaining 15% was evenly divided into the validation set and the test set. The hyperparameters of deep learning were set to 20 Epochs, 32 Batch size, and 0.001 Learning rate. The weight adjustment factor λ of each task i Optimization is performed through grid search with a step size of 0.01 to obtain the best performance of multiple tasks on the validation set. The weight adjustment factor λ of the main task citation intent classification task 0 Set to 1, the weight adjustment factor λ of the citation value classification task 1 Set to 0.1, the weight adjustment factor λ for the citation position classification task 2 It was set to 0.08. In order to avoid the accidental results, the experiment was repeated 20 times and the average value was recorded as the final result.

[0110] In order to evaluate the multi-task citation intent classification system (MTCIC) proposed in this study, compared with the main baselines and ablation models on the ACL-ARC dataset, the control experimental method is as follows:

[0111] Experimental method 1: Random Forest by Jurgens et al. This experimental method uses the random forest algorithm, and the feature set includes pattern-based features, topic modeling features, citation graph features, section title and section position features.

[0112] Experimental method 2: BiLSTM-Attn-ELMo by Cohan et al. This experimental method uses ELMo and Glove word vectors as input to a BiLSTM network with an attention mechanism, and optimizes the network using only the loss function of the main task of citation intent classification.

[0113] Experimental method 3: Structural-Scaffold by Cohan et al. This experimental method adopts the Structural-Scaffold multi-task structure based on hard constraints, uses ELMo and Glove word vectors as input, and optimizes the network using a multi-task loss function.

[0114] Experimental Method 4: SciBERT Finetune by Beltagy et al. This baseline uses the fine-tuned SciBERT pre-trained model as input to the fully connected layer of the citation intent classification task.

[0115] Experimental method 5: Multi-task citation intent classification system (MTCIC). The system proposed in this invention takes SciBERT and heterogeneous features as input, processes multiple classification tasks such as citation intent classification task, citation value classification task and citation position classification task at the same time, and outputs the result of recording citation intent classification.

[0116] Experimental method 6: Multi-task citation intent classification system removes heterogeneous feature sets (MTCIC withoutfeatures), with only SciBERT as input. Only SciBERT is used as input, and multiple classification tasks such as citation intent classification task, citation value classification task and citation location classification task are processed at the same time, and the output records the results of citation intent classification.

[0117] Experimental method 7: Multi-task citation intent classification system removes the citation location classification task (MTCIC without section-task). Taking SciBERT and heterogeneous features as input, it simultaneously processes the multi-classification tasks of citation intent classification task and citation value classification task, and outputs the results of recording citation intent classification.

[0118] Experimental method 8: Multi-task citation intent classification system removes the citation value classification task (MTCIC withoutworthiness-task). Taking SciBERT and heterogeneous features as input, it simultaneously processes the multi-classification tasks of citation intent classification task and citation location classification task, and outputs the result of recording citation intent classification.

[0119] Experimental method 9: Multi-task citation intent classification system removes two auxiliary tasks (MTCIC without any auxiliary task). Take SciBERT and heterogeneous features as input, only process the citation intent classification task, and output the result of recording the citation intent classification.

[0120] Experimental method 10: Multi-task citation intent classification system without multi-head attention mechanism (MTCIC withoutmulti-head). Taking SciBERT and heterogeneous features as input, it simultaneously processes multi-classification tasks such as citation intent classification task, citation value classification task and citation position classification task, removes the multi-task citation intent classification system (MTCIC) without multi-head attention, and outputs the result of recording citation intent classification.

[0121] The experimental results of the above experiments on the ACL-ARC dataset are shown in Table 2. The highest value of each indicator is indicated in bold.

[0122] Table 2 Comparison of citation intent classification results with baseline classification results

[0123]

[0124]

[0125] First, it is observed that the proposed multi-task citation intent classification system (MTCIC) has achieved significant improvements over existing methods in the citation intent classification task. Overall, the multi-task citation intent classification system (MTCIC) significantly improves the classification performance of key indicators. Compared with the Random Forest of Jurgens et al. in Experimental Method 1, the Marco-F1 of the multi-task citation intent classification system (MTCIC) is improved by 21.18%, the recall rate (Macro Recal) is improved by 22.9%, and the precision rate (Marco Precision) is improved by 17.17%. Compared with the Structural-Scaffold of Cohan et al. in Experimental Method 3, the Marco-F1 of the multi-task citation intent classification system (MTCIC) is improved by 7.88%, the recall rate (Macro Recal) is improved by 10.3%, and the precision rate (Marco Precision) is improved by 0.77%. This result shows that compared with Structural-Scaffold based on word embedding representation and hard-constrained multi-task framework, the joint heterogeneous citation representation model with regularization constraints and multi-task framework proposed in this paper provide useful information for citation intent classification.

[0126] In addition, compared with the SciBERT fine-tuned model of BiLSTM-Attn-ELMo of Cohan et al. in Experimental Method 2, the Marco-F1 of the Multi-Task Citation Intent Classification System (MTCIC) has obvious advantages, proving that the improvement in citation intent classification performance mainly comes from the heterogeneous features and regularization constrained multi-task framework, rather than from the SciBERT pre-trained language model.

[0127] In the ablation experiment results in Table 2, in Experimental Method 6, the multi-task citation intent classification system removes heterogeneous feature sets (MTCIC without features). After removing the heterogeneous feature sets, Marco-F1 drops from 75.78% to 73.21%, indicating that heterogeneous features provide useful information for citation intent prediction. In Experimental Method 10, the multi-task citation intent classification system removes the multi-head attention mechanism (MTCIC without multi-head). After removing the multi-head attention mechanism, Marco-F1 drops from 75.78% to 74.43%, indicating that the multi-head attention mechanism helps to extract implicit information in the feature space that is beneficial to citation intent classification. In the task ablation experiment, in the experimental method 7 multi-task citation intent classification system, after removing the citation position classification task (MTCIC without section-task), Marco-F1 dropped from 75.78% to 73.27%; in the experimental method 8 multi-task citation intent classification system, after removing the citation value classification task (MTCIC withoutworthiness-task), Marco-F1 dropped from 75.78% to 73.07%; in the experimental method 9 multi-task citation intent classification system, after removing two auxiliary tasks (MTCIC without any auxiliary task), Marco-F1 dropped from 75.78% to 72.75% when both auxiliary tasks were removed. This shows that the two auxiliary tasks, the citation position classification task and the citation value classification task, respectively provide useful supplementary information for the citation intent classification task.

[0128] Since the distribution of the six citation intent categories in the ACL-ARC dataset is obviously unbalanced, as shown in Table 1, in the experimental results, Micro F1, precision and recall are greatly affected by the categories with more instances. In order to observe the performance of MTCIC on each category, the classification results on all citation intent category instances in the ACL-ARC dataset are recorded in this embodiment, as shown in Table 3, where the two highest Micro F1 scores on each category are shown in bold. It can be observed that compared with the baseline model, the multi-task citation intent classification system proposed in this embodiment obtains higher Marco F1 values ​​in most categories, including the "Background" category and some categories with fewer instances, such as "Extentsion", "Future work / Future", "Motivation" and "Use". While most baseline models, such as Cohan et al.’s Structural-Scaffold, achieve higher F1 in the “Background” category, they have significantly lower Marco-F1 in categories with fewer instances, such as “Extentsion”, “Future work / Future”, “Motivation”, and “Use”.

[0129] Table 3 Experimental results on various categories

[0130]

[0131] The top two Micro F1 scores are in bold, p: precision, R: recall, F1: Micro F1.

[0132] In the ablation model, when the multi-task citation intent classification system removes the heterogeneous feature set (MTCIC withoutfeatures), the Marco-F1 of the "Extentsion", "Future work / Future", and "Use" categories is significantly reduced, proving that heterogeneous features are helpful for citation intent recognition. When the multi-task citation intent classification system removes the citation value classification task (MTCIC without worthiness-task), the Marco-F1 of the category with few instances is significantly reduced. When the multi-task citation intent classification system removes the citation location classification task (MTCIC without section-task), the Marco-F1 scores of the "Future work / Future", "Motivation", and "Use" categories are significantly reduced. When the multi-head attention mechanism is removed from the multi-task citation intent classification system (MTCICwithout multi-head), the Marco-F1 of the "Extentsion", "Future work / Future" and "Use" categories is significantly reduced, which shows that the multi-head attention mechanism improves the recognition ability of citation intent categories with few instances.

[0133] In order to observe the performance of the multi-task citation intent classification system when dealing with unbalanced data sets, the confusion matrices of the multi-task citation intent classification system (MTCIC), the multi-task citation intent classification system removing heterogeneous feature sets (MTCIC withoutfeatures), and the multi-task citation intent classification system removing two auxiliary tasks (MTCIC without any auxiliarytask) are selected, as shown in Figure 3-5 As shown. Figure 3-5It can be seen that there are still some instances of “Background” and “Compare” that are misclassified. Several instances of “Background”, “Compare”, and “Future work / Future” are misclassified as “Use”, while the erroneous instances of “Motivation” are mainly misclassified as “Compare” or “Background”. When the MTCIC without features is used, “Background” is more easily confused with “Extentsion”, “Future work / Future”, or “Motivation”; there is more confusion between “Compare” and “Background”. Instances of “Future work / Future” are more likely to be misclassified as “Background”; instances of “Use” are more often misclassified as “Background”, “Compare”, or “Future work / Future”. When the multi-task citation intent classification system removes the two auxiliary tasks (MTCIC withoutany auxiliary task), the true positive rate of “Background”, “Compare”, and “Use” is lower, and “Use” is misclassified as “Extentsion” or “Future work / Future”. However, the single-task model of MTCIC performs better on the “Extentsion” category, probably because the information from the two auxiliary tasks slightly interferes with the judgment of the “Extentsion” category.

[0134] In order to gain a deeper understanding of how the multi-task mechanism helps the Multi-Task Citation Intent Classification system (MTCIC) improve citation intent classification, we examined the attention weights obtained for each word of the input instance, and studied the instance “We will examine the worst-case complexity of interpretation as well as generation to shed some light on the hypothesis that vague descriptions are more difficult to process than others because they involve a comparison between objects (Beun and Cremers 1998, Krahmer and Theune 2002)”, whose true label is “Background”.

[0135] Figure 6 The heatmap of attention weights of input instances in the multi-task MTCIC system and the MTCIC system without any auxiliary task is shown. Figure 4 It can be seen that the multi-task MTCIC pays more attention to the words around "generation toshed" and "comparison between objects" and predicts the true label for this instance, which is reasonable. However, the MTCIC without any auxiliary task pays the most attention to "examine the worst-case", so it is incorrectly predicted as the "Compare" label. Since the only difference between the two models is whether the auxiliary tasks are included, this shows that the auxiliary tasks provide useful information for the main task of citation intent classification.

[0136] The embodiments of the present invention are given for the purpose of illustration and description, and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.

Claims

1. A method for constructing a multi-task citation intent classification system, characterized in that: The steps include: S1. Obtain a main task corpus of a main task and auxiliary task corpora of multiple auxiliary tasks, input the main task corpus into the SciBERT pre-trained language model, and obtain the SciBERT pre-trained language model representation of the main task corpus; input the multiple auxiliary task corpora into the SciBERT pre-trained language model respectively, and obtain the SciBERT pre-trained language model representation of the multiple auxiliary task corpora; S2. Extract the heterogeneous feature set of the main task corpus, and fuse the extracted heterogeneous feature set with the SciBERT pre-trained language model representation of the main task corpus obtained in step S1 to obtain Input intent ; Take the SciBERT pre-trained language model representation of the auxiliary task corpus obtained in step S1 as input auxiliary ; S3. Input obtained by soft parameter sharing constraints and BiLSTM learning step S2 intent The implicit features in ; Input is obtained through BiLSTM learning step S2 auxiliary Implicit features in S4. After the learning of the implicit features of the main task and the implicit features of the auxiliary task in step S3 is completed, the main task uses a multi-head attention mechanism to calculate the implicit feature weight of the main task, and the auxiliary task uses a self-attention mechanism to calculate the implicit feature weight of the auxiliary task; S5. multi-task joint learning, using sparse classification cross entropy as the loss function, regularizing the main task and the auxiliary task, obtaining the weight adjustment factors of the main task and the auxiliary task, and obtaining a citation intent classification system; In step S1, the main task is a citation intention classification task, and the main task corpus is a citation intention classification task corpus; the auxiliary tasks include a citation position classification task and a citation value classification task, the citation position classification task is used to classify the title of the paragraph where the citation is located, and the citation value classification task is used to classify whether a sentence includes a citation, and the auxiliary task corpus includes a citation position classification task corpus and a citation value classification task corpus, and the citation position classification task corpus and the citation value classification task corpus do not contain citation intention classification annotations; In step S2, the Input intent As shown in formula (1), Input intent =(SciBERT(S),feature ij (S)) (1) Among them, Input intent is the input of the main task citation intent classification task, SciBERT(S) is the SciBERT pre-trained language model representation of the main task corpus, i is the number of sentences, j is the number of words in the sentence, and feature ij (S) is the heterogeneous feature set of the main task corpus; feature ij As shown in formula (2), feature ij =cat([Onehot j (pos j ,pos_list),pattern j ,tfidf ij ,senti j ]) (2) Among them, Onehot j (pos j ,pos_list) is the part-of-speech tagging feature represented by a one-hot vector, pattern j is the syntactic structure feature, tfidf ij represents the TF-IDF value of word j in sentence i, senti j It is the weighted embedding representation of the domain sentiment word vector multiplied by the TF-IDF vector.

2. According to the method for constructing a multi-task citation intent classification system according to claim 1, it is characterized in that: pattern j It includes the following six syntactic structures: i) reference + verb + verb [past / present / third person / past / past participle]; ii) verb [past / gerund / third person] + verb [gerund / past participle]; iii) Verb [all forms] + (Adverb [Comparative / Superlative]) + Verb + (Adverb [Comparative / Superlative]) + Past Participle; iv) modal verb + (adverb [comparative / superlative]) + verb + (adverb [comparative / superlative]) + past participle; v) (adverb [comparative / superlative]) + personal pronoun + (adverb [comparative / superlative]) + verb [all forms]; vi) gerund + (proper noun + juxtaposition conjunction + proper noun), tfidf ij The calculation of is shown in (3), where d is an instance in the citation corpus, and f j,d is the frequency of word j in instance d, N is the number of instances in the citation corpus, and n j is the number of instances containing word j.

3. The method for constructing a multi-task citation intent classification system according to claim 1, characterized in that: In step S2, the input of the citation position classification task of the auxiliary task is shown in formula (4): Input section =SciBERT(S section ')(4) Among them, Input section As the input for the citation position classification task, SciBERT (S section ') is the SciBERT pre-trained language model representation for the citation position classification task; The input of the auxiliary task of citation position classification is shown in formula (5): Input worth =SciBERT(S worth ')(5) Among them, Input worth As the input of the citation value classification task, SciBERT (S worth ') is the SciBERT pre-trained language model representation for the citation value classification task.

4. The method for constructing a multi-task citation intent classification system according to claim 1, characterized in that: In step S4, the main task divides the input into multiple subspaces during the attention weight calculation process, calculates the subspace weights through the weight parameter matrix of multiple subspace query vectors, subspace key information and subspace word embedding vectors, and shares the training parameters in these multiple subspaces to obtain the main task implicit feature weights; The auxiliary task maps the input to the auxiliary task query vector, the auxiliary task key information and the auxiliary task word embedding vector during the attention weight calculation process, generates an auxiliary task attention weight map by the dot product of the auxiliary task query vector and the auxiliary task key information, and then generates an attention weighted feature by the dot product of the auxiliary task attention weight map and the auxiliary task word embedding vector. The dot product is adjusted by the adjustment factor of the smoothed Softmax function to obtain the auxiliary task implicit feature weight.

5. The method for constructing a multi-task citation intent classification system according to claim 1, characterized in that: In step S5, the citation value classification task uses the Sigmoid function as the activation function of multi-task joint learning, and the citation intent classification task and the citation position classification task use the Softmax function as the activation function of multi-task joint learning.

6. A method for constructing a multi-task citation intent classification system according to claim 1, characterized in that: In step S5, sparse classification cross entropy is used as the loss function and regularization constraints are performed, as shown in formula (6), Among them, p i is the probability that the instance belongs to category i, is the probability that the instance does not belong to category i, w i is the parameter weight matrix of the fully connected multi-task learning layer, The joint loss function L(w) is shown in formula (7): Among them, J j (w) is the loss function of task j, λ i It is the weight adjustment factor of each task.

7. A multi-task citation intent classification system, characterized in that: The method for constructing a multi-task citation intent classification system according to any one of claims 1 to 6 is adopted, comprising an input module, a representation module, a context feature learning module, an attention module, a multi-task joint learning module and a citation intent output module; The input module is used to input a data set and transmit the data set to the representation module, the input module includes a main task input module and an auxiliary task input module, the main task input module is used to input a main task data set, and the auxiliary task input module is used to input an auxiliary task data set; The representation module is used to represent the input data set to obtain input data and transmit the input data to the context feature learning module. The representation module includes a main task representation module and an auxiliary task representation module. The main task representation module obtains the main task input data by fusing the heterogeneous feature set with the SciBERT pre-trained language model, and the auxiliary task representation module obtains the auxiliary task input data by using the SciBERT pre-trained language model. The context feature learning module is used to learn the input data to obtain implicit features and transmit the implicit features to the attention module. The context feature learning module includes a main task context feature learning module and an auxiliary task context feature learning module. The main task context feature learning module learns the main task implicit features through soft parameter sharing constraints and BiLSTM, and the auxiliary task context feature learning module learns the auxiliary task implicit features through BiLSTM. The attention module is used to calculate the task weight according to the implicit feature and transmit the calculated task weight to the multi-task joint learning module, the attention module includes a main task attention module and an auxiliary task attention module, the main task attention module calculates the main task weight through a multi-head attention mechanism, and the auxiliary task attention module calculates the auxiliary task weight through a self-attention mechanism; The multi-task joint learning module is used to use sparse classification cross entropy as a loss function, regularize the main task and the auxiliary task, obtain the weight adjustment factors of the main task and the auxiliary task, and transmit the weight adjustment factors to the citation intention output module; The citation intent output module outputs a citation intent classification based on the weight adjustment factor.

8. An application of the system as claimed in claim 7 in citation intent classification, characterized in that: The data set to be classified for citation intent is input into the multi-task citation intent classification system to obtain citation intent classification data.

Citation Information

Patent Citations

  • Citation intention classification method based on multi-task bilateral branch network

    CN114328923A

  • Image classification method and system based on cross-modal semantic representation learning and fusion

    CN114898156A