A method for judging the difficulty of English reading materials based on text feature fusion

By combining pre-trained language models and LSTM to extract semantic and grammatical features of English reading materials and statistical information characteristics, the problem of difficulty in accurately judging the difficulty of English reading materials in the prior art is solved, and higher accuracy and efficiency are achieved.

CN115630140BActive Publication Date: 2025-06-27YUNNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211364247.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-06-27
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately judge the difficulty of English reading materials, especially in complex information environments, where traditional methods lack good generalization capabilities.

Method used

Using a method based on text feature fusion, semantic features are extracted through pre-trained language models, syntactic features are extracted in combination with LSTM, and statistical information features are counted. Finally, the difficulty score is output through the full connection layer and the sigmoid layer.

Benefits of technology

It improves the accuracy and efficiency of the difficulty judgment of English reading materials, performs better than traditional methods and is more robust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630140B_ABST
    Figure CN115630140B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for judging the difficulty of English reading materials based on text feature fusion, belonging to the related field of natural language processing. First, for the English reading material dataset, the input English text is encoded, and the encoded result is input into a pre-trained language model to calculate a feature vector containing semantic information. Then, part-of-speech tagging is performed on the English text, and the obtained part-of-speech sequence is input into an LSTM to calculate a feature vector containing syntactic information. For the factors affecting the difficulty of English reading materials, relevant factors are statistically analyzed and feature extraction is performed. All the obtained feature vectors are concatenated and then input into a fully connected layer, and finally a value between 0 and 1 is output through sigmoid to represent the difficulty. The present invention can effectively judge the difficulty of English reading materials and better assist various adaptive learning services in English teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for judging the difficulty of English reading materials based on text feature fusion, belonging to the technical field of natural language processing. Background Art

[0002] English, as a second language widely learned, and reading, as an important part of English learning, how to accurately judge the difficulty of English reading materials so that people with different English levels can receive education suitable for their own English levels and further promote personalized learning is particularly important.

[0003] In the early 20th century, research on measuring the difficulty of English reading materials emerged. Until now, the research on judging the difficulty of English reading materials has been the core issue concerned by relevant researchers at home and abroad. Therefore, numerous researchers have conducted a large number of studies on the factors affecting the difficulty of English reading materials, summarized many influencing factors, and produced many formulas for calculating the difficulty of English reading materials. These formulas have been helping people choose appropriate English texts for a long time. However, with the continuous development of informatization, the generated texts are becoming increasingly complex, and the method of formulating rules is usually relatively simple and does not have good generalization ability, so good results cannot be obtained.

[0004] With the continuous development of language models, in October 2018, Google proposed the BERT (Bidirectional Encoder Representation from Transformers) model, which has brought the development of the natural language processing field into a new stage. BERT is a pre-trained language model. Unlike traditional language models that only use unidirectional language models or shallow splicing of two unidirectional language models for training, it uses MLM (masked language model) to train bidirectional Transformers to generate deep bidirectional language representations and performs excellently in 11 different natural language processing (NLP) tests. Many scholars have achieved good results in combining BERT for other tasks in the field of natural language processing. This way of transferring the already trained model to a new model for training is called transfer learning. Considering that most tasks are somewhat related, passing the learned parameters to the new model in a certain way can greatly improve the efficiency of the model. As one of the methods of transfer learning, fine-tuning can further improve the learning time of the model and reduce the cost of model training by freezing the convolutional layers in the pre-trained model and training other convolutional layers and fully connected layers. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for judging the difficulty of English reading materials based on text feature fusion, so as to improve the accuracy and efficiency of judging the difficulty of English reading materials.

[0006] By summarizing the views of linguists on the factors affecting the difficulty of English reading materials and considering the advantages of pre-trained language models in natural language processing tasks, the present invention proposes a method for judging the difficulty of English reading materials based on text feature fusion, which fuses multiple text features and uses deep learning technology to judge the difficulty of English reading materials.

[0007] The technical solution of the present invention is: a method for judging the difficulty of English reading materials based on text feature fusion. First, for the English reading material dataset, the input English text is encoded, and the encoded information is input into a pre-trained language model that has been trained to obtain a feature vector containing semantic information; then the input text is part-of-speech tagged, and the obtained part-of-speech sequence is input into an LSTM to obtain a feature vector containing syntactic information; the factors affecting the difficulty of English reading materials are statistically analyzed and embedded, all features are concatenated and then input into a fully connected layer, and finally a value between 0 and 1 is output through a sigmoid layer to represent the difficulty.

[0008] The specific steps for judging the English reading difficulty are as follows:

[0009] Step1: Use a pre-trained language model to extract the semantic features of the text.

[0010] First, for the English reading material dataset (experiments are carried out using the Newsela dataset and the self-collected dataset), the input English text is encoded, and the encoded information is input into a pre-trained language model that has been trained to obtain a feature vector containing semantic information.

[0011] The specific process is to first extract information such as words, sentence positions, and word positions in the sentence for one-hot encoding, input it into the pre-trained language model, and obtain the semantic feature vector. The pre-trained model of the present invention selects the Bert model.

[0012] Step2: Extract syntactic information features.

[0013] The text is part-of-speech tagged, and the obtained part-of-speech sequence is input into an LSTM to obtain a feature vector containing syntactic information.

[0014] Step3: Extract statistical information features.

[0015] Statistically analyze the factors affecting the difficulty of English reading materials and perform embedded representation on them. Concatenate all features and input them into a fully connected layer, and finally output a value between 0 and 1 representing the difficulty after passing through the sigmoid layer.

[0016] Step4: Difficulty prediction.

[0017] Output a value between 0 and 1 representing the difficulty after passing through the sigmoid layer.

[0018] The specific content of Step1 is as follows:

[0019] Step1.1: Assume that the currently input English text is S t , S t contains n words, S t = {w1, w2, …, w i , …, w n}, where w i represents the i-th word.

[0020] The Bert model usually adds [CLS] at the beginning of a sentence to indicate the start of a paragraph and adds [SEP] in the middle of two sentences to separate them.

[0021] The transformed sentence is S BERT = {[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP]}.

[0022] Step1.2: Set the maximum length of S BERT to M. If the length of S t is less than M, then add [PAD] to S BERT to complete the padding. After the padding operation, S BERT is:

[0023] S BERT = {[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]}

[0024] If the length of S t is greater than M, then truncate and discard the subsequent content. After the truncation operation, S BERT is:

[0025] S BERT = {[CLS], w1, w2, …, [SEP], …, w M-2 , w M-1 , wM ,[SEP]}

[0026] Step1.3: Embed each content in S BERT , that is:

[0027]

[0028] Among them, D BERT represents the embedding dimension set by the pre-trained language model.

[0029] Step1.4: Perform sentence position encoding on the content in S BERT , that is:

[0030] S segmentembedding = {E A , E A , E A , E B , E B , E B , E B , …, E i , E i}

[0031] Among them, E A represents the first sentence, E B represents the second sentence, and so on for subsequent sentences, E i represents the i-th sentence.

[0032] Step1.5: Perform word position encoding on the content in S BERT , that is:

[0033] S positionembedding = {E1, E2, E3, …, E i , …, E n-2 , E n-1 , E n , …, E M}

[0034] Among them, E i represents the position encoding of the i-th word,

[0035] Step1.6: Input S embedding , S segmrntembedding , S positionembedding into the pre-trained language model (default is BERT) to obtain the feature vector O BERT output by the last layer, that is:

[0036]

[0037] Step1.7: There are multiple schemes for selecting sentence vectors. For example: 1) Take X [CLS] as the sentence vector. 2) Perform average pooling on O BERT and take the result. 3) Perform max pooling on O BERT and take the result. 4) Further extract features from the result of O BERT using CNN. 5) Input the result of O BERT into LSTM to extract features. In the task of the present invention, X [CLS] is selected as the sentence vector.

[0038] Specifically, the said Step2 is as follows:

[0039] Step2.1: For the input text S t ={w1, w2, w3, …, w n}, add [CLS] at the beginning of the sentence to indicate the start of a sentence, and add [SEP] in the middle of two sentences to separate them. The transformed sentence is:

[0040] S sen ={[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]}

[0041] Step2.2: Perform part-of-speech tagging on S sen to obtain:

[0042] S POS ={[SPACE], [PRP], [VBP], [NNP], …, [RB], [JJ], [SPACE], …, [PAD]}

[0043] Among them, [SPACE] represents [CLS] and [SEP], [PRP] represents a pronoun, [VBP] represents a verb, [NNP] represents a noun, [RB] represents an adverb of degree, and [JJ] represents an adjective.

[0044] Step2.3: Perform embedding representation on S POS to obtain E POS , that is:

[0045]

[0046] Among them, D POS represents the embedding dimension of the part-of-speech token.

[0047] Step2.4: Input E posInput the LSTM and take the output result O of the last layer pos As the feature vector (i.e., syntactic feature) of the sentence part-of-speech sequence, where

[0048] In Step2, the syntactic features of the sentence are mainly calculated. Syntax and vocabulary are the keys to differentiating the difficulty of English texts. Therefore, the complexity of syntax needs to be considered. The present invention takes the part-of-speech sequence of the sentence as the input, uses LSTM to learn the features of the sequence, so as to realize the vectorized representation of syntax and input it into the neural network for subsequent step calculations. In the existing methods, syntactic information is mainly obtained by counting the number of keywords and the co-occurrence of keywords. This method cannot fully represent the sequence information. Therefore, using LSTM in this invention can better learn syntactic features.

[0049] The specific content of Step3 is as follows:

[0050] Since the factors affecting the difficulty level of English reading materials, in addition to semantics and syntax, factors such as sentence length, the number of prepositions, and average word length also need to be considered as influencing factors. Then, these factors are statistically counted and encoded and input into the model. After adding the above information, the model converges faster during training, and at the same time, the robustness of the model is further improved.

[0051] The specific steps are as follows:

[0052] Step3.1: Count the sentence length and perform an embedding operation: For sentence S t ={w1, w2, …, w n}, then the sentence length embedding is

[0053] where L indicates that this vector is the embedding of the sentence length, n represents the number of words, and D represents the embedding dimension.

[0054] Step3.2: Count the number of prepositions and perform an embedding operation: For sentence S t ={w1, w2, …, w n}, the preposition number embedding

[0055] where P indicates that this vector is the embedding of the preposition number, * represents the specific quantity, and D represents the embedding dimension.

[0056] Step3.3: Count the average word length and perform an embedding operation: For sentence S t ={w1, w2, …, w1}, the preposition number embedding

[0057] where A indicates that this vector is the embedding of the average word length, * represents the specific quantity, and D represents the embedding dimension.

[0058] Step 3.4: Concatenate as the statistical information of the sentence:

[0059]

[0060] Among them,

[0061] The specific content of Step 4 is as follows:

[0062] Step 4.1: Concatenate the semantic feature X [CLS] , the syntactic feature O POS , and the statistical information feature O STA , then input the concatenated result into the fully connected layer, and then input it into the sigmoid layer to predict the result and output:

[0063]

[0064] Step 4.2: Calculate the loss:

[0065]

[0066] Among them, y ic represents the true category of sample i, taking 1 if it is equal to c, and taking 0 if it is not equal. p ic represents the predicted probability that the observed sample i belongs to category c;

[0067] Step 4.3: Use Adam to optimize the loss, aiming to minimize the loss. When the loss reaches the minimum, the model reaches the best effect.

[0068] This part concatenates the above three features and inputs them into the neural network, and uses the sigmoid function to limit the output between [0, 1], so as to realize the difficulty judgment.

[0069] The beneficial effects of the present invention are as follows: When the present invention judges the difficulty of English texts, it comprehensively considers features such as the semantic information, syntactic information, and statistical information of the texts. Compared with traditional methods, the present invention takes into account the importance of the semantic information of English texts, uses LSTM to learn the syntactic information of the texts, and at the same time inputs the traditional statistical information into the neural network for calculation. Thus, a difficulty judgment model with better effect and stronger robustness than traditional methods is obtained. Brief Description of the Drawings

[0070] Figure 1 is the step flow chart of the present invention. Detailed Embodiments

[0071] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0072] Example 1: As Figure 1 shown, a method for judging the difficulty of English reading materials based on text feature fusion. First, for the English reading material dataset, the input English text is encoded, and the encoded information is input into a pre-trained language model that has been trained to obtain a feature vector containing semantic information. Then, part-of-speech tagging is performed on the English text, and the obtained part-of-speech sequence is input into an LSTM to obtain a feature vector containing syntactic information. The factors affecting the difficulty of English reading materials are statistically analyzed and embedded, and all features are concatenated and then input into a fully connected layer. Finally, after passing through a sigmoid layer, a numerical value from 0 to 1 is output to represent the difficulty.

[0073] Suppose there is a set A of existing English reading materials, and there are N pieces of English reading material data in the set, then A = {S1, S2, S3, …, S N}, where S i represents the i-th English reading material text in the English reading material set. The specific steps for judging the English reading difficulty are as follows:

[0074] Step1: The pre-trained model of the present invention selects the Bert model. The pre-trained language model part is mainly used to learn the semantic information of the text. Three features are required for inputting into the pre-trained language model, namely the feature of each word, the sentence position feature, and the word position feature, and the three features are extracted.

[0075] Step2: Syntactic feature extraction.

[0076] Step3: Statistical information feature extraction.

[0077] Step4: Difficulty prediction.

[0078] The specific content of the said Step1 is as follows:

[0079] Step1.1: Suppose the currently input English text is S t , S t contains n words, S t = {w1, w2, …, w i , …, w n}, where w i represents the i-th word.

[0080] The Bert model usually adds [CLS] at the beginning of a sentence to represent the start of a paragraph, and adds [SEP] in the middle of two sentences to separate the sentences.

[0081] The transformed sentence is S BERT = {[CLS], w1, w2, …, [SEP], …, w n-2 , wn-1 , w n , [SEP]}。

[0082] Step1.2: Set the maximum length of S BERT to M. If the length of S t is less than M, then pad S BERT with [PAD] until it reaches M. After padding, S BERT = {[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]}.

[0083] If the length of S t is greater than M, then truncate and discard the subsequent content. After truncation, S BERT = {[CLS], w1, w2, …, [SEP], …, w M-2 , w M-1 , w M , [SEP]}.

[0084] Step1.3: Perform embedding encoding on each content in S BERT , that is:

[0085]

[0086] where D BERT represents the embedding dimension set by the pre-trained language model.

[0087] Step1.4: Perform sentence position encoding on the content in S BERT , that is:

[0088] S segmentembedding = {E A , E A , E A , E B , E B , E B , E B , …, E i , E i}

[0089] where E A represents the first sentence, E B represents the second sentence, and so on for subsequent sentences, where E i represents the i-th sentence.

[0090] Step1.5: For S BERTPerform word position encoding on the content in, that is:

[0091] S positionembedding ={E1, E2, E3, …, E i , …, E n-2 , E n-1 , E n , …, E M}

[0092] where E i represents the position encoding of the i-th word,

[0093] Step1.6: Input S embedding , S segmrntembedding , S positionembedding into the pre-trained language model (defaulting to BERT) to obtain the feature vector O BERT of the last layer output, that is:

[0094]

[0095] Step1.7: There are multiple schemes for selecting sentence vectors. For example: 1) Take X [CLS] as the sentence vector. 2) Perform average pooling on O BERT and take the result. 3) Perform max pooling on O BERT and take the result. 4) Further extract features from the result of O BERT using CNN. 5) Input the result of O BERT into LSTM to extract features. In the task of the present invention, X [CLS] is selected as the sentence vector.

[0096] The specific content of the said Step2 is as follows:

[0097] Step2.1: For the input text S t ={w1, w2, w3, …, w m}, add [CLS] at the beginning of the sentence to indicate the start of a sentence, and add [SEP] in the middle of two sentences to separate the sentences. The transformed sentence is:

[0098] S sen ={[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]}

[0099] Step2.2: Perform part-of-speech tagging on S sen to obtain:

[0100] S POS={[SPACE],[PRP],[VBP],[NNP],…,[RB],[JJ],[SPACE],…,[PAD]}

[0101] Among them, [SPACE] represents [CLS] and [SEP], [PRP] represents pronouns, [VBP] represents verbs, [NNP] represents nouns, [RB] represents adverbs of degree, and [JJ] represents adjectives.

[0102] Step2.3: Embed S POS to obtain E POS , that is:

[0103]

[0104] Among them, D POS represents the embedding dimension of the part-of-speech token.

[0105] Step2.4: Input E pos into the LSTM and take the output result O of the last layer pos as the feature vector of the sentence grammar information, where

[0106] The specific content of Step3 is as follows:

[0107] Since, in addition to semantics and grammar, factors such as sentence length, number of prepositions, and average word length also need to be considered as influencing factors for the difficulty level of English reading materials, these factors are statistically counted and encoded and then input into the model. The specific steps are as follows:

[0108] Step3.1: Statistically count the sentence length and perform an embedding operation: For the sentence S t ={w1, w2, …, w n}, the sentence length embedding is

[0109] Among them, L indicates that this vector is the embedding of the sentence length, n represents the number of words, and D represents the embedding dimension.

[0110] Step3.2: Statistically count the number of prepositions and perform an embedding operation: For the sentence S t ={w1, w2, …, w n}, the embedding of the number of prepositions

[0111] Among them, P represents that this vector is the embedding of the number of prepositions, * represents the specific quantity, and D represents the embedding dimension.

[0112] Step3.3: Statistically count the average word length and perform an embedding operation: For the sentence S t= {w1, w2, …, w n}, preposition number embedding

[0113] Among them, A represents that this vector is the embedding of the average word length, * represents a specific quantity, and D represents the embedding dimension.

[0114] Step3.4: Concatenate as the statistical information of the sentence:

[0115]

[0116] Among them,

[0117] The specific content of Step4 is as follows:

[0118] Step4.1: Concatenate the semantic feature X [CLS] , the syntactic feature O POS , and the statistical information feature O STA , then input them into the fully connected layer, and then input them into the sigmoid layer to predict the result and output:

[0119]

[0120] Step4.2: Calculate the loss:

[0121]

[0122] Among them, y ic represents the true category of sample i. If it is equal to c, take 1; if not, take 0. p ic represents the predicted probability that the observed sample i belongs to category c.

[0123] Step4.3: Use Adam to optimize the loss. The purpose is to make the loss reach the minimum. When the loss reaches the minimum, the model reaches the best effect.

[0124] In this embodiment, two English reading material datasets CEFR and Newsela with difficulty level markings, and a dataset CEED manually constructed by the present invention are selected. Among them, CEFR and CEED are publicly available graded English reading text datasets, and the Newsela dataset is a non-publicly available graded English reading text dataset (which can be applied for on the Newsela website). Basic data statistics are performed on the three datasets, and the statistical results are shown in Table 1. Among them, Num represents the number of texts contained in the dataset, and Class represents the number of grade categories.

[0125] Table 1 Basic Information of Datasets

[0126]

[0127] (1) CEFR consists of 1,493 English texts, which are labeled according to the levels of the Common European Framework of Reference for Languages (CEFR): A1, A2, B1, B2, C1, and C2, with the difficulty increasing from A1 to C2. The English texts in the dataset are taken from free online resources, including the British Council, ESLFast, and the CNN Daily Mail dataset. The content of the English texts includes conversations, descriptions, short stories, newspaper stories, and other articles.

[0128] (2) CEED is collected from 469 reading test questions from English exams such as the high school entrance examination, the national college entrance examination, CET-4, CET-6, TEM-4, and TEM-8. The difficulty levels are classified as follows: the difficulty level of the high school entrance examination is denoted as Z, the national college entrance examination as G, CET-4 as S, CET-6 as L, TEM-4 as E, and TEM-8 as B. The difficulty increases from the high school entrance examination to TEM-8.

[0129] (3) Newsela consists of 10,722 English texts, which are classified according to the standards of US K12 education. Each English text is labeled with a number from 2 to 12, with the difficulty increasing from 2 to 12.

[0130] The present invention organizes the English texts in the dataset as follows: First, each English text is read paragraph by paragraph; second, the corresponding difficulty level of each paragraph is labeled; third, a difficulty label is added to each paragraph; fourth, the number of words, the number of prepositions, and the average word length of each paragraph are calculated, and finally, they are organized into a csv file. The number of paragraphs contained in the organized datasets is as follows: CEFR contains 12,096 paragraphs, Newsela contains 227,971 paragraphs, and CEED contains 3,381 paragraphs.

[0131] To better obtain the difficulty coefficient in subsequent experiments, corresponding difficulty labels are added to the extracted paragraphs. In the CEFR dataset, the difficulty labels for A1, A2, B1, and B2 are set to 0, and the difficulty labels for C1 and C2 are set to 1. In the Newsela dataset, the difficulty labels for levels greater than or equal to 6 are set to 1, and the difficulty labels for levels less than 6 are set to 0. In the CEED dataset, due to the similarity in classification, the present invention divides the dataset into three subsets. The data for the high school entrance examination and the national college entrance examination are divided into one subset, abbreviated as CEED-EE; the data for CET-4 and CET-6 are divided into one subset, abbreviated as CEED-CET; the data for TEM-4 and TEM-8 are divided into one subset, abbreviated as CEED-TEM. Among them, the difficulty labels for the high school entrance examination, CET-4, and TEM-4 are set to 0, and the difficulty labels for the national college entrance examination, CET-6, and TEM-8 are set to 1. The number of positive and negative samples contained in each organized dataset is shown in Table 2.

[0132] Table 2: Number of positive and negative samples

[0133]

[0134] In this invention, several classic pre-trained language models for the Fill-mask task in recent years, such as Bert, Bart, xlnet, roberta, xlm-roberta, were selected for testing and compared with CNN, LSTM, and BiLSTM. In terms of parameter settings, PyTorch version 1.10 was used, and an NVIDIA GeForce RTX 2080Ti GPU was used. All pre-trained models were obtained from Huggingface. The selection of hyperparameters is as follows: Batchsize takes {16, 32, 64}, the learning rate takes {1e-3, 1e-4, 1e-5}, and the word embedding dimension takes 768. Different models were experimented on different datasets, and the experimental results are as follows:

[0135] Table 3: Experimental results of different models in CEFR and Newsela

[0136]

[0137] As can be seen from Table 3, in both datasets, the method of this invention (when using BERT as the pre-trained language model) is optimal in the three metrics AUC, ACC, RMSE and the results in the two datasets. In the CEFR dataset, in terms of the AUC, ACC, and RMS metrics, the method of this invention is higher than the second place. The AUC has increased by 5.81%, the ACC has increased by 7.02%, and the RMSE has decreased by 5.14%. In the Newsela dataset, in terms of the AUC, ACC, and RMS metrics, the method of this invention is also higher than the second place. The AUC has increased by 1.63%, the ACC has increased by 1.04%, and the RMSE has decreased by 1.15%. When the dataset is small (CEFR dataset), the pre-trained language model only needs less data to perform better.

[0138] Table 4: Results of different pre-trained language models in CEFR and Newsela

[0139]

[0140] As shown in Table 4, the present invention compares the effects of different pre-trained language models, which respectively improve and enhance BERT for different tasks. From the results, the BERT model can achieve the best results in the CEFR dataset. In terms of the AUC, ACC, and RMS metrics, the BERT model is higher than the second place. The AUC has increased by 0.35%, the ACC has increased by 0.24%, and the RMSE has decreased by 0.92%. The XLNet model can achieve the best results in the Newsela dataset. Compared with BERT, the AUC has increased by 0.37%, the ACC has increased by 0.58%, and the RMSE has decreased by 0.60%. However, the overall gap between these pre-trained models is not large, but the results are all better than those of CNN and LSTM.

[0141] Table 5: Experimental results of different models in CEED

[0142]

[0143]

[0144] As can be seen from Table 5, in the two datasets, the method of the present invention (when using BERT as the pre-trained language model) is the best in the three metrics of AUC, ACC, and RMSE and in the three datasets. In the CEED-EE dataset, in terms of the AUC, ACC, and RMS metrics, the method of the present invention is higher than the second place. The AUC has increased by 8.20%, the ACC has increased by 4.71%, and the RMSE has decreased by 7.05%. In the CEED-CET dataset, in terms of the AUC, ACC, and RMS metrics, the method of the present invention is also higher than the second place. The AUC has increased by 5.32%, the ACC has increased by 3.77%, and the RMSE has decreased by 1.95%. On the CEED-TEM dataset, compared with the second place, the AUC and ACC have increased by 9.09% and 12.5% respectively, and the REMSE has decreased by 8.51%.

[0145] Table 6: Results of different pre-trained language models in CEED

[0146]

[0147] As shown in Table 6, the present invention also compares the effects of different pre-trained language models in the CEED dataset. Generally speaking, RoBERTa can achieve good results in all three subsets of CEED. Compared with BERT in the CEED-EE dataset, the AUC is increased by 5.06%, the ACC is increased by 8.49%, and the RMSE is decreased by 13.1%. Compared with BERT in the CEED-CET dataset, the AUC is increased by 3.99%, the ACC is increased by 11.32%, and the RMSE is decreased by 11.20%. Compared with BERT in the CEED-TEM dataset, the AUC is increased by 3.17%, the ACC is increased by 3.12%, and the RMSE is decreased by 4.35%. Generally speaking, these pre-trained language models are superior to CNN and LSTM in all three metrics.

[0148] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for judging the difficulty of English reading materials based on text feature fusion, characterized in that: Step1: First, for the English reading material dataset, encode the input English text, and input the encoded information into a pre-trained language model that has been trained to obtain a feature vector containing semantic information; Step2: Perform part-of-speech tagging on the text, and input the obtained part-of-speech sequence into LSTM to obtain a feature vector containing syntactic information; Step3: Statistical information feature extraction; statistically analyze the factors affecting the difficulty of English reading materials and perform embedding representation on them, and splice all features and input them into a fully connected layer; Step4: Finally, output a value from 0 to 1 through the sigmoid layer to represent the difficulty, and complete the difficulty judgment; The specific content of Step3 is as follows: Step 3.1: Statistically analyze the sentence length and perform an embedding operation: For sentence S t ={w1, w2, …, w n}, the sentence length embedding is Among them, L represents the embedding of the sentence length of this vector, * represents a specific quantity, and D represents the embedding dimension; Step 3.2: Count the number of prepositions and perform an embedding operation: For sentence S t ={w1, w2, …, w n}, the embedding of the number of prepositions Among them, P represents the embedding of the number of prepositions of this vector, * represents a specific quantity, and D represents the embedding dimension; Step 3.3: Statistically analyze the average word length and perform an embedding operation: For sentence S t ={w1, w2, …, w n}, the average word length embedding Among them, A represents the embedding of the average word length of this vector, * represents a specific quantity, and D represents the embedding dimension; Step3.4: Combine as the statistical information of the sentence: Among them, 2. The method for judging the difficulty of English reading materials based on text feature fusion according to claim 1, wherein The specific content of Step1 is as follows: Step1.1: Assume that the currently input English text is S t , S t contains n words, S t = {w1, w2, …, w i , …, w n}, where w i represents the i-th word; The transformed sentence is S BERT ={[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP]}; Step1.2: Set the maximum length of S BERT to M. If the length of S t is less than M, then add [PAD] to S BERT to complete the padding. After the padding operation, S BERT is: S BERT = {[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]} If S t has a length greater than M, then truncate and discard the subsequent content. After the truncation operation, S BERT is as follows: S BERT = {[CLS], w1, w2, …, [SEP], …, w M-2 , w M-1 , w M , [SEP]} Step1.3: Embed and encode each content in S BERT as follows: Among them, D BERT represents the embedding dimension set by the pre-trained language model; Step1.4: Perform sentence position encoding on the content in S BERT as follows: S segment embedding = {E A , E A , E A , E B , E B , E B , E B , …, E i , E i} Among them, E A represents the first sentence, E B represents the second sentence, and so on for subsequent sentences, E i represents the i-th sentence; Step1.5: Perform word position encoding on the content in S BERT as follows: S position embedding = {E1, E2, E3, …, E i , …, E n-2 , E n-1 , E n , …, E M} Among them, E i represents the position encoding of the i-th word, Step1.6: Input S embedding , S segment embedding , S position embedding into the pre-trained language model to obtain the feature vector O BERT output by the last layer, that is: Step1.7: Select X [CLS] as the sentence vector.

3. The method for judging the difficulty of English reading materials based on text feature fusion according to claim 1, wherein The specific content of Step2 is as follows: Step2.1: For the input text S t ={w1, w2, w3, …, w n}, add [CLS] at the beginning of the sentence to indicate the start of a sentence, and add [SEP] in the middle of two sentences to separate them. The transformed sentence is: S sen = {[CLS], w1, w2, …, [SEP], …, w n-2 , w n-1 , w n , [SEP], …, [PAD]} Step2.2: For S sen perform part-of-speech tagging to obtain: S POS = {[SPACE], [PRP], [VBP], [NNP], …, [RB], [JJ], [SPACE], …, [PAD]} Among them, [SPACE] represents [CLS] and [SEP], [PRP] represents a pronoun, [VBP] represents a verb, [NNP] represents a noun, [RB] represents an adverb of degree, and [JJ] represents an adjective; Step 2.3: Embed S POS to obtain E POS , that is: Among them, D POS represents the embedding dimension of the part-of-speech token; Step2.4: Input E pos into the LSTM, and take the output result O of the last layer pos as the feature vector of the sentence grammar information, where 4. The method for judging the difficulty of English reading materials based on text feature fusion according to claim 1, characterized in that The specific content of Step4 is as follows: Step4.1: Concatenate the semantic feature X [CLS] , the syntactic feature O POS , and the statistical information feature O STA , then input the concatenated result into the fully connected layer, and then input it into the sigmoid layer to predict the result and output: Step4.2: Calculate the loss: Among them, y ic represents the true category of sample i, taking 1 if it is equal to c and 0 if it is not, and p ic represents the predicted probability that the observed sample i belongs to category c; Step4.3: Use Adam to optimize the loss, with the aim of minimizing the loss. When the loss reaches the minimum, the model reaches the best effect.

Citation Information

Patent Citations

  • Method for automatically predicting difficulty of reading and understanding test questions by introducing multi-text relationship

    CN113536808A

  • Deep learning based method and device for chinese semantics analysis

    WO2018028077A1