English language ability prediction system and method based on LSTM algorithm
Through the English language ability prediction system based on LSTM algorithm, the existing evaluation methods are solved, and accurate prediction and long-term accuracy of the CEFR level of the tester are achieved, which is suitable for a wide range of promotional applications.
Patent Information
- Application Number
- CN202510079702.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing English language proficiency assessment methods have problems such as inaccurate assessment, limited environmental impact and test counts, making it difficult to accurately predict the CEFR level of the tester.
The English language ability prediction system based on LSTM algorithm is adopted to achieve the prediction of the tester's CEFR corpus, preprocessing and feature extraction, construction and training of the LSTM recurrent neural network model.
It realizes long-term accurate prediction of the tester's English language ability CEFR grade, improves the accuracy of the prediction results, and is suitable for a wide range of promotional applications.
Smart Images

Figure CN120068857A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of algorithm-assisted English education, and particularly relates to an English language ability prediction system and method based on the LSTM algorithm. Background Art
[0002] The full name of CEFR is the Common European Framework of Reference for Languages, which is a standard system designed to describe language abilities and levels and is widely used in the fields of language teaching and assessment. This framework comprehensively describes the ability of language users to communicate using the four skills of listening, speaking, reading, and writing, and is widely used in various language tests and certifications, such as IELTS, TOEFL, etc., providing a unified standard for the assessment and certification of language abilities.
[0003] Existing English language abilities such as CEFR levels are usually obtained through tests, and the abilities of test takers in each single dimension of vocabulary, grammar, reading, listening, speaking, and writing are obtained through tests. The disadvantages of this method are as follows: (1) The test can only estimate the ability of the test taker as a whole through the score, but in the actual test process, evaluation tools such as test papers may not be able to fully and accurately reflect the true language ability of the test taker. (2) The test environment, such as test equipment and invigilators, may affect the performance of learners. An unsatisfactory test environment may cause the test taker to not be able to fully exert their language ability, thus affecting the objectivity of the test results. (3) There are restrictions on the number of tests. For example, test takers of KET and PET can only register for 1 to 4 times per year.
[0004] The purpose of the present invention is to provide an English language ability prediction system and method. By using this system and method for English ability prediction, the CEFR level can be accurately predicted, and test takers can formulate a more targeted learning plan according to their actual level and target level, thereby more effectively improving their language ability. Summary of the Invention
[0005] The purpose of the present invention is to provide an English language ability prediction system and method based on the LSTM algorithm to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above purpose, the present invention provides an English language ability prediction system and method based on the LSTM algorithm, including the following steps:
[0007] S1. Establish a CEFR corpus, where the CEFR corpus includes CEFR vocabulary and CEFR levels;
[0008] S2. Preprocess the CEFR corpus and complete feature extraction, and use the features as input feature vectors;
[0009] S3. Construct and train the LSTM recurrent neural network model for the CEFR corpus to obtain the trained LSTM recurrent neural network model;
[0010] S4. Predict the CEFR level of the tester according to the trained LSTM recurrent neural network model to obtain the prediction result of the tester's English language ability.
[0011] Furthermore, in step S1, it is necessary to ensure the correct matching relationship between the CEFR vocabulary and the CEFR level in the CEFR corpus. Among them, the CEFR levels are divided into A1 (beginner), A2 (elementary), B1 (intermediate), B2 (upper intermediate), C1 (advanced), and C2 (proficient).
[0012] Furthermore, adjust and supplement the CEFR corpus in step S1, and divide it into a training data set, a validation set, and a test set, and construct and train the LSTM recurrent neural network model through the training data set, the validation set, and the test set.
[0013] Furthermore, in step S2, the preprocessing steps include:
[0014] Step S21: Traverse the CEFR vocabulary, set all the vocabulary as target vocabulary, and associate it with the corresponding CEFR level;
[0015] Step S22: Set word vectors for the target vocabulary and the corresponding CEFR level;
[0016] Step S23: Use the values of the word vectors as input feature vectors.
[0017] Furthermore, in step S2, the feature extraction method is:
[0018] Train the word2vec model through the CBOW method, and extract the text feature representation of the CEFR vocabulary through the word2vec model;
[0019] Extract features from the speech data of spoken English to obtain prosodic features and MFCC cepstral features, and use the prosodic features and MFCC cepstral features as the feature representation of the speech data of the CEFR vocabulary.
[0020] Furthermore, the construction and training of the LSTM recurrent neural network model for the CEFR corpus in step S3 includes:
[0021] Initialize the LSTM model, train the LSTM model, and evaluate the LSTM model.
[0022] Furthermore, the initialization of the LSTM model includes:
[0023] Iteratively train a pre - constructed LTSM model using the input feature vectors described in S2.
[0024] Further, the training of the LSTM model includes:
[0025] Mark multiple said training data sets in chronological order, and use the training data set with earlier time to train the initialized LTSM model to obtain a candidate prediction model;
[0026] Stack the training data sets on the training data set with earlier time in chronological order to train the candidate prediction model.
[0027] Further, the evaluation of the LSTM model specifically includes:
[0028] Based on the test set, test the trained candidate prediction model;
[0029] If the test effect meets the preset threshold, generate a final LSTM model.
[0030] The present invention also provides an English language ability prediction system based on the LSTM algorithm. The system includes:
[0031] A data collection module for collecting CEFR vocabulary and CEFR level data to obtain a CEFR corpus;
[0032] A model training module for dividing the initial training data set into multiple training data sets, and iteratively training a pre - constructed LTSM model using the multiple training data sets to obtain multiple candidate prediction models and a final prediction model;
[0033] A candidate prediction model for predicting the CEFR level of a tester according to the multiple candidate prediction models and the final prediction model to obtain a deviation data training set; training a pre - constructed deviation LTSM model using the deviation data training set to obtain a trained candidate prediction model;
[0034] A prediction module for predicting the CEFR level according to the final prediction model and the trained candidate prediction model to obtain an initial prediction result and a prediction deviation; correcting the initial prediction result using the prediction deviation to obtain the final CEFR level value of the tester.
[0035] The technical effects of the present invention are as follows: A prediction system and method for English language ability based on the LSTM algorithm provided by the present invention, with CEFR as the background and reliance, mainly through the evaluation of vocabulary, constructs a model through a recurrent neural network with long short-term memory units, and converts the mastery of vocabulary into a predicted value of the CEFR level of the tester's English language ability. Among them, vocabulary is the most direct and easily obtainable rating data for the CEFR level. Therefore, the method can adapt to the changing characteristics of the CEFR level of the tester, and then achieve long-term and accurate advance prediction of the CEFR level of the tester. At the same time, it can greatly improve the accuracy of the prediction results, accurately distinguish the vocabulary ability required for the test, and the tester can conduct targeted training based on this to improve the language learning efficiency; and the method has strong versatility, low cost, wide applicability, and is suitable for wide promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The exemplary embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0037] Figure 1 is an exemplary flowchart of a method for predicting English language ability based on the LSTM algorithm of the present invention;
[0038] Figure 2 is an exemplary structural schematic diagram of a system for predicting English language ability based on the LSTM algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following description with reference to the drawings is for the purpose of facilitating a comprehensive understanding of the various embodiments of this application defined by the claims and their equivalent content. These embodiments include various specific details for the purpose of understanding, but these are only considered exemplary. Therefore, those skilled in the art can understand that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of this application. In addition, for the purpose of briefly and clearly describing this application, this application will omit the description of well-known functions and structures.
[0040] The terms and phrases used in the following specification and claims are not limited to their literal meanings, but are only for the purpose of clearly and consistently understanding this application. Therefore, for those skilled in the art, it can be understood that providing a description of the various embodiments of this application is only for the purpose of illustration, rather than limiting this application defined by the appended claims and their equivalents.
[0041] Next, in combination with the accompanying drawings in some embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0042] Refer to Figure 1 , which representatively shows a schematic flowchart of a method for predicting English language ability based on the LSTM algorithm proposed by the present invention. In the implementation manner of this embodiment, a method for predicting English language ability based on the LSTM algorithm proposed by the present invention is described by taking its application to English language ability prediction as an example. It is easy for those skilled in the art to understand that in order to apply the relevant designs of the present invention to the prediction of other languages, various modifications, additions, substitutions, deletions or other changes are made to the following specific implementation manners, and these changes are still within the scope of the principle of the English language ability prediction proposed by the present invention.
[0043] As Figure 1 shown, in this embodiment, a system and method for predicting English language ability based on the LSTM algorithm proposed by the present invention will be described in detail below in combination with the above-mentioned accompanying drawings for each main component, process relationship and functional relationship of the method for predicting English language ability based on the LSTM algorithm proposed by the present invention.
[0044] The embodiments of the present application provide a method for predicting English language ability based on the LSTM algorithm. To facilitate the understanding of the embodiments of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0045] As Figure 1 shown, in this implementation manner, a method for predicting English language ability based on the LSTM algorithm mainly includes the following steps:
[0046] S1. Establish a CEFR corpus, where the CEFR corpus includes CEFR vocabulary and CEFR levels;
[0047] S2. Preprocess the CEFR corpus and complete feature extraction, and use the features as input feature vectors;
[0048] S3. Construct and train an LSTM recurrent neural network model for the CEFR corpus to obtain a trained LSTM recurrent neural network model;
[0049] S4. Predict the CEFR level of the tester according to the trained LSTM recurrent neural network model to obtain the prediction result of the tester's English language ability.
[0050] Specifically: In the CEFR corpus described in step S1, it is necessary to ensure that the matching relationship between the CEFR vocabulary and the CEFR levels is correct. Among them, the CEFR levels are divided into A1 (beginner), A2 (elementary), B1 (intermediate), B2 (upper intermediate), C1 (advanced), and C2 (proficient). Adjust and supplement the CEFR corpus described in step S1, divide it into a training data set, a validation set, and a test set, and construct and train the LSTM recurrent neural network model through the training data set, the validation set, and the test set.
[0051] Among them, in step S2, the preprocessing steps include:
[0052] Step S21: Traverse the CEFR vocabulary, set all the vocabulary as target vocabulary, and associate it with the corresponding CEFR level;
[0053] Step S22: Set word vectors for the target vocabulary and the corresponding CEFR levels;
[0054] Step S23: Use the values of the word vectors as input feature vectors.
[0055] Specifically, in step S2, the feature extraction method is:
[0056] Train a word2vec model through the CBOW method, and extract the text feature representation of the CEFR vocabulary through the word2vec model;
[0057] Extract features from the speech data of spoken English to obtain prosodic features and MFCC cepstral features, and use the prosodic features and MFCC cepstral features as the feature representation of the speech data of the CEFR vocabulary.
[0058] Among them, constructing and training the LSTM recurrent neural network model of the CEFR corpus in step S3 includes: initializing the LSTM model, training the LSTM model, and evaluating the LSTM model.
[0059] Among them, initializing the LSTM model in step S3 includes:
[0060] Iteratively train a pre-constructed LTSM model using the input feature vectors described in S2.
[0061] Training the LSTM model in step S3 includes:
[0062] Mark the multiple training data sets in chronological order, and use the training data set with an earlier time to train the initialized LTSM model to obtain a candidate prediction model;
[0063] Train the candidate prediction model by stacking training data sets in chronological order on the training data set with an earlier time.
[0064] Evaluating the LSTM model specifically includes:
[0065] Based on the test set, test the trained candidate prediction model; if the test effect meets the preset threshold, generate a final LSTM model.
[0066] In one embodiment, perform CEFR level prediction based on the final LSTM model and the candidate prediction model to obtain an initial prediction result and a prediction deviation, including:
[0067] Perform CEFR level prediction based on the final LSTM model to obtain an initial prediction result;
[0068] Perform CEFR level prediction based on the candidate prediction model to obtain a prediction deviation; the prediction deviation is the deviation between the predicted CEFR level and the actual CEFR level.
[0069] This method has a feature: First, the tester issues the pronunciation and Chinese interpretation of CEFR vocabulary. This language ability prediction system is based on the pronunciation and Chinese interpretation of CEFR vocabulary issued by the tester. By identifying the vocabulary pronunciation and Chinese interpretation, an assessment of whether the tester has mastered the vocabulary is achieved.
[0070] Specifically, preprocess the tester's speech, extract feature parameters from it, compare them with the features of the CEFR corpus, and perform matching according to certain similarity criteria and make a determination, so as to achieve an assessment of whether the tester has mastered the vocabulary.
[0071] Next, find the CEFR level matched by the vocabulary and mark the vocabulary.
[0072] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0073] In one embodiment, as Figure 2As shown, an English language proficiency prediction system based on the LSTM algorithm is provided, including: a data collection module 101, a model training module 102, a candidate prediction model 103, and a prediction module 104, where:
[0074] The data collection module 101 is used to collect vocabulary and CEFR level data to obtain a CEFR corpus;
[0075] The model training module 102 is used to divide the initial training data set into multiple training data sets, and iteratively train a pre-constructed LTSM model using the multiple training data sets to obtain multiple candidate prediction models and a final prediction model;
[0076] The candidate prediction model 103 is used to predict the CEFR level of a tester based on the multiple candidate prediction models and the final prediction model to obtain a deviation data training set; use the deviation data training set to train a pre-constructed deviation LTSM model to obtain a trained deviation LTSM model;
[0077] The prediction module 104 is used to perform CEFR level prediction based on the final prediction model and the trained deviation LTSM model to obtain an initial prediction result and a prediction deviation; use the prediction deviation to correct the initial prediction result to obtain the final CEFR level value of the tester.
[0078] In one embodiment, the model training module 102 is further used to divide the initial training data set into multiple training data sets, and iteratively train a pre-constructed LTSM model using the multiple training data sets to obtain multiple candidate prediction models and a final prediction model, including:
[0079] Divide the initial training data set into multiple training data sets in chronological order, and use the training data set with the earliest time to train the pre-constructed LTSM model to obtain a candidate prediction model;
[0080] Stack the training data sets on the training data set with the earliest time in chronological order to train the candidate prediction model. A candidate prediction model will be obtained in each round of training, and finally the final prediction model will be obtained.
[0081] In one embodiment, the candidate prediction model 103 is further used to predict the current traffic flow of the area to be predicted based on the multiple candidate prediction models and the final prediction model to obtain a deviation data training set.
[0082] In summary, the English language proficiency prediction system and method based on the LSTM algorithm provided by the present invention are based on the CEFR. Mainly through the evaluation of vocabulary, a model is constructed by a recurrent neural network with long short-term memory units, and the mastery of vocabulary is converted into a predicted value of the CEFR level of the tester's English language proficiency. Among them, vocabulary is the most direct and easily obtainable rating data for the CEFR level. Therefore, the method can adapt to the changing characteristics of the tester's CEFR level, and thus achieve long-term and accurate advance prediction of the tester's CEFR level. At the same time, it can also greatly improve the accuracy of the prediction results, accurately distinguish the vocabulary capabilities required for testing, and testers can carry out targeted training based on this to improve language learning efficiency. And the method has strong versatility, low cost, wide applicability, and is suitable for wide promotion and application.
[0083] It should be noted that the above embodiments are merely used as examples, and the present application is not limited to such examples, but can be variously changed.
[0084] It should be noted that in this specification, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0085] Finally, it should also be noted that the above series of processes not only include the processes executed in time series in the order described herein, but also include processes executed in parallel or separately, rather than in time order.
[0086] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by hardware related to computer program instructions. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The above-disclosed are only some preferred embodiments of the present application, and the scope of rights of the present application cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments and make equivalent changes according to the claims of the present application, which still fall within the scope covered by the invention.
Claims
1. A method for predicting English language proficiency based on LSTM algorithm, characterized in that: The steps include: S1. Establish a CEFR corpus, wherein the CEFR corpus includes CEFR vocabulary and CEFR levels; S2, preprocessing the CEFR corpus and performing feature extraction, and using the features as input feature vectors; S3, constructing and training an LSTM recurrent neural network model of the CEFR corpus to obtain a trained LSTM recurrent neural network model; S4. Predict the test taker's CEFR level according to the trained LSTM recurrent neural network model to obtain a prediction result of the test taker's English language proficiency.
2. The English language proficiency prediction method based on LSTM algorithm according to claim 1, characterized in that: In step S1, the CEFR corpus needs to ensure that the CEFR vocabulary and the CEFR levels match correctly, wherein the CEFR levels are divided into A1 (beginner), A2 (elementary), B1 (intermediate), B2 (upper intermediate), C1 (advanced) and C2 (proficient).
3. The English language proficiency prediction method based on LSTM algorithm according to claim 1, characterized in that: The CEFR corpus in step S1 is adjusted and supplemented, and divided into a training data set, a validation set and a test set, and the LSTM recurrent neural network model is constructed and trained through the training data set, the validation set and the test set.
4. The English language proficiency prediction method based on LSTM algorithm according to claim 1, characterized in that: In step S2, the preprocessing step includes: Step S21: traverse the CEFR vocabulary, set all vocabulary as target vocabulary, and associate them with corresponding CEFR levels; Step S22: Setting word vectors for the target vocabulary and the corresponding CEFR level; Step S23: using the value of the word vector as an input feature vector.
5. The English language proficiency prediction method based on LSTM algorithm according to claim 1, characterized in that: In step S2, the feature extraction method is: Training a word2vec model using the CBOW method, and extracting text feature representations of the CEFR vocabulary using the word2vec model; Feature extraction is performed on the speech data of spoken English to obtain prosodic features and MFCC cepstrum features, and the prosodic features and MFCC cepstrum features are used as feature representations of the speech data of the CEFR vocabulary.
6. The English language proficiency prediction method based on LSTM algorithm according to claim 1, characterized in that: The LSTM recurrent neural network model constructed and trained for the CEFR corpus in step S3 includes: Initialize the LSTM model, train the LSTM model, and evaluate the LSTM model.
7. The English language proficiency prediction method based on LSTM algorithm according to claim 4 or 6, characterized in that: The initialization of the LSTM model includes: The LTSM model is pre-built by iterative training using the input feature vector described in S2.
8. The English language proficiency prediction method based on LSTM algorithm according to claim 3 or 6, characterized in that: The training LSTM model comprises: Marking the plurality of training data sets in chronological order, using the earlier training data sets to train the initialized LTSM model, and obtaining a candidate prediction model; The candidate prediction model is trained by stacking the training data set on the earlier training data set in chronological order.
9. The English language proficiency prediction method based on LSTM algorithm according to claim 3 or 6, characterized in that: The evaluation of the LSTM model specifically includes: Based on the test set, testing the trained candidate prediction model; If the test result meets the preset threshold, the final LSTM model is generated.
10. An English language proficiency prediction system based on LSTM algorithm, characterized in that: The system is used to implement any one of the English language proficiency prediction methods based on the LSTM algorithm described in claims 1-9, and the system includes: The data collection module is used to collect CEFR vocabulary and CEFR level data to obtain the CEFR corpus; A model training module, used to divide the initial training data set into multiple training data sets, and iteratively train a pre-built LTSM model using the multiple training data sets to obtain multiple candidate prediction models and a final prediction model; A candidate prediction model is used to predict the CEFR level of the test subject according to the multiple candidate prediction models and the final prediction model to obtain a deviation data training set; and a pre-built deviation LTSM model is trained using the deviation data training set to obtain a trained candidate prediction model; The prediction module is used to predict the CEFR level according to the final prediction model and the trained candidate prediction model to obtain an initial prediction result and a prediction deviation; and to correct the initial prediction result using the prediction deviation to obtain the final CEFR level value of the test subject.