CNN algorithm-based English composition automatic scoring system generation method
Through the generation method of the automatic scoring system of English composition based on CNN algorithm, the problems of efficiency and accuracy of the automatic scoring system of English composition in the prior art are solved, and efficient and accurate automatic scoring effect is achieved.
Patent Information
- Application Number
- CN202510079703.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to build an efficient and accurate automatic scoring system for English compositions, and face challenges such as data collection and preprocessing, feature engineering, model design, evaluation and optimization.
The generation method of the English composition automatic scoring system based on CNN algorithm is adopted, including data preprocessing, training and evaluation of CNN models. The specific steps include collecting and preprocessing the English composition sample data, using CNN for sequence modeling and feature extraction, training the CNN model, and performing model evaluation.
By unifying the data format, reducing noise, and extracting local and global features, efficient, accurate and automatic scoring of English compositions is achieved, reducing the time-consuming and subjective impact of manual scoring.
Smart Images

Figure CN120068858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and computer-aided education technology. Specifically, it relates to a method for generating an automatic English composition scoring system based on the CNN algorithm. Background Art
[0002] In today's globalized educational context, English writing ability has become one of the important indicators for measuring students' comprehensive language application ability. With the booming development of online education, how to efficiently and accurately evaluate students' English composition level has become an urgent problem to be solved in the field of educational technology. The traditional manual scoring method is not only time-consuming and laborious, but also easily affected by subjective factors, making it difficult to ensure the consistency and fairness of scoring. Therefore, it is particularly important to develop an automatic English composition scoring system based on artificial intelligence technology. In recent years, deep learning technology has made remarkable progress in the field of natural language processing. Among them, the convolutional neural network (CNN), due to its powerful feature extraction ability and good generalization performance, has been widely applied to tasks such as text classification and sentiment analysis. Applying CNN to automatic English composition scoring can make full use of its local feature capture ability to conduct a detailed analysis of the vocabulary, sentence patterns, paragraph structures, etc. in the composition, and then achieve automatic scoring. However, to build an efficient and accurate automatic English composition scoring system based on the CNN algorithm, there are still many challenges, such as data collection and preprocessing, feature engineering, model design, evaluation and optimization, etc. Summary of the Invention
[0003] In view of this, in view of the deficiencies of the prior art, the present invention proposes a method for generating an automatic English composition scoring system based on the CNN algorithm, aiming to solve at least one of the problems raised in the above background art.
[0004] The present invention provides a method for generating an automatic English composition scoring system based on the CNN algorithm, including the following steps:
[0005] S1. Collect English composition sample data and perform data preprocessing on the English composition sample data;
[0006] S2. Use CNN to perform sequence modeling on the composition sample data after data preprocessing and perform feature extraction;
[0007] S3. The composition sample data after feature extraction is used as training data and input into the CNN model, and the CNN model is trained;
[0008] S4. Evaluate the trained CNN model.
[0009] In some embodiments, the English composition sample data in step S1 is to collect English composition samples from multiple schools and educational institutions. The English composition samples cover different topics, styles, and difficulty levels, and each English composition sample has an expert score or a teacher score.
[0010] In some embodiments, the data preprocessing includes: removing personal information and special characters from the English composition samples; splitting the English composition samples into words or phrases; converting all English composition samples to lowercase, and counting all unique words that appear.
[0011] In some embodiments, the NLTK library is used to remove personal information and special characters from the English composition samples.
[0012] In some embodiments, step S2 uses a pre-trained Word2Vec model to convert the words in the English composition samples into vector representations of a fixed dimension. For new words or out-of-vocabulary words, random initialization is used for processing.
[0013] In some embodiments, based on the word vector representations of the English composition samples, CNN is used to perform sequence modeling on the English composition samples to capture local features. The CNN extracts n-gram features from the English composition samples through convolutional layers and reduces the feature dimension through pooling layers.
[0014] In some embodiments, an LSTM layer is added on the basis of the CNN to extract global features of the English composition samples.
[0015] In some embodiments, a CNN model is built using the TensorFlow or PyTorch framework.
[0016] In some embodiments, MSE is selected as the loss function in TensorFlow or PyTorch, and the Adam optimizer is used for model training;
[0017] The English composition sample dataset is divided into a training set, a validation set, and a test set. The training set is used for model training, and the validation set is used to monitor the model performance. When the performance on the validation set reaches a preset level, training is stopped and the model performance is evaluated on the test set.
[0018] In some embodiments, the evaluation metric functions in the Scikit-learn library are used to calculate the performance metrics of the model on the test set, and the performance of different models is compared; an online English composition scoring system is developed using Flask or Django, and users can upload compositions on this platform and obtain instant scoring feedback.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: By removing personal information, special characters, and converting the text to lowercase, the data format is unified and noise is reduced. Splitting into words or phrases helps the model better understand the text structure, and counting unique words lays the foundation for subsequent feature extraction and model training. With the help of the pre-trained Word2Vec model, words are mapped into a vector space of a fixed dimension, which not only preserves the semantic relationships between words but also accelerates the model training process. The handling strategy for out-of-vocabulary words ensures the robustness of the model. Combining the advantages of CNN and LSTM can capture local features (such as n-gram) in the text and learn global context information, which is crucial for understanding and evaluating the quality of English compositions. Using deep learning frameworks such as TensorFlow or PyTorch provides a high degree of flexibility and scalability. Choosing MSE as the loss function and Adam optimizer is because they can provide good convergence speed and stability in most cases. By dividing the dataset into training set, validation set, and test set, the model training process can be effectively monitored, overfitting can be prevented, and the true performance of the model can be evaluated on an independent test set. Using the evaluation metric functions in the Scikit-learn library can quantify the performance of the model, facilitating comparison and selection between different models. The developed online scoring system provides a convenient platform for users to upload compositions and obtain instant feedback, which not only improves the user experience but also makes the system more acceptable and usable by educational institutions and individual users. This English composition automatic scoring system based on the CNN algorithm can not only be applied to schools and educational institutions to help teachers reduce the scoring burden but also be used in language learning applications, online education platforms, etc., with broad market potential and application value.
[0020] The above general description and the following detailed description are exemplary and explanatory only and do not limit the present disclosure.
[0021] Other features and aspects of the present disclosure will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a method for generating an English composition automatic scoring system based on the CNN algorithm provided by an embodiment of the present invention. Detailed implementation manners
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0025] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.
[0026] The terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0027] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0028] As described in the background art, in today's globalized educational context, English writing ability has become one of the important indicators to measure students' comprehensive language application ability. With the booming development of online education, how to efficiently and accurately evaluate students' English composition level has become an urgent problem to be solved in the field of educational technology. The traditional manual scoring method is not only time-consuming and laborious, but also easily affected by subjective factors, making it difficult to ensure the consistency and fairness of scoring. Therefore, it is particularly important to develop an automatic English composition scoring system based on artificial intelligence technology. In recent years, deep learning technology has made remarkable progress in the field of natural language processing. Among them, the convolutional neural network (CNN) has been widely used in tasks such as text classification and sentiment analysis due to its powerful feature extraction ability and good generalization performance. Applying CNN to automatic English composition scoring can make full use of its local feature capture ability to carefully analyze the vocabulary, sentence patterns, paragraph structures, etc. in the composition, and then achieve automatic scoring. However, to build an efficient and accurate automatic English composition scoring system based on the CNN algorithm, there are still many challenges, such as data collection and preprocessing, feature engineering, model design, evaluation and optimization, etc.
[0029] To address the above issues, a method for generating an automatic English composition scoring system based on the CNN algorithm proposed in this application unifies the data format and reduces noise by removing personal information, special characters, and converting the text to lowercase. Splitting into words or phrases helps the model better understand the text structure, while counting unique words lays the foundation for subsequent feature extraction and model training. With the help of the pre-trained Word2Vec model, words are mapped into a vector space of a fixed dimension, which not only preserves the semantic relationships between words but also accelerates the model training process. The handling strategy for out-of-vocabulary words ensures the robustness of the model. Combining the advantages of CNN and LSTM can capture both local features (such as n-gram) in the text and learn global context information, which is crucial for understanding and evaluating the quality of English compositions. Using deep learning frameworks such as TensorFlow or PyTorch provides a high degree of flexibility and scalability. Choosing MSE as the loss function and the Adam optimizer is because they can provide good convergence speed and stability in most cases. By dividing the dataset into training set, validation set, and test set, the model training process can be effectively monitored, overfitting can be prevented, and the true performance of the model can be evaluated on an independent test set. Using the evaluation metric functions in the Scikit-learn library can quantify the model's performance, facilitating comparison and selection between different models. The developed online scoring system provides a convenient platform for users to upload compositions and obtain instant feedback, which not only improves the user experience but also makes the system more acceptable and usable by educational institutions and individual users. This automatic English composition scoring system based on the CNN algorithm can not only be applied to schools and educational institutions to help teachers reduce the scoring burden but also be used in language learning applications, online education platforms, etc., with broad market potential and application value.
[0030] Refer to Figure 1 As shown, a method for generating an automatic English composition scoring system based on the CNN algorithm according to an embodiment of the present application includes the following steps:
[0031] S1. Collect English composition sample data and perform data preprocessing on the English composition sample data;
[0032] S2. Use CNN to perform sequence modeling on the composition sample data after data preprocessing and perform feature extraction;
[0033] S3. The composition sample data after feature extraction is input into the CNN model as training data, and the CNN model is trained;
[0034] S4. Evaluate the trained CNN model.
[0035] In some specific embodiments, the English composition sample data in step S1 is collected from multiple schools and educational institutions. The English composition samples cover different topics, styles, and difficulty levels, and each English composition sample has an expert score or a teacher score.
[0036] It should be understood that multiple representative schools and educational institutions that can provide high-quality English composition samples are selected, including but not limited to public schools, private schools, language training institutions, and international schools. All grades from the upper grades of primary school to high school are covered, and it is ensured that there is a sufficient sample size for each grade. In addition, the works of students from different classes should be included to increase the diversity of the data. Collect the English composition parts in regular homework, mid-term / final exams, and entrance exams (such as the high school entrance examination, college entrance examination, or standardized tests like TOEFL and IELTS) to cover a wider range of scenarios and difficulty levels. Before starting any form of data collection, written consent from the students and their parents or guardians must be obtained to ensure compliance with the requirements of relevant laws and regulations. Convert the paper-based compositions into digital format through a scanner, or directly extract text information from electronic documents. For handwritten manuscripts, OCR technology can be used for recognition and conversion into editable text files. Ensure that the selected composition topics cover a wide range of subject areas, such as personal experiences, social phenomena, technological developments, cultural differences, etc., and at the same time, pay attention to balancing different types of writing tasks, such as argumentative essays, expository essays, narrative essays, etc. Carefully check each composition and manually delete all information that may disclose the student's identity, such as full name, student ID number, contact information, etc. Use regular expressions or other text processing tools to automatically detect and remove non-standard characters, such as extra spaces, tab characters, line break characters, etc.
[0037] In some specific embodiments, the data preprocessing includes: removing personal information and special characters from the English composition samples; splitting the English composition samples into words or phrases; converting all English composition samples to lowercase, and counting all the unique words that appear.
[0038] In some specific embodiments, the NLTK library is used to remove personal information and special characters from the English composition samples.
[0039] It should be understood that after word segmentation, meaningful features can be more easily extracted, which helps to improve the performance of the model, allows the model to learn more detailed language patterns, improves the accuracy of scoring, and the word segmentation results facilitate manual review and understanding of the model's decision-making process. Eliminating the influence of case differences makes different case forms of the same word be regarded as the same word, thus reducing the size of the vocabulary and improving statistical efficiency. It reduces the complexity of subsequent processing because most natural language processing tasks are case-insensitive. By counting the number of unique words, the language diversity and complexity of the composition can be evaluated. Knowing the most frequently used words can help determine which words should be given priority to be included in the model's vocabulary, thus saving storage space and computing resources. The optimization strategy based on frequent items can improve the model training speed and prediction efficiency. These data preprocessing steps not only help to improve the accuracy and efficiency of the English composition automatic scoring system, but also enhance the robustness and interpretability of the system.
[0040] In some specific embodiments, step S2 uses a pre-trained Word2Vec model to convert the words in the English composition sample into vector representations of a fixed dimension. For new words or out-of-vocabulary words, random initialization is used for processing.
[0041] It should be understood that by using a pre-trained Word2Vec model, a large number of general vocabulary and semantic relationships that it has learned can be utilized, which helps the model better understand and process new and unseen words. The Word2Vec model can capture the semantic relationships between words, making similar words close to each other in the vector space. This is particularly important for understanding the context and implicit meaning in the text, especially in composition scoring, which can more accurately evaluate the student's language application ability and creativity. For new words or out-of-vocabulary words, random initialization can ensure that these words have an initial representation in the model, thus allowing the model to learn their vector representations during the training process. This method increases the model's adaptability to new words and avoids performance degradation caused by vocabulary limitations. Since the Word2Vec model can generate high-dimensional vectors to represent words, these vectors can be used to calculate the similarity between words, thus helping the automatic scoring system to evaluate the composition quality more consistently and objectively. Using a pre-trained Word2Vec model can reduce the time and computing resources required to train the model from scratch because the pre-trained model already contains a large amount of language knowledge and only needs to be fine-tuned on specific tasks.
[0042] In some specific embodiments, based on the word vector representation of the English composition sample, CNN is used to perform sequence modeling on the English composition sample to capture local features. CNN extracts n-gram features in the English composition sample through convolutional layers and reduces the feature dimension through pooling layers.
[0043] It should be understood that CNN can automatically learn and capture local features in the text, such as phrases, word combinations, etc., which are crucial for understanding the semantics and emotions of the text. Through the convolutional layer, CNN can extract n-gram features of different lengths, which help the model understand the context information and language patterns in the text. The pooling layer can reduce the dimension of the feature map, thereby reducing the complexity and computational amount of the model while retaining the most important feature information. By capturing local features and extracting n-gram features, CNN can better generalize to unseen data and improve the performance of the model on new data. Since CNN can extract features from multiple perspectives (such as different n-gram lengths), it is insensitive to small changes in the input data, enhancing the robustness of the model. Compared with traditional recurrent neural networks (RNNs), CNN does not require complex time-step iterations when processing sequence data, so the model structure is simpler and easier to implement and optimize.
[0044] In some specific embodiments, an LSTM layer is added on the basis of CNN to extract global features of English composition samples.
[0045] In some specific embodiments, a CNN model is built using the TensorFlow or PyTorch framework.
[0046] In some specific embodiments, MSE is selected as the loss function in TensorFlow or PyTorch, and the Adam optimizer is used for model training;
[0047] The English composition sample dataset is divided into a training set, a validation set, and a test set. The training set is used for model training, and the validation set is used to monitor the model performance. When the performance on the validation set reaches the preset level, the training is stopped and the model performance is evaluated on the test set.
[0048] In some specific embodiments, the evaluation metric functions in the Scikit-learn library are used to calculate the performance metrics of the model on the test set, and the performance of different models is compared. An online English composition scoring system is developed using Flask or Django, and users can upload compositions on this platform and obtain instant scoring feedback.
[0049] It should be understood that the LSTM (Long Short-Term Memory) layer can capture long-term dependencies in the text, which is crucial for understanding the overall structure and theme of the composition. TensorFlow and PyTorch are two powerful deep learning frameworks that provide rich tools and libraries to simplify the model construction, training, and deployment processes. MSE (Mean Squared Error) is a commonly used regression loss function suitable for scoring tasks, while the Adam optimizer usually offers a faster convergence speed and better generalization ability due to its adaptive learning rate and momentum characteristics. Dividing the dataset into training set, validation set, and test set can ensure a reasonable evaluation of the model's performance on unseen data and avoid overfitting. Monitoring the model performance using the validation set and stopping the training when a preset level is reached can prevent the model from being overtrained. At the same time, evaluating the model performance using the test set can ensure that the model has good generalization ability. Using the evaluation metric functions in the Scikit-learn library to calculate the model performance metrics and comparing the performance of different models helps to select the best model structure. Developing an online scoring system using Web frameworks such as Flask or Django can provide a convenient platform for users to upload compositions and obtain instant scoring feedback. The online scoring system can provide real-time feedback to help students promptly understand their writing levels, thus promoting learning and improvement. Systems developed using modern Web frameworks usually have good scalability and maintainability, facilitating the addition of new features or upgrades in the future.
[0050] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A method for generating an automatic scoring system for English composition based on CNN algorithm, characterized in that: The following steps are involved: S1. Collect English composition sample data and perform data preprocessing on the English composition sample data; S2. Use CNN to perform sequence modeling on the composition sample data for data preprocessing and perform feature extraction; S3, the composition sample data after feature extraction is input into the CNN model as training data, and the CNN model is trained; S4. Perform model evaluation on the trained CNN model.
2. The method for generating an automatic scoring system for English composition based on a CNN algorithm according to claim 1, characterized in that: The English composition sample data in step S1 are English composition samples collected from multiple schools and educational institutions. The English composition samples cover different topics, styles and difficulty levels, and each English composition sample has an expert rating or a teacher rating.
3. The method for generating an automatic scoring system for English composition based on a CNN algorithm according to claim 2 is characterized in that: The data preprocessing includes: removing personal information and special characters from English composition samples; splitting the English composition samples into words or phrases; converting all English composition samples into lowercase letters, and counting all unique words that appear.
4. The method for generating an automatic scoring system for English composition based on a CNN algorithm according to claim 3 is characterized in that: Use the NLTK library to remove personal information and special characters from English composition samples.
5. The method for generating an automatic scoring system for English composition based on a CNN algorithm according to claim 4 is characterized in that: In step S2, the pre-trained Word2Vec model is used to convert the words in the English composition sample into vector representations of fixed dimensions. For new words or unregistered words, random initialization is used for processing.
6. The method for generating an automatic scoring system for English composition based on a CNN algorithm according to claim 5 is characterized in that: Based on the word vector representation of English composition samples, CNN is used to perform sequence modeling on English composition samples to capture local features. CNN extracts n-gram features in English composition samples through convolutional layers and reduces feature dimensions through pooling layers.
7. The method for generating an automatic scoring system for English composition based on CNN algorithm according to claim 6 is characterized in that: An LSTM layer is added on the basis of CNN to extract the global features of English composition samples.
8. The method for generating an automatic scoring system for English composition based on CNN algorithm according to claim 7 is characterized in that: Use TensorFlow or PyTorch framework to build CNN model.
9. The method for generating an automatic scoring system for English composition based on CNN algorithm according to claim 8, characterized in that: Select MSE as the loss function in TensorFlow or PyTorch and use the Adam optimizer for model training; The English composition sample dataset is divided into training set, validation set and test set. The training set is used for model training, and the validation set is used to monitor model performance. When the performance on the validation set reaches the preset level, the training is stopped and the model performance is evaluated on the test set.
10. The method for generating an automatic scoring system for English composition based on CNN algorithm according to claim 9, characterized in that: Use the evaluation metric function in the Scikit-learn library to calculate the performance indicators of the model on the test set and compare the performance of different models. Use Flask or Django to develop an online English composition scoring system, where users can upload their compositions and get instant scoring feedback.