Intelligent Composition Correction and Improvement System and Method

Through language model and cluster analysis, multi-dimensional composition scores were determined, and the learning paths were dynamically adjusted in combination with the Q-learning algorithm, which solved the problems of single-dimensional scoring and user personalized needs in traditional composition correction, achieved more comprehensive composition scoring and personalized learning path design, and improved writing ability.

CN119886155BActive Publication Date: 2025-07-22BEIJING CETEN EDUCATION TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510370218.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-22
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

In the essay correction, the existing technology often focuses on single-dimensional scoring, ignores the overall structure and logical coherence of the composition, and fails to fully consider the personalized needs of users, resulting in limited learning effects.

Method used

The language model is used to extract composition features, combine cluster analysis and K nearest neighbor regression model to determine multi-dimensional scores, create personalized learning paths, and dynamically adjust the learning paths through the Q-learning algorithm, supporting multiple composition input methods and real-time monitoring.

Benefits of technology

Multi-dimensional composition scoring is realized, which can more comprehensively reflect the quality of composition and generate personalized learning paths based on user history to improve the pertinence and effectiveness of writing skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886155B_ABST
    Figure CN119886155B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent composition correction and improvement system and method, which relates to the fields of natural language processing and educational technology. It includes receiving a user's composition, converting the composition into a processable format and performing preprocessing, using a language model to extract features from the preprocessed composition, collecting the user's writing history at the same time, using the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, calculating the feature values of each composition scoring dimension, comprehensively scoring the composition based on the feature values of each composition scoring dimension, creating a learning template, analyzing the user's writing history, and generating a personalized learning path and practice tasks in combination with the overall composition score. By using the DBSCAN cluster analysis method to select the composition scoring dimensions, the present invention overcomes the limitations of traditional single-dimensional scoring and can more comprehensively reflect the overall quality of the composition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and educational technology, and particularly to an intelligent composition correction and improvement system and method. Background Art

[0002] With the rapid development of information technology, various intelligent tools have gradually been introduced in the field of education to improve teaching efficiency and students' learning effects. Especially in the aspects of composition correction and writing ability improvement, the traditional manual correction method is not only time-consuming and laborious, but also difficult to provide comprehensive and personalized feedback. With the progress of machine learning, especially deep learning technology, training models using large-scale text data can more accurately understand the content of compositions and provide more detailed scoring and feedback.

[0003] Although the existing methods have improved the accuracy of scoring to a certain extent, they still face several challenges. First, the existing methods often focus on single-dimensional scoring while ignoring the overall structure and logical coherence of the composition. Second, most models fail to fully consider the personalized needs of users and cannot dynamically adjust the learning path and practice tasks according to the user's historical writing records, resulting in limited learning effects. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an intelligent composition correction and improvement system and method to solve the problem that the existing methods often focus on single-dimensional scoring while ignoring the overall structure and logical coherence of the composition.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In the first aspect, the present invention provides an intelligent composition correction and improvement method, which includes,

[0008] Receiving a user's composition, converting the composition into a processable format and performing preprocessing, extracting features of the preprocessed composition using a language model, and simultaneously collecting the user's writing history records;

[0009] Using the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, calculating the feature values of each composition scoring dimension, and comprehensively scoring the composition based on the feature values of each dimension;

[0010] Creating a learning template, analyzing the user's writing history records, and generating personalized learning paths and practice tasks in combination with the overall composition score;

[0011] Performing real-time monitoring during the user's practice tasks and dynamically adjusting the user's learning path using the Q-learning algorithm.

[0012] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, the following steps are included: receiving a user's composition, converting the composition into a processable format and performing preprocessing, specifically including the following steps:

[0013] The user inputs the composition in three ways: text, voice, and image. The text input is set as the default input method. The user directly types the composition text through a web page. For voice input, ASR is called to convert the voice signal into composition text. For uploading a handwritten composition picture, OCR is used to extract the composition text.

[0014] The preprocessing includes removing stop words, word segmentation, and part-of-speech tagging.

[0015] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, the following steps are included: using a language model to extract features from the preprocessed composition, and at the same time collecting the user's writing history records, specifically including the following steps:

[0016] Using BERT to extract the semantic features of the preprocessed composition.

[0017] The semantic features include lexical information, grammatical information, article structure, and sentiment tendency.

[0018] Using principal component analysis to compress the semantic features into low-dimensional features, forming a low-dimensional feature vector.

[0019] The user's writing history records are all previously submitted compositions and the corresponding scoring results and feedback comments for all compositions.

[0020] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, the following steps are included: using the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, specifically including the following steps:

[0021] Using DBSCAN to cluster the low-dimensional feature vectors, obtaining N clustering clusters and the clustering centers of each clustering cluster.

[0022] Inputting the low-dimensional feature vectors into the K-nearest neighbor regression model to obtain the predicted values of the clustering clusters to which the composition belongs.

[0023] Using SHAP value analysis to calculate the importance of each feature in the low-dimensional feature vectors of each clustering center, obtaining the features that contribute the most to each clustering center.

[0024] According to the meaning and contribution degree of the features, generating a scoring dimension name for each clustering center, and finally obtaining the scoring dimensions and the descriptions corresponding to the scoring dimensions, forming a set of scoring dimensions.

[0025] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, wherein: calculating the eigenvalue of each composition scoring dimension, and comprehensively scoring the composition based on the eigenvalues of each composition scoring dimension, specifically including the following steps,

[0026] Obtain the eigenvalue of each composition scoring dimension through the methods of direct reference, combination, and use of external models, and normalize it;

[0027] Set adjustment parameters that can control the slope and position of the scoring curve;

[0028] Use the Sigmoid function, and calculate according to the normalized eigenvalue of the composition scoring dimension and the adjustment parameters to obtain the score of each composition scoring dimension;

[0029] According to the scores of each composition scoring dimension, use the geometric mean and combine with the exponential function to calculate and obtain the overall score of the composition.

[0030] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, wherein: create a learning template, analyze the user's writing history record, and combine with the overall composition score to generate a personalized learning path and practice tasks, specifically including the following steps,

[0031] Calculate the average value of the scoring results in the user's writing history record, and obtain the part with the lowest performance of the user's composition according to the average value;

[0032] Set learning goals and learning paths according to the part with the lowest performance of the user's composition and the overall composition score;

[0033] The learning path refers to a series of gradually progressive practice tasks;

[0034] Create a learning template containing practice tasks, add classification labels and difficulty levels to each practice task, and combine them into a template library;

[0035] The learning template containing practice tasks includes the goals of setting tasks, completion criteria, and recommended learning resources;

[0036] By calculating the matching degree between the user's learning path and the templates in the template library, select the template with the highest matching degree as the practice task.

[0037] As a preferred embodiment of the intelligent composition correction and improvement method of the present invention, wherein: conduct real-time monitoring during the user's practice tasks and dynamically adjust the learning path, specifically including the following steps,

[0038] Whenever the user completes a practice task, automatically collect the submitted content and related metadata;

[0039] The metadata includes the completion time and the proportion of correct answers;

[0040] Calculate the user's task performance score based on the collected content and metadata, and map the task performance score to an immediate reward;

[0041] Define the state space and action space according to the learning objective, and initialize the Q-value table;

[0042] The state space includes the current practice task type and difficulty level, and the action space includes increasing the practice intensity and decreasing the practice frequency;

[0043] Based on the immediate reward, use the Q-learning algorithm to update the Q-value table until the Q-value no longer changes during iteration, indicating that the optimal action selection strategy has been found.

[0044] In a second aspect, the present invention provides an intelligent composition correction and improvement system, including,

[0045] A composition collection module that receives the user's composition, converts the composition into a processable format and performs preprocessing, extracts features from the preprocessed composition using a language model, and simultaneously collects the user's writing history;

[0046] A composition scoring module that uses the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, calculates the feature values of each composition scoring dimension, and comprehensively scores the composition based on the feature values of each dimension;

[0047] A composition improvement module that creates a learning template, analyzes the user's writing history, and generates a personalized learning path and practice tasks in combination with the overall composition score;

[0048] A learning adjustment module that monitors the user in real time during the practice task and dynamically adjusts the user's learning path using the Q-learning algorithm.

[0049] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the intelligent composition correction and improvement method described in the first aspect of the present invention is implemented.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the intelligent composition correction and improvement method described in the first aspect of the present invention is implemented.

[0051] The beneficial effects of the present invention are as follows: By using the DBSCAN clustering analysis method to cluster low-dimensional feature vectors, groups with similar features in the composition can be identified, thereby determining multiple scoring dimensions. This method not only overcomes the limitations of traditional single-dimensional scoring but also can more comprehensively reflect the overall quality of the composition. In addition, by creating a learning template and combining the user's historical writing records and the overall score of the current composition, personalized learning paths and practice tasks can be generated. This customized design fully considers the unique needs and development stages of each student, ensuring that the provided practice tasks are both challenging and can effectively improve writing ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the intelligent composition correction and improvement method in Embodiment 1.

[0054] Figure 2 It is a schematic diagram of the overall score of the composition in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be made in conjunction with the accompanying drawings of the specification.

[0056] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0057] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.

[0058] Embodiment 1, referring to Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides an intelligent composition correction and improvement method, including the following steps:

[0059] S1. Receive the user's composition, convert it into a processable format, and perform preprocessing.

[0060] Specifically, it includes the following steps:

[0061] S1.1. To improve the convenience and flexibility of user interaction, three different composition input methods are supported, namely text input, voice input, and image input.

[0062] Text input is the default input method. Users can directly type the composition content through the web interface. This method is simple and direct, suitable for most users.

[0063] Voice input means that the user inputs the composition by voice. Using automatic speech recognition (ASR) technology, the user's voice signal is converted into text in real time. This method is especially suitable for those who are inconvenient to type or prefer oral expression, such as children or students with special needs.

[0064] Image input means that the user uploads a picture of a handwritten composition, and the text is extracted from the picture through optical character recognition (OCR) technology. This input method is very suitable for those users who are used to writing compositions by hand, especially for the situation of correcting paper-based homework.

[0065] The received composition is uniformly converted into a processable plain text format.

[0066] Furthermore, supporting multiple input methods (text, voice, image) greatly improves user-friendliness. Different types of users can choose the most suitable input method according to their preferences, thus reducing the usage threshold. For example, visually impaired users may prefer to use voice input, while users who are used to writing by hand can choose image input.

[0067] S1.2. For the composition converted into plain text format, stop words are removed. Stop words refer to words that frequently appear in natural language processing but contribute little to semantic understanding, such as "de", "shi", "zai", etc. Removing these words can reduce noise and improve the effect of feature extraction; then the continuous text is segmented into meaningful word units. For example, "I like reading" will be segmented into: I / like / reading. This step is particularly important for non-space-separated languages such as Chinese because they lack clear word boundaries; finally, part-of-speech tagging is performed. Part-of-speech tagging is the process of assigning a grammatical role to each word, such as noun, verb, adjective, etc. This process helps to further understand the sentence structure and semantic relationship.

[0068] Furthermore, removing stop words reduces the interference of meaningless words on the model, making the key information more prominent and improving the accuracy of feature extraction. Word segmentation provides clear lexical boundaries for subsequent semantic understanding and feature extraction, which is particularly important when dealing with non-space-separated languages such as Chinese. Part-of-speech tagging can better understand the sentence structure and semantic relationships, thus more accurately evaluating the quality of the composition. For example, when judging whether a sentence is complete or has grammar errors, part-of-speech tagging provides important evidence.

[0069] S2. Use a language model to extract features from the preprocessed composition and collect the user's writing history at the same time.

[0070] Specifically, it includes the following steps:

[0071] S2.1. After completing the preprocessing, use the language model BERT to extract the lexical information, grammatical information, article structure, and sentiment tendency features of the composition. BERT is a deep learning model based on the Transformer architecture, which can capture bidirectional context information in the text, thus providing a richer semantic representation.

[0072] Lexical information means that BERT generates context-sensitive embedding vectors for each word through its multi-layer Transformer encoder structure. It can not only recognize the meaning of the word itself but also understand its meaning in a specific context.

[0073] Grammatical information means that by analyzing the lexical relationships in the sentence, BERT can identify complex grammatical structures, such as subject-predicate-object relationships, clause nesting, etc., which is crucial for evaluating the grammatical correctness and complexity of the composition.

[0074] Article structure means that BERT can help identify the logical connections between paragraphs and the overall structure of the article. For example, it can distinguish introductions, main texts, and supporting arguments, which helps to evaluate the organization and coherence of the composition.

[0075] Sentiment tendency means that BERT can be used for sentiment analysis to identify the sentiment tendency expressed in the composition (such as positive, negative, or neutral), which is of great significance for evaluating the sentiment expression ability of the composition and the author's attitude.

[0076] S2.1. To reduce the computational complexity and improve the efficiency of subsequent analysis, use the principal component analysis (PCA) technique to perform dimensionality reduction on the high-dimensional semantic features extracted by BERT. PCA is a commonly used linear dimensionality reduction method that can convert high-dimensional features into low-dimensional features and concatenate them into low-dimensional feature vectors while trying to retain the main information of the original data. Its expression is:

[0077] ;

[0078] Among them, represents the low-dimensional feature vector, represents principal component analysis, represents the matrix of the original high-dimensional features, represents the number of dimensions after dimensionality reduction, usually choosing = 50 to retain the main information.

[0079] Furthermore, through the PCA technique, the high-dimensional semantic features extracted by BERT are compressed into low-dimensional feature vectors, significantly reducing the computational complexity of subsequent analysis. This not only improves the running efficiency but also makes large-scale data processing possible. Although dimensionality reduction is performed, PCA can retain the main information of the original data, ensuring the accuracy of feature extraction. The low-dimensional feature vectors provide a solid foundation for subsequent clustering analysis and scoring models.

[0080] S2.2. Collect the user's writing history records. Each time a user submits an essay, it will be scored, and the scoring results will be stored in the database. These scores not only reflect the quality of the essay but also provide a basis for formulating personalized learning paths.

[0081] S3. Use the method of clustering analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the essay, calculate the feature values of each essay scoring dimension, and comprehensively score the essay based on the feature values of each essay scoring dimension.

[0082] Specifically, it includes the following steps:

[0083] S3.1. Adopt DBSCAN for the low-dimensional feature vector to perform clustering, output several clustering clusters, each clustering cluster represents a potential scoring dimension, and obtain the center of each clustering cluster, that is, the clustering center, and its expression is:

[0084] ;

[0085] Among them, represents the output clustering cluster, represents the DBSCAN algorithm in clustering analysis, represents the neighborhood radius, represents the minimum number of samples.

[0086] Assume that DBSCAN generates 3 main clustering clusters, corresponding to narrative coherence, emotional expression intensity, and view diversity respectively.

[0087] Furthermore, through DBSCAN clustering, different types of composition samples can be identified and divided into multiple clusters. Each cluster represents a specific writing style or ability level, thus enabling multi-dimensional composition evaluation. Compared with the single-dimensional scoring method, this method can more comprehensively reflect the quality of the composition.

[0088] S3.2. Use the K-nearest neighbor regression model as the interpretation model and input the low-dimensional feature vector into the K-nearest neighbor regression model to obtain the prediction result of which cluster the composition belongs to. Its expression is:

[0089] ;

[0090] where represents the predicted value of the KNN model, that is, the prediction result of which cluster the composition belongs to, represents the K-nearest neighbor regression model;

[0091] For the low-dimensional feature vectors of each cluster center, use the SHAP value formula to calculate the importance of the features in each low-dimensional feature vector. Its expression is:

[0092] ;

[0093] where represents the importance of the th feature, represents the set of all features, that is, all dimensions of the low-dimensional feature vector, represents the feature subset, which is a part of the feature combination selected from the set of all features, that is, the feature combination currently being considered, represents the feature subset , that is, the number of features included, represents the total number of features.

[0094] According to the SHAP value analysis results, find the features that contribute the most to each cluster center. For example: If the important feature of a certain cluster center is the semantic similarity between adjacent sentences, then this cluster may correspond to narrative coherence. If the important feature of a certain cluster center is the distribution of sentiment scores, then this cluster may correspond to the intensity of emotional expression.

[0095] Furthermore, the SHAP value is a technique for interpreting the model output. Through SHAP value analysis, it is possible to clarify which features have the greatest impact on the scoring of each cluster center.

[0096] According to the features with the highest importance and the meaning of the features in the SHAP value analysis results, generate a scoring dimension name for each cluster center.

[0097] For example, there are 3 clustering centers. The feature with the greatest contribution from clustering center 1 is (semantic similarity between adjacent sentences), which is interpreted as narrative coherence and reflects the fluency of the composition in terms of the timeline or logic; the feature with the greatest contribution from clustering center 2 is (emotional score distribution), which is interpreted as the intensity of emotional expression and reflects the performance of the composition in emotional transmission. The feature with the greatest contribution from clustering center 3 is (entropy value of theme distribution), which is interpreted as the diversity of viewpoints and reflects the richness of the theme distribution in the composition.

[0098] Finally, the obtained scoring dimensions and their corresponding explanations are formed into a set of scoring dimensions , for example:

[0099] ;

[0100] Furthermore, according to the features and importance of each clustering center, specific names and descriptions are generated for each scoring dimension. This customized design of scoring dimensions not only makes the scoring more transparent and understandable, but also provides strong support for the formulation of personalized learning paths.

[0101] S3.3. If a certain scoring dimension directly corresponds to a certain dimension in the low-dimensional feature vector, the value of that dimension can be directly used as the feature value.

[0102] If a certain scoring dimension needs to be represented by integrating multiple features, the feature value can be calculated by means of weighted summation or non-linear combination. For example, for the diversity of viewpoints, it is necessary to combine the entropy value of theme distribution and the number of themes to obtain.

[0103] If a certain scoring dimension requires more complex calculations, an external model can be called to calculate the feature value. For example, for the intensity of emotional expression, the TextBlob model can be used to extract the emotional score.

[0104] Normalize the feature value of each scoring dimension to the range [0, 1] for subsequent calculations.

[0105] S3.4. Set the slope and the adjustment parameter that controls the position of the scoring curve. A larger value will make the scoring change more sensitively, while the adjustment parameter determines the position of the scoring curve.

[0106] Use the Sigmoid function and calculate according to the normalized feature values of the composition scoring dimensions and the adjustment parameter to obtain the score of each composition scoring dimension. Its expression is:

[0107] ;

[0108] Among them, represents the score of the th composition scoring dimension, represents the eigenvalue corresponding to the th composition scoring dimension, represents the adjustment parameter that controls the slope of the scoring curve, represents the adjustment parameter that controls the position of the scoring curve;

[0109] Furthermore, the Sigmoid function restricts the scoring value between 0 and 1, ensuring the rationality of the scoring result, facilitating the subsequent comprehensive scoring calculation, making the scoring result smoother, avoiding the influence of extreme values, and enhancing the stability of the scoring.

[0110] According to the scores of each composition scoring dimension, the overall score of the composition is calculated and obtained by using the geometric mean and combining with the exponential function. Its expression is:

[0111] ;

[0112] Among them, represents the overall score of the composition, represents the number of composition scoring dimensions, represents the weight of the th dimension.

[0113] Furthermore, using the geometric mean and combining with the exponential function to calculate the overall score of the composition ensures an equal consideration of the scores of each dimension and avoids the excessive influence of too high or too low scores in a certain dimension on the overall score.

[0114] S4. Create a learning template, analyze the user's writing history records, and generate a personalized learning path and practice tasks in combination with the overall composition score.

[0115] Specifically, it includes the following steps:

[0116] S4.1. Create a learning template that contains practice tasks. Each template details the goal of the task, the completion criteria, and the recommended learning resources, adds classification labels to each practice task, and finally combines them into a learning template library.

[0117] For example, add classification labels (such as vocabulary expansion, grammar correction, etc.) and difficulty levels (beginner, intermediate, advanced) to each practice task.

[0118] The goal of the task refers to clarifying the goal of each task. For example, through this task, you will learn how to effectively use synonyms in writing.

[0119] Define the completion criteria for each task. For example, successfully replace at least 5 common words with more advanced synonyms.

[0120] Provide relevant learning resources such as online courses, reference books, or video tutorials to help users better understand and complete the tasks.

[0121] Organize all the practice tasks according to classification tags and difficulty levels to form a template library.

[0122] Further explain that as the user's writing level improves, the difficulty of the tasks in the template library can be adjusted to meet the changing learning needs of the user.

[0123] S4.2. Extract all the compositions submitted by the user in the past and their corresponding scoring results from the database. For each scoring dimension, calculate the average score of all the composition scores.

[0124] For example, if the user's scores for lexical richness in the past five compositions are 0.7, 0.6, 0.8, 0.5, and 0.6 respectively, then the average score for this dimension is 0.64.

[0125] Compare the average scores of each dimension and identify the lowest-performing part. For example, if the average score for lexical richness is 0.64 and the average score for logical coherence, another scoring dimension, is 0.9, then it can be determined that lexical richness is the area where the user most needs to improve.

[0126] Based on the identified weak links and the overall score of the composition, set specific learning goals and a progressive learning path for the user.

[0127] First, set learning goals. For example, if the user has a low average score in lexical richness, set the goal to improve the score for this dimension to above 0.8.

[0128] Second, design a series of progressive tasks according to the set goals. These tasks should be from simple to complex, gradually guiding the user to master the required skills. For example, the initial task may be to identify and use synonyms to replace common words, and subsequent tasks may involve more complex word collocations and sentence structures.

[0129] Further explain that according to the user's performance and overall score, set scientific and reasonable learning goals for them and design a series of progressive learning paths. This kind of planning not only takes into account the short-term progress needs but also lays a solid foundation for long-term development.

[0130] Finally, by calculating the matching degree between the user's current learning path and the templates in the template library, select the template with the highest matching degree as the practice task.

[0131] Further explanation, by calculating the matching degree between the user's learning path and the templates in the template library, the most suitable practice tasks can be customized for each user. This method ensures that the provided tasks are both challenging and can effectively promote the user's progress.

[0132] S5. Monitor in real time when the user is performing practice tasks and dynamically adjust the learning path.

[0133] Specifically, it includes the following steps:

[0134] Collect the practice tasks submitted by the user and related metadata

[0135] Calculate the user's task performance score based on the collected practice tasks and metadata. Its expression is:

[0136] ;

[0137] Among them, represents the performance score of the user in this practice task, represents the adjustment coefficient, which is used to consider the influence of additional factors such as creativity or complexity.

[0138] Directly use the calculated task performance score as an immediate reward .

[0139] In order to perform dynamic adjustment using the Q-learning algorithm, it is necessary to define the state space and action space according to the learning goal, and at the same time set the initial values of the Q-value table to 0. This also means that at the beginning, there is no preference for any action in any state.

[0140] The state space includes the current practice task type and difficulty level. For example, the state can be represented as a grammar correction task, intermediate difficulty.

[0141] The action space includes possible action options, such as increasing the practice intensity (providing more complex tasks) or reducing the practice frequency (giving more review time). These actions are designed to help the user gradually improve their abilities or consolidate the knowledge they have learned.

[0142] According to the immediate reward, use the Q-learning algorithm to update the Q-value table until the Q-value no longer changes in the iteration, indicating that the optimal action selection strategy has been found. Its expression is:

[0143] ;

[0144] Among them, represents the updated Q-value when taking action in state , represents the Q-value in state Take an action below The current Q-value of Denote the learning rate Denote the immediate reward Denote the discount factor Denote taking an action in the next state The Q-value of

[0145] When the Q-value no longer changes during iteration, it means that the optimal action selection strategy has been found. At this time, the learning path of the user can be dynamically adjusted according to this strategy to ensure that the provided practice tasks always meet the actual needs and development stage of the user.

[0146] Furthermore, by continuously exploring the optimal action selection strategy using the Q-learning algorithm, the most effective learning path can be found in a complex learning environment. As the user's writing level improves, the strategy is dynamically adjusted to ensure that the provided practice tasks always meet the actual needs and development stage of the user, supporting their long-term development.

[0147] This embodiment also provides an intelligent composition correction and improvement system, including:

[0148] A composition collection module that receives the user's composition, converts the composition into a processable format and preprocesses it, extracts features of the preprocessed composition using a language model, and simultaneously collects the user's writing history;

[0149] A composition scoring module that determines the scoring dimensions of the composition using the method of cluster analysis and the K-nearest neighbor regression model, calculates the feature values of each composition scoring dimension, and comprehensively scores the composition based on the feature values of each dimension;

[0150] A composition improvement module that creates a learning template, analyzes the user's writing history, and generates a personalized learning path and practice tasks in combination with the overall composition score;

[0151] A learning adjustment module that monitors in real time when the user is performing practice tasks and dynamically adjusts the user's learning path using the Q-learning algorithm.

[0152] This embodiment also provides a computer device applicable to the intelligent composition correction and improvement method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent composition correction and improvement method as proposed in the above embodiment.

[0153] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or buttons, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0154] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing intelligent composition correction and improvement as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0155] In summary, the present invention: By using the DBSCAN clustering analysis method to cluster low-dimensional feature vectors, it can identify groups with similar features in the composition, thereby determining multiple scoring dimensions. This method not only overcomes the limitations of traditional single-dimensional scoring but also can more comprehensively reflect the overall quality of the composition. In addition, by creating a learning template and combining the user's historical writing records and the overall score of the current composition, personalized learning paths and practice tasks can be generated. This customized design fully considers the unique needs and development stages of each student, ensuring that the provided practice tasks are both challenging and can effectively improve writing ability.

[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An intelligent method for correcting and improving compositions, characterized in that: including, receiving the user's composition, converting the composition into a processable format and performing preprocessing, using a language model to extract features from the preprocessed composition, and simultaneously collecting the user's writing history; using the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, calculating the feature values of each composition scoring dimension, and comprehensively scoring the composition based on the feature values of each composition scoring dimension; creating a learning template, analyzing the user's writing history, and generating a personalized learning path and practice tasks in combination with the overall score of the composition; performing real-time monitoring during the user's practice tasks and dynamically adjusting the user's learning path using the Q-learning algorithm; using the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, specifically including the following steps, using DBSCAN to cluster the low-dimensional feature vectors to obtain N clustering clusters and the clustering centers of each clustering cluster; inputting the low-dimensional feature vectors into the K-nearest neighbor regression model to obtain the predicted values of the clustering clusters to which the composition belongs; using SHAP value analysis to calculate the importance of each feature in the low-dimensional feature vectors of each clustering center to obtain the feature that contributes the most to each clustering center; generating a scoring dimension name for each clustering center according to the meaning and contribution degree of the feature, and finally obtaining the scoring dimensions and the descriptions corresponding to the scoring dimensions to form a scoring dimension set; the calculating the feature values of each composition scoring dimension and comprehensively scoring the composition based on the feature values of each dimension specifically includes the following steps, obtaining the feature values of each composition scoring dimension by directly referring to, combining, and using external models, and normalizing; setting adjustment parameters that can control the slope and position of the scoring curve; using the Sigmoid function and calculating according to the normalized composition scoring dimension feature values and adjustment parameters to obtain the scores of each composition scoring dimension; calculating the overall score of the composition according to the scores of each composition scoring dimension using the geometric mean and combining with the exponential function.

2. The intelligent composition correction and improvement method according to claim 1, characterized in that: receiving the user's composition, converting the composition into a processable format and performing preprocessing, specifically including the following steps, the user inputs the composition in three ways: text, voice, and image, the text input is set as the default input method, and the user directly types the composition text through the web page; for voice input, ASR is called to convert the voice signal into composition text; for uploading a handwritten composition picture, OCR is used to extract the composition text; the preprocessing includes removing stop words, word segmentation, and part-of-speech tagging.

3. The intelligent composition marking and improvement method according to claim 2, wherein: using a language model to extract features from the preprocessed composition and simultaneously collecting the user's writing history, specifically including the following steps, using BERT to extract the semantic features of the preprocessed composition; the semantic features include lexical information, grammatical information, article structure, and sentiment tendency; using principal component analysis to compress the semantic features into low-dimensional features to form low-dimensional feature vectors; the user's writing history is all the compositions submitted in the past and all the corresponding scoring results and feedback opinions of the compositions.

4. The intelligent composition correction and improvement method according to claim 3, characterized in that: Create a learning template, analyze the user's writing history, and generate a personalized learning path and practice tasks in combination with the overall composition score. The specific steps are as follows: Calculate the average of the scoring results in the user's writing history and obtain the part with the lowest performance of the user's composition based on the average. Set learning goals and a learning path according to the part with the lowest performance of the user's composition and the overall score of the composition. The learning path refers to a series of gradually progressive practice tasks. Create a learning template containing practice tasks, add classification labels and difficulty levels to each practice task, and combine them into a template library. The learning template containing practice tasks includes the goal of setting tasks, completion criteria, and recommended learning resources. By calculating the matching degree between the user's learning path and the templates in the template library, select the template with the highest matching degree as the practice task.

5. The intelligent composition correction and improvement method according to claim 4, characterized in that: Monitor in real time when the user is performing practice tasks and dynamically adjust the learning path. The specific steps are as follows: Whenever the user completes a practice task, automatically collect the submitted content and related metadata. The metadata includes the completion time and the proportion of correct answers. Calculate the task performance score of the user based on the collected content and metadata, and at the same time map the task performance score to an immediate reward. Define the state space and action space according to the learning goal and initialize the Q-value table. The state space includes the current practice task type and difficulty level, and the action space includes increasing the practice intensity and decreasing the practice frequency. Based on the immediate reward, use the Q-learning algorithm to update the Q-value table until the Q-value no longer changes in the iteration, indicating that the best action selection strategy has been found.

6. An intelligent composition correction and improvement system, based on the intelligent composition correction and improvement method according to any one of claims 1 to 5, characterized in that: Including: A composition collection module that receives the user's composition, converts the composition into a processable format and preprocesses it, extracts features from the preprocessed composition using a language model, and at the same time collects the user's writing history. A composition scoring module that uses the method of cluster analysis and the K-nearest neighbor regression model to determine the scoring dimensions of the composition, calculates the feature values of each composition scoring dimension, and comprehensively scores the composition based on the feature values of each dimension. A composition improvement module that creates a learning template, analyzes the user's writing history, and generates a personalized learning path and practice tasks in combination with the overall composition score. A learning adjustment module that monitors in real time when the user is performing practice tasks and dynamically adjusts the user's learning path using the Q-learning algorithm.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent composition correction and improvement method described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent composition correction and improvement method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic composition scoring method based on multi-stage learning

    CN115659954A

  • Intelligent language education system and method

    CN118135851A