English language ability assessment method and system
By analyzing the stylistic categories and sentence structures of students' English writing content and combining it with the Q-learning algorithm, we can identify logical relationship deviations and quantify the degree of contextual ambiguity. This solves the problem of low accuracy in logical and contextual ambiguity analysis in existing English language proficiency assessment methods, and achieves more accurate and personalized assessment.
Patent Information
- Application Number
- CN202511094450.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing English language proficiency assessment methods fail to comprehensively examine students' language abilities, especially in the analysis of logic and contextual ambiguity, which leads to large assessment errors.
By analyzing the stylistic categories and sentence structures of students' English writing content, identifying logical relationship deviations, and using the Q-learning algorithm to quantify the degree of contextual ambiguity, we design an English language proficiency assessment framework to achieve personalized assessment.
It improves the accuracy of the analysis of the logic and contextual ambiguity of English language proficiency, reduces assessment errors, provides more accurate and personalized language proficiency assessment, and can identify logical structure problems in students' writing and give targeted suggestions.
Smart Images

Figure CN120597003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of English language proficiency assessment, and in particular to an English language proficiency assessment method and system. Background Art
[0002] Computer technology can identify grammatical errors, irregular sentence structures, unclear logical relationships, and other issues in students' writing, and conduct in-depth analysis of the language proficiency of the written content. Especially in English writing assessments, factors such as contextual ambiguity, clarity of logical relationships, and appropriateness of stylistic structure have become important indicators for evaluating students' language proficiency. Previously, existing automated assessment methods typically focused on a single dimension, such as grammatical correctness, ignoring the more complex cognitive and contextual aspects of writing, and failing to comprehensively examine students' language proficiency. However, traditional English language proficiency assessment methods suffer from low accuracy in analyzing the logical and contextual ambiguity of the English language, resulting in large errors in English language proficiency assessments. Summary of the Invention
[0003] Based on this, it is necessary to provide an English language proficiency assessment method and system to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for evaluating English language proficiency is provided, comprising the following steps: Step S1: collecting English writing content texts uploaded by students through the terminal; analyzing the style categories of the English writing content texts uploaded by students, and then performing category sentence structure analysis to obtain the content style category sentence structure; Step S2: Identifying logical relationship deviations of English writing content texts based on content style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the degree of context ambiguity based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data; Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0005] Preferably, step S1 includes the following steps: Step S11: collecting English writing content texts uploaded by students through the terminal; Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text; Step S13: performing a stylistic analysis on the revised English writing content to obtain a content stylistic category; Step S14: performing a category sentence structure analysis on the revised English writing content text according to the content style category to obtain the content style category sentence structure.
[0006] Preferably, step S2 includes the following steps: Step S21: performing a content tense confusion structure analysis on the English writing content text based on the content style category sentence structure to obtain the content text tense confusion structure; Step S22: identifying logical relationship deviations of the English writing content text according to the tense disorder structure of the content text, thereby obtaining content logical relationship deviation data; Step S23: performing deviation clustering processing on the content logic relationship deviation data to obtain content logic relationship deviation clustering data; Step S24: quantifying the degree of context ambiguity of the English writing content text according to the content logical relationship deviation clustering data, thereby obtaining context ambiguity quantification data.
[0007] Preferably, step S22 includes the following steps: Step S221: extracting tense misuse / jump / missing structures in the tense confusion structure of the content text; Step S222: Detecting the English writing content text for paragraphs with invalid tense correspondence based on tense misuse / jump / missing structures, thereby obtaining paragraphs with invalid tense correspondence; Step S223: performing content time axis conflict analysis on the paragraphs with invalid tense correspondence, and obtaining content time axis conflict data between the paragraphs with invalid tense correspondence; Step S224: Identify the logic state fragmentation of the instant event based on the content time axis conflict data and the tense misuse / jump / missing structure between the tense echo failure paragraphs to obtain the logic state fragmentation data of the instant event; Step S225: performing logical relationship deviation identification on the English writing content text according to the instant event logical state segmentation data, thereby obtaining content logical relationship deviation data.
[0008] Preferably, step S224 includes the following steps: Perform content paragraph single-chain logical break analysis on the content timeline conflict data between paragraphs with failed temporal correspondence, and obtain a content paragraph single-chain logical break dataset; Based on the single-strand logical break dataset of content paragraphs and the temporal misuse / jump / missing structure, a single-strand logical break timeline directed graph is constructed. Perform break density calculation on the single-chain logical break time axis directed graph to obtain the logical break density; The logical disorder entropy value of the single-chain logical break time axis directed graph is calculated based on the logical break density to obtain the content logical disorder entropy value; Based on the logic fracture density and content logic disorder entropy, the logic state splitting of instant events is identified to obtain the logic state splitting data of instant events.
[0009] Preferably, step S24 includes the following steps: Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data; Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse; Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability; Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data; Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data; Step S246: quantifying the context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity quantification data.
[0010] Preferably, step S244 includes the following steps: Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content; Quantify the degree of semantic focus disorder in content subjective / objective context ambiguity data; Calculate the focus shift frequency variance in the semantic focus disorder degree; Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, content rhetoric dislocation analysis was conducted to obtain content semantic rhetoric dislocation data; The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus shift frequency variance to obtain the focus chain break ratio; Based on the variance of focus shift frequency and the ratio of focus chain breaks, a regression analysis of content context frame disorder was conducted to obtain context frame disorder data.
[0011] Preferably, step S3 includes the following steps: Step S31: performing logic learning on the context ambiguity degree quantified data to obtain context ambiguity degree learning data; Step S32: performing feature sampling on the context fuzziness degree learning data to obtain fuzziness degree learning feature sampling data; Step S33: Designing an English language proficiency assessment framework based on the fuzzy degree learning feature sampling data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; Step S34: Send the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0012] Preferably, the present invention further provides an English language proficiency assessment system for executing the above-mentioned English language proficiency assessment method, the English language proficiency assessment system comprising: The category sentence structure analysis module is used to collect English writing content texts uploaded by students through the terminal; analyze the style category of the English writing content texts uploaded by students, and then perform category sentence structure analysis to obtain the content style category sentence structure; The context ambiguity degree quantification module is used to identify logical relationship deviations in English writing content based on content style category sentence structure to obtain content logical relationship deviation clustering data; based on the content logical relationship deviation clustering data, the context ambiguity degree is quantified to obtain context ambiguity degree quantitative data; The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
[0013] The beneficial effect of the present invention is that by collecting English writing texts uploaded by students through a terminal and analyzing their writing style and sentence structure, a comprehensive understanding of students' writing characteristics and language usage habits can be obtained. The analysis of writing style helps identify students' expression methods in different writing scenarios, while the analysis of sentence structure can reveal the complexity and fluency of students' language expression. This analysis not only accurately divides the language structure of the writing content, but also lays the foundation for subsequent logical relationship identification and quantification of context ambiguity. Through the systematic analysis of student writing texts, data support can be provided for the assessment of English language proficiency. Based on the analysis of content style and sentence structure, logical relationship deviations are identified, and different logical deviations are identified through clustering algorithms, providing deeper data support for student writing analysis. By clustering these deviations, common logical structure problems in students' writing can be revealed, such as insufficient argumentation and unclear viewpoints, which often affect the overall logic and persuasiveness of the writing. At the same time, quantifying the degree of context ambiguity based on logical relationship deviation data can accurately assess the ambiguity and uncertainty of students' language expression in real contexts. The quantitative data of context ambiguity provides an objective basis for subsequent ability assessment, which can help the system more accurately assess students' language application ability in different contexts and make up for the shortcomings of traditional assessment methods in details. In the process of designing an English language ability assessment framework based on the Q-learning algorithm, the algorithm can dynamically adjust the assessment model according to the students' writing characteristics and the quantitative data of context ambiguity through continuous learning and optimization, thereby realizing personalized language ability assessment. The introduction of the Q-learning algorithm makes the assessment process not just a static scoring, but an intelligent and adaptive assessment system that can be optimized and adjusted at any time according to the progress of students' writing. This assessment framework can not only accurately identify students' problems in grammar, sentence structure, logical relationships, etc., but also provide specific feedback and improvement suggestions based on the characteristics of different students, providing students with more targeted learning guidance. Therefore, the present invention is an optimization process made to a traditional English language ability assessment method, which solves the problem that a traditional English language ability assessment method has low accuracy in analyzing the logic and context ambiguity of the English language, thereby causing large errors in English language ability assessment, improves the accuracy of analyzing the logic and context ambiguity of the English language, and reduces the errors in English language ability assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flowchart of the steps of an English language proficiency assessment method; Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG. Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG. DETAILED DESCRIPTION
[0015] See also Figures 1 to 3 , an English language proficiency assessment method, the method comprising the following steps: Step S1: collecting English writing content texts uploaded by students through the terminal; analyzing the style categories of the English writing content texts uploaded by students, and then performing category sentence structure analysis to obtain the content style category sentence structure; Step S2: Identifying logical relationship deviations of English writing content texts based on content style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the degree of context ambiguity based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data; Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0016] In the embodiment of the present invention, reference Figure 1 The above is a schematic flow chart of the steps of an English language proficiency assessment method of the present invention. In this example, the English language proficiency assessment method includes the following steps: Step S1: collecting English writing content texts uploaded by students through the terminal; analyzing the style categories of the English writing content texts uploaded by students, and then performing category sentence structure analysis to obtain the content style category sentence structure; In the embodiment of the present invention, the English writing text submitted by the student is received through a preset terminal interface, the terminal is connected to the Nginx reverse proxy server through the HTTPS protocol, and the proxy server forwards the data request to the Flask service interface module built in the back-end Python language. After receiving the English writing text, the service interface immediately calls the text preprocessing module for standardization. The preprocessing module uses regular expressions to clean control characters, non-language characters and redundant spaces, unifies the line break format and removes non-English language content, and then inputs the cleaned text into the line style category recognition module. This module is based on the Bayesian text classification algorithm, and the training corpus uses a public English writing genre corpus (such as COCA, BNC). It extracts 42 feature dimensions in five categories, including word frequency distribution, syntactic structure label distribution, transition word usage frequency, sentence length mean, and conjunction type frequency, as input vectors. After TF-IDF vectorization processing, MultinomialNaive is used. Bayes (Multinomial Naive Bayes) is used for genre prediction. The prediction results are divided into four categories: narrative, expository, argumentative, and descriptive. The prediction results and the text are sent to the category sentence structure analysis module. The sentence structure analysis module calls the depparse component of the dependency syntax analyzer Stanford CoreNLP to extract the subject-verb-object structure, modifiers, non-restrictive clauses, and adverbial position labels of each sentence. It analyzes the sentence usage frequency in each genre and counts the occurrence ratios of typical structures such as noun clauses, attributive clauses, passive voice, modal verb structures, and subjunctive mood. Finally, it generates the corresponding content genre category sentence structure vector for subsequent calls.
[0017] In another embodiment, the English writing content text uploaded by the student is obtained through the terminal. The terminal is a remote data input interface module of the teaching platform. The module has a structured data channel configuration function, which can parse the text content into paragraph structure and use the sentence vector encoding mechanism based on the BERT model to embed the uploaded text into a semantic vector. After the original corpus is segmented at the sentence level, the TextRazor API (i.e., a platform for natural language processing services) is first used to perform preliminary style recognition. By extracting indicators such as modifier density, syntactic structure ratio, and voice distribution, a style vector feature set is constructed. The Gunning Fog Index value is combined with the Flesch Reading The Ease score is used as an auxiliary parameter for genre identification to determine the genre category of the text. This process sets the classification label set into five categories: expository, argumentative, narrative, applied, and descriptive. For narrative and argumentative categories, the density of transition markers and the proportion of subjective emotional words are additionally extracted to enhance the accuracy of discrimination. After the genre classification is completed, the sentence structure of each sentence is analyzed based on the syntactic analysis tree constructed based on StanfordParser (an open source syntactic analysis tool). Sentence combination type labels (such as main-subordinate compound sentences, parallel sentences, imperative sentences, etc.) are extracted from the tree structure. Combined with the sentence-initial part-of-speech pattern and the sentence-final mood structure, the sentence structures of each category are statistically classified, ultimately forming a content genre category sentence structure containing feature vectors such as the proportion of sentence types in each paragraph, the frequency of conjunction types, and the distribution of tense patterns.
[0018] Step S2: Identifying logical relationship deviations of English writing content texts based on content style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the degree of context ambiguity based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data; In the embodiment of the present invention, according to the content style category sentence structure vector obtained in step S1, the logical relationship deviation identification operation of the English writing text is performed through the structural disorder detection module. The module has a built-in inter-sentence connection graph constructed by depth-first traversal. The nodes in the graph represent each sentence in the text. The edge weight is calculated by the semantic continuity of the inter-sentence conjunctions and the logical connection structure. If there is no clear causal or temporal connection clue between two adjacent sentences, the edge weight is set to 0. The graph further uses a minimum spanning tree-based method to identify isolated nodes and edge break areas. Isolated nodes are logically broken sentences. All broken sentences are clustered using the K-means clustering algorithm. The vector dimension includes sentence complexity. (number of subordinate clauses), information entropy value (frequency distribution of key words), cosine value of the angle between the subject vectors of the previous and next sentences, the default K value is set to 5, and it is adjusted according to the stability of the variance within the cluster. After the clustering is completed, the logical relationship deviation clustering data is obtained, and then the break density, sentence flatness rate, and logical connection vacancy rate in the cluster are input into the context fuzzy quantification module. The module uses a multidimensional normalization method to standardize the data to the [0,1] interval, and uses PCA principal component analysis to extract the top three main influencing factors, assigning weights of 0.4, 0.35, and 0.25 for weighted processing, and outputs the quantitative value as the quantitative data of the context fuzziness degree. The data format is a quantitative score and an index list of corresponding text segment numbers.
[0019] In another embodiment, the content style category sentence structure obtained in step S1 is used to implement a logical relationship deviation identification operation. This operation adopts a local adjacent paragraph comparison analysis method based on sentence connection logic. In the specific implementation, a paragraph-level semantic sequence alignment model is constructed, and a topic sentence semantic embedding vector is constructed for each two adjacent paragraphs. The word migration distance (Word Mover's Distance) is used to calculate the topic coherence between paragraphs. If the topic coherence of three consecutive paragraphs is lower than the set threshold of 0.38 and the sentence structure cross-paragraph consistency index is greater than 0.72, it is marked as a logical connection weakening area. Then, in these areas, the inter-sentence causal conjunction (such as because, therefore) omission detection algorithm is further used to identify inter-sentence logical breakpoints. If the logical conjunction is missing but there is a contradiction in semantic advancement or a sequence interruption, it is determined to be a logical relationship deviation point. Such deviation points are marked as logical break nodes in the sentence structure diagram. In the clustering process of all deviation nodes, a method based on DB is used. The SCAN density clustering algorithm uses the position of the deviation point in the text structure and the sentence category to which it belongs as the clustering feature dimension, and aggregates the areas with abnormal logical structure distribution in the text to form deviation clusters. Each cluster corresponds to the type and density interval of logical anomalies. Then, the contextual ambiguity of each logical relationship deviation cluster is quantified. First, the three indicators of the frequency of failed anaphora in each cluster, the frequency of broken causal chains, and the tense inconsistency rate are calculated, and weighted comprehensive scores are assigned with weights of 0.35, 0.4, and 0.25 respectively. The obtained ambiguity value is output in percentage. All ambiguity data constitute a contextual ambiguity quantification dataset.
[0020] Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0021] In the embodiment of the present invention, the Q-learning reinforcement learning algorithm is used to design an evaluation architecture for the above-mentioned context fuzziness degree quantitative data. In the specific operation, the state space is set to the distribution interval of the fuzziness degree quantitative score [0,100], which is divided into 20 state levels with a unit of 5. The action space is set to four categories of language ability grading behaviors (low level, medium level, good level, and excellent level). The reward function is defined as an inverse penalty function based on the matching accuracy between the real annotated English ability level label and the algorithm output. This function assigns a negative reward value to the result with a large prediction deviation. In the Q learning iteration process, the initial Q table is a zero matrix, and the ε-greedy strategy is adopted. For state behavior selection, the learning rate is set to 0.3, the discount factor is set to 0.9, and the number of iterations is set to 1000. In each round, samples of different ambiguity levels are grouped to train the learning path and update the Q-value table. Finally, a decision path mapping structure is constructed that automatically determines the language proficiency level after inputting quantitative data on the ambiguity level of the context. This is the English language proficiency assessment architecture. In this architecture, each level output node is bound to the feature combination pattern of the semantic ambiguity level and its corresponding probability threshold range. After the evaluation architecture is constructed, the architecture is written to the local memory interface registration module, which controls the return of the structured evaluation architecture data to the terminal execution module for the entire subsequent evaluation operation process.
[0022] Step S1 includes the following steps: Step S11: collecting English writing content texts uploaded by students through the terminal; Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text; Step S13: performing a stylistic analysis on the revised English writing content to obtain a content stylistic category; Step S14: performing a category sentence structure analysis on the revised English writing content text according to the content style category to obtain the content style category sentence structure.
[0023] In an embodiment of the present invention, the English writing content text submitted by students is received through the data upload interface module of the teaching platform. The interface module has an HTTP POST data access mechanism and implements asynchronous data reading based on a multi-channel cache strategy. After access, the data stream is first split into chunks using a paragraph segmentation logic based on sentence-end punctuation rules. Regular expression rules are used to match English end markers such as periods, question marks, and exclamation marks to perform preliminary sentence boundary recognition on the text. The paragraph position relationship of each sentence structure is retained in the cache through a numbered index. The input content is then converted into a standard byte stream using a UTF-8 encoding and decoding module to ensure consistency in subsequent character-level processing. The total length of the text content obtained in this stage is not less than 500 characters. A statistical strategy of dividing the average sentence length by 20 words is used to obtain a set of valid sentences of not less than 25 sentences, which serves as the basic corpus data set for subsequent processing. A spelling error correction operation is performed on the English writing content text obtained in step S11. This operation uses a spelling error detection method based on character-level edit distance calculation to compare each word in the text with a standard English dictionary library. The dictionary library contains Oxford Dictionary. The system uses a 3000 vocabulary list, the Coca corpus's top 10,000 words, and the Collins standard vocabulary set. It uses the Levenshtein distance algorithm to detect the minimum number of edit steps between the target word and the vocabulary entry. When the distance value is less than or equal to 2 and the target word appears less than or equal to 1 time in the context, it is marked as a low-confidence word. Such entries are required to participate in the context semantic similarity calculation with the three words that appear most frequently in the context window. The Jaccard similarity is used to compare the similarity of the context 3-gram phrases. When the similarity is greater than 0.68, the semantic nearest neighbor word is selected to replace the original word. At the same time, the original word and the replacement word are retained in the correction comparison table. All replacement actions are performed within a logical loop to complete the spelling correction of each sentence in turn. The total number of corrected words in the final output text is no less than 2% of the total number of words in the text, forming a complete English writing content correction text.
[0024] The revised text of the English writing content obtained in step S12 is subjected to stylistic category analysis, and a multi-dimensional text style vector encoding method is used for style modeling analysis. First, the main-slave compound structure of the text is annotated based on the sentence depth embedding index. The index obtains the complex sentence density score by calculating the difference between the average and maximum clause embedding depths. At the same time, the iconic conjunctions in the sentence, such as however, moreover, in contrast, and in The unit sentence frequency of "conclusion" and "conclusion" was calculated, and the weighted distribution of semantic polarity scores of subjective evaluation words at the end of sentences, such as "amazing," "terrible," and "impressive," was calculated. A feature vector group containing 12-dimensional style parameters, such as average sentence length, subject precedence rate, frequency of passive structure use, density of abstract nouns, and emotional color weight, was constructed. Then, a KNN-based unsupervised clustering method was used to determine the style classification. The K value was set to 5, and the style similarity between feature vectors was measured using Manhattan distance. The classification labels were divided into five standard styles: expository, argumentative, narrative, descriptive, and applied. Each style feature template was represented by an average vector. The argumentative type was required to meet the requirements of passive sentence frequency greater than 25% and connection logic frequency greater than 5 times per 100 words. Finally, each paragraph was assigned to the corresponding style based on the minimum distance between the style vector and the five types of templates, completing the style category analysis process. The style label and corresponding style vector weight data of each paragraph were output to form the content style category. Based on the content genre category obtained in step S13, the corresponding English writing content revision text is subjected to category sentence structure analysis. First, the Stanford syntax analyzer is used to perform dependency structure analysis on the text. The grammatical path information of the subject-verb-object structure is extracted from each sentence, and the adverb modification direction, the adjective clause nesting path, and the tense verb position vector are recorded. For narrative texts, the position of time adverbs and the distribution of verb tense changes are extracted. The tense distribution ratio curve is calculated and the presence of tense mutation points is determined. For expository texts, noun phrases are counted and the logical consistency of the subsequent modification structure is analyzed. For argumentative texts, the frequency of "if-then" and "although-yet" structures and the density of causal sentence combinations are counted. Combined with the frequency distribution of connectives, a sentence distribution matrix corresponding to each genre type is constructed. The sentence structure in each genre sample is encoded in vector form and merged to form a genre sentence set. This set is used for subsequent logical relationship identification and context consistency verification. The final output includes a sentence structure label set, a sentence category frequency vector, and a sentence distribution density matrix, forming a complete content genre category sentence structure dataset.
[0025] In this embodiment, reference Figure 2 The above is a flowchart of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Step S21: performing a content tense confusion structure analysis on the English writing content text based on the content style category sentence structure to obtain the content text tense confusion structure; Step S22: identifying logical relationship deviations of the English writing content text according to the tense disorder structure of the content text, thereby obtaining content logical relationship deviation data; Step S23: performing deviation clustering processing on the content logic relationship deviation data to obtain content logic relationship deviation clustering data; Step S24: quantifying the degree of context ambiguity of the English writing content text according to the content logical relationship deviation clustering data, thereby obtaining context ambiguity quantification data.
[0026] In an embodiment of the present invention, step S21 performs a tense confusion structure analysis operation on the English writing content text based on the content style category sentence structure obtained in step S14. In the specific implementation process, the corrected English text is first tense-tagged, and a lexical analysis algorithm based on tense recognition rules is used to identify the tense type of the predicate verb of each sentence. The recognition logic relies on rule tree matching technology to determine the tense type based on the combination relationship between the verb prototype and the auxiliary verb. The tense categories are set to eight categories: general present tense, general past tense, present perfect tense, past perfect tense, future tense, present continuous tense, past continuous tense, and future perfect tense. Each type of tense has a specific combination pattern in the sentence structure, among which The past tense is primarily composed of the past tense of the verb, while the perfect tense is composed of auxiliary verbs such as have or has followed by a past participle. All sentence stem verbs are annotated with tenses based on a rule set. Paragraph-level tense jumps are identified by combining sentence position indexes with contextual sentence consistency. When the frequency of tense switching between three adjacent sentences is greater than 2 and is not driven by a causal temporal structure, it is marked as a non-logical tense jump region. If tenses are mixed and alternated under the same subject within a paragraph, it is identified as a tense misuse structure. This analysis uses a linear temporal index vector to record the tense switching paths between sentences. Ultimately, a collection of all tense misuse, jump, and missing structures in the content text is integrated to form a tense confusion structure. According to the temporal chaos structure of the content text obtained in step S21, the logical relationship deviation recognition operation is performed on the English writing content text. The specific implementation process is to build a semantic advancement relationship graph based on the recognized chaos structure. The graph uses sentence sequences as nodes and event time sequence as directed edges. The event state words at the beginning and end of the sentence (such as begin, end, continue, finish) are used to build time sequence anchor points, and all inconsistent timeline paths are marked as time advancement breakpoints. Then, the logical consistency of the connectives between the sentences corresponding to these breakpoints is further verified. The connective dependency matching method is used to retrieve the sentence head connection component and compare it with the previous and next sentences. The semantic relationship is mapped and checked. When the conjunction does not conform to the semantic logic of the sentence (for example, "although" connects two essentially identical sentences), it is determined to be a logical relationship deviation point. In the analysis process, the nesting depth of inter-sentence propositions is introduced as the basis for redundant logical judgment. If the logical advancement in a sentence is unclear due to excessive nesting of conditional sentences and concession sentences, it is regarded as a logical deviation caused by structural complexity. In the logical deviation identification operation, the three-valued logic backoff method is used to disassemble the nested logical relationship to form a logical relationship chain, and the broken parts are marked for deviation. Finally, the content logical relationship deviation data consisting of all inter-sentence relationships with structural logical errors is output.
[0027] Deviation clustering is performed on the content logic relationship deviation data obtained in step S22. First, a three-dimensional feature space is constructed for all deviation points based on their position in the text, deviation type and temporal association. Each deviation point is represented as a vector. The first dimension of the vector is the normalized value of the sentence sequence number, the second dimension is the deviation type number (such as missing conjunctions are marked as 1, tense jumps are marked as 2, and structural nested redundancy is marked as 3), and the third dimension is the context consistency measure of the corresponding temporal chaos structure. The consistency measure is calculated based on the cosine similarity between the sentence vector and the adjacent sentence vectors. The density-based clustering algorithm DBSCAN is used to perform density clustering on the above three-dimensional deviation vector set. The minimum sample number is set to 4, and the ε radius distance is set to 0.3. The final clustering output is several logical deviation clusters. Each cluster contains a group of deviation points that have an association relationship in the semantic structure, syntactic structure and time advancement path. Through clustering, high-density areas of logical errors can be identified at the paragraph level to form content logic relationship deviation clustering data. The context ambiguity degree is quantified based on the content logic relationship deviation clustering data obtained in step S23. First, the subject-predicate structure consistency analysis is performed on the sentences in each logic relationship deviation cluster, and the expressions with ambiguity or conflict in the semantic direction between the subject, predicate and object are extracted, and the frequency of each conflict type is recorded. Secondly, the ambiguity strength of the connectives in each cluster is scored, and the frequency of known fuzzy connectives (such as since, while, as) in the Brown corpus is used as the semantic fuzziness index benchmark. The ambiguity strength is calculated by the word meaning fuzzy context overlap rate indicator. If the overlap rate exceeds 0.6, it is considered a highly ambiguous connective. The pronoun reference failure rate in the paragraph is further counted. If it appears in two or more consecutive sentences, the ambiguity strength is calculated. If a pronoun has no clear referent, it is counted as the frequency of contextual reference confusion. At the same time, the word meaning jump rate in the cluster is analyzed, and the contextual semantic category of each noun is counted. If the contextual semantic label of the noun switches more than two category levels within two sentences (for example, jumping from the event class to the entity class), it is determined to be a semantic jump. Finally, a weighted comprehensive score is given to various fuzzy features in the cluster, with the subject-predicate structure fuzziness weight being 0.3, the conjunction ambiguity weight being 0.25, the reference failure weight being 0.2, and the word meaning jump weight being 0.25. The contextual fuzziness score of the cluster is obtained by weighted averaging, and the scores of all clusters are then aggregated to form a quantitative data set of contextual fuzziness, which is used in the subsequent design of the English language proficiency assessment architecture.
[0028] Step S22 includes the following steps: Step S221: extracting tense misuse / jump / missing structures in the tense confusion structure of the content text; Step S222: Detecting the English writing content text for paragraphs with invalid tense correspondence based on tense misuse / jump / missing structures, thereby obtaining paragraphs with invalid tense correspondence; Step S223: performing content time axis conflict analysis on the paragraphs with invalid tense correspondence, and obtaining content time axis conflict data between the paragraphs with invalid tense correspondence; Step S224: Identify the logic state fragmentation of the instant event based on the content time axis conflict data and the tense misuse / jump / missing structure between the tense echo failure paragraphs to obtain the logic state fragmentation data of the instant event; Step S225: performing logical relationship deviation identification on the English writing content text according to the instant event logical state segmentation data, thereby obtaining content logical relationship deviation data.
[0029] In the embodiment of the present invention, based on the obtained tense chaotic structure data of the content text, the tense misuse, tense jump and tense missing structure therein are explicitly extracted. The operation process is to first construct a tense identification sequence based on part-of-speech tagging for all sentences, and adopt the rule of auxiliary verb + predicate verb combination structure to identify English tense. The auxiliary verb set includes is, are, was, were, have, has, had, will, would, etc. The predicate verb is identified based on word form transformation to determine the past tense and past participle. Sentences without any auxiliary verb identification and the predicate verb is in the original form are marked as tense missing. Sentences where the auxiliary verb and predicate do not match the English standard tense composition rules are marked as tense misuse. For example, a sentence contains has went or will All "eating" structures are recorded as misused structures. In cross-sentence comparison, when the switch from future tense to past tense or from perfect tense to present tense occurs twice or more in three consecutive sentences under the same subject, and there is no supporter of the logical explanation of time adverbial, it is marked as tense jump. This extraction operation is completed using a joint judgment mechanism based on the syntactic tree structure and the nesting of tense rules. Each error type is located as a different structural conflict node in the sentence structure tree. All structural annotations are recorded in the index table for subsequent logical analysis. Step S222 detects paragraphs with failed tense correspondence in the English writing content based on the extracted tense misuse, jump, and missing structures. The detection method is to calculate the tense consistency coefficient between sentences in a paragraph on a paragraph basis. This coefficient is based on the pairing of tense label sequences between sentences and is defined as the ratio of the number of sentences with the same tense in a paragraph to the total number of sentences in the paragraph. If this value is lower than 0.35 and there are two or more marked tense-incorrect sentences in the paragraph, the paragraph is marked as a paragraph with failed tense correspondence. Further statistics are collected to see whether the chronological sequence of actions or events in the paragraph contains a mixed path switching from past to present and then back to past. For example, if in an event narrative, the sentence switches from "he had eaten" to "he eats" and then back to "he went", it is determined to be a typical tense correspondence failure path. All such paragraphs are uniformly filed in a collection of failed tense correspondence paragraphs for subsequent content timeline conflict analysis.
[0030] A content timeline conflict analysis is performed on the paragraph with tense-response failure obtained in step S222. First, an event timestamp is established for each event sentence. The main action verbs in each sentence are identified using an event dictionary. The event dictionary includes common action verbs such as arrive, leave, eat, decide, start, and finish. The relative temporal position of the event is determined by combining temporal adverbials in the sentence such as yesterday, last week, soon, and already. An event directed graph is constructed and a path consistency check is performed on it. When a path in which the same subject entity jumps forward and then returns backward in the time path occurs, that is, a path loop structure exists in the event graph, it is marked as a timeline conflict. This conflict is detected by detecting whether there is a loop with a length greater than 2 in the event graph. The nodes in the loop represent event actions, and the edges represent the direction of time advancement. For example, start→eat→arrive→start constitutes a conflict loop. The system marks the paragraph as a paragraph with a timeline conflict. All conflict path data and their structural features are encoded and recorded to form content timeline conflict data. Based on the above-obtained content timeline conflict data and the original tense misuse, jump, and missing structures, this paper identifies logical state fragmentation of immediate events in English writing content. The specific method is to identify whether there is semantic discontinuity in the event state change path at the same time point or within a very short period of time. The detection step first extracts the state verbs used in all events, including state-indicating words such as become, remain, stay, and feel, and analyzes their logical continuity paths. If there is an event state "he was staying home" followed by the next sentence "he had gone outside", it indicates that the state transition is inconsistent and there is no intermediary state connection. The state jumps directly from the static state to the completed movement state, which is a fragmented path. In addition, during the analysis process, a state transition diagram is constructed for all events showing state fragmentation. State fragmentation behavior is identified by determining whether there are unconnected isolated state nodes or jump-connected cross-state transition paths in the diagram. The occurrence time, state type, and sentence index of all identified state fragmentation segments are extracted and merged to form the immediate event logical state fragmentation data.According to the instant event logical state segmentation data obtained in step S224, the logical relationship deviation of the English writing content text is identified. The identification operation is to use the sentence pairs in the segmentation data as input, and perform the semantic adaptability judgment of the connectives and the proposition logic matching test in the inter-sentence logical chain. First, the connective semantic classification dictionary is used to divide the connectives into four categories: causal, conditional, transitional and parallel. The connective semantics are reversely verified for the sentence pairs before and after the state segmentation. For example, if the state of the previous sentence is "he was sick" and the state of the subsequent sentence is "he went to work", the connective is "although", then it is considered to be structurally reasonable. If the connective is "because", then it is considered to be semantically mismatched and marked as a logical relationship deviation. Secondly, the subject consistency and the main clause and subordinate clause logical strength relationship in the segmentation segment are verified by proposition logic. When the main sentence is "he wanted to rest" and the subordinate clause is "heran ten miles", the logic is not sound, and the logical relationship strength score is less than 0.2 (based on the normalized score of the inner product between sentence vectors). Such sentence groups are classified as logical deviation units. The system associates all such logically abnormal sentence groups with their corresponding deviation reasons (inappropriate conjunctions, proposition contradictions, and tense conflicts) to form content logical relationship deviation data for subsequent cluster analysis and context fuzzy quantification.
[0031] Step S224 includes the following steps: Perform content paragraph single-chain logical break analysis on the content timeline conflict data between paragraphs with failed temporal correspondence, and obtain a content paragraph single-chain logical break dataset; Based on the single-strand logical break dataset of content paragraphs and the temporal misuse / jump / missing structure, a single-strand logical break timeline directed graph is constructed. Perform break density calculation on the single-chain logical break time axis directed graph to obtain the logical break density; The logical disorder entropy value of the single-chain logical break time axis directed graph is calculated based on the logical break density to obtain the content logical disorder entropy value; Based on the logic fracture density and content logic disorder entropy, the logic state splitting of instant events is identified to obtain the logic state splitting data of instant events.
[0032] In an embodiment of the present invention, based on the content timeline conflict data between the tense-echo failure paragraphs obtained in the preceding step S223, a content paragraph single-chain logical break analysis is performed on the time advancement path between the paragraphs. This step first establishes an event sequence set for each paragraph marked as having a timeline conflict, extracts the action verbs in each event sentence in the paragraph as event nodes, and establishes an initial linear connection path between the event nodes according to the sentence sequence number, that is, constructs an initial single-chain event sequence, wherein each event node contains four types of parameters, namely, the event verb form, the subject word, the intra-sentence time adverbial word and its sentence sequence index number. Subsequently, a logical connection check is completed by determining the temporal logical consistency relationship between adjacent event nodes. The logical consistency check relies on the event sequence semantic template for determination. The template construction rule is to identify the state starting point and end point of each event based on the verb semantic role labeling technology. If the end state of the previous event is semantically discontinuous with the start state of the next event or there is a logical conflict, it is marked as a logical breakpoint. For example, the previous event is "he stayed at home" and the next event is "he had arrived at School" has no intermediary process events and the time adverbials conflict, which constitutes a single-chain logical break. All breakpoints are uniformly archived and recorded according to the position number of the original event sequence. Finally, a single-chain logical structure diagram consisting of alternating logically continuous segments and broken segments is constructed for each paragraph. The broken paragraph index set, break type identifier, and semantic difference between events before and after the break are extracted to form the content paragraph single-chain logical break dataset.
[0033] On the basis of the single-chain logical break dataset of the content paragraphs obtained in the above steps, the tense misuse, jump and missing structures extracted previously are combined to construct a single-chain logical break timeline directed graph. During the construction process, an event point set is first constructed for all event sentences in the paragraph. The event point takes the main predicate verb in the sentence as the central node, and adds meta-attributes such as time adverbial label, sentence sequence number and paragraph number. All event points are connected in sequence within the paragraph to form an initial linear path graph. Then, the direction edge is constructed by analyzing the time advancement direction between events. The direction of the direction edge is determined based on the temporal sequence relationship between the previous and next events. If the time adverbial in the sentence is past, next morning, then, etc. with advancement logic, the edge direction is from front to back. If there is a semantic backtracking structure such as "earlier that day" or "before", the direction edge is from front to back. That" is determined to be a reverse edge. Subsequently, all event nodes with misused, skipped, or missing temporal markers are embedded in the graph and treated as candidate break nodes. If the edge direction between such a node and its adjacent nodes is inconsistent or the edge weight is negative (calculated based on the temporal difference of the sentence vector), the edge is set as a break edge. Finally, the above nodes and edges form a complete directed graph structure, which has the event sequence path, the break edge path, and all temporal abnormality marker point information. A separate timeline directed graph structure is established at each paragraph level, and each graph number corresponds one-to-one with the paragraph number. After completing the construction of the above directed graph, the fracture density calculation is performed on the single-chain logical fracture timeline directed graph corresponding to each paragraph. The calculation operation first counts the total number of fracture edges B and the total number of event nodes N in the graph, and calculates the preliminary fracture rate parameter as B / N. Then, the fracture distribution concentration correction factor is introduced to perform weighted adjustment on whether the fracture edges are concentrated in a specific graph segment. The correction factor calculation method is to convolve the density function of each fracture edge distribution on the graph node path. If the concentration of the fracture edge on the continuous path exceeds a certain threshold (set to 0.65), the fracture density weighting coefficient is multiplied by the inverse concentration function. The value is scaled and further normalized based on the ratio of the maximum continuous normal path length to the total path length in the graph. If the maximum normal path segment is less than 40% of the total length, the abnormal weighted factor penalty mechanism is triggered, and the final fracture density is multiplied by a coefficient of 1.3 for correction. The final fracture density value is expressed as a comprehensive expression of the proportion of fracture structures and the degree of continuity loss in the graph. The higher the fracture density, the more serious the interruption of the paragraph event flow in the process of logical advancement. The value is finally output together with the graph structure number to form a logical fracture density sequence table for subsequent context entropy value calculation and immediate logical state split recognition operation.
[0034] The fracture density calculation operation is performed on the basis of the established single-chain logical fracture timeline directed graph. The operation takes the directed edge set marked as fracture edge in the graph structure as the core input. First, the total number of nodes N and the number of fracture edges B in the graph are counted. Then, the node span evaluation operation is performed on each fracture edge, that is, the distance D between the starting node and the ending node connected by the fracture edge on the path in the graph is calculated, and the D values of all fracture edges are summed to form the total fracture span value S. Then, for each fracture node, the local fracture edge density within the third-order adjacent range around it is calculated, which is recorded as the local fracture density vector Ld=[b1,b2,...,bn], where b i represents the ratio of the number of third-order internal fracture edges around the i-th fracture node to the total number of edges. The Pearson correlation between this vector and the edge density distribution vector of the entire graph is then calculated. If the correlation coefficient is lower than 0.35, it indicates that the fracture distribution is uneven, and a penalty factor coefficient α=1.2 needs to be introduced to perform weighted correction on the overall fracture density. The final fracture density value is calculated as B / N×α×log(S+1). The fracture density results corresponding to all segments are organized into a vector M=[m1,m2,...,mn] for subsequent logical disorder entropy value calculation to ensure that each fracture density value includes the influence of path length, fracture aggregation degree, and structural distribution consistency. The obtained logical fracture density vector M performs logical disorder entropy calculation on the single-chain logical fracture timeline directed graph. The calculation is based on the event path sequence in the graph. First, each event path is sequenced and the node events in all legal event paths are numbered in order according to the order of the time adverbials in the sentence to form a time sequence T=[t1,t2,...,tk], where tk represents the position number of the kth event on the time axis relative to the entire article. At the same time, all the broken edge node sets are extracted from the directed graph, and their relative positions in the sequence T are marked as the break point set D=[d1,d2,...,dn]. Then, the entropy value is calculated based on the distribution discreteness of the break point set in the sequence T. Operation, using the Shannon entropy calculation formula to take the probability of each break point in the sequence as input to calculate the sequence entropy value H of the overall logical path. The H value represents the degree of uncertainty of the break in the event advancement path. The H value is further adjusted together with the break density value mi to form the entropy weight adjustment factor θ=H×mi. For each paragraph directed graph, its θ value is calculated separately and summarized to form a logical disorder entropy value sequence θ=[θ1,θ2,...,θn]. If the θ value of a segment exceeds the mean of the sequence plus 1.5 times the standard deviation, it is marked as a high logical disorder segment. Finally, each paragraph graph corresponds to an entropy value record. Its logic is that the break distribution entropy and the structural break density jointly determine the disorder degree of the logical advancement of events in the segment.Based on the above-obtained logical fracture density sequence M and logical disorder entropy value sequence θ, an immediate event logical state split recognition operation is performed. During the recognition process, the single-chain logical fracture timeline directed graph structure corresponding to all paragraphs with logical disorder entropy values higher than the threshold is first extracted. Then, all events containing "state change" verbs in all nodes in the graph are marked and screened. The state change verb set includes words such as become, remain, stay, appear, feel, turn, get, etc. This type of word represents an explicit change in the state of an entity or situation. In the graph structure, this type of node is marked as a state node. Subsequently, a continuity analysis operation is performed on the path between any two adjacent state nodes. If there is a path with less than 3 nodes and no process action verb nodes in the middle (such as go, m ove, do, arrive, etc.), it is judged as a direct state jump path. Combined with the number of broken edges contained in the path, if the path contains at least one broken edge and the overall path entropy contribution value is greater than the average entropy value of the paragraph, it is marked as a state split path. The state split path is judged by the difference between the logical semantic vectors of the event description. If the inner product of the previous and next event description vectors is less than 0.15, the split judgment strength is further strengthened. Finally, all state jump segments in all paths that meet the above three conditions are recorded as immediate event logic state split data. Each split path record contains parameters such as the previous and next state event subjects, verbs, sentence order numbers, connecting logical words, path length, fracture density weights, and local entropy contribution coefficients, which are used for subsequent logical relationship deviation analysis and context fuzziness modeling calls.
[0035] Step S24 includes the following steps: Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data; Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse; Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability; Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data; Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data; Step S246: quantifying the context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity quantification data.
[0036] In an embodiment of the present invention, a causal fault analysis is performed on the English writing content text based on the content logic relationship deviation clustering data obtained in the preceding step S23. In the operation, a set of deviation fragments marked as "logical jump class" is first extracted from the clustering data, and the sentence index range in the corresponding paragraph of each deviation fragment is located. Then, a sentence causal relationship mapping structure diagram is constructed in the paragraph. The structure diagram uses each sentence as a node, and the edges between sentences indicate the existence of clear causal logic. The identification of causal edges is based on the joint judgment of conjunction rule matching and sentence meaning relationship analysis. The conjunction rule set includes because, so, therefore, as a result, since, due To, consequently, hence, and other conjunctions or connecting structures that explicitly indicate causal relationships. The sentence meaning relationship analysis adopts the subject-predicate structure analysis method within the sentence to determine whether the action of the latter sentence is directly triggered by the subject action of the previous sentence. The judgment is completed by determining whether the result of the semantic verb of the subject action and the prerequisite of the latter action form a causal connection boundary in the concept network. If a sentence has no obvious causal conjunction with its previous sentence, and the subject behavior logic between the sentence meanings does not have a triggering relationship, it is marked as a causal fault point. This point records its paragraph number, sentence number, conjunction missing status, subject-predicate structure vocabulary, and logical jump mark, which together constitute the content causal relationship fault data. This step requires relying on the syntactic dependency tree structure parsing results and the verb trigger-type semantic comparison function to compare the relationship between sentences one by one. The semantic comparison function used is constructed based on the trigger-result pair rule defined in the verb semantic role labeling tool.After obtaining the content causal relationship fault data, the text is analyzed for reference logic disorder and the pronoun ambiguity frequency is calculated. The specific operation is to first extract all pronoun examples from the sentence segments located by the fault data. The target pronoun range includes subjective and objective personal pronouns (he, she, they, it, him, her, them), and demonstrative pronouns (this, that, these, those). When each pronoun is extracted, its sentence number, position within the sentence, grammatical role and context coreference target candidate set are recorded. Subsequently, based on the coreference resolution algorithm, each pronoun is tried to match the noun phrase entity in the previous text. The entity matching process prioritizes the shortest distance. Semantic consistency and grammatical role matching are prioritized as hierarchical matching strategies. The candidate coreference entities for each pronoun match are sorted by confidence score. If there is a preferred coreference item with a confidence score lower than 0.5, and the pronoun is at the beginning of a sentence or the first sentence of a paragraph in the text, it is marked as an "ambiguous reference point". Each ambiguous reference point records its corresponding pronoun, sentence position number, candidate confidence sequence, grammatical role, whether there are similar entities in the two sentences before and after the text, and other information. Then, for each paragraph of text, the ratio of the number of ambiguous reference points to the total number of sentences is counted and defined as the frequency of pronoun misuse. This frequency is included in the text context ambiguity index system, and the final output is a pronoun misuse frequency mapping table corresponding to the paragraph number.
[0037] Based on the aforementioned causal fault data, a word meaning confusion probability estimation and analysis operation is carried out. This step is based on the keyword phrases within the sentences in the fault paragraphs, and semantic polysemous words in noun and verb phrases are extracted to calculate the ambiguity probability. The specific process is to first use the part-of-speech tagging system to tag each sentence, and select all the core verbs and the noun components of the sentence subject and object to form a context phrase pair. Then, the central word in the phrase pair is called by the semantic dictionary to count its interpretation number. If the interpretation number is greater than 3, it is judged to be a highly polysemous term. Subsequently, the context sliding window method is used to construct the context sentence vectors before and after each term. The vector dimension adopts a 300-dimensional word embedding model, and the extraction window size is set to two sentences before and after. On this basis, the context vector and the word meaning vector center are calculated. The cosine similarity between them is used to score the degree of adaptation of each interpretation, and the difference between the maximum similarity gap and the second largest value is selected as the ambiguity identification distance indicator. If the distance is less than 0.15, it is considered that the word has a high confusion probability in the context. The confusion probability is normalized by the standard deviation of the cosine similarity distribution between all interpretations, and the normalized value is multiplied by the word sense number weight factor to obtain the final word sense confusion probability value. All high confusion probability terms are recorded and a word sense confusion information table consisting of fields such as term position index, grammatical role, number of interpretations, contextual semantic offset value, and final confusion probability is generated. Finally, the average confusion probability value of each paragraph is counted to form a content word sense confusion probability vector, which is used for subsequent semantic focus disorder and context ambiguity modeling operations.
[0038] The specific processing method for content context framework disorder regression analysis based on content word meaning confusion probability, pronoun misuse frequency and content causal relationship fault data is as follows: first, a sentence order and semantic mapping matrix is constructed for each English writing content text. The matrix is numbered with sentence order as the horizontal axis and the vertical axis is the semantic centroid vector of the corresponding sentence. The semantic centroid vector is extracted using the context aggregation word embedding mean processing method, and the dimension is set to 300 dimensions. Each sentence vector is represented by the mean of the content word vector after removing stop words. Then, each sentence is arranged in the order of the original text to form a sentence semantic flow vector group S=[v1,v2,...,vn]. Then, the content causal relationship fault data obtained in the previous step are combined to locate all sentence pair number intervals with logical jump points, and regression slope mutation detection is performed within the interval. A five-point sliding window is used for local linear Fitting, calculate the main direction of the mean of the semantic centroid vector of each five sentences, and judge whether the semantic flow direction has a sudden change by comparing the angle of the fitting slope vectors of adjacent windows. If the angle exceeds 45 degrees, it is marked as a local frame disorder node. Then, the mixed disturbance coefficient β is established by combining the word meaning confusion probability and the pronoun misuse frequency in each sentence segment. This coefficient is defined as the weighted average of the mean word meaning confusion probability and the pronoun misuse frequency of the segment, with weights set to 0.6 and 0.4. The disturbance coefficient β is calculated at all marked local frame disorder nodes and compared with the average semantic consistency value of the adjacent area. If the ratio exceeds 1.5 times the standard deviation, it is finally marked as a context frame disorder point. The paragraph number, sentence order position, semantic slope jump angle, disturbance coefficient, and semantic flow residual value of the point are recorded, and finally a context frame disorder data table is formed for the next stage to call. After obtaining the disordered data of the context framework, the core subject concept offset perception operation is performed. This step uses the standard subject concept vocabulary under the style category as a reference to identify the full text subject concept. The specific process is to first extract the corresponding subject concept word set according to the content style category. The set is formed by manual expert annotation. Each concept word contains the corresponding word meaning and related extended word family. Then, the subject word scan is performed in each text. Each sentence is matched with the word in the set and its occurrence frequency and semantic similarity value are recorded. The semantic similarity is calculated using the cosine similarity method. The word embedding mean vector in the sentence is calculated with the concept word vector. If the similarity is greater than 0, the semantic similarity is calculated. .65 is recorded as a main theme coverage. The main theme concept density of the entire text is defined as the ratio of all main theme coverage times to the number of sentences. Then, a local main theme sparse analysis is performed on the main theme coverage density of all paragraphs where the frame disorder points are located. If the density is lower than 0.6 times the average density of the entire text, and the change in the main theme term frequency between the two paragraphs exceeds 50%, it is marked as a main theme concept offset point. The main theme coverage missing value of this point and the main theme term distribution offset in the adjacent area together constitute the main theme concept offset degree. The final output main theme concept offset degree data includes fields such as the offset segment number, the main theme density within the segment, the main theme change value of the adjacent paragraphs, and the number of main theme coverage interruption sentences.After obtaining the context frame disorder data and the main concept deviation degree data, the context fuzziness degree quantification operation is performed. This step establishes a multi-factor fuzziness aggregation function based on the sentence segment. The function input is three sets of core parameters, namely the context frame disorder index F, the main concept deviation intensity value P and the semantic coherence attenuation coefficient D. The semantic coherence attenuation coefficient is obtained by normalizing the semantic cosine similarity reduction between the two sentences before and after the context frame disorder point. The specific calculation method is to extract the semantic embedding vector means of the two sentences before and after, which are A and B respectively, and calculate the cosine similarity S=cos (A, B), and then normalize the average semantic similarity of the entire text to form a difference index D. The final fuzziness level M = F × 0.5 + P × 0.3 + D × 0.2, which is defined as the context fuzziness level of the current paragraph. This value is calculated for all paragraphs separately and summarized to form a fuzziness distribution vector. The fuzziness level quantification data structure includes five fields: paragraph number, context framework disorder, main concept offset value, semantic coherence attenuation value, and fuzziness level aggregation value. A complete fuzziness level quantification map is constructed for each text as input for subsequent language proficiency assessment architecture learning module calls.
[0039] Step S244 includes the following steps: Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content; Quantify the degree of semantic focus disorder in content subjective / objective context ambiguity data; Calculate the focus shift frequency variance in the semantic focus disorder degree; Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, content rhetoric dislocation analysis was conducted to obtain content semantic rhetoric dislocation data; The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus shift frequency variance to obtain the focus chain break ratio; Based on the variance of focus shift frequency and the ratio of focus chain breaks, a regression analysis of content context frame disorder was conducted to obtain context frame disorder data.
[0040] In the embodiment of the present invention, a fuzzy analysis of the subjective context and objective context of the content is performed based on the probability of confusion of content word meanings, the frequency of misuse of pronouns and the content causal relationship fault data. The specific operation is to first extract the paragraph area with logical jumps or causal logic chain breaks based on the sentence number marked in the content causal relationship fault data, and calculate the context fuzziness score by combining the word meaning confusion probability value and the pronoun misuse frequency of each sentence in the area. The score calculation process first extracts the coverage frequency of the first person and emotional words as subjective marking indicators according to the sentence unit, and uses the emotional vocabulary composed of adjectives and adverbs to match their distribution positions. The subjective context fuzziness is defined as the product of the mean probability of word meaning confusion and the proportion of the first person and emotional words. Weighted aggregation is performed, where the weight of the probability of word confusion is 0.5, the weight of the first-person frequency ratio is 0.3, and the weight of the emotional word coverage is 0.2. Similarly, after excluding sentences with subjective expressions in the objectivity analysis, the logical chain density of technical verbs and factual description phrases is statistically analyzed. The objective context fuzziness is defined as the inverse of the causal chain node density and the inter-sentence logical coherence, where the causal chain node density is obtained by the ratio of the number of logical chains in each paragraph to the total number of sentences, and the inter-sentence logical coherence is expressed by the mean cosine similarity of adjacent sentence vectors. Finally, the subjective and objective context fuzzy data of all paragraphs retain the subjective fuzzy score and the objective fuzzy score respectively and associate the paragraph position with the vocabulary statistics to form the subjective and objective context fuzzy dataset structure. In order to quantify the degree of semantic focus disorder in the subjective and objective context fuzzy data mentioned above, the operation process is to take the paragraph as the minimum analysis unit, establish a sentence focus word identification mechanism, perform subject-verb-object structure analysis on each sentence, extract the subject and object as candidate focus entities, and then screen them through TF-IDF value and word vector center clustering, retain the semantic core words in the sentence, and construct a focus word sequence stream F=[f1,f2,...,fn] for each paragraph. This sequence is used to analyze the flow path of semantic focus in the paragraph. On this basis, statistics are counted to see whether the focus jumps in three consecutive sentences, that is, the subject or object in the sentence no longer reappears in the later sentences. If the sentence changes continuously for more than two times, it is marked as a focus breakpoint, and the focus retention rate in the paragraph is calculated. This value is the ratio of the number of repeated focus words to the total number of focus words. The lower the focus retention rate, the higher the degree of focus disorder. The focus disorder index is generated by combining the subjective and objective context fuzzy scores. The index is defined as the aggregate value of subjective fuzzy score × 0.4 + objective fuzzy score × 0.3 + focus retention rate × 0.3. Finally, the focus disorder index is extracted from all paragraphs to form a semantic focus disorder degree dataset. The data fields include segment number, focus word change path, break position distribution, focus retention rate and focus disorder index.In order to calculate the focus shift frequency variance in the degree of semantic focus disorder, the specific operation is to number and encode the focus word transfer trajectory of each paragraph in each text, encode each focus change event as a jump, and construct a complete focus shift sequence vector G=[g1,g2,...,gm], where gi is the relative position difference of the i-th focus change, and the unit is set as the sentence sequence distance. The variance statistical analysis of the sequence vector is performed to calculate the focus shift frequency variance. If the focus jump frequency in an article is highly concentrated, the variance is small. If the jump frequency is high in the paragraph, the variance is small. Uneven distribution of focus or concentration in local areas will show high variance values. To prevent deviations in short texts, focus sequence normalization is performed on texts with a length of less than 200 words, that is, the focus jump interval is divided by the total number of sentences to maintain scale consistency. Finally, the focus shift frequency variance and the focus disorder index are jointly output to form a semantic focus shift stability dataset, in which fields include segment number, focus sequence, focus change distance mean, focus change variance, standardized variance value, etc., which are used for further aggregation into the context frame disorder indicator in subsequent regression analysis. Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, the content rhetorical dislocation analysis is carried out. The specific operation is to use the degree of semantic focus disorder as a measure of the stability of focus transfer within a paragraph, cross-match the causal chain missing area marked in the causal relationship fault data with the focus disorder area and construct a dislocation segment mapping matrix, number each paragraph of text according to sentence order to form a sentence vector sequence, and quantitatively represent the semantic distance between the subject and the object in each pair of adjacent sentences. The cosine similarity of the word vector is used to calculate the subject semantic continuity and object pointing consistency. If the angle between the subject vectors of adjacent sentences is greater than 1.2rad and the object semantic deviation is greater than 0. The verb logical framework of the previous sentence is marked as a rhetorical jump point, and the word meaning confusion probability threshold of the paragraph where the jump point is located is further set to 0.4. Words exceeding this threshold are identified as semantic rhetorical breakpoints when they appear in the subject or object position. The ratio of the number of breakpoints to the total number of sentences in each paragraph is recorded as the rhetorical dislocation rate of the paragraph. At the same time, the focus words before and after the jump point are marked, and a rhetorical focus migration table is constructed to record the semantic path mutation. Finally, the rhetorical dislocation rate, jump point distribution, focus deviation angle, and word meaning confusion weight of each paragraph are extracted to form the content semantic rhetorical dislocation data. The data fields include paragraph number, breakpoint position, break type, focus semantic change path and rhetorical offset weight sequence.
[0041] The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus transfer frequency variance. The specific operation is to first construct a directed linked list structure based on the inter-sentence recurrence of the focus word sequence F=[f1,f2,…,fn] in each text. If a focus word does not appear three times in a row in the linked list, it is marked as a broken link point. The ratio of the total number of break points in all focus chains to the total length of the linked list is recorded as the initial value of the focus chain break. At the same time, the threshold T is set with reference to the focus transfer frequency variance. In the experiment, the sample focus variance with an average of 1.5 is taken. If the current If the text focus variance is greater than T, it is considered that the break is not a random transfer but a semantic chain disconnection. The break count weight is recalibrated and the broken chain points are screened again. The broken chain points of all paragraphs are accumulated and then the ratio with the total number of chains is used to form the final focus chain break ratio. The focus chain break ratio calculation formula includes the rhetorical jump density factor R, which is defined as the product of the break point frequency and the rhetorical dislocation rate in a unit paragraph. The break ratio field output is the segment number, broken chain point position, average chain length, variance level and ratio result, which constitute the focus chain break ratio retrieval result set. In order to conduct a regression analysis of content context frame disorder based on the focus shift frequency variance and focus chain break ratio, the specific operation is to establish a multivariate linear regression structure, set the context frame disorder degree Y as the target variable, introduce the focus shift frequency variance X1 and the focus chain break ratio X2 as explanatory variables, and initialize the regression weight parameters to 0.5 and 0.5 respectively. The manually annotated context stability level L in 50 training sample texts is selected as the target control, and the X1 and X2 values of each text are substituted into the regression equation to optimize the least squares error and adjust the regression parameters. Finally, the regression coefficient is stabilized at X1 weight 0.6 and X2 weight 0.4. The corresponding paragraph is input into the new sample text. The X1 and X2 values are substituted into the regression function, and the Y value is calculated and then compared with the level interval in the training set to identify the level of context disorder in this paragraph. The higher the Y value, the more chaotic the frame. The context frame disorder data is output in paragraph units. The fields include paragraph number, X1 value, X2 value, regression prediction value Y, context disorder level label and participation factor contribution evaluation, which are used as key indicator basis for subsequent main concept deviation degree analysis and context ambiguity degree quantification operation. Among them, the level interval in the training set, the Y value range Y∈[0,0.3) corresponds to "stable", Y∈[0.3,0.7) corresponds to "mild disorder", and Y∈[0.7,1.0] corresponds to "severe disorder".
[0042] In this embodiment, reference Figure 3 The above is a flowchart of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Step S31: performing logic learning on the context ambiguity degree quantified data to obtain context ambiguity degree learning data; Step S32: performing feature sampling on the context fuzziness degree learning data to obtain fuzziness degree learning feature sampling data; Step S33: Designing an English language proficiency assessment framework based on the fuzzy degree learning feature sampling data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; Step S34: Send the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0043] In the embodiment of the present invention, the quantitative data of context ambiguity is subjected to logic learning processing. In this process, the main concept deviation degree, semantic rhetoric dislocation frequency, focus chain break ratio and reference ambiguity weight in the quantitative data of context ambiguity are used as core variables to construct a 4-dimensional fuzzy feature vector set. ,in is the frequency of context focus jump, is the number of causal fault levels, is the semantic confusion rate, To refer to conflicting values, a co-occurrence matrix between fuzzy variables is constructed using logical reduction method. In the same text and The frequency of co-occurrence is calculated, and the principal component normalization is performed after constructing the M matrix. The collinear interference value is eliminated and the feature group with a feature synergy factor greater than 0.7 is retained. The feature importance weighting mechanism based on structural entropy is used to perform discrete value Boolean learning on each group of co-occurrence factors to form a feature pattern sequence set. , each set of sequences represents a typical fuzzy structure configuration, thus constructing a contextual fuzzy degree logic learning sample set , and use this set in the subsequent feature sampling task. Feature sampling is performed on the context fuzziness learning data. In this stage, the stratified multi-segment sampling method is used to sample the logical learning sample set. The extraction is performed paragraph by paragraph. First, the sample set is divided into low ambiguity group, medium ambiguity group, and high ambiguity group according to the degree of subject deviation. Each group is further divided into intervals according to the focus jump density. 20% of the paragraphs in each interval are randomly selected as sampling samples. After sampling, the feature vectors in each paragraph are maximum entropy encoded to construct a feature training set D with context ambiguity level as the label. Each data structure is ⟨feature group sequence, ambiguity level label>, where the ambiguity level label is divided into 5 levels from 1 to 5. There are 600 samples in total, 120 in each group. The training set is normalized by standard deviation, with the mean set to 0 and the standard deviation set to 1, forming a fuzzy degree learning feature sampling dataset. , the dataset contains fields such as focus chain break ratio, jump point frequency, ambiguity density, logical offset value, fuzzy level label and other 5-dimensional structures. Based on the Q-learning algorithm, the English language proficiency assessment architecture is designed for the fuzzy degree learning feature sampling data. In this stage, the Q-learning state space S and action space A are constructed, where the state space S is the 5-dimensional context feature combination in the sampling data, and the action space A is the corresponding ability assessment level division action set {A1, A2, A3, A4, A5}. Initially, the Q value matrix is set to a full zero matrix, and the ε-greedy strategy is used to control the exploration rate, where the initial value of ε is set to 0.2, the discount factor γ is set to 0.9, and the learning rate α is set to 0.1. In the iterative process, the state space S is set to 0.2, the discount factor γ is set to 0.9, and the learning rate α is set to 0.1. Extract the current context fuzzy feature combination and select the action The evaluation grade assigned to the current state is used to calculate the immediate feedback value based on the difference between the evaluation grade of the current action and the manually labeled grade in the training set. As a reward signal, The update strategy is used to modify the Q value matrix. Each round of iteration is updated 100 times. The total number of training rounds is set to 500 rounds. The final convergence condition is that the change of the Q value matrix is less than After completion, the final Q-value matrix is extracted and a Q-strategy function Q: S→A is constructed, mapping the state input to the language proficiency level. This completes the design and output of the English language proficiency assessment architecture. To send the English language proficiency assessment architecture to the terminal for execution, this step first quantifies the contextual ambiguity of the user-uploaded text, obtains a 5-dimensional feature combination as the current state input, and inputs it into the trained Q-strategy function Q: S→A. The corresponding language proficiency level output is extracted based on the action index corresponding to the maximum Q-value in the Q-value matrix corresponding to the current state. The language proficiency level is divided into five levels, from A1 to A5, where A1 represents extremely low context stability and A5 represents extremely high context stability. The output level is accompanied by the Q-value weight that matches the current input state vector to the level for evaluation and interpretation. The level is finally presented through the terminal interface.
[0044] The present invention also provides an English language proficiency assessment system for executing the above-mentioned English language proficiency assessment method, the English language proficiency assessment system comprising: The category sentence structure analysis module is used to collect English writing content texts uploaded by students through the terminal; analyze the style category of the English writing content texts uploaded by students, and then perform category sentence structure analysis to obtain the content style category sentence structure; The context ambiguity degree quantification module is used to identify logical relationship deviations in English writing content based on content style category sentence structure to obtain content logical relationship deviation clustering data; based on the content logical relationship deviation clustering data, the context ambiguity degree is quantified to obtain context ambiguity degree quantitative data; The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
[0045] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating English language proficiency, characterized in that: The following steps are involved: Step S1: collecting English writing content texts uploaded by students through the terminal; Analyze the stylistic categories of the English writing content uploaded by students, and then conduct category sentence structure analysis to obtain the content stylistic category sentence structure; Wherein, step S2 includes: Step S21: performing a content tense confusion structure analysis on the English writing content text based on the content style category sentence structure to obtain the content text tense confusion structure; Step S22: identifying logical relationship deviations of the English writing content text according to the tense disorder structure of the content text, thereby obtaining content logical relationship deviation data; Step S23: performing deviation clustering processing on the content logic relationship deviation data to obtain content logic relationship deviation clustering data; Wherein, step S24 is specifically as follows: Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data; Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse; Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability; Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data; Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data; Step S246: quantifying the context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity degree quantification data; Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
2. The English language proficiency assessment method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting English writing content texts uploaded by students through the terminal; Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text; Step S13: performing a stylistic analysis on the revised English writing content to obtain a content stylistic category; Step S14: performing a category sentence structure analysis on the revised English writing content text according to the content style category to obtain the content style category sentence structure.
3. The English language proficiency assessment method according to claim 1, characterized in that: Step S22 includes the following steps: Step S221: extracting tense misuse / jump / missing structures in the tense confusion structure of the content text; Step S222: Detecting the English writing content text for paragraphs with invalid tense correspondence based on tense misuse / jump / missing structures, thereby obtaining paragraphs with invalid tense correspondence; Step S223: performing content time axis conflict analysis on the paragraphs with invalid tense correspondence, and obtaining content time axis conflict data between the paragraphs with invalid tense correspondence; Step S224: Identify the logic state fragmentation of the instant event based on the content time axis conflict data and the tense misuse / jump / missing structure between the tense echo failure paragraphs to obtain the logic state fragmentation data of the instant event; Step S225: performing logical relationship deviation identification on the English writing content text according to the instant event logical state segmentation data, thereby obtaining content logical relationship deviation data.
4. The English language proficiency assessment method according to claim 3, characterized in that: Step S224 includes the following steps: Perform content paragraph single-chain logical break analysis on the content timeline conflict data between paragraphs with failed temporal correspondence, and obtain a content paragraph single-chain logical break dataset; Based on the single-strand logical break dataset of content paragraphs and the temporal misuse / jump / missing structure, a single-strand logical break timeline directed graph is constructed. Perform break density calculation on the single-chain logical break time axis directed graph to obtain the logical break density; The logical disorder entropy value of the single-chain logical break time axis directed graph is calculated based on the logical break density to obtain the content logical disorder entropy value; Based on the logic fracture density and content logic disorder entropy, the logic state splitting of instant events is identified to obtain the logic state splitting data of instant events.
5. The English language proficiency assessment method according to claim 1, wherein: Step S244 includes the following steps: Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content; Quantify the degree of semantic focus disorder in content subjective / objective context ambiguity data; Calculate the focus shift frequency variance in the semantic focus disorder degree; Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, content rhetoric dislocation analysis was conducted to obtain content semantic rhetoric dislocation data; The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus shift frequency variance to obtain the focus chain break ratio; Based on the variance of focus shift frequency and the ratio of focus chain breaks, a regression analysis of content context frame disorder was conducted to obtain context frame disorder data.
6. The English language proficiency assessment method according to claim 1, wherein: Step S3 includes the following steps: Step S31: performing logic learning on the context ambiguity degree quantified data to obtain context ambiguity degree learning data; Step S32: performing feature sampling on the context fuzziness degree learning data to obtain fuzziness degree learning feature sampling data; Step S33: Designing an English language proficiency assessment framework based on the fuzzy degree learning feature sampling data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; Step S34: Send the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
7. An English language proficiency assessment system, characterized in that: For executing the English language proficiency assessment method according to claim 1, the English language proficiency assessment system comprises: The category sentence structure analysis module is used to collect English writing content texts uploaded by students through the terminal; analyze the style category of the English writing content texts uploaded by students, and then perform category sentence structure analysis to obtain the content style category sentence structure; The context ambiguity degree quantification module is used to identify logical relationship deviations in English writing content based on content style category sentence structure to obtain content logical relationship deviation clustering data; and to quantify the context ambiguity degree based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data; The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
Citation Information
Patent Citations
Language evaluation generation method based on reinforcement learning
CN110532555A
Multi-dimensional English composition scoring method and device and readable storage medium
CN113836894A
Automatic Mongolian speech quality evaluation method based on hierarchical transfer learning
CN116434778A
Intelligent English writing reviewing system and use method thereof
CN120235141A
Chinese article evaluation method and system and computer reading medium
TW579470B
Cited By
Pretrained classification model-based interpretable text classification method and system
CN121071155A
Explainable text classification method and system based on pre-trained classification model
CN121071155B
Method and system for generating intelligent insight report based on AI large model
CN121328504A
Intelligent analysis and text optimization method and system for external financial and economic words
CN121543600A