A method and system for assessing english language proficiency
By analyzing the stylistic categories and sentence structures of students' English writing content and combining it with the Q-learning algorithm, we can identify logical relationship deviations and quantify the degree of contextual ambiguity. This solves the problem of low accuracy in logical and contextual ambiguity analysis in existing English language proficiency assessment methods, and achieves more accurate and personalized assessment.
Patent Information
- Application Number
- CN202511094450.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing English language proficiency assessment methods fail to comprehensively examine students' language abilities, especially in the analysis of logic and contextual ambiguity, which leads to large assessment errors.
By analyzing the stylistic categories and sentence structures of the English writing content uploaded by students, identifying logical relationship deviations, and using the Q-learning algorithm to quantify the degree of context ambiguity, we design an English language proficiency assessment framework and achieve personalized assessment.
It improves the accuracy of the analysis of the logic and contextual ambiguity of English language proficiency, reduces assessment errors, provides a more accurate and personalized language proficiency assessment, and can identify logical structure problems in students' writing and give targeted improvement suggestions.
Smart Images

Figure CN120597003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of English language proficiency assessment, and in particular to an English language proficiency assessment method and system. Background Art
[0002] Computer technology can identify grammatical errors, irregular sentence structures, unclear logical relationships, and other issues in students' writing, and conduct in-depth analysis of the language proficiency of the written content. Especially in English writing assessments, factors such as contextual ambiguity, clarity of logical relationships, and appropriateness of stylistic structure have become important indicators for evaluating students' language proficiency. Previously, existing automated assessment methods typically focused on a single dimension, such as grammatical correctness, ignoring the more complex cognitive and contextual aspects of writing, and failing to comprehensively examine students' language proficiency. However, traditional English language proficiency assessment methods suffer from low accuracy in analyzing the logical and contextual ambiguity of the English language, resulting in large errors in English language proficiency assessments. Summary of the Invention
[0003] Based on this, it is necessary to provide an English language proficiency assessment method and system to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for evaluating English language proficiency is provided, comprising the following steps:
[0005] Step S1: collecting English writing content texts uploaded by students through the terminal; analyzing the style categories of the English writing content texts uploaded by students, and then performing category sentence structure analysis to obtain the content style category sentence structure;
[0006] Step S2: Identifying logical relationship deviations of English writing content texts based on content style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the degree of context ambiguity based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data;
[0007] Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0008] Preferably, step S1 includes the following steps:
[0009] Step S11: collecting English writing content texts uploaded by students through the terminal;
[0010] Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text;
[0011] Step S13: performing genre category analysis on the English writing content revision text to obtain a content genre category;
[0012] Step S14: performing category sentence pattern structure analysis on the English writing content revision text according to the content genre category to obtain a content genre category sentence pattern structure.
[0013] Preferably, step S2 comprises the following steps:
[0014] Step S21: performing content tense confusion structure analysis on the English writing content text based on the content genre category sentence pattern structure to obtain a content text tense confusion structure;
[0015] Step S22: performing logical relationship deviation identification on the English writing content text according to the content text tense confusion structure to obtain content logical relationship deviation data;
[0016] Step S23: performing deviation clustering processing on the content logical relationship deviation data to obtain content logical relationship deviation clustering data;
[0017] Step S24: performing context fuzziness degree quantification on the English writing content text according to the content logical relationship deviation clustering data to obtain context fuzziness degree quantification data.
[0018] Preferably, step S22 comprises the following steps:
[0019] Step S221: extracting tense misuse / skipping / missing structures in the content text tense confusion structure;
[0020] Step S222: performing sentence tense echo failure paragraph detection on the English writing content text according to the tense misuse / skipping / missing structures to obtain content sentence tense echo failure paragraphs;
[0021] Step S223: performing content time axis conflict analysis on the content sentence tense echo failure paragraphs to obtain content time axis conflict data between the tense echo failure paragraphs;
[0022] Step S224: performing immediacy event logical state fragmentation identification based on the content time axis conflict data between the tense echo failure paragraphs and the tense misuse / skipping / missing structures to obtain immediacy event logical state fragmentation data;
[0023] Step S225: performing logical relationship deviation identification on the English writing content text according to the immediacy event logical state fragmentation data to obtain content logical relationship deviation data.
[0024] Preferably, step S224 comprises the following steps:
[0025] Perform content paragraph single-chain logical break analysis on the content timeline conflict data between paragraphs with failed temporal correspondence, and obtain a content paragraph single-chain logical break dataset;
[0026] Based on the single-strand logical break dataset of content paragraphs and the temporal misuse / jump / missing structure, a single-strand logical break timeline directed graph is constructed.
[0027] Perform break density calculation on the single-chain logical break time axis directed graph to obtain the logical break density;
[0028] The logical disorder entropy value of the single-chain logical break time axis directed graph is calculated based on the logical break density to obtain the content logical disorder entropy value;
[0029] Based on the logic fracture density and content logic disorder entropy, the logic state splitting of instant events is identified to obtain the logic state splitting data of instant events.
[0030] Preferably, step S24 includes the following steps:
[0031] Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data;
[0032] Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse;
[0033] Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability;
[0034] Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data;
[0035] Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data;
[0036] Step S246: quantifying the degree of context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity degree quantification data.
[0037] Preferably, step S244 includes the following steps:
[0038] Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content;
[0039] quantifying a semantic focus disorder degree in the context ambiguity data;
[0040] calculating a focus shift frequency variance in the semantic focus disorder degree;
[0041] performing content rhetoric dislocation analysis based on the semantic focus disorder degree, content causal relationship fault data, and content word sense confusion probability, to obtain content semantic rhetoric dislocation data;
[0042] performing focus chain breakage ratio searching calculation according to the content semantic rhetoric dislocation data and the focus shift frequency variance, to obtain a focus chain breakage ratio;
[0043] performing content context framework disorder regression analysis based on the focus shift frequency variance and the focus chain breakage ratio, to obtain context framework disorder data.
[0044] Preferably, the step S3 comprises the following steps:
[0045] Step S31: performing logical learning on the context ambiguity degree quantification data, to obtain context ambiguity degree learning data;
[0046] Step S32: performing feature sampling on the context ambiguity degree learning data, to obtain ambiguity degree learning feature sampling data;
[0047] Step S33: performing English language ability evaluation architecture design on the ambiguity degree learning feature sampling data based on a Q-learning algorithm, to obtain an English language ability evaluation architecture;
[0048] Step S34: sending the English language ability evaluation architecture to a terminal, to perform English language ability evaluation.
[0049] Preferably, the present application further provides an English language ability evaluation system for performing the English language ability evaluation method as described above, which comprises:
[0050] a category sentence structure analysis module, configured to collect English writing content text uploaded by a student through a terminal, analyze the category of the English writing content text uploaded by the student, and then perform category sentence structure analysis, to obtain a content category sentence structure;
[0051] a context ambiguity degree quantification module, configured to perform logical relationship deviation identification on the English writing content text based on the content category sentence structure, to obtain content logical relationship deviation clustering data, and perform context ambiguity degree quantification according to the content logical relationship deviation clustering data, to obtain context ambiguity degree quantification data;
[0052] The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
[0053] The beneficial effect of the present invention is that by collecting English writing texts uploaded by students through a terminal and analyzing their writing style and sentence structure, a comprehensive understanding of students' writing characteristics and language usage habits can be obtained. The analysis of writing style helps identify students' expression methods in different writing scenarios, while the analysis of sentence structure can reveal the complexity and fluency of students' language expression. This analysis not only accurately divides the language structure of the writing content, but also lays the foundation for subsequent logical relationship identification and quantification of context ambiguity. Through the systematic analysis of student writing texts, data support can be provided for the assessment of English language proficiency. Based on the analysis of content style and sentence structure, logical relationship deviations are identified, and different logical deviations are identified through clustering algorithms, providing deeper data support for student writing analysis. By clustering these deviations, common logical structure problems in students' writing can be revealed, such as insufficient argumentation and unclear viewpoints, which often affect the overall logic and persuasiveness of the writing. At the same time, quantifying the degree of context ambiguity based on logical relationship deviation data can accurately assess the ambiguity and uncertainty of students' language expression in real contexts. The quantitative data of context ambiguity provides an objective basis for subsequent ability assessment, which can help the system more accurately assess students' language application ability in different contexts and make up for the shortcomings of traditional assessment methods in details. In the process of designing an English language ability assessment framework based on the Q-learning algorithm, the algorithm can dynamically adjust the assessment model according to the students' writing characteristics and the quantitative data of context ambiguity through continuous learning and optimization, thereby realizing personalized language ability assessment. The introduction of the Q-learning algorithm makes the assessment process not just a static scoring, but an intelligent and adaptive assessment system that can be optimized and adjusted at any time according to the progress of students' writing. This assessment framework can not only accurately identify students' problems in grammar, sentence structure, logical relationships, etc., but also provide specific feedback and improvement suggestions based on the characteristics of different students, providing students with more targeted learning guidance. Therefore, the present invention is an optimization process made to a traditional English language ability assessment method, which solves the problem that a traditional English language ability assessment method has low accuracy in analyzing the logic and context ambiguity of the English language, thereby causing large errors in English language ability assessment, improves the accuracy of analyzing the logic and context ambiguity of the English language, and reduces the errors in English language ability assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a step flow chart of an English language ability evaluation method;
[0055] Figure 2 It is Figure 1 The detailed implementation step flow chart of step S2 in the embodiment;
[0056] Figure 3 It is Figure 1 The detailed implementation step flow chart of step S3 in the embodiment. DETAILED DESCRIPTION
[0057] Please refer to Figures 1 to 3 An English language ability evaluation method, the method comprises the following steps:
[0058] Step S1: collecting the English writing content text uploaded by students through a terminal; analyzing the text style category of the English writing content text uploaded by students, and then performing category sentence structure analysis to obtain the content text style category sentence structure;
[0059] Step S2: identifying the logical relationship deviation of the English writing content text based on the content text style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the context fuzziness degree according to the content logical relationship deviation clustering data, thereby obtaining context fuzziness degree quantification data;
[0060] Step S3: designing the English language ability evaluation architecture based on the Q-learning algorithm to quantify the context fuzziness degree, thereby obtaining the English language ability evaluation architecture; sending the English language ability evaluation architecture to the terminal to perform English language ability evaluation.
[0061] In the embodiment of the application, reference Figure 1 The step flow chart of an English language ability evaluation method, in this example, the English language ability evaluation method comprises the following steps:
[0062] Step S1: collecting the English writing content text uploaded by students through a terminal; analyzing the text style category of the English writing content text uploaded by students, and then performing category sentence structure analysis to obtain the content text style category sentence structure;
[0063] In the embodiment of the present invention, the English writing text submitted by the student is received through a preset terminal interface, the terminal is connected to the Nginx reverse proxy server through the HTTPS protocol, and the proxy server forwards the data request to the Flask service interface module built in the back-end Python language. After receiving the English writing text, the service interface immediately calls the text preprocessing module for standardization. The preprocessing module uses regular expressions to clean control characters, non-language characters and redundant spaces, unifies the line break format and removes non-English language content, and then inputs the cleaned text into the line style category recognition module. This module is based on the Bayesian text classification algorithm, and the training corpus uses a public English writing genre corpus (such as COCA, BNC). It extracts 42 feature dimensions in five categories, including word frequency distribution, syntactic structure label distribution, transition word usage frequency, sentence length mean, and conjunction type frequency, as input vectors. After TF-IDF vectorization processing, MultinomialNaive is used. Bayes (Multinomial Naive Bayes) is used for genre prediction. The prediction results are divided into four categories: narrative, expository, argumentative, and descriptive. The prediction results and the text are sent to the category sentence structure analysis module. The sentence structure analysis module calls the depparse component of the dependency syntax analyzer Stanford CoreNLP to extract the subject-verb-object structure, modifiers, non-restrictive clauses, and adverbial position labels of each sentence. It analyzes the sentence usage frequency in each genre and counts the occurrence ratios of typical structures such as noun clauses, attributive clauses, passive voice, modal verb structures, and subjunctive mood. Finally, it generates the corresponding content genre category sentence structure vector for subsequent calls.
[0064] In another embodiment, the English writing content text uploaded by the student is obtained through the terminal. The terminal is a remote data input interface module of the teaching platform. The module has a structured data channel configuration function, which can parse the text content into paragraph structure and use the sentence vector encoding mechanism based on the BERT model to embed the uploaded text into a semantic vector. After the original corpus is segmented at the sentence level, the TextRazor API (i.e., a platform for natural language processing services) is first used to perform preliminary style recognition. By extracting indicators such as modifier density, syntactic structure ratio, and voice distribution, a style vector feature set is constructed. The Gunning Fog Index value is combined with the Flesch Reading The Ease score is used as an auxiliary parameter for genre identification to determine the genre category of the text. This process sets the classification label set into five categories: expository, argumentative, narrative, applied, and descriptive. For narrative and argumentative categories, the density of transition markers and the proportion of subjective emotional words are additionally extracted to enhance the accuracy of discrimination. After the genre classification is completed, the sentence structure of each sentence is analyzed based on the syntactic analysis tree constructed based on StanfordParser (an open source syntactic analysis tool). Sentence combination type labels (such as main-subordinate compound sentences, parallel sentences, imperative sentences, etc.) are extracted from the tree structure. Combined with the sentence-initial part-of-speech pattern and the sentence-final mood structure, the sentence structures of each category are statistically classified, ultimately forming a content genre category sentence structure containing feature vectors such as the proportion of sentence types in each paragraph, the frequency of conjunction types, and the distribution of tense patterns.
[0065] Step S2: Identifying logical relationship deviations of English writing content texts based on content style category sentence structure to obtain content logical relationship deviation clustering data; quantifying the degree of context ambiguity based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data;
[0066] In the embodiment of the present invention, according to the content style category sentence structure vector obtained in step S1, the logical relationship deviation identification operation of the English writing text is performed through the structural disorder detection module. The module has a built-in inter-sentence connection graph constructed by depth-first traversal. The nodes in the graph represent each sentence in the text. The edge weight is calculated by the semantic continuity of the inter-sentence conjunctions and the logical connection structure. If there is no clear causal or temporal connection clue between two adjacent sentences, the edge weight is set to 0. The graph further uses a minimum spanning tree-based method to identify isolated nodes and edge break areas. Isolated nodes are logically broken sentences. All broken sentences are clustered using the K-means clustering algorithm. The vector dimension includes sentence complexity. (number of subordinate clauses), information entropy value (frequency distribution of key words), cosine value of the angle between the subject vectors of the previous and next sentences, the default K value is set to 5, and it is adjusted according to the stability of the variance within the cluster. After the clustering is completed, the logical relationship deviation clustering data is obtained, and then the break density, sentence flatness rate, and logical connection vacancy rate in the cluster are input into the context fuzzy quantification module. The module uses a multidimensional normalization method to standardize the data to the [0,1] interval, and uses PCA principal component analysis to extract the top three main influencing factors, assigning weights of 0.4, 0.35, and 0.25 for weighted processing, and outputs the quantitative value as the quantitative data of the context fuzziness degree. The data format is a quantitative score and an index list of corresponding text segment numbers.
[0067] In another embodiment, the content style category sentence structure obtained in step S1 is used to implement a logical relationship deviation recognition operation. The operation uses a local adjacent paragraph comparison analysis method based on sentence connection logic. In a specific implementation, a paragraph-level semantic coherence alignment model is constructed to build a semantic embedding vector for each adjacent two paragraphs. The Word Mover's Distance is used to calculate the theme coherence between the paragraphs. If the theme coherence of three consecutive paragraphs is lower than a set threshold of 0.38 and the sentence structure consistency index across the paragraphs is greater than 0.72, it is marked as a weak logical connection area. Further, an inter-sentence causal connection word (such as because, therefore) missing detection algorithm is used to identify the inter-sentence logical breakpoints in these areas. If the logical connection word is missing but the semantics has a contradiction or a coherence interruption, it is determined as a logical relationship deviation point. This type of deviation point is marked in the sentence structure graph in the form of a logical break node. In the clustering process of all deviation nodes, the DBSCAN density clustering algorithm is used to take the position of the deviation point in the text structure and the sentence category it belongs to as the clustering feature dimensions. The abnormal distribution area of the logical structure in the text is aggregated to form a deviation cluster. Each cluster corresponds to the type and density interval of the logical abnormality occurrence. Then, the context fuzziness quantification processing is performed on each logical relationship deviation cluster. First, the frequency of reference structure back-reference failure, the frequency of causal relationship chain breakage, and the temporal inconsistency rate in each cluster are calculated. The weighted values are 0.35, 0.4, and 0.25, respectively, for weighted comprehensive scoring. The fuzzy degree value is output in percentage. All fuzzy degree data constitute the context fuzziness quantification data set.
[0068] Step S3: Based on the Q-learning algorithm, the context fuzziness quantification data is used to design an English language ability evaluation architecture, so as to obtain an English language ability evaluation architecture. The English language ability evaluation architecture is sent to a terminal to perform English language ability evaluation.
[0069] In the embodiment of the present invention, the Q-learning reinforcement learning algorithm is used to design an evaluation architecture for the above-mentioned context fuzziness degree quantitative data. In the specific operation, the state space is set to the distribution interval of the fuzziness degree quantitative score [0,100], which is divided into 20 state levels with a unit of 5. The action space is set to four categories of language ability grading behaviors (low level, medium level, good level, and excellent level). The reward function is defined as an inverse penalty function based on the matching accuracy between the real annotated English ability level label and the algorithm output. This function assigns a negative reward value to the result with a large prediction deviation. In the Q learning iteration process, the initial Q table is a zero matrix, and the ε-greedy strategy is adopted. For state behavior selection, the learning rate is set to 0.3, the discount factor is set to 0.9, and the number of iterations is set to 1000. In each round, samples of different ambiguity levels are grouped to train the learning path and update the Q-value table. Finally, a decision path mapping structure is constructed that automatically determines the language proficiency level after inputting quantitative data on the ambiguity level of the context. This is the English language proficiency assessment architecture. In this architecture, each level output node is bound to the feature combination pattern of the semantic ambiguity level and its corresponding probability threshold range. After the evaluation architecture is constructed, the architecture is written to the local memory interface registration module, which controls the return of the structured evaluation architecture data to the terminal execution module for the entire subsequent evaluation operation process.
[0070] Step S1 includes the following steps:
[0071] Step S11: collecting English writing content texts uploaded by students through the terminal;
[0072] Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text;
[0073] Step S13: performing a stylistic analysis on the revised English writing content to obtain a content stylistic category;
[0074] Step S14: performing a category sentence structure analysis on the revised English writing content text according to the content style category to obtain the content style category sentence structure.
[0075] In an embodiment of the present invention, the English writing content text submitted by students is received through the data upload interface module of the teaching platform. The interface module has an HTTP POST data access mechanism and implements asynchronous data reading based on a multi-channel cache strategy. After access, the data stream is first split into chunks using a paragraph segmentation logic based on sentence-end punctuation rules. Regular expression rules are used to match English end markers such as periods, question marks, and exclamation marks to perform preliminary sentence boundary recognition on the text. The paragraph position relationship of each sentence structure is retained in the cache through a numbered index. The input content is then converted into a standard byte stream using a UTF-8 encoding and decoding module to ensure consistency in subsequent character-level processing. The total length of the text content obtained in this stage is not less than 500 characters. A statistical strategy of dividing the average sentence length by 20 words is used to obtain a set of valid sentences of not less than 25 sentences, which serves as the basic corpus data set for subsequent processing. A spelling error correction operation is performed on the English writing content text obtained in step S11. This operation uses a spelling error detection method based on character-level edit distance calculation to compare each word in the text with a standard English dictionary library. The dictionary library contains Oxford Dictionary. The system uses a 3000 vocabulary list, the Coca corpus's top 10,000 words, and the Collins standard vocabulary set. It uses the Levenshtein distance algorithm to detect the minimum number of edit steps between the target word and the vocabulary entry. When the distance value is less than or equal to 2 and the target word appears less than or equal to 1 time in the context, it is marked as a low-confidence word. Such entries are required to participate in the context semantic similarity calculation with the three words that appear most frequently in the context window. The Jaccard similarity is used to compare the similarity of the context 3-gram phrases. When the similarity is greater than 0.68, the semantic nearest neighbor word is selected to replace the original word. At the same time, the original word and the replacement word are retained in the correction comparison table. All replacement actions are performed within a logical loop to complete the spelling correction of each sentence in turn. The total number of corrected words in the final output text is no less than 2% of the total number of words in the text, forming a complete English writing content correction text.
[0076] The English writing content revision text obtained in step S12 is subjected to style category analysis, and a multi-dimensional text style vector coding method is used for style modeling analysis. First, the main and subordinate composite structures of the text are labeled based on the sentence depth nesting index. The index obtains the complex sentence density score by calculating the difference between the average and maximum embedding depth of subordinate sentences. Meanwhile, the unit sentence frequency of signifying conjunctions such as however, moreover, in contrast, in conclusion, etc. is counted, and the semantic polarity score weighted distribution of subjective evaluation words such as amazing, terrible, impressive, etc. at the end of the sentence is calculated. A feature vector group containing 12-dimensional style parameters such as average sentence length, subject antecedent rate, passive structure usage frequency, abstract noun density, and emotional color weight is constructed. Then, an unsupervised clustering method based on KNN is used for style classification. The K value is set to 5, and the Manhattan distance is used to measure the style similarity between feature vectors. The classification labels are explanation class, argumentative class, narrative class, description class, and application class, a total of five standard styles. The argumentative type needs to meet the conditions that the passive sentence frequency ratio is greater than 25% and the connection logic frequency is greater than 5 times per hundred words. Finally, the corresponding style is attributed according to the minimum distance between the style vector of each paragraph and the five type templates, and the style category analysis process is completed. The output of each paragraph belonging to the style label and the corresponding style vector weight data constitutes the content style category. According to the content style category obtained in step S13, the corresponding English writing content revision text is subjected to category sentence structure analysis. First, the Stanford Parser is used to analyze the dependency structure of the text. In each sentence, the syntax path information of the subject-predicate-object structure is extracted, and the adverb modification direction, adjective clause nesting path, and tense verb position vector are recorded. For narrative texts, the position of temporal adverbs and the distribution of verb tenses are extracted. The tense distribution ratio curve is calculated and whether there is a tense mutation point is judged. For explanatory texts, noun phrases are counted and the logical consistency of their post-modification structures is analyzed. For argumentative texts, the frequencies of "if-then" and "although-yet" structures and the causal sentence combination density are counted. Combined with the frequency distribution of conjunctions, the sentence structure distribution matrix corresponding to each style type is constructed. The sentence structure in each style sample is encoded in the form of a vector and merged to form a style sentence structure set. This set is used for subsequent logical relationship identification and context consistency verification. The final output includes the sentence structure label set, the sentence category frequency vector, and the sentence distribution density matrix, which constitute the complete content style category sentence structure data set.
[0077] In this embodiment, reference is made to Figure 2 The detailed implementation steps of step S2 are shown in the flowchart. In this embodiment, the detailed implementation steps of step S2 include:
[0078] Step S21: performing a content tense confusion structure analysis on the English writing content text based on the content style category sentence structure to obtain the content text tense confusion structure;
[0079] Step S22: identifying logical relationship deviations of the English writing content text according to the tense disorder structure of the content text, thereby obtaining content logical relationship deviation data;
[0080] Step S23: performing deviation clustering processing on the content logic relationship deviation data to obtain content logic relationship deviation clustering data;
[0081] Step S24: quantifying the degree of context ambiguity of the English writing content text according to the content logical relationship deviation clustering data, thereby obtaining context ambiguity quantification data.
[0082] In the embodiment of the present application, step S21 performs tense confusion structure analysis operation on the English writing content text based on the content style category sentence structure obtained in step S14. In the specific implementation process, first, the modified English text is subjected to tense annotation processing. A morphological analysis algorithm based on tense recognition rules is used to identify the tense type of the predicate verb of each sentence. The identification logic relies on rule tree matching technology to determine the tense type according to the combination relationship between the verb prototype and the auxiliary verb. There are eight types of tense categories: present tense, past tense, present perfect tense, past perfect tense, future tense, present progressive tense, past progressive tense, and future perfect tense. Each type of tense has a specific combination mode in the sentence structure. In the past tense, the past tense of the verb is used as the main structure. In the perfect tense, the auxiliary verb is have or has and is followed by the past participle. Based on the rule set, the main verb of all sentences is annotated with the tense. At the same time, paragraph-level tense jump identification is performed in combination with the sentence position index and the context sentence consistency. When the frequency of tense switching between adjacent three sentences is greater than 2 and is not driven by causal time structure, it is marked as a non-logical tense jump area. If the same subject object is used with alternating tenses in a paragraph, it is determined to be a tense misuse structure. This analysis process uses a linear time sequence index vector to record the tense switching path between sentences. Finally, the set of all tense misuse, jump, and missing structures in the content text is integrated to form the content text tense confusion structure. According to the content text tense confusion structure obtained in step S21, logical relationship deviation identification operation is performed on the English writing content text. In the specific implementation process, a semantic advancement relationship graph is constructed based on the identified confusion structure. The graph uses sentence sequences as nodes and event time sequence order as directed edges. Time sequence anchor points are constructed using sentence head and tail event state words (such as begin, end, continue, and finish). All inconsistent timeline paths are marked as time advancement breakpoints. Further, the connection word logical consistency between the sentences corresponding to these breakpoints is verified. A connection word dependency relationship matching method is used to search for the first connection component of the sentence, and the semantic relationship of the previous and next sentences is mapped and checked. When the connection word and the sentence semantic logic are inconsistent (for example, although connects two essentially identical sentences), it is determined to be a logical relationship deviation point. In the analysis process, the inter-sentence proposition nesting depth is introduced as a redundant logic judgment basis. If a sentence has excessive nested conditional sentences or concession sentences, causing logical advancement to be ambiguous, it is considered to be a logical deviation caused by structural complexity. A three-value logic backtracking method is used in the logical deviation identification operation to disassemble and process the nested logical relationship, form a logical relationship chain, and mark the broken parts for deviation. Finally, the content logical relationship deviation data composed of all sentence relationship structures with logical errors are output.
[0083] Deviation clustering is performed on the content logic relationship deviation data obtained in step S22. First, a three-dimensional feature space is constructed for all deviation points based on their position in the text, deviation type and temporal association. Each deviation point is represented as a vector. The first dimension of the vector is the normalized value of the sentence sequence number, the second dimension is the deviation type number (such as missing conjunctions are marked as 1, tense jumps are marked as 2, and structural nested redundancy is marked as 3), and the third dimension is the context consistency measure of the corresponding temporal chaos structure. The consistency measure is calculated based on the cosine similarity between the sentence vector and the adjacent sentence vectors. The density-based clustering algorithm DBSCAN is used to perform density clustering on the above three-dimensional deviation vector set. The minimum sample number is set to 4, and the ε radius distance is set to 0.3. The final clustering output is several logical deviation clusters. Each cluster contains a group of deviation points that have an association relationship in the semantic structure, syntactic structure and time advancement path. Through clustering, high-density areas of logical errors can be identified at the paragraph level to form content logic relationship deviation clustering data. The context ambiguity degree is quantified based on the content logic relationship deviation clustering data obtained in step S23. First, the subject-predicate structure consistency analysis is performed on the sentences in each logic relationship deviation cluster, and the expressions with ambiguity or conflict in the semantic direction between the subject, predicate and object are extracted, and the frequency of each conflict type is recorded. Secondly, the ambiguity strength of the connectives in each cluster is scored, and the frequency of known fuzzy connectives (such as since, while, as) in the Brown corpus is used as the semantic fuzziness index benchmark. The ambiguity strength is calculated by the word meaning fuzzy context overlap rate indicator. If the overlap rate exceeds 0.6, it is considered a highly ambiguous connective. The pronoun reference failure rate in the paragraph is further counted. If it appears in two or more consecutive sentences, the ambiguity strength is calculated. If a pronoun has no clear referent, it is counted as the frequency of contextual reference confusion. At the same time, the word meaning jump rate in the cluster is analyzed, and the contextual semantic category of each noun is counted. If the contextual semantic label of the noun switches more than two category levels within two sentences (for example, jumping from the event class to the entity class), it is determined to be a semantic jump. Finally, a weighted comprehensive score is given to various fuzzy features in the cluster, with the subject-predicate structure fuzziness weight being 0.3, the conjunction ambiguity weight being 0.25, the reference failure weight being 0.2, and the word meaning jump weight being 0.25. The contextual fuzziness score of the cluster is obtained by weighted averaging, and the scores of all clusters are then aggregated to form a quantitative data set of contextual fuzziness, which is used in the subsequent design of the English language proficiency assessment architecture.
[0084] Step S22 includes the following steps:
[0085] Step S221: extracting tense misuse / jump / missing structures in the tense confusion structure of the content text;
[0086] Step S222: Perform sentence tense echo failure paragraph detection on the English writing content text according to the tense misuse / jump / missing structure, so as to obtain a content sentence tense echo failure paragraph;
[0087] Step S223: Perform content timeline conflict analysis on the content sentence tense echo failure paragraph, so as to obtain content timeline conflict data between the tense echo failure paragraphs;
[0088] Step S224: Perform immediacy event logical state fragmentation identification based on the content timeline conflict data between the tense echo failure paragraphs and the tense misuse / jump / missing structure, so as to obtain immediacy event logical state fragmentation data;
[0089] Step S225: Perform logical relationship deviation identification on the English writing content text according to the immediacy event logical state fragmentation data, so as to obtain content logical relationship deviation data.
[0090] In the embodiment of the present application, based on the obtained content text time confusion structure data, the temporal misuse, temporal jump and temporal missing structure existing therein are explicitly extracted. The operation process is as follows. Firstly, a time identification sequence based on part-of-speech tagging is constructed for all sentences. A rule is used to identify the English tense, which is a combination of auxiliary verbs and predicate verbs. The auxiliary verb set includes is, are, was, were, have, has, had, will, would, etc. The past tense and past participle are determined based on the recognition of the word form transformation. If there is no auxiliary verb in the sentence and the predicate verb is in the original form, it is marked as a temporal missing. If the auxiliary verb and the predicate verb in the sentence do not meet the standard English tense structure rules, it is marked as a temporal misuse. For example, the structure of has went or will eating in the sentence is recorded as a misuse structure. In the cross-sentence comparison, if the same subject appears twice or more in the next three sentences, the tense is switched from the future tense to the past tense or from the completed tense to the ongoing tense, and there is no time adverbial logical explanation support, it is marked as a temporal jump. The extraction operation is completed using a joint judgment mechanism based on the nested structure of the syntactic tree and the tense rules. Each error type is located as a different structure conflict node on the sentence structure tree. All structure annotations are recorded in the index table for subsequent logical analysis. In step S222, according to the extracted temporal misuse, jump and missing structure, the sentence tense echo failure paragraph detection of the English writing content text is performed. The detection method is to calculate the temporal consistency coefficient between sentences in a paragraph. The coefficient is based on the pairing of the tense label sequence between sentences. It is defined as the ratio of the number of sentences with the same tense in a paragraph to the total number of sentences in the paragraph. If the value is less than 0.35 and there are two or more sentences with marked temporal errors in the paragraph, the paragraph is marked as a temporal echo failure paragraph. Further, it is determined whether the mixed path of switching from the past to the present and then to the past appears in the order of action or event occurrence in the paragraph. For example, in an event description, it is switched from "he had eaten" to "he eats" and then back to "he went". It is determined as a typical temporal echo failure path. All such paragraphs are uniformly archived in the temporal echo failure paragraph set for subsequent content time axis conflict analysis.
[0091] A content timeline conflict analysis is performed on the paragraph with tense-response failure obtained in step S222. First, an event timestamp is established for each event sentence. The main action verbs in each sentence are identified using an event dictionary. The event dictionary includes common action verbs such as arrive, leave, eat, decide, start, and finish. The relative temporal position of the event is determined by combining temporal adverbials in the sentence such as yesterday, last week, soon, and already. An event directed graph is constructed and a path consistency check is performed on it. When a path in which the same subject entity jumps forward and then returns backward in the time path occurs, that is, a path loop structure exists in the event graph, it is marked as a timeline conflict. This conflict is detected by detecting whether there is a loop with a length greater than 2 in the event graph. The nodes in the loop represent event actions, and the edges represent the direction of time advancement. For example, start→eat→arrive→start constitutes a conflict loop. The system marks the paragraph as a paragraph with a timeline conflict. All conflict path data and their structural features are encoded and recorded to form content timeline conflict data. Based on the above-obtained content timeline conflict data and the original tense misuse, jump, and missing structures, this paper identifies logical state fragmentation of immediate events in English writing content. The specific method is to identify whether there is semantic discontinuity in the event state change path at the same time point or within a very short period of time. The detection step first extracts the state verbs used in all events, including state-indicating words such as become, remain, stay, and feel, and analyzes their logical continuity paths. If there is an event state "he was staying home" followed by the next sentence "he had gone outside", it indicates that the state transition is inconsistent and there is no intermediary state connection. The state jumps directly from the static state to the completed movement state, which is a fragmented path. In addition, during the analysis process, a state transition diagram is constructed for all events showing state fragmentation. State fragmentation behavior is identified by determining whether there are unconnected isolated state nodes or jump-connected cross-state transition paths in the diagram. The occurrence time, state type, and sentence index of all identified state fragmentation segments are extracted and merged to form the immediate event logical state fragmentation data.The logical relationship deviation of the English writing content text is identified according to the instant event logical state fragmentation data obtained in step S224. The identification operation is to take the sentence pair in the fragmentation data as input, perform the judgment of the semantic adaptability of the conjunction and the matching test of the propositional logic in the inter-sentence logical chain, first, use the conjunction semantic classification dictionary to classify the conjunction into four categories: cause and effect, condition, transition and parallel. The conjunction semantic reverse verification is performed on the sentence pair before and after the state fragmentation. If the previous sentence state is "he was sick" and the latter sentence state is "he went to work", and the conjunction is "although", it is considered as a reasonable structure. If the conjunction is "because", it is considered as semantic mismatch, and is marked as logical relationship deviation. Secondly, the propositional logic verification is performed on the subject consistency and the logical strength relationship between the main sentence and the subordinate sentence in the fragmentation fragment. When the main sentence is "he wanted to rest" and the subordinate sentence is "heran ten miles", the logic is not correct, and the logical relationship strength score is less than 0.2 (according to the normalized score of the inner product between sentence vectors). The sentence group is classified as a logical deviation unit. The system corresponds the logical abnormal sentence group and the corresponding deviation reason (inappropriate conjunction, propositional contradiction, and time conflict) one by one to form the content logical relationship deviation data for subsequent clustering analysis and context fuzzy quantization.
[0092] Step S224 includes the following steps:
[0093] The content paragraph single-chain logical break analysis is performed on the content time axis conflict data between the time echo invalid paragraphs, and the content paragraph single-chain logical break data set is obtained.
[0094] Based on the content paragraph single-chain logical break data set and the time misuse / skipping / missing structure, the single-chain logical break time axis directed graph is constructed.
[0095] The break density of the single-chain logical break time axis directed graph is calculated, and the logical break density is obtained.
[0096] According to the logical break density, the logical disordered entropy value of the single-chain logical break time axis directed graph is calculated, and the content logical disordered entropy value is obtained.
[0097] Based on the logical break density and the content logical disordered entropy value, the instant event logical state fragmentation is identified, and the instant event logical state fragmentation data is obtained.
[0098] In an embodiment of the present invention, based on the content timeline conflict data between the tense-echo failure paragraphs obtained in the preceding step S223, a content paragraph single-chain logical break analysis is performed on the time advancement path between the paragraphs. This step first establishes an event sequence set for each paragraph marked as having a timeline conflict, extracts the action verbs in each event sentence in the paragraph as event nodes, and establishes an initial linear connection path between the event nodes according to the sentence sequence number, that is, constructs an initial single-chain event sequence, wherein each event node contains four types of parameters, namely, the event verb form, the subject word, the intra-sentence time adverbial word and its sentence sequence index number. Subsequently, a logical connection check is completed by determining the temporal logical consistency relationship between adjacent event nodes. The logical consistency check relies on the event sequence semantic template for determination. The template construction rule is to identify the state starting point and end point of each event based on the verb semantic role labeling technology. If the end state of the previous event is semantically discontinuous with the start state of the next event or there is a logical conflict, it is marked as a logical breakpoint. For example, the previous event is "he stayed at home" and the next event is "he had arrived at School" has no intermediary process events and the time adverbials conflict, which constitutes a single-chain logical break. All breakpoints are uniformly archived and recorded according to the position number of the original event sequence. Finally, a single-chain logical structure diagram consisting of alternating logically continuous segments and broken segments is constructed for each paragraph. The broken paragraph index set, break type identifier, and semantic difference between events before and after the break are extracted to form the content paragraph single-chain logical break dataset.
[0099] On the basis of the single-chain logical break dataset of the content paragraphs obtained in the above steps, the tense misuse, jump and missing structures extracted previously are combined to construct a single-chain logical break timeline directed graph. During the construction process, an event point set is first constructed for all event sentences in the paragraph. The event point takes the main predicate verb in the sentence as the central node, and adds meta-attributes such as time adverbial label, sentence sequence number and paragraph number. All event points are connected in sequence within the paragraph to form an initial linear path graph. Then, the direction edge is constructed by analyzing the time advancement direction between events. The direction of the direction edge is determined based on the temporal sequence relationship between the previous and next events. If the time adverbial in the sentence is past, next morning, then, etc. with advancement logic, the edge direction is from front to back. If there is a semantic backtracking structure such as "earlier that day" or "before", the direction edge is from front to back. That" is determined to be a reverse edge. Subsequently, all event nodes with misused, skipped, or missing temporal markers are embedded in the graph and treated as candidate break nodes. If the edge direction between such a node and its adjacent nodes is inconsistent or the edge weight is negative (calculated based on the temporal difference of the sentence vector), the edge is set as a break edge. Finally, the above nodes and edges form a complete directed graph structure, which has the event sequence path, the break edge path, and all temporal abnormality marker point information. A separate timeline directed graph structure is established at each paragraph level, and each graph number corresponds one-to-one with the paragraph number. After completing the construction of the above directed graph, the fracture density calculation is performed on the single-chain logical fracture timeline directed graph corresponding to each paragraph. The calculation operation first counts the total number of fracture edges B and the total number of event nodes N in the graph, and calculates the preliminary fracture rate parameter as B / N. Then, the fracture distribution concentration correction factor is introduced to perform weighted adjustment on whether the fracture edges are concentrated in a specific graph segment. The correction factor calculation method is to convolve the density function of each fracture edge distribution on the graph node path. If the concentration of the fracture edge on the continuous path exceeds a certain threshold (set to 0.65), the fracture density weighting coefficient is multiplied by the inverse concentration function. The value is scaled and further normalized based on the ratio of the maximum continuous normal path length to the total path length in the graph. If the maximum normal path segment is less than 40% of the total length, the abnormal weighted factor penalty mechanism is triggered, and the final fracture density is multiplied by a coefficient of 1.3 for correction. The final fracture density value is expressed as a comprehensive expression of the proportion of fracture structures and the degree of continuity loss in the graph. The higher the fracture density, the more serious the interruption of the paragraph event flow in the process of logical advancement. The value is finally output together with the graph structure number to form a logical fracture density sequence table for subsequent context entropy value calculation and immediate logical state split recognition operation.
[0100] The fracture density calculation operation is performed on the basis of the established single-chain logical fracture timeline directed graph. The operation takes the directed edge set marked as fracture edge in the graph structure as the core input. First, the total number of nodes N and the number of fracture edges B in the graph are counted. Then, the node span evaluation operation is performed on each fracture edge, that is, the distance D between the starting node and the ending node connected by the fracture edge on the path in the graph is calculated, and the D values of all fracture edges are summed to form the total fracture span value S. Then, for each fracture node, the local fracture edge density within the third-order adjacent range around it is calculated, which is recorded as the local fracture density vector Ld=[b1,b2,...,bn], where b i represents the ratio of the number of third-order internal fracture edges around the i-th fracture node to the total number of edges. The Pearson correlation between this vector and the edge density distribution vector of the entire graph is then calculated. If the correlation coefficient is lower than 0.35, it indicates that the fracture distribution is uneven, and a penalty factor coefficient α=1.2 needs to be introduced to perform weighted correction on the overall fracture density. The final fracture density value is calculated as B / N×α×log(S+1). The fracture density results corresponding to all segments are organized into a vector M=[m1,m2,...,mn] for subsequent logical disorder entropy value calculation to ensure that each fracture density value includes the influence of path length, fracture aggregation degree, and structural distribution consistency. The obtained logical fracture density vector M performs logical disorder entropy calculation on the single-chain logical fracture timeline directed graph. The calculation is based on the event path sequence in the graph. First, each event path is sequenced and the node events in all legal event paths are numbered in order according to the order of the time adverbials in the sentence to form a time sequence T=[t1,t2,...,tk], where tk represents the position number of the kth event on the time axis relative to the entire article. At the same time, all the broken edge node sets are extracted from the directed graph, and their relative positions in the sequence T are marked as the break point set D=[d1,d2,...,dn]. Then, the entropy value is calculated based on the distribution discreteness of the break point set in the sequence T. Operation, using the Shannon entropy calculation formula to take the probability of each break point in the sequence as input to calculate the sequence entropy value H of the overall logical path. The H value represents the degree of uncertainty of the break in the event advancement path. The H value is further adjusted together with the break density value mi to form the entropy weight adjustment factor θ=H×mi. For each paragraph directed graph, its θ value is calculated separately and summarized to form a logical disorder entropy value sequence θ=[θ1,θ2,...,θn]. If the θ value of a segment exceeds the mean of the sequence plus 1.5 times the standard deviation, it is marked as a high logical disorder segment. Finally, each paragraph graph corresponds to an entropy value record. Its logic is that the break distribution entropy and the structural break density jointly determine the disorder degree of the logical advancement of events in the segment.Based on the above-obtained logical fracture density sequence M and logical disorder entropy value sequence θ, an immediate event logical state split recognition operation is performed. During the recognition process, the single-chain logical fracture timeline directed graph structure corresponding to all paragraphs with logical disorder entropy values higher than the threshold is first extracted. Then, all events containing "state change" verbs in all nodes in the graph are marked and screened. The state change verb set includes words such as become, remain, stay, appear, feel, turn, get, etc. This type of word represents an explicit change in the state of an entity or situation. In the graph structure, this type of node is marked as a state node. Subsequently, a continuity analysis operation is performed on the path between any two adjacent state nodes. If there is a path with less than 3 nodes and no process action verb nodes in the middle (such as go, m ove, do, arrive, etc.), it is judged as a direct state jump path. Combined with the number of broken edges contained in the path, if the path contains at least one broken edge and the overall path entropy contribution value is greater than the average entropy value of the paragraph, it is marked as a state split path. The state split path is judged by the difference between the logical semantic vectors of the event description. If the inner product of the previous and next event description vectors is less than 0.15, the split judgment strength is further strengthened. Finally, all state jump segments in all paths that meet the above three conditions are recorded as immediate event logic state split data. Each split path record contains parameters such as the previous and next state event subjects, verbs, sentence order numbers, connecting logical words, path length, fracture density weights, and local entropy contribution coefficients, which are used for subsequent logical relationship deviation analysis and context fuzziness modeling calls.
[0101] Step S24 includes the following steps:
[0102] Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data;
[0103] Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse;
[0104] Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability;
[0105] Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data;
[0106] Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data;
[0107] Step S246: quantifying the context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity quantification data.
[0108] In an embodiment of the present invention, a causal fault analysis is performed on the English writing content text based on the content logic relationship deviation clustering data obtained in the preceding step S23. In the operation, a set of deviation fragments marked as "logical jump class" is first extracted from the clustering data, and the sentence index range in the corresponding paragraph of each deviation fragment is located. Then, a sentence causal relationship mapping structure diagram is constructed in the paragraph. The structure diagram uses each sentence as a node, and the edges between sentences indicate the existence of clear causal logic. The identification of causal edges is based on the joint judgment of conjunction rule matching and sentence meaning relationship analysis. The conjunction rule set includes because, so, therefore, as a result, since, due To, consequently, hence, and other conjunctions or connecting structures that explicitly indicate causal relationships. The sentence meaning relationship analysis adopts the subject-predicate structure analysis method within the sentence to determine whether the action of the latter sentence is directly triggered by the subject action of the previous sentence. The judgment is completed by determining whether the result of the semantic verb of the subject action and the prerequisite of the latter action form a causal connection boundary in the concept network. If a sentence has no obvious causal conjunction with its previous sentence, and the subject behavior logic between the sentence meanings does not have a triggering relationship, it is marked as a causal fault point. This point records its paragraph number, sentence number, conjunction missing status, subject-predicate structure vocabulary, and logical jump mark, which together constitute the content causal relationship fault data. This step requires relying on the syntactic dependency tree structure parsing results and the verb trigger-type semantic comparison function to compare the relationship between sentences one by one. The semantic comparison function used is constructed based on the trigger-result pair rule defined in the verb semantic role labeling tool.After obtaining the content causal relationship fault data, the text is analyzed for reference logic disorder and the pronoun ambiguity frequency is calculated. The specific operation is to first extract all pronoun examples from the sentence segments located by the fault data. The target pronoun range includes subjective and objective personal pronouns (he, she, they, it, him, her, them), and demonstrative pronouns (this, that, these, those). When each pronoun is extracted, its sentence number, position within the sentence, grammatical role and context coreference target candidate set are recorded. Subsequently, based on the coreference resolution algorithm, each pronoun is tried to match the noun phrase entity in the previous text. The entity matching process prioritizes the shortest distance. Semantic consistency and grammatical role matching are prioritized as hierarchical matching strategies. The candidate coreference entities for each pronoun match are sorted by confidence score. If there is a preferred coreference item with a confidence score lower than 0.5, and the pronoun is at the beginning of a sentence or the first sentence of a paragraph in the text, it is marked as an "ambiguous reference point". Each ambiguous reference point records its corresponding pronoun, sentence position number, candidate confidence sequence, grammatical role, whether there are similar entities in the two sentences before and after the text, and other information. Then, for each paragraph of text, the ratio of the number of ambiguous reference points to the total number of sentences is counted and defined as the frequency of pronoun misuse. This frequency is included in the text context ambiguity index system, and the final output is a pronoun misuse frequency mapping table corresponding to the paragraph number.
[0109] Based on the aforementioned causal fault data, a word meaning confusion probability estimation and analysis operation is carried out. This step is based on the keyword phrases within the sentences in the fault paragraphs, and semantic polysemous words in noun and verb phrases are extracted to calculate the ambiguity probability. The specific process is to first use the part-of-speech tagging system to tag each sentence, and select all the core verbs and the noun components of the sentence subject and object to form a context phrase pair. Then, the central word in the phrase pair is called by the semantic dictionary to count its interpretation number. If the interpretation number is greater than 3, it is judged to be a highly polysemous term. Subsequently, the context sliding window method is used to construct the context sentence vectors before and after each term. The vector dimension adopts a 300-dimensional word embedding model, and the extraction window size is set to two sentences before and after. On this basis, the context vector and the word meaning vector center are calculated. The cosine similarity between them is used to score the degree of adaptation of each interpretation, and the difference between the maximum similarity gap and the second largest value is selected as the ambiguity identification distance indicator. If the distance is less than 0.15, it is considered that the word has a high confusion probability in the context. The confusion probability is normalized by the standard deviation of the cosine similarity distribution between all interpretations, and the normalized value is multiplied by the word sense number weight factor to obtain the final word sense confusion probability value. All high confusion probability terms are recorded and a word sense confusion information table consisting of fields such as term position index, grammatical role, number of interpretations, contextual semantic offset value, and final confusion probability is generated. Finally, the average confusion probability value of each paragraph is counted to form a content word sense confusion probability vector, which is used for subsequent semantic focus disorder and context ambiguity modeling operations.
[0110] The specific processing method for content context framework disorder regression analysis based on content word meaning confusion probability, pronoun misuse frequency and content causal relationship fault data is as follows: first, a sentence order and semantic mapping matrix is constructed for each English writing content text. The matrix is numbered with sentence order as the horizontal axis and the vertical axis is the semantic centroid vector of the corresponding sentence. The semantic centroid vector is extracted using the context aggregation word embedding mean processing method, and the dimension is set to 300 dimensions. Each sentence vector is represented by the mean of the content word vector after removing stop words. Then, each sentence is arranged in the order of the original text to form a sentence semantic flow vector group S=[v1,v2,...,vn]. Then, the content causal relationship fault data obtained in the previous step are combined to locate all sentence pair number intervals with logical jump points, and regression slope mutation detection is performed within the interval. A five-point sliding window is used for local linear Fitting, calculate the main direction of the mean of the semantic centroid vector of each five sentences, and judge whether the semantic flow direction has a sudden change by comparing the angle of the fitting slope vectors of adjacent windows. If the angle exceeds 45 degrees, it is marked as a local frame disorder node. Then, the mixed disturbance coefficient β is established by combining the word meaning confusion probability and the pronoun misuse frequency in each sentence segment. This coefficient is defined as the weighted average of the mean word meaning confusion probability and the pronoun misuse frequency of the segment, with weights set to 0.6 and 0.4. The disturbance coefficient β is calculated at all marked local frame disorder nodes and compared with the average semantic consistency value of the adjacent area. If the ratio exceeds 1.5 times the standard deviation, it is finally marked as a context frame disorder point. The paragraph number, sentence order position, semantic slope jump angle, disturbance coefficient, and semantic flow residual value of the point are recorded, and finally a context frame disorder data table is formed for the next stage to call. After obtaining the disordered data of the context framework, the core subject concept offset perception operation is performed. This step uses the standard subject concept vocabulary under the style category as a reference to identify the full text subject concept. The specific process is to first extract the corresponding subject concept word set according to the content style category. The set is formed by manual expert annotation. Each concept word contains the corresponding word meaning and related extended word family. Then, the subject word scan is performed in each text. Each sentence is matched with the word in the set and its occurrence frequency and semantic similarity value are recorded. The semantic similarity is calculated using the cosine similarity method. The word embedding mean vector in the sentence is calculated with the concept word vector. If the similarity is greater than 0, the semantic similarity is calculated. .65 is recorded as a main theme coverage. The main theme concept density of the entire text is defined as the ratio of all main theme coverage times to the number of sentences. Then, a local main theme sparse analysis is performed on the main theme coverage density of all paragraphs where the frame disorder points are located. If the density is lower than 0.6 times the average density of the entire text, and the change in the main theme term frequency between the two paragraphs exceeds 50%, it is marked as a main theme concept offset point. The main theme coverage missing value of this point and the main theme term distribution offset in the adjacent area together constitute the main theme concept offset degree. The final output main theme concept offset degree data includes fields such as the offset segment number, the main theme density within the segment, the main theme change value of the adjacent paragraphs, and the number of main theme coverage interruption sentences.After obtaining the context frame disorder data and the main concept deviation degree data, the context fuzziness degree quantification operation is performed. This step establishes a multi-factor fuzziness aggregation function based on the sentence segment. The function input is three sets of core parameters, namely the context frame disorder index F, the main concept deviation intensity value P and the semantic coherence attenuation coefficient D. The semantic coherence attenuation coefficient is obtained by normalizing the semantic cosine similarity reduction between the two sentences before and after the context frame disorder point. The specific calculation method is to extract the semantic embedding vector means of the two sentences before and after, which are A and B respectively, and calculate the cosine similarity S=cos (A, B), and then normalize the average semantic similarity of the entire text to form a difference index D. The final fuzziness level M = F × 0.5 + P × 0.3 + D × 0.2, which is defined as the context fuzziness level of the current paragraph. This value is calculated for all paragraphs separately and summarized to form a fuzziness distribution vector. The fuzziness level quantification data structure includes five fields: paragraph number, context framework disorder, main concept offset value, semantic coherence attenuation value, and fuzziness level aggregation value. A complete fuzziness level quantification map is constructed for each text as input for subsequent language proficiency assessment architecture learning module calls.
[0111] Step S244 includes the following steps:
[0112] Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content;
[0113] Quantify the degree of semantic focus disorder in content subjective / objective context ambiguity data;
[0114] Calculate the focus shift frequency variance in the semantic focus disorder degree;
[0115] Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, content rhetoric dislocation analysis was conducted to obtain content semantic rhetoric dislocation data;
[0116] The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus shift frequency variance to obtain the focus chain break ratio;
[0117] Based on the variance of focus shift frequency and the ratio of focus chain breaks, a regression analysis of content context frame disorder was conducted to obtain context frame disorder data.
[0118] In the embodiment of the present application, the content subjective context and objective context fuzzy analysis is performed according to the content word sense confusion probability, pronoun misuse frequency and content causality fault data, and the specific operation is as follows: first, the paragraph area with logical jump or causality logical chain rupture is extracted according to the sentence segment number marked in the content causality fault data, and the context fuzziness score is calculated by combining the word sense confusion probability value and the pronoun misuse frequency of each sentence in the area. The score calculation process first extracts the first person and emotional word coverage frequency in each sentence as the subjectivity marker index according to the sentence unit, matches the distribution position by using the emotional word library composed of adjectives and adverbs, and the subjective context fuzziness is defined as the weighted aggregation of the word sense confusion probability mean value and the first person and emotional word proportion, wherein the word sense confusion probability weight is 0.5, the first person frequency proportion weight is 0.3, and the emotional word coverage degree weight is 0.2. Similarly, in the objective analysis, after excluding the sentences with subjective expression, the logical chain density of technical verbs and fact description phrases is counted, the objective context fuzziness is defined as the reciprocal value of the causality chain node density and the inter-sentence logical coherence, the causality chain node density is obtained by the ratio of the number of logical chains in each paragraph to the total number of sentences, and the inter-sentence logical coherence is expressed by the mean value of the cosine similarity of adjacent sentence vectors. Finally, the subjective and objective context fuzzy data of all paragraphs are respectively retained as the subjective fuzzy score and the objective fuzzy score, and are associated with the paragraph position and the word statistical value to form the subjective and objective context fuzzy data set structure. In order to quantify the semantic focus disorder degree in the above-mentioned subjective and objective context fuzzy data, the operation process is as follows: taking the paragraph as the minimum analysis unit, a focus word recognition mechanism is established in each sentence, the subject and object in each sentence are extracted as candidate focus entities through subject-predicate-object structure analysis, and the semantic core word items are retained through TF-IDF value and word vector center clustering, a focus word sequence flow F=[f1, f2,..., fn] is constructed for each paragraph, the sequence is used to analyze the flow path of the semantic focus in the paragraph, and then it is counted whether the focus jumps in the next sentence if the subject or object in the sentence does not reappear, and if the focus jumps more than twice, it is marked as a focus rupture point, and the focus retention rate in the paragraph is calculated, which is the ratio of the number of repeated focus words to the total number of focus words. The lower the focus retention rate, the higher the focus disorder degree. Then, the focus disorder index is generated by combining the subjective and objective context fuzzy scores, which is defined as the aggregation value of the subjective fuzzy score x 0.4 + the objective fuzzy score x 0.3 + the focus retention rate x 0.3. Finally, the focus disorder index is extracted from all paragraphs to form the semantic focus disorder degree data set, and the data fields include the segment number, the focus word change path, the rupture position distribution, the focus retention rate and the focus disorder index.In order to calculate the focus shift frequency variance in the degree of semantic focus disorder, the specific operation is to number and encode the focus word transfer trajectory of each paragraph in each text, encode each focus change event as a jump, and construct a complete focus shift sequence vector G=[g1,g2,...,gm], where gi is the relative position difference of the i-th focus change, and the unit is set as the sentence sequence distance. The variance statistical analysis of the sequence vector is performed to calculate the focus shift frequency variance. If the focus jump frequency in an article is highly concentrated, the variance is small. If the jump frequency is high in the paragraph, the variance is small. Uneven distribution of focus or concentration in local areas will show high variance values. To prevent deviations in short texts, focus sequence normalization is performed on texts with a length of less than 200 words, that is, the focus jump interval is divided by the total number of sentences to maintain scale consistency. Finally, the focus shift frequency variance and the focus disorder index are jointly output to form a semantic focus shift stability dataset, in which fields include segment number, focus sequence, focus change distance mean, focus change variance, standardized variance value, etc., which are used for further aggregation into the context frame disorder indicator in subsequent regression analysis. Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, the content rhetorical dislocation analysis is carried out. The specific operation is to use the degree of semantic focus disorder as a measure of the stability of focus transfer within a paragraph, cross-match the causal chain missing area marked in the causal relationship fault data with the focus disorder area and construct a dislocation segment mapping matrix, number each paragraph of text according to sentence order to form a sentence vector sequence, and quantitatively represent the semantic distance between the subject and the object in each pair of adjacent sentences. The cosine similarity of the word vector is used to calculate the subject semantic continuity and object pointing consistency. If the angle between the subject vectors of adjacent sentences is greater than 1.2rad and the object semantic deviation is greater than 0. The verb logical framework of the previous sentence is marked as a rhetorical jump point, and the word meaning confusion probability threshold of the paragraph where the jump point is located is further set to 0.4. Words exceeding this threshold are identified as semantic rhetorical breakpoints when they appear in the subject or object position. The ratio of the number of breakpoints to the total number of sentences in each paragraph is recorded as the rhetorical dislocation rate of the paragraph. At the same time, the focus words before and after the jump point are marked, and a rhetorical focus migration table is constructed to record the semantic path mutation. Finally, the rhetorical dislocation rate, jump point distribution, focus deviation angle, and word meaning confusion weight of each paragraph are extracted to form the content semantic rhetorical dislocation data. The data fields include paragraph number, breakpoint position, break type, focus semantic change path and rhetorical offset weight sequence.
[0119] The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus transfer frequency variance. The specific operation is to first construct a directed linked list structure based on the inter-sentence recurrence of the focus word sequence F=[f1,f2,…,fn] in each text. If a focus word does not appear three times in a row in the linked list, it is marked as a broken link point. The ratio of the total number of break points in all focus chains to the total length of the linked list is recorded as the initial value of the focus chain break. At the same time, the threshold T is set with reference to the focus transfer frequency variance. In the experiment, the sample focus variance with an average of 1.5 is taken. If the current If the text focus variance is greater than T, it is considered that the break is not a random transfer but a semantic chain disconnection. The break count weight is recalibrated and the broken chain points are screened again. The broken chain points of all paragraphs are accumulated and then the ratio with the total number of chains is used to form the final focus chain break ratio. The focus chain break ratio calculation formula includes the rhetorical jump density factor R, which is defined as the product of the break point frequency and the rhetorical dislocation rate in a unit paragraph. The break ratio field output is the segment number, broken chain point position, average chain length, variance level and ratio result, which constitute the focus chain break ratio retrieval result set. In order to conduct a regression analysis of content context frame disorder based on the focus shift frequency variance and focus chain break ratio, the specific operation is to establish a multivariate linear regression structure, set the context frame disorder degree Y as the target variable, introduce the focus shift frequency variance X1 and the focus chain break ratio X2 as explanatory variables, and initialize the regression weight parameters to 0.5 and 0.5 respectively. The manually annotated context stability level L in 50 training sample texts is selected as the target control, and the X1 and X2 values of each text are substituted into the regression equation to optimize the least squares error and adjust the regression parameters. Finally, the regression coefficient is stabilized at X1 weight 0.6 and X2 weight 0.4. The corresponding paragraph is input into the new sample text. The X1 and X2 values are substituted into the regression function, and the Y value is calculated and then compared with the level interval in the training set to identify the level of context disorder in this paragraph. The higher the Y value, the more chaotic the frame. The context frame disorder data is output in paragraph units. The fields include paragraph number, X1 value, X2 value, regression prediction value Y, context disorder level label and participation factor contribution evaluation, which are used as key indicator basis for subsequent main concept deviation degree analysis and context ambiguity degree quantification operation. Among them, the level interval in the training set, the Y value range Y∈[0,0.3) corresponds to "stable", Y∈[0.3,0.7) corresponds to "mild disorder", and Y∈[0.7,1.0] corresponds to "severe disorder".
[0120] In this embodiment, reference Figure 3 The above is a flowchart of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include:
[0121] Step S31: performing logic learning on the context ambiguity degree quantified data to obtain context ambiguity degree learning data;
[0122] Step S32: performing feature sampling on the context fuzziness degree learning data to obtain fuzziness degree learning feature sampling data;
[0123] Step S33: Designing an English language proficiency assessment framework based on the fuzzy degree learning feature sampling data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework;
[0124] Step S34: Send the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
[0125] In the embodiment of the present invention, the quantitative data of context ambiguity is subjected to logic learning processing. In this process, the main concept deviation degree, semantic rhetoric dislocation frequency, focus chain break ratio and reference ambiguity weight in the quantitative data of context ambiguity are used as core variables to construct a 4-dimensional fuzzy feature vector set. ,in is the frequency of context focus jump, is the number of causal fault levels, is the semantic confusion rate, To refer to conflicting values, a co-occurrence matrix between fuzzy variables is constructed using logical reduction method. In the same text and The frequency of co-occurrence is calculated, and the principal component normalization is performed after constructing the M matrix. The collinear interference value is eliminated and the feature group with a feature synergy factor greater than 0.7 is retained. The feature importance weighting mechanism based on structural entropy is used to perform discrete value Boolean learning on each group of co-occurrence factors to form a feature pattern sequence set. , each set of sequences represents a typical fuzzy structure configuration, thus constructing a contextual fuzzy degree logic learning sample set , and use this set in the subsequent feature sampling task. Feature sampling is performed on the context fuzziness learning data. In this stage, the stratified multi-segment sampling method is used to sample the logical learning sample set. The extraction is performed paragraph by paragraph. First, the sample set is divided into low ambiguity group, medium ambiguity group, and high ambiguity group according to the degree of subject deviation. Each group is further divided into intervals according to the focus jump density. 20% of the paragraphs in each interval are randomly selected as sampling samples. After sampling, the feature vectors in each paragraph are maximum entropy encoded to construct a feature training set D with context ambiguity level as the label. Each data structure is ⟨feature group sequence, ambiguity level label>, where the ambiguity level label is divided into 5 levels from 1 to 5. There are 600 samples in total, 120 in each group. The training set is normalized by standard deviation, with the mean set to 0 and the standard deviation set to 1, forming a fuzzy degree learning feature sampling dataset. , the dataset contains fields such as focus chain break ratio, jump point frequency, ambiguity density, logical offset value, fuzzy level label and other 5-dimensional structures. Based on the Q-learning algorithm, the English language proficiency assessment architecture is designed for the fuzzy degree learning feature sampling data. In this stage, the Q-learning state space S and action space A are constructed, where the state space S is the 5-dimensional context feature combination in the sampling data, and the action space A is the corresponding ability assessment level division action set {A1, A2, A3, A4, A5}. Initially, the Q value matrix is set to a full zero matrix, and the ε-greedy strategy is used to control the exploration rate, where the initial value of ε is set to 0.2, the discount factor γ is set to 0.9, and the learning rate α is set to 0.1. In the iterative process, the state space S is set to 0.2, the discount factor γ is set to 0.9, and the learning rate α is set to 0.1. Extract the current context fuzzy feature combination and select the action The evaluation grade assigned to the current state is used to calculate the immediate feedback value based on the difference between the evaluation grade of the current action and the manually labeled grade in the training set. As a reward signal, The update strategy is used to modify the Q value matrix. Each round of iteration is updated 100 times. The total number of training rounds is set to 500 rounds. The final convergence condition is that the change of the Q value matrix is less than After completion, the final Q-value matrix is extracted and a Q-strategy function Q: S→A is constructed, mapping the state input to the language proficiency level. This completes the design and output of the English language proficiency assessment architecture. To send the English language proficiency assessment architecture to the terminal for execution, this step first quantifies the contextual ambiguity of the user-uploaded text, obtains a 5-dimensional feature combination as the current state input, and inputs it into the trained Q-strategy function Q: S→A. The corresponding language proficiency level output is extracted based on the action index corresponding to the maximum Q-value in the Q-value matrix corresponding to the current state. The language proficiency level is divided into five levels, from A1 to A5, where A1 represents extremely low context stability and A5 represents extremely high context stability. The output level is accompanied by the Q-value weight that matches the current input state vector to the level for evaluation and interpretation. The level is finally presented through the terminal interface.
[0126] The present invention also provides an English language proficiency assessment system for executing the above-mentioned English language proficiency assessment method, the English language proficiency assessment system comprising:
[0127] The category sentence structure analysis module is used to collect English writing content texts uploaded by students through the terminal; analyze the style category of the English writing content texts uploaded by students, and then perform category sentence structure analysis to obtain the content style category sentence structure;
[0128] The context ambiguity degree quantification module is used to identify logical relationship deviations in English writing content based on content style category sentence structure to obtain content logical relationship deviation clustering data; based on the content logical relationship deviation clustering data, the context ambiguity degree is quantified to obtain context ambiguity degree quantitative data;
[0129] The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
[0130] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating English language proficiency, characterized in that: The following steps are involved: Step S1: collecting English writing content texts uploaded by students through the terminal; Analyze the stylistic categories of the English writing content uploaded by students, and then conduct category sentence structure analysis to obtain the content stylistic category sentence structure; Wherein, step S2 includes: Step S21: performing a content tense confusion structure analysis on the English writing content text based on the content style category sentence structure to obtain the content text tense confusion structure; Step S22: identifying logical relationship deviations of the English writing content text according to the tense disorder structure of the content text, thereby obtaining content logical relationship deviation data; Step S23: performing deviation clustering processing on the content logic relationship deviation data to obtain content logic relationship deviation clustering data; Wherein, step S24 is specifically as follows: Step S241: performing a causal relationship fault analysis on the English writing content text based on the content logical relationship deviation clustering data to obtain content causal relationship fault data; Step S242: performing a reference logic disorder analysis on the English writing content text based on the content causal relationship fault data, and calculating the frequency of pronoun reference ambiguity, thereby obtaining the frequency of pronoun misuse; Step S243: estimating the word meaning confusion probability of the English writing content text based on the content causal relationship fault data to obtain the content word meaning confusion probability; Step S244: performing content context frame disorder regression analysis based on the content word meaning confusion probability, pronoun misuse frequency, and content causal relationship fault data to obtain context frame disorder data; Step S245: performing core subject concept shift perception based on the contextual framework disorder data to generate subject concept shift degree data; Step S246: quantifying the context ambiguity of the English writing content text according to the context frame disorder data and the main concept deviation degree data, thereby obtaining context ambiguity degree quantification data; Step S3: Designing an English language proficiency assessment framework for the context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; and sending the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
2. The English language proficiency assessment method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting English writing content texts uploaded by students through the terminal; Step S12: correcting spelling errors in the English writing content uploaded by the student to obtain a corrected English writing content text; Step S13: performing a stylistic analysis on the revised English writing content to obtain a content stylistic category; Step S14: performing a category sentence structure analysis on the revised English writing content text according to the content style category to obtain the content style category sentence structure.
3. The English language proficiency assessment method according to claim 1, characterized in that: Step S22 includes the following steps: Step S221: extracting tense misuse / jump / missing structures in the tense confusion structure of the content text; Step S222: Detecting the English writing content text for paragraphs with invalid tense correspondence based on tense misuse / jump / missing structures, thereby obtaining paragraphs with invalid tense correspondence; Step S223: performing content time axis conflict analysis on the paragraphs with invalid tense correspondence, and obtaining content time axis conflict data between the paragraphs with invalid tense correspondence; Step S224: Identify the logic state fragmentation of the instant event based on the content time axis conflict data and the tense misuse / jump / missing structure between the tense echo failure paragraphs to obtain the logic state fragmentation data of the instant event; Step S225: performing logical relationship deviation identification on the English writing content text according to the instant event logical state segmentation data, thereby obtaining content logical relationship deviation data.
4. The English language proficiency assessment method according to claim 3, characterized in that: Step S224 includes the following steps: Perform content paragraph single-chain logical break analysis on the content timeline conflict data between paragraphs with failed temporal correspondence, and obtain a content paragraph single-chain logical break dataset; Based on the single-strand logical break dataset of content paragraphs and the temporal misuse / jump / missing structure, a single-strand logical break timeline directed graph is constructed. Perform break density calculation on the single-chain logical break time axis directed graph to obtain the logical break density; The logical disorder entropy value of the single-chain logical break time axis directed graph is calculated based on the logical break density to obtain the content logical disorder entropy value; Based on the logic fracture density and content logic disorder entropy, the logic state splitting of instant events is identified to obtain the logic state splitting data of instant events.
5. The English language proficiency assessment method according to claim 1, wherein: Step S244 includes the following steps: Based on the probability of confusion of content meanings, the frequency of misuse of pronouns and the data of content causal relationship fault, the subjective / objective context fuzziness analysis of the content was conducted to obtain the subjective / objective context fuzziness data of the content; Quantify the degree of semantic focus disorder in content subjective / objective context ambiguity data; Calculate the focus shift frequency variance in the semantic focus disorder degree; Based on the degree of semantic focus disorder, content causal relationship fault data and the probability of content word meaning confusion, content rhetoric dislocation analysis was conducted to obtain content semantic rhetoric dislocation data; The focus chain break ratio is retrieved and calculated based on the content semantic rhetoric dislocation data and the focus shift frequency variance to obtain the focus chain break ratio; Based on the variance of focus shift frequency and the ratio of focus chain breaks, a regression analysis of content context frame disorder was conducted to obtain context frame disorder data.
6. The English language proficiency assessment method according to claim 1, wherein: Step S3 includes the following steps: Step S31: performing logic learning on the context ambiguity degree quantified data to obtain context ambiguity degree learning data; Step S32: performing feature sampling on the context fuzziness degree learning data to obtain fuzziness degree learning feature sampling data; Step S33: Designing an English language proficiency assessment framework based on the fuzzy degree learning feature sampling data based on the Q-learning algorithm, thereby obtaining an English language proficiency assessment framework; Step S34: Send the English language proficiency assessment framework to the terminal to perform English language proficiency assessment.
7. An English language proficiency assessment system, characterized in that: For executing the English language proficiency assessment method according to claim 1, the English language proficiency assessment system comprises: The category sentence structure analysis module is used to collect English writing content texts uploaded by students through the terminal; analyze the style category of the English writing content texts uploaded by students, and then perform category sentence structure analysis to obtain the content style category sentence structure; The context ambiguity degree quantification module is used to identify logical relationship deviations in English writing content based on content style category sentence structure to obtain content logical relationship deviation clustering data; and to quantify the context ambiguity degree based on the content logical relationship deviation clustering data to obtain context ambiguity degree quantitative data; The evaluation architecture design module is used to design an English language proficiency evaluation architecture for context ambiguity quantification data based on the Q-learning algorithm, thereby obtaining an English language proficiency evaluation architecture; and sending the English language proficiency evaluation architecture to the terminal to perform English language proficiency evaluation.
Citation Information
Patent Citations
Language evaluation generation method based on reinforcement learning
CN110532555A
Multi-dimensional English composition scoring method and device and readable storage medium
CN113836894A