An intelligent rewriting method for reading comprehension proposition materials based on large models
Through intelligent rewriting methods based on big models, the core theme and context information of English reading comprehension topic materials are automatically identified and simplified, and the problem of inefficiency of traditional manual operations is solved, high-quality and diverse proposition generation is achieved, and personalized feedback and dynamic model adjustment capabilities are improved.
Patent Information
- Application Number
- CN202411565207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Traditional English reading comprehension propositions and material revision rely on manual operations, which are inefficient and have repetitive problems, making it difficult to ensure the diversity and innovation of content, and lacks personalized feedback and dynamic model adjustment capabilities.
Using a large model-based intelligent rewriting method, a text representation model is constructed by preprocessing English chapters, a GPT big model is used to automatically identify core themes and context information, vocabulary, sentence structure and grammatical complexity is evaluated, and adaptive simplification and rewrite according to the junior high school English syllabus and the middle school entrance examination question standards.
It improves teaching efficiency, ensures the quality and diversity of generated content, avoids the problems of semantic breakage and logic loss in the simplification process of traditional tools, and achieves highly personalized proposition generation.
Smart Images

Figure CN119514554B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of English, and particularly to an intelligent rewriting method for reading comprehension proposition materials based on a large model. Background Art
[0002] With the rapid development of artificial intelligence and natural language processing technologies, the application of intelligent tools in the education field has gradually become popular. In English teaching, reading comprehension, as an important link for cultivating students' language application ability and thinking ability, has become the focus of attention of teachers and students. However, in the traditional teaching process, the proposition, rewriting, and evaluation of English reading comprehension mainly rely on manual operations by teachers, which not only has low efficiency but also has various limitations.
[0003] In the prior art, the proposition and material rewriting of English reading comprehension are usually completed manually by teachers or teaching and research teams. Teachers need to select appropriate article materials according to the curriculum standards and manually design or adjust the vocabulary, sentence patterns, and grammar of the test questions to make them meet the comprehension levels and examination requirements of students in different grades. Although this method can ensure the consistency of test question quality and teaching objectives to a certain extent, it also exposes significant problems in practical applications. First of all, the workload of manually designing proposition materials is huge. When adapting materials with different themes and difficulties, teachers often need to spend a lot of time and it is difficult to continue to be efficient under heavy teaching tasks. Secondly, due to manual writing, repetitive problems are likely to occur, and the proposition content may lack diversity and innovation, making it difficult to stimulate students' learning interests. In addition, the personal subjective judgment of teachers during proposition may lead to inconsistencies in vocabulary, sentence patterns, and grammar, and it is impossible to strictly ensure that all test questions meet the high school entrance examination proposition standards.
[0004] In recent years, some auxiliary education systems have begun to introduce intelligent tools to adjust English reading comprehension propositions and materials. Auxiliary education tools usually adapt the content based on keyword matching and syntactic analysis. However, such tools also have many deficiencies in practical applications. On the one hand, the system based on keyword matching can only perform surface replacement and cannot deeply understand the core theme and context logic of the materials, easily resulting in the lack of coherence or deviation from teaching objectives in the adapted content. On the other hand, the system usually lacks the adaptive adjustment ability for the junior high school English teaching syllabus and high school entrance examination standards and cannot generate personalized test question versions according to the needs of different students. In addition, the grammar and sentence pattern rewriting modules of existing systems have limited capabilities, are difficult to handle complex clause structures and advanced grammar expressions, and lack a perfect evaluation mechanism for the quality of the generated content. Summary of the Invention
[0005] An object of the present invention is to propose an intelligent rewriting method for reading comprehension proposition materials based on a large model, which improves teaching efficiency and ensures the quality and diversity of the generated content.
[0006] An intelligent rewriting method for reading comprehension proposition materials based on a large model according to an embodiment of the present invention includes the following steps:
[0007] S1. Receive an English passage input by a user, where the English passage includes clause structures, advanced vocabulary, and grammar expressions beyond the cognitive level of junior high school students;
[0008] S2. Preprocess the English passage, perform word segmentation, syntactic analysis, and context structure parsing, and construct a text representation model including lexical information, syntactic structure information, and context theme information;
[0009] S3. Use the pre-trained GPT large model to automatically identify and extract the core theme representation model and context information set of the English passage based on the constructed text representation model;
[0010] S4. Based on the core theme representation model and context information set, the GPT large model automatically evaluates the vocabulary, sentence patterns, and grammar complexity of the English passage, and generates a text feature set including vocabulary difficulty, sentence pattern complexity, and grammar complexity;
[0011] S5. The GPT large model adaptively simplifies the English passage according to the junior high school English teaching syllabus and the middle school entrance examination proposition standards, replaces advanced vocabulary with common vocabulary that conforms to the junior high school English teaching syllabus based on the context information and difficulty evaluation, converts complex clause structures into simple sentences or compound sentences, and adjusts complex grammar to a grammar structure that conforms to the middle school entrance examination proposition standards based on grammar associations;
[0012] S6. The GPT large model evaluates the grammar correctness, semantic integrity, and logical coherence of the rewritten version of the English passage generated after simplification;
[0013] S7. Based on the evaluation results and the user's feedback, the feedback learning module of the GPT large model dynamically adjusts the simplification rules and generation strategies, and the GPT large model adaptively adjusts the simplification rules and generation strategies according to the type of English passage input by the user, the context relevance, and the proposition difficulty requirements;
[0014] S8. Output the English text simplified by the GPT large model and conforming to the middle school entrance examination proposition standards to the user or the education system.
[0015] Optionally, the S1 specifically includes:
[0016] S11. Receive an English passage T input by a user, where the English passage includes multiple sentences S i The sentence set S = {S1, S2,..., S n}, and each sentence S i includes clause structures, advanced vocabulary, and grammar expressions beyond the cognitive level of junior high school students;
[0017] S12. For each sentence S in the input English passage T i perform syntactic analysis, separate the main clause and subordinate clauses, and construct a set of subordinate clause structures:
[0018] C i = {c i1 , c i2 , …, c im};
[0019] where c ij represents the j-th subordinate clause in sentence S i ;
[0020] S13. Extract a set of advanced vocabulary from each sentence S i :
[0021] V i = {v i1 , v i2 , …, v ik};
[0022] where v ij represents the j-th advanced vocabulary in sentence S i , and advanced vocabulary refers to words that do not match the junior high school English syllabus or have a low word frequency;
[0023] S14. According to the syntactic structure and vocabulary complexity of each sentence S i , construct a set of grammatical expressions beyond the cognitive level:
[0024] G i = {g i1 , g i2 , …, g il};
[0025] where g ij represents that the j-th grammatical structure in sentence S i is beyond the cognitive level of junior high school students;
[0026] S15. Merge the structures, vocabulary, and grammatical features of all sentences in the received English passage T to construct an English passage matrix M T :
[0027] Sentence number Number of subordinate clauses Number of advanced vocabulary Number of grammatical structures beyond the level
[0028]
[0029] where m i , k i , l i respectively represent sentence S iThe number of clauses, the number of advanced vocabulary, and the number of grammatical structures beyond the cognitive level of students.
[0030] Optionally, the S2 specifically includes:
[0031] S21. Perform word segmentation on the received English passage T = {S1, S2,..., S n} to decompose each sentence S i into a sequence of words:
[0032]
[0033] where L i represents the total number of words in sentence S i , and w ij represents the j-th word in sentence S i . Perform part-of-speech tagging and lexical attribute extraction on each word w ij to obtain a lexical feature vector:
[0034] F ij = [pos(w ij ), lemma(w ij ), freq(w ij ), complexity(w ij )];
[0035] where pos(w ij ) is the part of speech of word w ij , lemma(w ij ) is the prototype form of word w ij , freq(w ij ) is the frequency of occurrence of word w ij in the junior high school English syllabus, and complexity(w ij ) is the complexity score of word w ij , reflecting its difficulty for junior high school students;
[0036] S22. Perform syntactic analysis on each sentence S i to construct a grammatical dependency tree:
[0037] D i = (N i , E i );
[0038] where is the set of nodes, each node n ij corresponds to word w ij , and E i = {(n ip , n iq , r ipq)∣p, q ∈ [1, L i} is the edge set, r ipq represents the syntactic relation type from node n ip to node n iq ;
[0039] Calculate the syntactic complexity index C i (S syntax ) of sentence S i :
[0040]
[0041] where weight(r ipq ) is the complexity weight of syntactic relation r ipq , reflecting the difficulty level of the syntactic structure in the junior high school English syllabus;
[0042] S23. Perform context structure analysis on the English passage T, and use the topic model to extract the topic distribution of the English passage to construct the topic distribution vector of the English passage:
[0043] θ T = [θ1, θ2,..., θ K ;
[0044] where K is the preset number of topics, and θ k represents the weight of the English passage T on the k-th topic; at the same time, construct the topic word distribution matrix Φ = [φ kj , where φ kj represents the probability of the word w j in the k-th topic;
[0045] Based on the topic distribution vector θ T and the topic word distribution matrix Φ, construct the context topic representation vector C T ;
[0046] S24. Based on the lexical feature set F i , the syntactic dependency tree D i and the context topic representation vector C T , construct the comprehensive text representation model R T of the English passage:
[0047] R T = (F T , G T , C′ T );
[0048] where, is the lexical feature set of all sentences in the English passage, G T = (N T , ET ) is the global grammar graph of the English passage, C′ T = α1·C T + β1·Embed(T) is the weighted context topic representation vector, α1 and β1 are balance coefficients, and Embed(T) is the semantic embedding vector of the English passage T.
[0049] Optionally, the S3 specifically includes:
[0050] S31. Use the pre-trained GPT large model to process the comprehensive text representation model R of the constructed English passage T to automatically identify and extract the core theme and context information of the English passage;
[0051] S32. Calculate the theme relevance S of the English passage T based on the lexical feature set F T and the context topic representation vector C′ topic (T):
[0052]
[0053] where w j is the j-th word in the lexical feature set F T , θ T is the theme distribution vector of the English passage, and φ kj is the element in the theme-word distribution matrix, representing the probability of the word w j in the k-th theme;
[0054] S33. Use the syntactic dependency tree G T =(N T , E T ) to analyze the syntactic relationship between sentences and construct the sentence connectivity matrix L T =[l ij :
[0055]
[0056] where l ij represents the connectivity between sentence S i and sentence S j ;
[0057] S34. Extract the main idea M of the article T and the set of key information points K T ={k1, k2,..., k m} through the semantic understanding module of the GPT large model, where k i represents the i-th key information point;
[0058] S35. Combine with the main idea M T , the set of key information points K T , the theme relevance S topic (T) and the sentence connection degree matrix L T , construct the core theme representation model C core :
[0059] C core =(M T , K T , S topic (T), L T );
[0060] S36. Based on the core theme representation model C core Analyze the logical relationship, narrative structure and semantic association before and after between paragraphs of the article, and generate the context information set:
[0061]
[0062] Among them, C context represents the context information set, which is the quantitative representation result of the logical relationship and semantic association before and after between paragraphs of the article. Sim(S i , S j ) represents the semantic similarity between sentence S i and sentence S j , measures the semantic overlap degree between two sentences, and is used to capture the potential information connection between sentences in an English text. L T [i, j] is an element in the sentence connection degree matrix, representing the direct syntactic relationship strength between sentence S i and S j . S topic (S i , S j ) is the theme relevance between sentence S i and S j . Rel(S i , M T ) is the correlation function between sentence S i and the main idea M T . α2 and β2 are the balance coefficients of semantic similarity and theme relevance, and λ is the adjustment coefficient of theme relevance. is the logistic regression function, which is used to map the comprehensive result of sentence connection degree and theme relevance to the interval [0, 1].
[0063] Optionally, the S4 specifically includes:
[0064] S41. Based on the core theme representation model C core and the context information set C context , analyze the lexical features F of the textT 、The global syntax graph G of the English passage T and the context theme information C′ T ;
[0065] S42. Use the GPT large model to evaluate the lexical difficulty D vocab (T) of the English passage, and calculate the lexical complexity score for each word w j ∈F T :
[0066]
[0067] where complexity(w j ) represents the complexity of the word w j , freq(w j ) is the frequency of w j in the junior high school English syllabus, and max(freq(F T )) is the maximum frequency of all words;
[0068] S43. Calculate the sentence pattern complexity D syntax (S i ) based on the syntactic dependency tree;
[0069] S44. Use the context information set C context and the theme relevance S topic (T) to calculate the grammar complexity D grammar (T) of the English passage:
[0070]
[0071] where n is the number of sentences, and Rel(S i , M T ) represents the relevance between the sentence S i and the main idea M T , and γ and δ are the balance coefficients between the sentence pattern complexity and the relevance;
[0072] S45. Generate the text feature set F text (T) by integrating the lexical, sentence pattern, and grammar complexities:
[0073] F text (T) = (D vocab (T), D syntax (T), D grammar (T)).
[0074] Optionally, the S5 specifically includes:
[0075] S51. According to the lexical feature set F T and the text feature set Ftext (T) Identify the set V of advanced vocabulary in the English passage H ={v1, v2, …, v k}, where represents vocabulary beyond the junior high school English syllabus;
[0076] S52. Use the vocabulary replacement function Replace(v j , V C ) to replace each advanced vocabulary v j ∈V H with a common vocabulary v′ j ∈V C :
[0077]
[0078] where Sim(v j , v) is the semantic similarity between the vocabulary v j and the common vocabulary v, freq(v) is the frequency of v in the junior high school English syllabus, and V C is the set of common vocabulary;
[0079] S53. Based on the syntactic dependency tree G T identify the set C of complex clause structures in the English passage H ={c1, c2, …, c m};
[0080] S54. Use the clause simplification function Simplify(c i ) to convert each complex clause c i ∈C H into a simple sentence or a compound sentence. The set of simplified sentences is represented as:
[0081] S′={s′1, s′2, …, s′ m};
[0082] where s′ i is the simplified form of the clause c i , retaining the semantic core of the original sentence but reducing the syntactic complexity;
[0083] S55. Use the context information set C context and the sentence connectivity matrix L T to adjust the grammatical connections of the simplified sentences S' to meet the high school entrance examination proposition standards;
[0084] S56. Define the grammar adjustment function Adjust(S′) to optimize the grammatical structure of the sentences based on the grammatical connection degree and context information:
[0085]
[0086] Among them, S″ i is the optimized form of the sentence s′ i The adjustment coefficient is λ, ensuring that the optimized syntax meets the requirements of the teaching syllabus;
[0087] S57. Generate the simplified English passage T':
[0088] T′ = {S″1, S″2, …, S″ n}.
[0089] The beneficial effects of the present invention are as follows:
[0090] (1) The present invention proposes a multi-dimensional feature evaluation algorithm by combining the vocabulary feature set, syntactic dependency tree and context theme information. By comprehensively calculating the vocabulary complexity, sentence pattern complexity and grammar complexity, it can automatically adjust the material content according to the junior high school English teaching syllabus and the high school entrance examination proposition standard. Compared with the traditional keyword replacement, the present invention innovatively adopts the context relevance calculation function and the logistic regression function to accurately replace the complex syntax and advanced vocabulary in the material, converts the complex clauses into simple sentences or compound sentences and optimizes the grammar relevance, so as to ensure that the rewritten material meets the teaching standard and maintains semantic coherence, and can effectively avoid the problems of semantic break and logical loss in the simplification process of traditional tools, improving the accuracy and applicability of the generated content.
[0091] (2) The present invention designs a multi-layer adaptive proposition generation mechanism based on the core theme model. By analyzing the main idea, key information points and the grammar relevance matrix between sentences of the original text, it can generate multiple versions of proposition materials according to different educational goals, and realizes the full-process adaptive adjustment from the input text to the proposition generation by using the semantic similarity and theme relevance as adjustment parameters. Different from the common fixed-template propositions in existing tools, the present invention can flexibly generate proposition versions suitable for different difficulties and themes, realizing the diversification and innovation of proposition content.
[0092] (3) The present invention integrates a dynamic feedback and real-time optimization module, enabling the GPT large model to continuously optimize the generation effect according to user feedback in actual applications. Teachers or educational systems evaluate the generated materials and propositions during use, and the system inputs these feedbacks as training data into the model to automatically adjust the generation strategy and algorithm parameters. Through this closed-loop optimization mechanism, the present invention can gradually improve the adaptability of the model to different student groups and realize highly personalized proposition generation, solving the problems of lack of personalized feedback and model dynamic adjustment in the prior art, and providing more efficient teaching support for teachers and students. Description of the Drawings
[0093] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0094] Figure 1 It is a flowchart of an intelligent rewriting method for reading comprehension proposition materials based on a large model proposed by the present invention. Detailed implementation manners
[0095] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0096] Reference Figure 1 , an intelligent rewriting method for reading comprehension proposition materials based on a large model, includes the following steps:
[0097] S1. Receive the English passage input by the user. The English passage includes clause structures, advanced vocabulary, and grammar expressions beyond the cognitive level of junior high school students;
[0098] S2. Preprocess the English passage, perform word segmentation, syntactic analysis, and context structure parsing, and construct a text representation model containing lexical information, syntactic structure information, and context theme information;
[0099] S3. Use the pre-trained GPT large model to automatically identify and extract the core theme representation model and context information set of the English passage based on the constructed text representation model;
[0100] S4. Based on the core theme representation model and context information set, the GPT large model automatically evaluates the vocabulary, sentence patterns, and grammar complexity of the English passage, and generates a text feature set containing vocabulary difficulty, sentence pattern complexity, and grammar complexity;
[0101] S5. The GPT large model adaptively simplifies the English passage according to the junior high school English teaching syllabus and the middle school entrance examination proposition standard. Based on the context information and difficulty evaluation, it replaces advanced vocabulary with common vocabulary that meets the junior high school English teaching syllabus, converts complex clause structures into simple sentences or compound sentences, and adjusts complex grammar to a grammar structure that meets the middle school entrance examination proposition standard based on grammar association;
[0102] S6. The GPT large model evaluates the grammar correctness, semantic integrity, and logical coherence of the rewritten version of the English passage generated after simplification;
[0103] S7. Based on the evaluation results and user feedback, the feedback learning module of the GPT large model dynamically adjusts the simplification rules and generation strategies. The GPT large model adaptively adjusts the simplification rules and generation strategies according to the type of English passage input by the user, the context relevance, and the proposition difficulty requirements.
[0104] S8. Output the English text simplified by the GPT large model and meeting the middle school entrance examination proposition standards to the user or the education system.
[0105] In this embodiment, S1 specifically includes:
[0106] S11. Receive the English passage T input by the user. The English passage includes multiple sentences S i The sentence set S = {S1, S2,..., S n}, and each sentence S i includes clause structures, advanced vocabulary, and grammar expressions beyond the cognitive level of junior high school students.
[0107] S12. Conduct grammatical analysis on each sentence S i in the input English passage T, separate the main clause and subordinate clauses, and construct a subordinate clause structure set:
[0108] C i = {c i1 , c i2 ,..., c im};
[0109] where c ij represents the jth subordinate clause in the sentence S i .
[0110] S13. Extract the advanced vocabulary set from each sentence S i :
[0111] V i = {v i1 , v i2 ,..., v ik};
[0112] where v ij represents the jth advanced vocabulary in the sentence S i . Advanced vocabulary refers to words that do not match the junior high school English syllabus or have a low word frequency.
[0113] S14. According to the syntactic structure and lexical complexity of each sentence S i , construct a set of grammar expressions beyond the cognitive level:
[0114] G i = {g i1 , g i2 ,..., g il};
[0115] Among them, g ij represents that the j-th grammatical structure in sentence S i exceeds the cognitive level of junior high school students;
[0116] S15. Merge the structures, vocabulary, and grammatical features of all sentences in the received English passage T to construct an English passage matrix M T :
[0117] Sentence number Clause number Number of advanced vocabulary Number of grammatical structures beyond the level
[0118]
[0119] Among them, m i , k i , l i respectively represent the number of clauses, the number of advanced vocabulary, and the number of grammatical structures beyond the cognitive level of students in sentence S i .
[0120] In this embodiment, S2 specifically includes:
[0121] S21. Perform word segmentation on the received English passage T = {S1, S2,..., S n}, and decompose each sentence S i into a vocabulary sequence:
[0122]
[0123] Among them, L i represents the total number of vocabulary in sentence S i , w ij represents the j-th vocabulary in sentence S i ; perform part-of-speech tagging and vocabulary attribute extraction on each vocabulary w ij to obtain a vocabulary feature vector:
[0124] F ij = [pos(w ij ), lemma(w ij ), freq(w ij ), complexity(w ij )];
[0125] Among them, pos(w ij ) is the part of speech of vocabulary w ij , lemma(w ij ) is the prototype form of vocabulary w ij , freq(w ij ) is the frequency of vocabulary w ijThe frequency of occurrence in the junior high school English teaching syllabus, complexity(w ij ) is the complexity score of vocabulary w ij , reflecting its difficulty for junior high school students;
[0126] S22. For each sentence S i Perform syntactic analysis to construct a grammatical dependency tree:
[0127] D i =(N i , E i );
[0128] Among them, is a set of nodes, and each node n ij corresponds to vocabulary w ij , E i ={(n ip , n iq , r ipq )∣p, q∈[1, L i} is a set of edges, and r ipq represents the type of grammatical relationship pointing from node n ip to node n iq ;
[0129] Calculate the syntactic complexity index C i (S syntax ) of sentence S i :
[0130]
[0131] Among them, weight(r ipq ) is the complexity weight of grammatical relationship r ipq , reflecting the difficulty level of the grammatical structure in the junior high school English teaching syllabus;
[0132] S23. Perform context structure analysis on the English passage T, and use the topic model to extract the topic distribution of the English passage to construct the topic distribution vector of the English passage:
[0133] θ T =[θ1, θ2,..., θ K ;
[0134] Among them, K is the preset number of topics, and θ k represents the weight of the English passage T on the k-th topic; at the same time, construct the topic word distribution matrix Φ = [φ kj , where φ kj represents the probability of vocabulary w j in the k-th topic;
[0135] Based on the topic distribution vector θT and the subject word distribution matrix Φ, construct the context theme representation vector C of the English text T ;
[0136] S24. Based on the lexical feature set F i , the syntactic dependency tree D i and the context theme representation vector C T , construct the comprehensive text representation model R of the English text T :
[0137] R T =(F T , G T , C′ T );
[0138] Among them, is the lexical feature set of all sentences in the English text, G T =(N T , E T ) is the global syntactic graph of the English text, C′ T =α1·C T +β1·Embed(T) is the weighted context theme representation vector, α1 and β1 are balance coefficients, and Embed(T) is the semantic embedding vector of the English text T.
[0139] In this embodiment, S3 specifically includes:
[0140] S31. Use the pre-trained GPT large model to process the constructed comprehensive text representation model R of the English text T to automatically identify and extract the core theme and context information of the English text;
[0141] S32. Based on the lexical feature set F T and the context theme representation vector C′ T , calculate the theme relevance S topic (T):
[0142]
[0143] Among them, w j is the j-th word in the lexical feature set F T , θ T is the theme distribution vector of the English text, φ kj is the element in the subject word distribution matrix, indicating the probability of the word w j in the k-th theme;
[0144] S33. Use the syntactic dependency tree G T =(N T, E T ), analyze the syntactic relationships between sentences and construct the sentence connectivity matrix L T = [l ij :
[0145]
[0146] where l ij represents the connectivity between sentence S i and sentence S j ;
[0147] S34. Extract the main idea M T and the set of key information points K T = {k1, k2,..., k m}, where k i represents the i-th key information point;
[0148] S35. Combine the main idea M T , the set of key information points K T , the topic relevance S topic (T), and the sentence connectivity matrix L T to construct the core topic representation model C core :
[0149] C core = (M T , K T , S topic (T), L T );
[0150] S36. Analyze the logical relationships, narrative structures, and semantic associations between paragraphs of the article based on the core topic representation model C core to generate the context information set:
[0151]
[0152] where C context represents the context information set, which is the quantitative representation result of the logical relationships and semantic associations between paragraphs of the article. Sim(S i , S j ) represents the semantic similarity between sentence S i and sentence S j , measuring the semantic overlap degree between two sentences and used to capture the potential information connections between sentences in an English text. L T [i, j] is an element in the sentence connectivity matrix, representing the direct syntactic relationship strength between sentence S i and S j , and S topic (Si , S j ) is the topic relevance degree between and S i and S j , Rel(S i , M T ) is the sentence S i 's correlation function with the main idea M T . α2 and β2 are the balance coefficients of semantic similarity and topic relevance degree, and λ is the adjustment coefficient of topic relevance degree. is the logistic regression function, which is used to map the comprehensive result of sentence connection degree and topic relevance degree to the interval [0, 1].
[0153] In this embodiment, S4 specifically includes:
[0154] S41. Based on the core topic representation model C core and the context information set C context , analyze the lexical features F T of the text, the global grammar graph G T of the English passage, and the context topic information C' T ;
[0155] S42. Use the GPT large model to evaluate the lexical difficulty D vocab (T) of the English passage, and calculate the lexical complexity score of each word w j ∈F T :
[0156]
[0157] Among them, complexity(w j ) represents the complexity of the word w j , freq(w j ) is the frequency of w j appearing in the junior high school English syllabus, and max(freq(F T )) is the maximum frequency of all words;
[0158] S43. Calculate the sentence pattern complexity D syntax (S i ) based on the syntactic dependency tree;
[0159] S44. Use the context information set C context and the topic relevance degree S topic (T) to calculate the grammar complexity D grammar (T) of the English passage:
[0160]
[0161] Among them, n is the number of sentences, Rel(Si , M T ) represents sentence S i 's relevance to the main idea M T , where γ and δ are the balance coefficients between the syntactic complexity and the relevance;
[0162] S45. Generate the text feature set F by integrating vocabulary, syntax, and grammar complexity text (T):
[0163] F text (T) = (D vocab (T), D syntax (T), D grammar )(T)).
[0164] In this embodiment, S5 specifically includes:
[0165] S51. Identify the advanced vocabulary set V in the English passage according to the vocabulary feature set F T and the text feature set F text (T), where V H = {v1, v2,..., v k}, and where represents the vocabulary that exceeds the junior high school English teaching syllabus;
[0166] S52. Use the vocabulary replacement function Replace(v j , V C ) to replace each advanced vocabulary v j ∈V H with a common vocabulary v' j ∈V C :
[0167]
[0168] where Sim(v j , v) is the semantic similarity between the vocabulary v j and the common vocabulary v, freq(v) is the occurrence frequency of v in the junior high school English teaching syllabus, and V C is the common vocabulary set;
[0169] S53. Identify the complex clause structure set C in the English passage based on the syntactic dependency tree G T = {c1, c2,..., c H}; m}
[0170] S54. Use the clause simplification function Simplify(c i ) to simplify each complex clause c i ∈C HConvert to simple sentences or compound sentences. The set of simplified sentences is represented as:
[0171] S′ = {s′1, s′2, …, s′ m};
[0172] where s′ i is the simplified form of clause c i , retaining the semantic core of the original sentence while reducing syntactic complexity;
[0173] S55. Utilize the context information set C context and the sentence connection degree matrix L T to adjust the grammatical associations of the simplified sentence S' to meet the middle school entrance examination proposition standards;
[0174] S56. Define the grammar adjustment function Adjust(S′) to optimize the grammatical structure of the sentence based on the grammatical association degree and context information:
[0175]
[0176] where S″ i is the optimized form of sentence s′ i , and λ is the adjustment coefficient to ensure that the optimized syntax meets the requirements of the teaching syllabus;
[0177] S57. Generate the simplified English passage T':
[0178] T′ = {S″1, S″2, …, S″ n}.
[0179] Example 1:
[0180] In May 2024, a junior high school in a certain city was preparing for a mock middle school entrance examination. The teacher team planned to select reading materials from public English resources and generate reading comprehension questions that met the requirements of the middle school entrance examination. During the screening process, the English teaching and research group chose a news article from China Daily titled "Climate Change and Global Cooperation". However, the content of the article was complex, containing multiple nested clauses and advanced vocabulary such as "mitigate", "emission", and "convergence". Based on experience, the teachers judged that the language difficulty of this article far exceeded the comprehension level of junior high school students and needed to be rewritten.
[0181] To solve the problems of high workload and grammatical consistency during the rewriting process, the teaching and research group decided to use the intelligent rewriting system based on the GPT large model of the present invention for processing. The following is the operation process of the system in a real scenario:
[0182] The teacher input the original text fragment into the system at 15:00 on May 10, 2024. The content is as follows:
[0183] “Efforts to mitigate climate change require convergence among nations to reduce carbon emissions and adopt sustainable practices. Failure to act in time may result in irreversible consequences.”;
[0184] The system first identified the advanced vocabulary in the article: “mitigate,” “convergence,” and “irreversible,” and marked them as words beyond the junior high school English teaching syllabus. The system also detected the nested clause structure in the sentence: “Efforts to mitigate climate change require convergence among nations to reduce carbon emissions and adopt sustainable practices,” which contains multiple verb phrases and conditional expressions, and the overall sentence structure is complex.
[0185] At 15:02 on May 10, 2024, the system automatically replaced the advanced vocabulary with words that are easier for students to understand according to the vocabulary list in the junior high school English syllabus.
[0186] Original word: “mitigate” → replaced with “reduce”;
[0187] Original word: “convergence” → replaced with “agreement”;
[0188] Original word: “irreversible” → replaced with “serious”;
[0189] At the same time, the system disassembled the complex clause structure into multiple short sentences to improve students' comprehension ability.
[0190] Original sentence: “Efforts to mitigate climate change require convergence among nations to reduce carbon emissions and adopt sustainable practices.”
[0191] Simplified: "Countries need to work together to reduce pollution. They should also find better ways to protect the environment."
[0192] The system generated three rewritten versions at 15:05 and selected the optimal version based on grammar, syntax, and context evaluation.
[0193] The system completed the automatic generation of reading comprehension questions at 15:08 on May 10, 2024, and output the following questions:
[0194] Question 1: What do countries need to do to fight climate change?
[0195] A. Use more resources;
[0196] B. Reduce pollution;
[0197] C. Ignore environmental issues;
[0198] D. Increase emissions;
[0199] Question 2: Why is it important for countries to work together?
[0200] A. To adopt new practices;
[0201] B. To agree on environmental laws;
[0202] C. To reduce serious problems;
[0203] D. To find new energy sources;
[0204] After generating the questions, the teacher team reviewed the questions output by the system at 15:15 and provided feedback to the system: It is hoped to further emphasize the specific measures for environmental protection. The system adjusted the options and descriptions in the questions in real time according to the feedback and generated a new version of the questions at 15:20.
[0205] Table 1 Comparison of Time and Efficiency between the System of the Present Invention and Traditional Manual Methods in the Process of Processing Reading Comprehension Proposition Materials
[0206] Task Traditional manual method The system of the present invention Material preparation and screening 2 hours 10 minutes Advanced vocabulary replacement and sentence pattern simplification 4 hours 3 minutes Question generation and multi-version output 2 hours 5 minutes Feedback optimization and version adjustment 1 hour 5 minutes Total time consumption 9 hours 23 minutes
[0207] As can be seen from Table 1 above, after using the present system for intelligent rewriting, the teacher team used the generated test questions for mock exams. The passing rate of the test questions generated by the system reached 92%, while the passing rate of the traditional proposition test paper was 84%. In addition, the teachers were satisfied with the diversity and grammar accuracy of the test questions generated by the system, with an average score of 9.1 points (out of 10), while the average score of the traditional test questions was 7.8 points.
[0208] Example 1 demonstrates the high efficiency and applicability of the present invention in actual teaching. Compared with the traditional method, the present invention significantly shortens the time for material processing and proposition generation, reduces the workload of teachers, and ensures the grammar consistency and diversity that meet the high school entrance examination standards of the test questions.
[0209] The present invention proposes a multi-dimensional feature evaluation algorithm by combining lexical feature sets, syntactic dependency trees, and context theme information. By comprehensively calculating lexical complexity, sentence pattern complexity, and grammar complexity, it can automatically adjust the material content according to the junior high school English teaching syllabus and high school entrance examination proposition standards. Compared with traditional keyword replacement, the present invention innovatively uses context relevance calculation functions and logistic regression functions to accurately replace complex syntax and advanced vocabulary in the material, converts complex clauses into simple sentences or coordinate sentences, and optimizes the grammar relevance, thereby ensuring that the rewritten material not only meets the teaching standards but also maintains semantic coherence, effectively avoiding the problems of semantic breakage and logical loss in the simplification process of traditional tools, and improving the accuracy and applicability of the generated content.
[0210] The present invention designs a multi-layer adaptive proposition generation mechanism based on the core theme model. By analyzing the main idea, key information points, and grammar relevance matrix between sentences of the original text, it can generate multiple versions of proposition materials according to different educational goals, and realizes the full-process adaptive adjustment from the input text to proposition generation by using semantic similarity and theme relevance as adjustment parameters. Different from the common fixed-template proposition in existing tools, the present invention can flexibly generate proposition versions suitable for different difficulties and themes, realizing the diversification and innovation of proposition content.
[0211] The present invention integrates a dynamic feedback and real-time optimization module, enabling the GPT large model to continuously optimize the generation effect according to user feedback in practical applications. During the use process, teachers or the education system evaluate the generated materials and propositions, and the system inputs these feedbacks as training data into the model to automatically adjust the generation strategy and algorithm parameters. Through this closed-loop optimization mechanism, the present invention can gradually improve the adaptability of the model to different student groups and achieve highly personalized proposition generation, solving the problems of lack of personalized feedback and model dynamic adjustment in the prior art, and providing more efficient teaching support for teachers and students.
[0212] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. An intelligent rewriting method for reading comprehension question materials based on a large model, characterized in that: The steps include: S1. receiving an English passage input by a user, wherein the English passage includes a clause structure, advanced vocabulary, and grammatical expressions beyond the cognitive level of junior high school students; S2. Preprocess the English text by performing word segmentation, syntactic analysis and context structure analysis, and construct a comprehensive text representation model that includes vocabulary information, syntactic structure information and context theme information; S3. Using the pre-trained GPT large model, based on the constructed comprehensive text representation model, automatically identify and extract the core topic representation model and context information set of the English passage; S4. Based on the core topic representation model and context information set, the GPT large model automatically evaluates the vocabulary, sentence structure and grammatical complexity of English passages, and generates a text feature set including vocabulary difficulty, sentence structure complexity and grammatical complexity; S5, GPT large model adaptively simplifies English passages according to the junior high school English teaching syllabus and the high school entrance examination test standards, replaces advanced vocabulary with common vocabulary that meets the junior high school English teaching syllabus based on contextual information and difficulty assessment, converts complex clause structures into simple sentences or parallel sentences, and adjusts complex grammar to a grammatical structure that meets the high school entrance examination test standards based on grammatical associations; S6. The GPT model evaluates the grammatical correctness, semantic completeness, and logical coherence of the rewritten English text generated after simplification. S7. Based on the evaluation results and user feedback, the feedback learning module of the GPT large model dynamically adjusts the simplification rules and generation strategies. The GPT large model adaptively adjusts the simplification rules and generation strategies according to the type of English passages input by users, the relevance of the context, and the difficulty requirements of the proposition; S8. Output the English text simplified by the GPT large model and meeting the standards for the high school entrance examination to the user or the education system.
2. According to claim 1, a method for intelligently rewriting reading comprehension question materials based on a large model is characterized in that: The S1 specifically includes: S11. Receive an English passage T input by a user, where the English passage includes a plurality of sentences S i The sentence set S = {S1, S2, ..., S n }, each sentence S i Includes clause structure, advanced vocabulary, and grammatical expressions beyond the cognitive level of junior high school students; S12, for each sentence S in the input English passage T i Perform grammatical analysis, separate the main clause and the subordinate clause, and construct a set of subordinate clause structures: Among them, c ij Represents sentence S i The j-th clause in ; S13. From each sentence S i Extract high-level vocabulary sets from: in, Represents sentence S i The first advanced vocabulary in the text refers to the vocabulary that does not match the junior high school English teaching syllabus or has a low frequency; S14. According to each sentence S i The syntactic structure and vocabulary complexity of the text are used to construct a set of grammatical expressions that exceed the cognitive level of junior high school students: in, Represents sentence S i The j2nd grammatical structure is beyond the cognitive level of junior high school students; S15. Merge the structure, vocabulary and grammatical features of all sentences in the received English text T to construct an English text matrix M T : Number of sentences Number of clauses Number of advanced vocabulary Number of grammatical structures that exceed the cognitive level of junior high school students 3. According to claim 2, a method for intelligently rewriting reading comprehension question materials based on a large model is characterized in that: The S2 specifically includes: S21, for the received English chapter T = {S1, S2, ..., S n } to perform word segmentation and separate each sentence S i Decomposed into a sequence of words: Among them, L i Represents sentence S i The total number of words, w ij Represents sentence S i The jth word in ij Perform part-of-speech tagging and vocabulary attribute extraction to obtain vocabulary feature vectors: F ij =[pos(w ij ),lemma(w ij ),freq(w ij ),complexity(w ij )]; Among them, pos(w ij ) is the vocabulary w ij Part of speech, lemma (w ij ) is the vocabulary w ij The prototype form of freq(w ij ) is the vocabulary w ij The frequency of occurrence in the junior high school English teaching syllabus, complexity (w ij ) is the vocabulary w ij The complexity score reflects its difficulty for junior high school students; S22. For each sentence S i Perform syntactic analysis to build a grammatical dependency tree: D i =(N i ,E i ); in, is a set of nodes, each node n ij Corresponding vocabulary w ij , E i ={(n ip ,n iq ,r ipq )|p,q∈[1,L i ]} is the edge set, r ipq Represents slave node n ip Points to node n iq grammatical relationship; Calculate the sentence S i The syntactic complexity index C syntax (S i ): Among them, weight(r ipq ) is the grammatical relation r ipq The complexity weight reflects the difficulty level of the grammatical structure in the junior high school English teaching syllabus; S23, perform context structure analysis on the English article T, use the topic model to extract the topic distribution of the English article and construct the topic distribution vector of the English article: i T =[θ1,θ2,…,θ K ]; Among them, K is the preset number of topics, θ k represents the weight of English article T on the kth topic; at the same time, construct the topic word distribution matrix in Represents the word w in the kth topic j3 probability; Based on the topic distribution vector θ T And the topic word distribution matrix Φ, construct the context topic representation vector C of the English chapter T ; S24, based on each sentence S in English passage T i The corresponding vocabulary feature set F i 、Grammar dependency tree D i and the contextual topic representation vector C T , construct a comprehensive text representation model R for English texts T : R T =(F T ,G T ,C′ T ); in, is the lexical feature set of all sentences in the English text, G T =(N T ,E T ) is the global grammatical graph of the English text, C′ T =α1·C T +β1·Embed(T) is the weighted context topic representation vector, α1 and β1 are balance coefficients, and Embed(T) is the semantic embedding vector of the English text T.
4. According to claim 1, a method for intelligently rewriting reading comprehension question materials based on a large model is characterized in that: The S3 specifically includes: S31. Comprehensive text representation model R of English chapters constructed using the pre-trained GPT large model T Processing, automatically identifying and extracting the core topics and contextual information of English passages; S32, based on the vocabulary feature set F of all sentences in the English text T and the weighted contextual topic representation vector C′ T , calculate the topic relevance S of the English chapter topic (T): Among them, w j4 is the lexical feature set F of all sentences in the English text T The j4th word in θ T is the topic distribution vector of the English text, is an element in the topic word distribution matrix, representing the vocabulary in the kth topic probability; S33. Using the global grammar graph G of English text T =(N T ,E T ), analyze the grammatical relationship between sentences and construct the sentence connectivity matrix L T =[l ij ]: Among them, l ij Represents sentence S i With sentence S j Connectivity S34. Extract the main idea of the article through the semantic understanding module of the GPT large model T And the key information point set K T ={k1,k2,…,k m }, where k i1 Indicates the i 1 Key information points; S35. Combine the main ideas M T , key information point set K T , topic relevance S topic (T) and the sentence connectivity matrix L T , build the core topic representation model C core : C core =(M T ,K T ,S topic (T),L T ); S36, based on the core topic representation model C core Analyze the logical relationship, narrative structure and semantic association between paragraphs in the article to generate a set of context information: Among them, C context Represents a set of context information, which is a quantitative representation of the logical relationship between paragraphs and the semantic association between the previous and next paragraphs. Sim(S i ,S j ) represents sentence S i and sentence S j The semantic similarity of L is used to measure the semantic overlap between two sentences and to capture the potential information connection between sentences in English texts. T [i,j]=l ij , which means sentence S i With S j The direct grammatical relationship strength between topic (S i ,S j ) is a sentence S i and S j The topic relevance between them, Rel(S i ,M T ) is a sentence S i With the main idea M T , α2 and β2 are the balance coefficients between semantic similarity and topic relevance, and λ is the adjustment coefficient of topic relevance. is a logistic regression function, which is used to map the comprehensive results of sentence connectivity and topic relevance into the interval [0,1], and n is the number of sentences in the English text.
5. According to claim 1, a method for intelligently rewriting reading comprehension question materials based on a large model is characterized in that: The S4 specifically includes: S41. Core topic representation model C core and context information set C context , analyze the vocabulary feature set F of the text T 、The global grammatical graph G of English text T and the weighted contextual topic representation vector C′ T ; S42. Using the GPT model to evaluate the vocabulary difficulty of English passages D vocab (T), calculate each word w j ∈F T The lexical complexity score of is: Among them, complexity (w j ) represents the word w j The complexity of freq(w j ) is w j The frequency of occurrence in the junior high school English teaching syllabus, max(freq(F T )) is the maximum frequency of all words; S43. Global grammar graph G based on English text T Calculate sentence complexity D syntax (S i ); S44, using context information set C context The degree of relevance to the topic of the English text S topic (T) Calculate the grammatical complexity of English texts D grammar (T): Where n is the number of sentences, Rel(S i ,M T ) represents sentence S i With the main idea M T , γ and δ are the balance coefficients between sentence complexity and correlation; S45. Generate text feature set F by integrating vocabulary, sentence structure and grammatical complexity of English texts text (T): F text (T)=(D vocab (T),D syntax (T),D grammar (T))。 6. The intelligent rewriting method for reading comprehension question materials based on a large model according to claim 1 is characterized in that: The S5 specifically includes: S51, according to the vocabulary feature set F T and text feature set F text (T), identify the high-level vocabulary set V in the English text H ={v1,v2,…,v k },in It indicates vocabulary beyond the junior high school English syllabus; S52, using the vocabulary replacement function Replace(v j ,V C ) Each advanced vocabulary v j ∈V H Replace with common vocabulary v′ that conforms to the junior high school English teaching syllabus j ∈V C : Among them, Sim(v j ,v) is the vocabulary v j and the semantic similarity of common vocabulary v, freq(v) is the frequency of v in the junior high school English teaching syllabus, V C A collection of common words; S53, Global grammar graph G based on English text T Identify complex clause structures in English texts C H ={c1,c2,…,c m }; S54、Use clauses to simplify functions Simplify(c i ) Each complex clause c i ∈C H Converted into simple sentences or parallel sentences, the simplified sentence set is expressed as: S′={s′1,s′2,…,s′ m }; Among them, s′ i For clause c i A simplified form that retains the semantic core of the original sentence but reduces syntactic complexity; S55, using context information set C context And the sentence connectivity matrix L T , adjust the grammatical association of the simplified sentence set S' to make it meet the standards of the high school entrance examination; S56. Define a grammar adjustment function Adjust(S′) to optimize the grammatical structure of the sentence based on grammatical relevance and context information: Among them, S″ i For sentence s′ i The optimized form, λ is the adjustment coefficient to ensure that the optimized syntax meets the syllabus requirements, L T [i,f] represents sentence S i With S f the strength of the direct grammatical relationship between S57. Generate a simplified English text T'.
Citation Information
Patent Citations
Auxiliary question setting system and method based on large model
CN118609437A
Examination test question generation method based on large model
CN118839003A