A progressive language learning content generation system based on a restricted expression space

CN122658147APending Publication Date: 2026-08-28CHUANGZHI YUNWEI (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610787547.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0005]固定教材或统一课程通常依据群体平均水平设计,难以匹配单个用户真实词汇能力,导致内容过难或过易,降低学习效率

Benefits of technology

[0020] This invention integrates content generation, vocabulary development, difficulty adjustment, native language assistance display, terminal adaptation, and continuous content supply into a unified closed-loop system, achieving a language learning model with high adaptability, high retention rate, and low content supply cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658147A_ABST
    Figure CN122658147A_ABST
Patent Text Reader

Abstract

The application discloses a kind of progressive language learning content generation systems based on limited expression space, including user state acquisition module, limited expression space construction module, content generation module, mother tongue mapping output module, familiar word progressive module, continuous supply module and terminal adaptation module.System constructs limited expression space according to user familiar word state, learning level, reading record and interactive behavior data, and constraint determination is executed in real time during generation process, so that the path construction of language content and the content generation limit are completed in limited expression space.The generated content is output as an auxiliary learning interface through mother tongue mapping, and user learning behavior is used to update familiar word state and reconstruct the next round of limited expression space.The limited expression space can also include target language presentation density, upper limit of single-screen vocabulary quantity and current visible area load threshold to control interface display intensity and achieve progressive learning supply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence content generation, natural language processing, educational informatization, personalized teaching and language learning assistance, and in particular to a language learning content generation system that dynamically generates progressive language learning content within a limited expression space based on user learning status data, and continuously updates the content based on user learning behavior. Background Technology

[0002] Existing language learning systems mainly include vocabulary memorization software, fixed textbook courses, question bank training systems, graded reading materials, and general content generation systems.

[0003] The aforementioned existing technologies generally suffer from the following problems:

[0004] (1) Mismatch between the difficulty of the learning content

[0005] Fixed textbooks or standardized courses are usually designed based on the average level of a group, which makes it difficult to match the actual vocabulary ability of an individual user, resulting in content that is too difficult or too easy, thus reducing learning efficiency.

[0006] (2) The new learning burden is uncontrollable

[0007] Existing general content generation systems typically generate content freely in an open expression space. The number of new words, sentence complexity, and semantic span in the generated text fluctuate greatly, which can easily cause an imbalance in the reading pressure on users.

[0008] (3) Low utilization rate of already mastered vocabulary

[0009] Existing systems often fail to effectively identify and reuse users’ existing vocabulary assets, resulting in a high rate of repetitive learning and making it difficult for users’ existing knowledge to be transformed into a continuous input advantage.

[0010] (4) Fragmented learning path

[0011] Vocabulary learning, reading training, content consumption, and ability assessment are usually scattered across different systems or functional modules, lacking a continuous closed-loop growth path.

[0012] (5) High content supply costs

[0013] Manually compiling graded reading materials, textbook content, or phased learning content is costly, slow to update, and has weak individual adaptability, making it difficult to meet the continuous learning needs of a large number of users.

[0014] (6) Existing graded reading systems lack real-time individual adjustment capabilities.

[0015] Existing graded content mostly adopts a fixed level system, which cannot be dynamically adjusted in real time based on an individual user's vocabulary familiarity, reading speed, recognition rate, and recent progress.

[0016] Therefore, there is a need for a language learning content generation system that can automatically generate appropriate content within a limited expression space based on the user's current learning status, and continuously evolve as the user grows. Summary of the Invention

[0017] This invention provides a progressive language learning content generation system based on a restricted expression space. By acquiring user learning status data, a corresponding individualized restricted expression space Ω is constructed, and language learning content is generated within the restricted expression space Ω.

[0018] The content generation process does not involve generating content in an open expression space first and then filtering it. Instead, the constraint judgment of the restricted expression space Ω is performed in real time during the generation process, so that the path construction, expression selection and output results of the language content are all limited to the restricted expression space Ω.

[0019] The present invention further dynamically updates the boundary parameters of the restricted expression space Ω based on user learning behavior data, so that the generated content continuously changes with the user's vocabulary familiarity, reading ability and learning progress, thereby forming an individualized, continuous and low-burden language learning path.

[0020] This invention integrates content generation, vocabulary development, difficulty adjustment, native language assistance display, terminal adaptation, and continuous content supply into a unified closed-loop system, achieving a language learning model with high adaptability, high retention rate, and low content supply cost. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0022] Figure 1 This is a schematic diagram of the overall structure of a progressive language learning content generation system based on a constrained expression space according to an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram illustrating the construction of the restricted expression space Ω according to an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of an optional content generation process according to an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of an optional word progression closed loop according to an embodiment of the present invention;

[0026] Figure 5This is a schematic diagram of an optional terminal adapter output structure according to an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] This application provides a progressive language learning content generation system based on a constrained expression space, such as... Figure 1 As shown, the system includes the following modules:

[0030] 1) User Status Acquisition Module

[0031] The data is used to acquire user learning status data, which includes at least one of the following: user's vocabulary familiarity status, learning level, reading history, reading speed, reading dwell time, continuous reading completion rate, historical correct recognition rate, user's interest topics, learning activity level, and terminal device information.

[0032] 2) Restricted Expression Space Construction Module

[0033] like Figure 2 As shown, the restricted expression space construction module is used to construct a corresponding restricted expression space Ω based on the user learning state data. The restricted expression space Ω includes at least one of the following constraint sets:

[0034] (1) Lexical constraints, including: user familiar word set, new word candidate set, upper limit of the number of new words in this round, word frequency level range and target part of speech distribution range.

[0035] (2) Grammatical constraints, including: sentence length range, tense complexity, number of clauses, range of sentence variation and grammatical structure level.

[0036] (3) Content constraints, including: text length, reading time, content theme, content genre, emotional style and interest tags.

[0037] (4) User behavior constraints, including: familiar word confirmation rate, reading speed range, continuous reading completion rate, recent learning activity and historical correct recognition rate.

[0038] The boundary parameters of the restricted expression space Ω are dynamically adjusted based on user learning behavior data. Furthermore, the boundary parameters of the restricted expression space Ω may also include at least one of the following: target language presentation density, maximum number of target language words per screen, maximum number of target language words displayed per line, native language auxiliary ratio, and current visible area load threshold. These boundary parameters not only limit the vocabulary, grammar, or topic range of the generated content but also limit the display intensity that the generated content can bear in the learning interface, ensuring that the generated content matches the user's current cognitive load and the terminal display environment.

[0039] 3) Content Generation Module

[0040] This is used to perform constraint determination of the restricted expression space Ω in real time during the language content generation process, and only allows the path construction and content generation of language content to be completed within the restricted expression space Ω.

[0041] The content generation module does not include a generation method that first generates content outside the restricted expression space Ω and then filters or maps it to the restricted expression space Ω.

[0042] The generated content includes at least one of the following: short reading passages, continuous stories, scene dialogues, news briefs, content on topics of interest, exam training materials, and high-frequency expression training content.

[0043] The content generation module can make new words appear repeatedly in the same learning content a preset number of times or a dynamic number of times, so as to enhance the user's memory reinforcement effect.

[0044] The content generation module generates content as follows: Figure 3 As shown.

[0045] 4) Native Language Mapping Output Module

[0046] This is used to convert target language content into a native language-assisted learning interface.

[0047] The output methods include at least one of the following: native language wrapping display, bilingual comparison display, familiar word hiding display, keyword prompt display, fade-out auxiliary display, click-to-expand auxiliary display, and segmented progressive display.

[0048] In one implementation, the native language mapping output module can send the generated content to the vocabulary replacement load control module. The vocabulary replacement load control module allocates the display position, number of words, or intensity of assistance for the target language vocabulary based on the user's familiarity with the words, the current visible area, the target language presentation density, and the replacement capacity. Therefore, the native language-assisted learning interface can not only present the generated content but also further control the display load of the target language vocabulary according to the user's current learning status.

[0049] 5) Familiar Word Progression Module

[0050] Used to update the user's familiarity with words based on the user's learning behavior.

[0051] The learning behaviors include at least one of the following: clicking to confirm mastery, continuous correct recognition, high-frequency natural reading, active repetition and use, multi-round stable recognition, and reading speed reaching a threshold.

[0052] When an improvement in a user's vocabulary proficiency is detected, the vocabulary range of the restricted expression space Ω is automatically expanded, and at least one of the following is adjusted: the upper limit of the number of new words, the proportion of the target language, and the proportion of native language assistance.

[0053] Familiar words progressive closed loop, such as Figure 4 As shown.

[0054] 6) Continuous supply module

[0055] It is used to reconstruct the restricted expression space Ω based on the updated user learning status data and generate the next round of learning content.

[0056] The continuous supply module dynamically adjusts the difficulty level, text length, and number of new words for the next round of learning content based on the user's historical correct recognition rate, the growth rate of familiar words, the reading completion rate, the reading speed, or the current learning level.

[0057] 7) Terminal adaptation module

[0058] This system is used to automatically adjust the output content density, layout structure, content segmentation structure per screen, target language display density per line, font size, or auxiliary display ratio based on the terminal device's screen size, display orientation, visible area size, or interaction method. In one implementation, the system can use the current visible area as the control unit for the output content density or native language auxiliary display ratio, ensuring the stability of the target language presentation load for the same long text in different screen areas. When the user slides to a new visible area, the system can re-determine the output density or auxiliary display ratio based on the new visible area size, text content distribution, and the user's learning status to avoid the target language content being too dense or too sparse in a local area. Specific terminal-adapted output structures are as follows: Figure 5 As shown.

[0059] Example 1 (Junior High School Users): The user's vocabulary level corresponds to 1500 words. System settings: Maximum number of new words = 6; Reading length = 180 characters; Content theme = Campus life; Native language assistance ratio = High. Generate a bilingual short article containing 6 new words, which appear 2 to 4 times in the same content. When the user's vocabulary level increases to 1550: The maximum number of new words is adjusted to 8; the native language assistance ratio decreases; the target language display ratio increases.

[0060] Example 2 (University Students): The user's vocabulary level corresponds to a vocabulary size of 3200. System settings: Maximum number of new words = 10; Reading length = 450 characters; Content theme = News commentary; Native language assistance ratio = Medium. Intermediate to advanced reading materials are generated, and the difficulty of the next round is dynamically adjusted based on the recognition rate.

[0061] Example 3 (Early Childhood Learning Users): The user's vocabulary level corresponds to 80 words. System settings: high-frequency repeated words; single-sentence structure; high proportion of native language; picture and text auxiliary mode. Generate early childhood learning reading content.

[0062] Example 4 (Mobile Terminal Adaptation): When the system detects that the screen size of the terminal device is small or in portrait mode, it automatically reduces the density of new words per screen and adjusts it to a three-segment layout output structure per screen, while limiting the display density of the target language in each line.

[0063] Compared with the prior art, the present invention has at least the following beneficial effects: (1) The difficulty of the content is matched with the user's ability in real time, avoiding the decline in learning efficiency caused by overly difficult or easy content; (2) The new learning burden is controllable, and the learning pace is stabilized by the new word window mechanism; (3) Familiar word assets continue to compound, and the mastered vocabulary can be reused continuously to improve reading efficiency; (4) A closed-loop growth system is formed, with content generation, learning behavior, familiar word growth and difficulty upgrade interconnected; (5) The cost of content supply is greatly reduced, and a large amount of adapted content is automatically generated through the limited expression space; (6) The long-term retention rate is improved, and the user's willingness to continue using the product is enhanced through the low-pain input learning path; (7) The mobile reading experience is improved, and the display stability and reading comfort under different device environments are improved through the terminal adaptation mechanism.

[0064] This invention also provides a method for generating progressive language learning content based on a constrained expression space, comprising the following steps:

[0065] S101 collects user learning status data, including the status of familiar words, learning level, reading records, and interaction behavior data.

[0066] In this embodiment, when a user enters the learning system through a learning terminal, the user status acquisition module first obtains the learning status data corresponding to the current user and establishes a set of status information corresponding to the current learning task. The learning status data can originate from the user's historical learning archives or from real-time interactive behavior data during the current learning process. Specifically, it reads the familiar word status, learning level, and historical reading records corresponding to the current user from the user's learning database, and simultaneously acquires the user's reading behavior data, interactive behavior data, and learning activity data generated within the most recent preset time period to form a description of the current user's learning status.

[0067] The "familiarity status" represents the set of target language vocabulary that the user has stably mastered. In some implementations, the mastery status of target vocabulary can be determined based on the number of consecutive correct recognitions, the number of times the user naturally passes through vocabulary during reading, the number of times the user actively confirms mastery, and historical test results. Vocabulary that meets preset mastery conditions is then included in the familiarity set. The "learning level" represents the user's current language ability stage, such as the beginner, basic, intermediate, or advanced stage. The "reading record" represents the reading content the user has already studied, including information such as reading topic, number of readings, reading completion status, and reading duration.

[0068] Furthermore, to improve the accuracy of subsequent content adaptation, user interest topic information and terminal device information are also acquired. The interest topic information reflects the user's long-term preferred content areas, such as campus life, travel and culture, sports, technology news, or exam preparation. The terminal device information reflects the display characteristics of the user's current device, including screen size, display orientation, resolution, and interaction methods. After acquiring the above data, the system forms a learning state data set corresponding to the current user and sends this learning state data set to the restricted expression space construction module, providing a data foundation for the subsequent construction of the restricted expression space Ω.

[0069] S102, calculate the current newly added word window and determine the difficulty level of the target content.

[0070] After obtaining user learning status data, the new word window and target content difficulty level are determined based on the user's current learning ability. Since there are significant differences in vocabulary reserves, reading ability, and learning speed among different users, this embodiment does not use a fixed configuration of the number of new words. Instead, it dynamically determines the allowable new learning burden for this round based on the user's current learning status.

[0071] Specifically, the system first calculates the user's reading completion rate, historical correct recognition rate, reading speed, and vocabulary growth rate within the most recent preset period. The reading completion rate reflects the stability of the user's completion of learning tasks; the historical correct recognition rate reflects the user's mastery of learned vocabulary; reading speed reflects the user's fluency in processing target language content; and the vocabulary growth rate reflects the recent trend in the user's learning ability. Based on this learning behavior data, the system assesses the user's current learning load capacity and determines the corresponding difficulty range based on the current learning level.

[0072] In this embodiment, the new word window is not simply a limit on the number of new words, but rather represents the range of new knowledge that can be introduced in the current learning round. For example, when the system determines that the user is currently in the basic learning stage and has a high recent reading completion rate, the range of the new word window can be appropriately expanded; when the system detects that the user's recent reading speed has decreased or the reading completion rate has continuously decreased, the range of the new word window will be reduced accordingly to avoid excessive increase in the learning burden. At the same time, the system simultaneously determines the difficulty level of the target content based on the determined new word window and generates corresponding difficulty control parameters. These difficulty control parameters will be used as boundary constraints in the subsequent construction of the restricted expression space Ω.

[0073] In one specific embodiment, when the system detects that a user's current vocabulary is approximately 1500 words and the average completion rate of the last ten reading tasks is over 90%, the new word window for this round can be set to 6 new words, and the target content level can be determined as the basic reading level. When the system detects that the user's vocabulary reaches 3200 words and the continuous reading completion rate and recognition rate remain at a high level, the new word window can be expanded to 10 new words, and higher-level reading content can be generated.

[0074] S103, construct the current restricted expression space Ω based on the user's learning status data.

[0075] After obtaining user learning status data and the difficulty level of the target content, the restricted expression space construction module constructs a corresponding restricted expression space Ω based on the current user's ability status. The restricted expression space Ω is used to limit the range of expressions allowed in this round of learning content, thereby ensuring that the subsequently generated content matches the user's current learning ability.

[0076] Specifically, a vocabulary constraint set is first constructed. This set includes a user-familiar vocabulary set, a candidate set of new words, an upper limit on the number of new words added in this round, a word frequency range, and a target part-of-speech distribution range. The user-familiar vocabulary set originates from the familiar vocabulary status data obtained in step S101; the candidate set of new words comes from a candidate word library corresponding to the current learning level; and the upper limit on the number of new words is limited by the new word window determined in step S102. Through the above processing, the system establishes the allowed vocabulary boundary range for the current round.

[0077] Subsequently, a set of grammatical constraints is constructed. This set of grammatical constraints includes constraints such as sentence length range, tense complexity, number of clauses, range of sentence variations, and grammatical structure level. For example, for basic-level users, sentence length can be limited to a preset range, and the frequency of complex clause structures can be restricted; for intermediate and advanced-level users, richer grammatical structures and more complex sentence variations are allowed.

[0078] After constructing the vocabulary and grammar constraints, a set of content constraints and a set of user behavior constraints are further constructed. The content constraint set is used to limit the content attributes of the reading content, such as the topic direction, content genre, text length, reading time, and mood style; the user behavior constraint set determines the dynamically adjusted boundaries based on behavioral indicators such as the user's recent reading speed, familiar word confirmation rate, continuous reading completion rate, and historical correct recognition rate.

[0079] After completing the above processing, the system merges the sets of lexical constraints, grammatical constraints, content constraints, and user behavior constraints to construct the restricted expression space Ω corresponding to the current user. Unlike the fixed hierarchical system in the prior art, the restricted expression space Ω in this embodiment is not a static parameter set, but a dynamic expression space that can be continuously updated as the user's learning behavior changes. In other words, different users correspond to different restricted expression spaces at the same time, and the boundary of the restricted expression space corresponding to the same user will also change at different learning stages, thereby achieving dynamic adaptation to the individual's ability state.

[0080] S104, Path construction and language content generation are performed within the restricted expression space Ω, and constraint determination is performed in real time during the generation process.

[0081] After constructing the restricted expression space Ω, the content generation module begins the language learning content generation process. This embodiment differs significantly from existing content generation systems. Existing technologies typically employ an open content generation approach, where complete content is generated first, and then the generated results are filtered, reduced, or replaced based on vocabulary or reading levels. In this embodiment, however, the content generation process is constrained by the restricted expression space Ω in real time from the very beginning. All content is generated within the restricted expression space Ω, eliminating the need for secondary filtering after generation.

[0082] In practice, an initial semantic node is first established based on the current content theme, and a restricted expression path tree is constructed around the initial semantic node. The restricted expression path tree describes the expression expansion path during the content generation process, and includes multiple candidate expression nodes and multiple candidate expansion directions. Each candidate expression node corresponds to a possible linguistic expression form that may appear in the current generation state.

[0083] When it is necessary to expand a candidate expression node, multiple candidate words capable of expressing the target semantics are first selected from the candidate vocabulary set, and corresponding candidate expression results are generated by combining them with the current grammatical structure. Subsequently, the various constraints in the restricted expression space Ω are invoked to perform real-time evaluation of the candidate expression results. The evaluation criteria include whether the number of newly added words exceeds the limit of the newly added word window, whether the candidate words belong to the allowed vocabulary set, whether the current sentence length exceeds the limit range, whether the grammatical structure meets the corresponding level requirements, whether the content theme is consistent with the target theme, and whether the current expression complexity meets the user behavior constraints.

[0084] Only when a candidate expression result simultaneously meets the above constraints is the corresponding candidate expression node allowed to continue expanding and generate the next layer of expression nodes; for candidate expression results that do not meet the constraints, the corresponding path expansion is directly terminated and does not enter the subsequent generation stage. In this way, the content generation path is kept within the restricted expression space Ω throughout the entire generation process, avoiding the expression boundary out-of-bounds problem caused by the open generation mode from the source.

[0085] In some implementations, to further ensure the coherence and learning value of the generated content, cumulative path constraints are also established during the expansion of the restricted expression path tree. This involves not only single-time constraint determination of the current candidate expression node, but also continuous statistics on the number of newly added words, cumulative sentence length, cumulative reading load, and content topic offset in the current generated path. When a path's current node meets the constraints, but its cumulative result is close to the boundary of the restricted expression space Ω, the subsequent expansion degrees of freedom are automatically reduced, and candidate expressions with lower complexity are prioritized for continued generation.

[0086] For example, when the system detects that the current content has already used five new words, and the maximum number of new words in this round is six, subsequent path expansion will prioritize selecting expressions from the user's familiar word set, rather than prioritizing the introduction of new candidate words. As another example, when the system detects that the current text length is approaching the preset reading length, it automatically reduces the complexity of subsequent sentences, allowing the remaining content to be expressed within the limited space. Through these cumulative constraint mechanisms, the final generated result can be more stably maintained within the constrained expression space Ω.

[0087] Furthermore, to improve the learning effect of new words, once a new word is identified for inclusion in the current learning content, the content generation module does not simply make it appear once, but registers the new word in the new word reinforcement list. Subsequently, during the subsequent path expansion process, it actively seeks suitable semantic positions for reusing the new word based on the current content context, so that the new word appears repeatedly in the same learning content a preset number of times or a dynamic number of times.

[0088] In this embodiment, the number of repetitions can be dynamically adjusted according to the user's learning level. For example, for users in the beginner stage, the same new word can be repeated three to five times; for users in the basic stage, it can be repeated two to four times; and for users in the intermediate and advanced stages, the number of repetitions is dynamically determined according to the actual content needs. Since the new words are repeated in a natural context, it can improve the user's ability to recognize target vocabulary and the long-term memory effect, while avoiding the learning burden caused by rote memorization.

[0089] After constructing and expanding the restricted expression path tree, the target language content that satisfies all constraints of the restricted expression space Ω is obtained. This target language content can take various forms, such as short reading passages, continuous stories, scene dialogues, news briefs, interest-themed content, exam training materials, or high-frequency expression training content. At this point, the content generation module completes the content generation process for this round of learning and sends the results to the native language mapping output module.

[0090] S105 outputs the generated content to the native language-assisted learning interface.

[0091] After obtaining the target language content, the native language mapping output module performs native language mapping processing on the target language content according to the current user's learning level and auxiliary learning needs, and generates a corresponding native language auxiliary learning interface.

[0092] In this embodiment, native language mapping does not simply translate the entire content into the user's native language. Instead, it provides progressive support to users while preserving the learning value of the target language. Specifically, it first identifies familiar and new words in the target language content based on the user's vocabulary familiarity level, and then employs different display strategies for different types of vocabulary.

[0093] For familiar words, the system prioritizes maintaining the original display format of the target language to enhance the user's ability to directly recognize already mastered vocabulary. For new words, corresponding native language prompts are configured based on the current level of assistance. For example, in the initial learning stage, a native language wrapping display method can be used, that is, the native language definition is displayed simultaneously after the target language vocabulary. As the user's learning ability improves, the assistance intensity is gradually reduced, switching to keyword prompts, click-to-expand assistance display, or fade-out assistance display.

[0094] Furthermore, in some implementations, a bilingual comparative display mode can also be adopted. The target language content and the native language content are displayed synchronously according to paragraph correspondence, allowing users to gradually establish a semantic mapping relationship between the target language and their native language while understanding the meaning of the content. When the system detects that the user has correctly identified the corresponding words multiple times consecutively, it can automatically reduce the display ratio of the native language and increase the display ratio of the target language, thereby forming a learning path that gradually reduces the reliance on the native language.

[0095] Meanwhile, the native language mapping output module can also dynamically adjust the page layout based on the display parameters output by the terminal adaptation module. For example, for users with weaker reading skills, the proportion of auxiliary information can be appropriately increased; for users with stronger reading skills, the continuous display of target language content is prioritized. Through these methods, users at different learning stages can obtain an auxiliary learning experience that matches their ability level.

[0096] S106 records user click behavior, reading speed, recognition results, dwell time, completion rate, and other behavioral data.

[0097] Once the target language content is displayed on the learning terminal, the system begins recording various behavioral data of the user during the learning process, and uses this behavioral data as an important basis for subsequent updates to the learning status.

[0098] The system monitors users' reading progress in real time, including start and end times, page dwell time, scrolling activity, and completion status. It also records user interactions during the learning process, such as the number of times they view definitions, click on helpful tips, actively expand on native language content, and mark content as mastered.

[0099] For learning content containing new words, the system further records the user's recognition results for the new words. For example, when a new word reappears, the system can determine the user's mastery of the new word by whether the user calls upon auxiliary prompts, whether they correctly complete the corresponding exercises, and whether they can complete the reading naturally. At the same time, the system also tracks indicators such as the user's reading speed, continuous reading completion rate, and overall learning activity level for the current learning task.

[0100] After collecting learning behavior data, the obtained data is structured and formed into a corresponding behavioral feedback data set. This behavioral feedback data set is then sent to the vocabulary progression module to update the user's vocabulary proficiency and growth status.

[0101] S107, Update user's familiarity with words and growth status.

[0102] After obtaining user behavior feedback data, the familiar word progression module updates the current familiar word status and growth status of the user based on the user's actual learning performance.

[0103] Specifically, the mastery status of newly added words involved in this round of learning is first assessed. For the same new word, a comprehensive judgment can be made based on its number of reads passed, number of correct recognitions, number of active uses, and number of times auxiliary prompts are invoked. When a new word is detected to continuously meet the preset mastery conditions, the new word is moved from the new word set to the user's familiar word set.

[0104] For example, in one implementation, when a newly added word can be correctly identified by the user for multiple consecutive learning cycles, and can be understood without calling auxiliary prompts during reading, it can be determined that the newly added word has entered a stable mastery state and is included in the familiar word status data.

[0105] In addition to vocabulary growth, the system also updates the user's overall growth status. This growth status includes changes in current vocabulary size, learning level, reading ability, and learning stability. When the system detects a continuous increase in a user's vocabulary size and a significant improvement in reading ability, it can trigger the expansion mechanism of the restricted expression space Ω.

[0106] Specifically, the system automatically expands the allowed vocabulary range and appropriately increases the upper limit on the number of new words. Simultaneously, it dynamically adjusts the target language and native language support ratios based on the user's current learning progress. For example, as the user's learning ability improves, the native language support ratio can be appropriately reduced while the target language display ratio is increased, thus promoting the user's gradual development of the ability to directly understand using the target language.

[0107] S108, reconstruct the next round of restricted expression space Ω based on the updated user learning status data, and generate subsequent learning content.

[0108] Once the familiar word progression module updates the user's familiar word status and growth status, the continuous supply module reconstructs the restricted expression space Ω corresponding to the next round of learning tasks based on the latest learning status, thus forming a continuously progressive learning loop.

[0109] First, the user status data updated in step S107 is read, and parameters such as the current number of familiar words, the growth rate of familiar words, the reading completion rate, the historical correct recognition rate, the reading speed, and the current learning level are extracted. Then, based on these parameters, the range of the new word window, the difficulty level of the target content, and the proportion of native language assistance are recalculated. Unlike the initial restricted expression space Ω constructed in step S103, the restricted expression space Ω constructed at this time already reflects the user's growth status after this round of learning; therefore, its boundary parameters will change accordingly.

[0110] For example, when the system detects that a user maintains a high reading completion rate over multiple consecutive learning cycles and the recognition rate of new words continues to improve, the range of candidate new words can be appropriately expanded, while the level of allowed grammatical structures can be increased. Conversely, when the system detects that a user's reading speed has decreased significantly recently or the frequency of auxiliary prompts has increased significantly, the new word window can be appropriately narrowed to reduce the complexity of the content, so that subsequent learning content returns to the user's acceptable learning load range.

[0111] Furthermore, in this embodiment, the continuous supply module adjusts not only the boundary parameters at the vocabulary level but also simultaneously adjusts the boundary parameters at the content level. For example, when a user is primarily interested in technology-related content, the weight of technology-related content is increased when reconstructing the restricted expression space Ω; when a user prefers reading about campus life or daily conversations, corresponding content is generated first. This ensures that subsequent learning content is both appropriate for the user's current ability level and maintains a high level of learning interest and sustained engagement.

[0112] After completing the reconstruction of the restricted expression space Ω, the system executes the restricted expression path tree generation process described in step S104 again, constructing the next round of learning content within the new restricted expression space Ω. Since the learning results of each round are fed back into the construction process of the next round of restricted expression space Ω, the system can form a closed-loop learning mechanism encompassing content generation, learning behavior collection, familiar word growth, space expansion, and content regeneration, ensuring that the user's ability improvement process keeps pace with the increase in the difficulty of the learning content.

[0113] Through the above processing, this application no longer adopts the supply model of fixed levels, fixed textbooks or fixed content libraries in traditional graded reading systems. Instead, it continuously reconstructs the limited expression space Ω according to the user's growth status and dynamically generates new learning content within the limited expression space Ω, thereby realizing a continuous, individualized and progressive language learning process.

[0114] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A progressive language learning content generation system based on a constrained expression space, characterized in that, include: The user status acquisition module is used to acquire user learning status data, which includes user vocabulary familiarity status, learning level, reading history and interaction behavior data. The restricted expression space construction module is used to construct a corresponding restricted expression space Ω based on the user learning state data. The restricted expression space Ω includes lexical constraints, grammatical constraints, content constraints, and user behavior constraints. The content generation module is used to perform constraint determination of the restricted expression space Ω in real time during the generation process, and only allows the path construction and content generation of language content to be completed within the restricted expression space Ω; The native language mapping output module is used to convert the generated content into a native language-assisted learning interface; The vocabulary familiarity module is used to update the user's vocabulary familiarity status based on the user's learning behavior; The continuous supply module is used to reconstruct the restricted expression space Ω based on the updated user learning status data and generate the next round of learning content; The content generation module does not include a generation method that first generates content outside the restricted expression space Ω and then filters or maps it to the restricted expression space Ω; the boundary parameters of the restricted expression space Ω are dynamically updated based on user learning behavior data.

2. The system according to claim 1, characterized in that, The vocabulary constraints include at least one of the following: a set of familiar user words, a set of candidate new words, an upper limit on the number of new words added in this round, a word frequency level range, and a target part-of-speech distribution range.

3. The system according to claim 1, characterized in that, The grammatical constraints include at least one of the following: sentence length range, tense complexity, number of clauses, range of sentence variation, and grammatical structure level.

4. The system according to claim 1, characterized in that, The user behavior constraints include at least one of the following: reading dwell time, continuous reading completion rate, familiar word confirmation rate, reading speed, historical correct recognition rate, and learning activity.

5. The system according to claim 1, characterized in that, After detecting an improvement in the user's familiarity with words, the familiarity word progression module automatically expands the vocabulary range of the restricted expression space Ω and adjusts at least one of the following: the upper limit of the number of new words, the proportion of the target language, and the proportion of the native language as an auxiliary language.

6. The system according to claim 1, characterized in that, The native language mapping output module adopts at least one of the following display methods: native language wrapping display, bilingual comparison display, familiar word hidden display, keyword prompt display, fade-out auxiliary display, and click-to-expand auxiliary display.

7. The system according to claim 1, characterized in that, The continuous supply module dynamically adjusts the difficulty level, text length, and number of new words for the next round of learning content based on the user's historical correct recognition rate, the growth rate of familiar words, the reading completion rate, or the reading speed.

8. The system according to claim 1, characterized in that, The system automatically adjusts the output content density, layout structure, segmentation structure of content per screen, display density of target language per line, font size, or auxiliary display ratio based on the terminal device's screen size, display orientation, visible area size, or interaction method.

9. The system according to claim 1, characterized in that, The boundary parameters of the restricted expression space Ω also include at least one of the following: target language presentation density, upper limit of target language vocabulary on a single screen, upper limit of target language display on a single line, native language auxiliary ratio, or current visible area load threshold; the system uses the boundary parameters to control the intensity of target language presentation of generated content in the native language assisted learning interface.

10. The system according to claim 1, characterized in that, The content generation module enables new words to appear repeatedly in the same learning content a preset number of times or dynamically, thereby enhancing the user's memory reinforcement effect.