English teaching story intelligent generation method and device, equipment and medium

By introducing a pragmatic repetition distribution constraint mechanism into the generation of English teaching stories, and combining semantic intent vectors and retrieval-enhanced generation techniques, the grammatical structure and positional distribution of target learning vocabulary are controlled. This solves the problems of insufficient pragmatic clarity and teaching guidance in the generated content in existing technologies, and achieves higher accuracy and teaching effectiveness.

CN120952157APending Publication Date: 2025-11-14ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511040367.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing English teaching story generation methods based on RAG technology lack cross-language transfer control capabilities, resulting in insufficient pragmatic clarity, contextual typicality, and grammatical diversity in the generated content, which affects the effectiveness and accuracy of teaching guidance.

Method used

A pragmatic repetition distribution constraint mechanism is introduced, which controls the frequency and position distribution of the grammatical structure of the target vocabulary in the generated content through semantic intent vectors. Combined with retrieval-enhanced generation technology, the usage status of vocabulary is dynamically tracked and the generation probability distribution is adjusted to ensure that vocabulary appears reasonably and repeatedly in the language generation process.

Benefits of technology

It significantly improved the accuracy of English teaching story generation and the effectiveness of teaching guidance, achieved the repeated appearance of diverse grammatical forms and reasonable contextual rhythms of target vocabulary, and enhanced language internalization ability and the readability of teaching stories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952157A_ABST
    Figure CN120952157A_ABST
Patent Text Reader

Abstract

The invention provides an English teaching story intelligent generation method and device, equipment and a medium, and relates to the technical field of the English teaching story intelligent generation method and device, and the method comprises the steps: obtaining the input content of a user; performing semantic analysis on the input content, and identifying a semantic intention vector and a target learning vocabulary of the user; based on the semantic intention vector, utilizing a retrieval enhancement generation technology to execute query retrieval in a preset corpus, and extracting text paragraphs from the query retrieval; constructing a retrieval result containing the target learning vocabulary based on the text paragraph; a pragmatic repeated distribution constraint mechanism is applied, and the pragmatic repeated distribution constraint mechanism is used for controlling the occurrence frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; and taking the semantic intention vector and the retrieval result as input, and generating target English story content in combination with a pragmatic repetitive distribution constraint mechanism. According to the method, the distribution of the target vocabularies can be controlled, so that the accuracy of English teaching story generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field of this application specifically relates to a method, device, equipment, and medium for intelligently generating English teaching stories. Background Technology

[0002] Currently, English teaching technologies based on large models generally adopt context-aware retrieval enhancement generation mechanisms. By semantically encoding user input, they perform context-related retrieval operations in large-scale corpora to extract fragments or sentence structures related to the learning objectives, and generate teaching content containing teaching words accordingly. The selection of teaching words is based on context relevance, vocabulary difficulty gradient, and language usage frequency, thereby achieving personalized English learning support tailored to individual needs.

[0003] RAG (Retrieval-Enhanced Generative Architecture) technology is an architecture that integrates generative language models and external knowledge retrieval mechanisms. It first encodes the input semantics, retrieves the most semantically relevant text fragments as contextual clues, and then the generative model reconstructs the language from the retrieval results to generate the target text. Existing multi-level large-scale model planning methods for generating English teaching stories based on RAG technology typically construct multi-layered intent parsing and target vocabulary planning modules. First, it identifies the user's language proficiency intent and target vocabulary set within their learning objectives. Then, it uses the RAG mechanism to retrieve plot fragments and language expressions matching the intent from a corpus. Combining language level and vocabulary mastery, it dynamically adjusts the complexity and language structure of the plot, guiding the target vocabulary to be naturally embedded into the story text in stages. It also applies positional control and syntactic diversification to recurring teaching target vocabulary, resulting in semantically consistent, highly targeted, and progressively advancing English teaching story content.

[0004] While existing multi-level large-scale model-based methods for generating English teaching stories using RAG technology have improved content relevance and teaching relevance by introducing retrieval enhancement mechanisms, they still suffer from insufficient modeling of the "vocabulary input—semantic transfer—language internalization" path in teaching tasks due to the lack of cross-language transfer control capabilities in the generation model itself. Traditional generative large-scale language models typically focus on the fluency and narrative logic of language when generating teaching stories, but neglect the distribution control of target vocabulary in pragmatic clarity, contextual typicality, and grammatical diversity. This results in generated content that, while natural and fluent, has limited teaching guidance value and fails to effectively achieve deep mastery of target vocabulary and language transfer. Consequently, the generated English teaching stories do not meet user needs and have low accuracy. Therefore, a method is needed to control the distribution of target vocabulary to improve the accuracy of English teaching story generation. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for intelligent generation of English teaching stories, which can control the distribution of target vocabulary, thereby improving the accuracy of English teaching story generation.

[0006] The first aspect of this application provides a method for intelligently generating English teaching stories, the method comprising: Get the user's input; The input content is semantically parsed to identify the user's semantic intent vector and target vocabulary. Based on the semantic intent vector, a search enhancement generation technique is used to perform a query search in a preset corpus and extract text paragraphs from it. Based on the text paragraph, construct search results containing the target vocabulary for learning; A pragmatic repetition distribution constraint mechanism is applied, wherein the pragmatic repetition distribution constraint mechanism is used to control the frequency and positional distribution of the grammatical structure of the target learning vocabulary in the generated content; The semantic intent vector and the retrieval results are used as input, and the target English story content is generated by combining the pragmatic repetition distribution constraint mechanism.

[0007] Based on the above technical solutions, preferably, the step of using retrieval enhancement generation technology based on the semantic intent vector to perform a query retrieval in a preset corpus and extract text paragraphs therefrom specifically includes: The semantic intent vector is mapped to the target dimension vector space using vectorization technology to form a query vector that can match the text data in the preset corpus, wherein the corpus contains the structured text data; A matching mechanism based on vector similarity calculation is used to calculate the similarity between the query vector and each text data in the preset corpus; Select candidate text data from multiple text data that have a similarity to the query vector that meets a preset threshold; The candidate text data is filtered by semantic relevance to obtain the text paragraphs.

[0008] Based on the above technical solutions, preferably, the pragmatic repetition distribution constraint mechanism is used to control the frequency and positional distribution of the grammatical structure of the target learning vocabulary in the generated content, specifically including: Set pragmatic repetition distribution constraint parameters for the target vocabulary; Dynamically track the cumulative occurrence count and current position distribution of the target vocabulary in the generated content; The cumulative occurrence count and the current position distribution are compared with the pragmatic repetition distribution constraint parameters in real time. When the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameters, the generation probability distribution is adjusted to improve the usage priority of the target learned vocabulary in subsequent texts.

[0009] Based on the above technical solutions, preferably, after the application repetition distribution constraint mechanism, the method further includes: Based on the context dependency control mechanism, it is determined whether the target vocabulary has already been semantically expressed in the previous segment. When the semantic expression has been closed, the continuous generation of the target vocabulary is suppressed through a probability cooling strategy.

[0010] Based on the above technical solutions, preferably, when the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameter, adjusting the generation probability distribution specifically includes: When the cumulative number of occurrences does not reach the preset occurrence frequency threshold, or when the current position distribution does not conform to the context distribution pattern, the word probability distribution of the current decoding step is reweighted. In the reweighting operation, an adaptive boosting factor is applied to the lexical items corresponding to the target learning vocabulary in the current vocabulary candidate list. The boosting factor is calculated based on the degree of usage absence of the target learning vocabulary, and the degree of absence is directly proportional to the increase in the generation probability.

[0011] Based on the above technical solutions, preferably, the step of using the semantic intent vector and the retrieval results as input, and combining them with the pragmatic repetition distribution constraint mechanism to generate target English story content, specifically includes: The text segments in the retrieval results are weighted based on the attention mechanism, and the text segments whose relevance to the current generated state meets the preset requirements are selected for language modeling. The semantic intent vector guides the direction of content generation; The pragmatic repetition distribution constraint mechanism is embedded in the generation process to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; When it is found that the cumulative occurrence count or current position distribution of the target learning vocabulary does not meet the pragmatic repetition distribution constraint parameter, the generation probability distribution is adjusted, and the target learning vocabulary is guided to be embedded in the target English story content in a grammatical structure that has not appeared before.

[0012] Based on the above technical solutions, preferably, the step of semantically parsing the input content to identify the user's semantic intent vector and target learning vocabulary specifically includes: Convert the input content into a feature vector; By performing semantic analysis on the input content, the semantic information of user needs can be obtained; Based on the semantic information of user needs, the target learning vocabulary is identified and extracted according to the frequency, contextual association and importance of different words in the input content and in the user's learning path. Convert the target vocabulary into vocabulary intent vectors; The semantic intent vector is constructed based on the feature vector and the lexical intent vector.

[0013] A second aspect of this application provides an intelligent English teaching story generation device, the device being used to execute an intelligent English teaching story generation method as described in any of the above-described methods, the device comprising an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire the user's input content; The processing module is used to perform semantic parsing on the input content and identify the user's semantic intent vector and target learning vocabulary; The processing module is used to perform a query retrieval in a preset corpus based on the semantic intent vector and using retrieval enhancement generation technology to extract text paragraphs from it; The processing module is used to construct retrieval results containing the target learning vocabulary based on the text paragraph; The processing module is used to apply a pragmatic repetition distribution constraint mechanism, wherein the pragmatic repetition distribution constraint mechanism is used to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; The output module is used to take the semantic intent vector and the retrieval results as input, and combine them with the pragmatic repetition distribution constraint mechanism to generate target English story content.

[0014] Based on the above technical solutions, preferably, the processing module is used to map the semantic intent vector to a target dimension vector space through vectorization technology to form a query vector that can match the text data in the preset corpus, wherein the corpus contains the text data that has been structured. The processing module is used to calculate the similarity between the query vector and each text data in the preset corpus using a matching mechanism based on vector similarity calculation. The processing module is used to select candidate text data from a plurality of text data that have a similarity to the query vector that meets a preset threshold. The processing module is used to filter the candidate text data based on semantic relevance to obtain the text paragraph.

[0015] Based on the above technical solutions, preferably, the processing module is used to set pragmatic repetition distribution constraint parameters for the target learning vocabulary; The processing module is used to dynamically track the cumulative occurrence and current position distribution of the target learning vocabulary in the generated content; The processing module is used to compare the cumulative occurrence count and the current position distribution with the pragmatic repetition distribution constraint parameter in real time. When the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameter, the module adjusts the generation probability distribution to improve the usage priority of the target learning vocabulary in subsequent texts.

[0016] Based on the above technical solutions, preferably, the processing module is used to determine whether the target learning vocabulary has been semantically expressed in a previous segment based on a context dependency control mechanism. When the semantic expression has been closed, the continuous generation of the target learning vocabulary is suppressed by a probability cooling strategy.

[0017] Based on the above technical solutions, preferably, the processing module is used to reweight the word probability distribution of the current decoding step when the cumulative number of occurrences does not reach the preset occurrence frequency threshold, or when the current position distribution does not conform to the context distribution pattern. The processing module is used to apply an adaptive boosting factor to the lexical items corresponding to the target learning vocabulary in the current vocabulary candidate list during the reweighting operation. The boosting factor is calculated based on the degree of usage absence of the target learning vocabulary, and the degree of absence is directly proportional to the increase in the generation probability.

[0018] Based on the above technical solutions, preferably, the processing module is used to assign weights to text segments in the retrieval results based on an attention mechanism, and select text segments whose relevance to the current generation state meets preset requirements for language modeling; The processing module is used to guide the generation direction of the generated content based on the semantic intent vector; The processing module is used to embed the pragmatic repetition distribution constraint mechanism during the generation process to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content. The output module is used to adjust the generation probability distribution and guide the target learning vocabulary to be embedded in the target English story content in a non-occurring grammatical structure form when it is found that the cumulative occurrence frequency or current position distribution of the target learning vocabulary does not meet the pragmatic repetition distribution constraint parameter.

[0019] Based on the above technical solutions, preferably, the output module is used to convert the input content into a feature vector; The output module is used to obtain user demand semantic information by performing semantic analysis on the input content; The output module is used to identify and extract the target learning vocabulary based on the semantic information of the user's needs, according to the frequency, contextual association and importance of different words in the input content and in the user's learning path. The output module is used to convert the target learning vocabulary into a vocabulary intent vector; The output module is used to construct the semantic intent vector based on the feature vector and the lexical intent vector.

[0020] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0021] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0022] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. This application introduces a pragmatic repetition distribution constraint mechanism and combines semantic intent vectors to implement precise control over the generated content. During the language generation process, it dynamically tracks the cumulative occurrence frequency and contextual position distribution of target learning vocabulary. When the preset pragmatic distribution requirements are not met, it adjusts the generation probability and embedding structure in real time to ensure that target learning vocabulary appears repeatedly in diverse grammatical forms and reasonable contextual rhythms. This achieves controllability of vocabulary distribution, significantly improves the coverage of target learning vocabulary and teaching guidance effect in the generated English teaching story content, and thus improves the accuracy of the generated results and language internalization ability.

[0023] 2. Based on semantic intent vectors, and utilizing retrieval-enhanced generation technology, queries are performed in a pre-set corpus to extract text paragraphs. This enables the acquisition of semantically relevant texts that meet user learning needs. Through vectorized matching and semantic relevance filtering, the selected text paragraphs are ensured to not only fit the context but also possess pedagogical value in their language structure, thereby providing accurate and highly relevant language materials for subsequent story generation.

[0024] 3. By applying a pragmatic repetition distribution constraint mechanism, the frequency and positional distribution of the grammatical structure of the target vocabulary in the generated content can be controlled. This enables the controllable output of the target vocabulary in the text. By setting repetition parameters and dynamically monitoring usage, it ensures that the vocabulary is evenly distributed, meets the frequency standard, and has diverse structures during the language generation process, effectively improving its pragmatic clarity and teaching reinforcement effect.

[0025] 4. Based on the context dependency control mechanism, it can determine whether the target vocabulary has completed semantic expression and suppress continuous generation through a probability cooling strategy. This can avoid excessive word stacking in phrases, improve the naturalness of language generation rhythm and the coherence of content progression, enhance the rhythm of vocabulary semantic transfer, and improve the readability and semantic hierarchical structure of the overall teaching story.

[0026] 5. The scheme of improving the priority of target vocabulary usage by adjusting the generation probability distribution can adaptively intervene in the generation probability of word units when vocabulary usage is detected to be missing. By reweighting and introducing probability enhancement factors, it ensures that target vocabulary appears in the generated content in a timely manner and has grammatical adaptability and contextual consistency, thereby enhancing the achievement rate of vocabulary teaching objectives.

[0027] 6. By using semantic intent vectors and retrieval results as inputs and combining them with a pragmatic repetition distribution constraint mechanism to generate target English story content, it can achieve accurate language generation under multi-source conditions. Through the integration of the retrieval paragraph context, semantic intent guiding the structural direction, and pragmatic constraints controlling the vocabulary distribution, it can construct personalized English story texts that are pedagogically targeted, have a strong emphasis on vocabulary transferability, and are natural in language.

[0028] 7. The steps of semantic parsing the input content and constructing a semantic intent vector can transform the user's free input into a structured expression target. Through feature extraction, vocabulary selection and semantic modeling, the target learning vocabulary and learning intent are accurately extracted, providing a clear semantic control foundation for the subsequent retrieval and generation process, and realizing the mapping of teaching tasks to executable generation paths. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating an intelligent method for generating English teaching stories disclosed in an embodiment of this application; Figure 2 This is a schematic diagram of a module of an intelligent English teaching story generation device disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0030] Explanation of reference numerals in the attached drawings: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation

[0031] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0032] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0033] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0034] While existing multi-level large-scale model planning methods for generating English teaching stories based on RAG technology possess context awareness and teaching relevance optimization mechanisms, enabling them to dynamically adjust teaching content according to user intent and language proficiency, the generation models lack the ability to model the "vocabulary input—semantic transfer—language internalization" path. This results in a failure to effectively control the pragmatic clarity, contextual typicality, and syntactic diversity of target vocabulary in the generated content. Consequently, while the generated content may be fluent, its teaching guidance is insufficient, and the language transfer and internalization effects of target vocabulary are poor. This, in turn, affects the accuracy of English teaching story generation and personalized learning outcomes. Therefore, it is urgent to introduce a target vocabulary distribution control mechanism to improve the quality of teaching story generation and language transfer capabilities.

[0035] This embodiment discloses an intelligent method for generating English teaching stories, referring to... Figure 1 This includes the following steps S110-S160: S110, Obtain user input.

[0036] The English teaching story intelligent generation method disclosed in this application is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running the English teaching story intelligent generation method. The server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0037] When receiving user input, a natural language access module can be used to receive and structure the user's free text. This process supports users inputting language data in various forms, such as statements, questions, keyword suggestions, or task instructions, without limiting the input length or expression style, thus adapting to the expression habits of different learners. The input content is directly converted into a standardized text string, preserving the syntactic hierarchy and segmentation boundary information in the original language order.

[0038] In implementation, an input channel can be set up to receive English teaching requests such as: "I want to learn the word 'absorb' and hope it appears in an English story about a science experiment." This input includes the target vocabulary word "absorb" and the contextual indication of the preferred context "science experiment." Upon receiving this input, lexicalization, part-of-speech tagging, entity extraction, and dependency analysis are performed sequentially, while preserving the semantic connections between words. For example, "absorb" is identified as a verb, possessing typical action attributes in an experimental context.

[0039] Next, a vocabulary filtering operation is performed. Based on the language usage frequency database, the teaching vocabulary grading standard, and the context co-occurrence pattern, it is determined whether the vocabulary has teaching value. If it meets the criteria, it is added to the learning vocabulary set as a target learning vocabulary. Further, a learning intent embedding is constructed based on the syntactic clues and semantic orientations in the input sentence. This embedding includes feature fields such as target learning vocabulary, context category, and expected style, which are then encoded into a semantic intent vector.

[0040] This semantic intent vector serves as the input condition for subsequent semantic retrieval and generative modeling, thereby constraining the consistency of teaching objectives and guiding contextual adaptability. This ensures that the generated content can revolve around the vocabulary to be learned and maintain the desired story context, thus ensuring that the language generation results have high teaching relevance and personalized learning value.

[0041] S120 performs semantic parsing on the input content to identify the user's semantic intent vector and target learning vocabulary.

[0042] In one possible implementation, semantic parsing of the input content is performed to identify the user's semantic intent vector and target learning vocabulary. Specifically, this includes: converting the input content into a feature vector; obtaining user demand semantic information through semantic analysis of the input content; identifying and extracting target learning vocabulary based on the frequency, contextual relevance, and importance of different words in the input content and in the user's learning path, based on the user demand semantic information; converting the target learning vocabulary into a vocabulary intent vector; and constructing a semantic intent vector based on the feature vector and the vocabulary intent vector.

[0043] Specifically, converting input content into feature vectors refers to the operation of embedding user-input natural language text into a semantic encoding model. This process employs a multi-layer semantic encoder based on the Transformer architecture. Each word in the original text is used as input, and through nested computation of positional encoding, contextual interaction, and attention mechanisms, the entire sentence or paragraph-level input is mapped into a vector representation with a fixed dimension. This feature vector not only contains word order, syntactic structure, and semantic information, but also retains pragmatic features such as emotional tone and contextual cues. For example, if a user inputs "I want to study the word absorb in a science-themed story," the generated feature vector will represent their learning objective, lexical focus, and story type preference in a high-dimensional semantic space, serving as the foundation for subsequent learning intent modeling and contextual retrieval.

[0044] Semantic analysis of the input content yields user demand semantic information. This involves structurally decomposing the feature vectors after semantic encoding using a semantic annotator and intent recognition module to identify the core language task types and learning scenario preferences within the input text. Semantic analysis includes action verb analysis (e.g., "learn" or "study" indicating learning intent), contextual keyword extraction (e.g., "science-themed" indicating story background preference), and contextual co-structure recognition (e.g., the dependency structure between words and scenarios). The resulting user demand semantic information is a structured set of semantic descriptions, containing semantic metadata such as task type, vocabulary target, context category, and expression style, used to support target vocabulary selection and content generation control.

[0045] Based on the semantic information of user needs, the system identifies and extracts target learning vocabulary according to the frequency, contextual relevance, and importance of different words in the input content and in the user's learning path. This means that the system evaluates the pedagogical value of all words appearing in the input text and extracts a set of learning vocabulary that can be used for generation process control through a joint decision based on three indicators. Word frequency reflects the salience of the word in the current input; contextual relevance is used to assess the semantic fit between the word and the user's intent; and the importance of the learning path is dynamically evaluated based on the user's language proficiency, vocabulary acquisition history, and the setting of the teaching task. For example, in the aforementioned input, "absorb" is identified as a target learning vocabulary and enters the vocabulary control list because it is explicitly mentioned and has typical scientific semantic expression, satisfying multiple conditions such as significant word frequency, contextual matching, and high semantic weight.

[0046] Converting target vocabulary into lexical intent vectors refers to independently encoding the identified target vocabulary to form lexical-level control instructions that guide the generation process. These lexical intent vectors not only include the semantic embedding of the word itself but also integrate the semantic role, preference structure, and teaching priority carried by the word in user input, forming a priori control over how the target vocabulary is used in the generated content. For example, if "absorb" is identified as a core learning word in the aforementioned input, its lexical intent vector will explicitly require that the word appear as a verb in the story, be contextualized within experimental scenarios, and possess diverse grammatical usage requirements.

[0047] Constructing a semantic intent vector based on feature vectors and lexical intent vectors involves fusing the overall semantic representation corresponding to the user input with the control vector of the target vocabulary to form a high-dimensional composite vector that guides the entire process of retrieval and generation. During the fusion process, a multi-head attention mechanism is used to structurally align the feature vectors and lexical intent vectors, and the semantic dependencies between them are encoded through a fusion layer. The generated semantic intent vector simultaneously possesses language task orientation, vocabulary generation constraints, and context adaptability. This semantic intent vector is then used as the main control signal input to the retrieval enhancement generation process, ensuring that subsequent content generation closely adheres to the target vocabulary, responds to user preferences, and maintains semantic consistency and alignment with teaching objectives.

[0048] S130, based on semantic intent vectors, uses retrieval enhancement generation technology to perform a query retrieval in a preset corpus and extract text paragraphs from it.

[0049] In one possible implementation, based on semantic intent vectors, a retrieval enhancement generation technique is used to perform a query retrieval in a preset corpus to extract text segments. Specifically, this includes: mapping the semantic intent vectors to a target dimension vector space using vectorization technology to form a query vector that can match text data in the preset corpus, which contains structured text data; calculating the similarity between the query vector and each piece of text data in the preset corpus using a matching mechanism based on vector similarity calculation; selecting candidate text data from multiple text data whose similarity to the query vector meets a preset threshold; and filtering the candidate text data based on semantic relevance to obtain text segments.

[0050] Specifically, mapping semantic intent vectors to a target-dimensional vector space using vectorization techniques to form query vectors capable of matching text data in a pre-defined corpus involves further adjusting the semantic intent vectors, which already incorporate features such as semantic tasks, target vocabulary, and contextual preferences, to a vector space consistent with the corpus's vector index structure. This ensures computational feasibility and retrieval efficiency in the matching operation. Vectorization techniques typically employ projection functions. , to transform high-dimensional semantic intent vector Projected onto a target-dimensional vector space consistent with the corpus vector structure. ,in The dimensions must be consistent with the vectorized dimensions of each text segment in the corpus. For example, for a corpus vector indexing system built on BERT, the input vector needs to be adjusted to 768 dimensions to form a standard query vector. .

[0051] A matching mechanism based on vector similarity calculation is employed to calculate the similarity between the query vector and each text data in a pre-defined corpus. This involves measuring the degree of similarity between the query vector and the semantic vector corresponding to each text segment in the corpus to determine their semantic relevance. Commonly used similarity calculation methods include cosine similarity and Euclidean distance, where cosine similarity is defined as: in Represents the query vector. Indicates the first in the corpus The embedding vectors of each text segment are calculated. This result measures semantic alignment and serves as a ranking metric for the relevance of text segments in the corpus. A result closer to 1 indicates a greater semantic similarity between the text data and the query vector.

[0052] Selecting candidate text data from multiple text datasets that meet a preset threshold in similarity to the query vector refers to setting a similarity threshold based on the aforementioned similarity calculation results. The process involves selecting multiple text segments that meet a set of semantic consistency criteria with the query vector, and using these segments as a candidate text dataset. The selection criteria are typically set as follows: if a certain text segment... If the similarity is greater than or equal to a preset threshold, it is included in the candidate set. In practical applications, the similarity threshold is generally set above 0.75 to ensure that the selected candidate text data has a high degree of semantic consistency with the semantic intent vector. For example, when the query vector expresses "the way the verb 'absorb' is expressed in the context of scientific experiments", paragraphs containing expressions such as "absorb light" or "absorb energy during reaction" are selected first.

[0053] Selecting text paragraphs from candidate text data by semantic relevance refers to the process of filtering candidate text data sets. Each paragraph in the search results undergoes further semantic analysis to determine whether it meets the usage requirements of the target vocabulary at the lexical, syntactic, and pragmatic levels. The best-matching set of paragraphs is then selected as the final retrieval output. This semantic relevance assessment includes two core processes: first, determining whether the target vocabulary is used in a teachable manner within the paragraph, meaning it must appear as a core component with clear context; and second, determining whether the linguistic context of the paragraph aligns with the context described in the semantic intent vector. For example, if the target vocabulary is "absorb" and the intent is "scientific experiments," paragraphs containing "absorb radiation during chemical reactions" are prioritized while those containing "absorb the news" are discarded. The final selected set of text paragraphs will be used as the retrieval results and input into the subsequent language generation process, providing contextual corpus support for the generation of the target English story content.

[0054] S140, construct search results containing the target learning vocabulary based on text paragraphs.

[0055] First, for the selected set of text paragraphs, the target vocabulary recognition operation is performed paragraph by paragraph. This operation is based on a lexical-level positive matching and word form restoration mechanism, detecting whether any variant form of the target vocabulary (such as root words, tense changes, nominalized forms, etc.) exists in the text paragraphs. This ensures that even if the text does not use the basic form of the target vocabulary, the semantically consistent actual expression can still be identified. For example, if the target vocabulary is "absorb", then the appearance of "absorbing", "absorbed", and "absorption" in the paragraph is considered to meet the condition of the target vocabulary's presence.

[0056] Next, the paragraphs containing the target vocabulary are further analyzed to extract the grammatical role (e.g., subject, predicate, object, modifier) ​​and syntactic structure (e.g., declarative, interrogative, passive) of the target vocabulary. This step is performed jointly by a dependency parser and a grammatical tag extractor, aiming to determine whether the target vocabulary has pedagogical value in language expression. Paragraphs where the target vocabulary is a core syntactic component and has a clear pragmatic context are prioritized for retention, while paragraphs where the target vocabulary appears marginalized, has unclear grammatical dependencies, or is ambiguous in context are filtered out. For example, paragraphs like "the material absorbs infrared radiation efficiently" are retained, while paragraphs like "a discussion about absorption followed" are discarded.

[0057] Then, contextual diversity and structural coverage assessments are performed across multiple candidate paragraphs to ensure that the target vocabulary is presented in diverse ways in the final search results. This assessment includes calculating lexical contextual variability, analyzing the distribution of syntactic structure types, and measuring sentence complexity to guarantee that the selected paragraphs provide a cross-scenario, cross-structural language transfer foundation for the target vocabulary in subsequent generation. For example, paragraphs in which the target vocabulary appears in active, passive, and subordinate clause structures are prioritized to avoid monotonous expression during language generation.

[0058] Finally, the set of text paragraphs that meet the selection criteria is output as the search results. These results, while satisfying semantic relevance, further ensure that the target vocabulary has significant visibility, grammatical clarity, and pragmatic typicality at the linguistic content level. This provides directly embeddable semantic support for the subsequent generation process and a distribution benchmark for the pragmatic repetition distribution constraint mechanism. Through this process, the constructed search results not only meet the user's learning intent but also ensure that the generated story content revolves around the target vocabulary, thereby improving the teaching effectiveness and accuracy of the generated story content.

[0059] S150, applying the repetitive distribution constraint mechanism.

[0060] In one possible implementation, a pragmatic repetition distribution constraint mechanism is applied to control the frequency and positional distribution of the grammatical structure of the target vocabulary in the generated content. Specifically, this includes: setting pragmatic repetition distribution constraint parameters for the target vocabulary; dynamically tracking the cumulative occurrence count and current positional distribution of the target vocabulary in the generated content; comparing the cumulative occurrence count and current positional distribution with the pragmatic repetition distribution constraint parameters in real time; and adjusting the generation probability distribution to increase the priority of the target vocabulary in subsequent texts when the cumulative occurrence count or current positional distribution does not meet the pragmatic repetition distribution constraint parameters.

[0061] Specifically, setting pragmatic repetition distribution constraint parameters for target vocabulary refers to defining a set of constraint indicators for each target vocabulary word before content generation begins, based on teaching objectives and language generation strategies. These constraints limit its usage behavior in the generated text. These parameters include minimum occurrence frequency, maximum occurrence interval, syntactic structure type coverage requirements, and contextual distribution location. For example, the minimum occurrence frequency of the target vocabulary word "absorb" can be set to 3, requiring it to appear in at least one active voice, one passive voice, and one subordinate clause structure, and distributed across the beginning, middle, and end paragraphs of the story to form a rhythmically reasonable and formally diverse pragmatic structure. This parameter set constitutes the baseline strategy for controlling the semantic transfer and language internalization path of the target vocabulary words, and is embedded in the decoding logic as a language scheduling basis during the generation process.

[0062] Dynamically tracking the cumulative occurrence and current position distribution of target vocabulary in generated content refers to recording the actual occurrence of target vocabulary in each segment in real time during text generation and constructing a dynamic vocabulary usage trajectory. This tracking mechanism includes two parallel modules: a vocabulary frequency recorder, which counts the total number of occurrences of the target vocabulary in the generated text; and a context distribution mapper, which records the distribution of the target vocabulary in different paragraphs, sentence structures, and grammatical structures. This information is updated after each decoding step, forming a usage record matrix synchronized with the evolution of the generated content. For example, if "absorb" only appears in the first paragraph as a passive sentence in the first five segments of the generated text, then the current cumulative occurrence is 1, the current position distribution is in the first segment set, and the grammatical structure is simple.

[0063] The cumulative occurrence count and current position distribution are compared in real time with pragmatic repetition distribution constraint parameters. When the cumulative occurrence count or current position distribution does not meet the pragmatic repetition distribution constraint parameters, the generation probability distribution is adjusted to increase the usage priority of the target vocabulary in subsequent texts. This involves comparing the tracked actual usage behavior with the preset target parameters item by item. When there are substandard situations such as insufficient usage frequency, incomplete grammatical structure coverage, or uneven paragraph distribution, a probability intervention mechanism is activated to actively guide the target vocabulary into subsequent generated content. This mechanism reconstructs the lexical decoding probability distribution of the language generation model. Specifically, it identifies the lexical items corresponding to the target vocabulary in the current decoding candidate set and calculates their original generation probabilities. Add a dynamic boosting factor , obtain the adjusted probability ,in Adaptive calculation based on the degree of missing information, such as: in, This represents the current cumulative number of occurrences. To preset the minimum number of occurrences, This is the adjustment coefficient. This method ensures that, without compromising the naturalness of the language, the target vocabulary is prioritized and pushed into the generation path, while maintaining the diversity of its pragmatic structure and the rationality of its contextual logic, thus effectively achieving the teaching objective of repetitive distribution of the target vocabulary.

[0064] In one possible implementation, when the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameters, the generation probability distribution is adjusted. Specifically, this includes: when the cumulative occurrence count does not reach a preset occurrence frequency threshold, or the current position distribution does not conform to the context distribution pattern, a reweighting operation is performed on the word probability distribution of the current decoding step; in the reweighting operation, an adaptive boosting factor is applied to the lexical items corresponding to the target learned words in the current word candidate list. The boosting factor is calculated based on the degree of usage absence of the target learned words, and the degree of absence is proportional to the boosting magnitude of the generation probability.

[0065] Specifically, when the cumulative occurrence count falls below the preset frequency threshold, or the current position distribution does not conform to the context distribution pattern, a reweighting operation is performed on the word probability distribution of the current decoding step. This means that during the decoding stage of the language generation model, the probability distribution in the candidate word set corresponding to the word to be generated is dynamically adjusted to prioritize the output of the target word. Specifically, at each generation step, the model outputs a candidate word list based on the preceding context, with each word having a corresponding initial generation probability value. If the cumulative occurrence count of the current target word is lower than the set minimum frequency threshold, or its occurrence in the context is concentrated in a specific segment (e.g., only appearing at the beginning or end), it is considered not to meet the distribution diversity or rhythmic rationality required by the pragmatic repetition distribution constraint parameters. At this time, the probability reweighting module is activated to redistribute the generation probabilities of all words in the current generation candidate space. The core is to increase the probability weight of the word item corresponding to the target word and appropriately scale the remaining words to maintain overall normalization. For example, if "absorb" does not appear in the middle segment, the system will adjust its original generation probability at the current decoding position from... Upgraded to a new value calculated using weighting. This makes it competitive to enter the generation sequence.

[0066] In the reweighting operation, an adaptive boosting factor is applied to the lexical items corresponding to the target learning vocabulary in the current vocabulary candidate list. This boosting factor is calculated based on the degree of absence of the target learning vocabulary, and the degree of absence is directly proportional to the increase in generation probability. This means that a dynamic adjustment mechanism automatically calculates the increase in generation probability based on the degree of absence of the target learning vocabulary in the current generated content, forming a differentiated generation priority control strategy. The degree of absence is defined as the proportional gap between the current cumulative occurrence count of the target learning vocabulary and its preset minimum occurrence count, denoted as: in This indicates the number of times the target vocabulary appears in the generated text. This indicates the preset minimum number of occurrences. The boost factor is set as follows: in This is a control coefficient used to regulate the intensity of the boost. The final adjusted generation probability is: For example, if the target vocabulary word "absorb" appears only once in the text, while the preset target is three times, then the degree of absence... If set ,but In the current decoding step, its probability is changed from Upgraded to This significantly increases the likelihood of the target vocabulary being selected for generation. This mechanism ensures that target vocabulary is distributed and introduced as needed during language generation, thereby meeting the requirements of the teaching pace and improving the efficiency of language internalization and the diversity of contextual presentation.

[0067] In one possible implementation, after using a pragmatic repetition distribution constraint mechanism to control the frequency and positional distribution of the grammatical structure of the target learned vocabulary in the generated content, the method further includes: based on a context dependency control mechanism, determining whether the target learned vocabulary has already completed semantic expression in a previous segment; when the semantic expression has been closed, suppressing the continuous generation of the target learned vocabulary through a probability cooling strategy.

[0068] Specifically, during context tracking, the occurrence position, dependency syntax structure, and semantic function of the target vocabulary in each generated segment are recorded. Dependency syntax analysis identifies the position of the target vocabulary within a sentence, such as subject, predicate, object, or modifier, and analyzes its semantic combination relationship with contextual vocabulary, such as agent, state receiver, and condition constraint. Furthermore, the logical connection between the sentence containing the target vocabulary and the preceding and following segments is analyzed, such as whether it leads to the result of an event or forms a complete plot progression unit.

[0069] The criteria for determining semantic closure include the following categories: First, the target vocabulary has become the main predicate, constituting the core action of the sentence and completing the event description; second, it forms a stable collocation structure or phrase with other linguistic components, such as a fixed phrase or functional block; third, the context in the current segment has developed a complete information expression chain around the target vocabulary, for example, by semantically defining it through multiple modifiers; fourth, the segment has reached a logical conclusion in terms of plot, such as forming a causal chain, describing a final state, or an emotional conclusion. When any one of these conditions is met, the semantic expression of the target vocabulary is considered closed.

[0070] When the semantic expression is closed, a probability cooling strategy is activated to suppress the continuous generation of the target learned vocabulary. This strategy introduces a control factor during the decoding phase of the language generation model to temporarily lower the generation probability of the target learned vocabulary, preventing its recurrence within a short period. Specifically, after the semantic expression of the target learned vocabulary is closed... In each generation step, a cooling coefficient is applied to the probability value of the corresponding lexical unit in the vocabulary candidate list. Let its new generation probability be: in Represents the original probability. The value is typically between 0.2 and 0.5. Cooling window length The settings can be dynamically adjusted based on the current text complexity, adaptive context density, and pacing requirements. For example, in describing an experimental scenario, if "absorb" has already appeared as the predicate of the main clause and completed a description of an energy absorption behavior, the model will suppress its regeneration in the following 2 to 3 sentence segments, prioritizing the advancement of other content.

[0071] The cooling strategy can be restarted in conjunction with the contextual tension adjustment mechanism. When generated content enters a new semantic unit or a new scene transition node, the generation priority of the target vocabulary is reactivated to support its re-embedding in the new context. Through the joint implementation of the context dependency control mechanism and the probabilistic cooling strategy, it ensures that the target vocabulary is not overused and repetitive, while maintaining its periodic semantic recurrence rhythm in the text, thereby enhancing the pedagogical guidance, content diversity, and contextual integration effect of language generation.

[0072] S160 takes the semantic intent vector and retrieval results as input and combines them with a pragmatic repetition distribution constraint mechanism to generate target English story content.

[0073] In one possible implementation, the semantic intent vector and retrieval results are used as input, and a pragmatic repetition distribution constraint mechanism is combined to generate target English story content. Specifically, this includes: weighting text paragraphs in the retrieval results based on an attention mechanism, and selecting text paragraphs whose relevance to the current generation state meets preset requirements for language modeling; guiding the generation direction of the generated content based on the semantic intent vector; embedding a pragmatic repetition distribution constraint mechanism during the generation process to control the frequency and positional distribution of the grammatical structure of the target learning vocabulary in the generated content; and when it is found that the cumulative occurrence count or current positional distribution of the target learning vocabulary does not meet the pragmatic repetition distribution constraint parameters, adjusting the generation probability distribution and guiding the target learning vocabulary to be embedded in the target English story content in a grammatical structure that has not yet appeared.

[0074] Specifically, weighting text segments in the retrieval results based on an attention mechanism, and selecting text segments whose relevance to the current generated state meets preset requirements for language modeling, refers to dynamically evaluating the semantic relevance between each retrieved text segment and the current decoding state when constructing the context representation of the language generation model, and assigning different attention weights to each text segment based on the relevance. The attention mechanism quantifies the degree of language contribution to the currently generated content by calculating the attention score between the query vector (i.e., the representation of the current generated state) and the vector of each segment in the retrieval results. The specific calculation method usually adopts dot product attention, i.e.: in Generate a state vector for the current state. For the first Vector representation of each paragraph The vector dimension is used. All scores are normalized into attention weights using the Softmax function, and the paragraphs are then weighted and fused to form a context vector for current word generation. For example, if the user intends to generate story content related to scientific experiments, and the target vocabulary word is "absorb," the paragraph with higher weight might be "the material absorbs radiation under laboratory conditions" rather than other irrelevant paragraphs.

[0075] Guiding the generation direction of content based on semantic intent vectors refers to embedding semantic intent vectors as generation conditions into the control structure of a language generation model. This adjusts the output tendency of each decoding round, ensuring that the generated content maintains semantic consistency with the user's target task. Semantic intent vectors integrate semantic information from multiple dimensions, including target vocabulary, story style, emotional tone, and thematic background. In the language model, they participate as conditional vectors in the context modeling of each layer of the generative network. Common implementation methods include concatenating the semantic intent vector into the input embedding of each decoding step at the input of the Transformer decoder, or dynamically adjusting the influence weight between the semantic intent vector and the current positional state in the attention module through gating mechanisms. For example, if the semantic intent vector points to "laboratory, energy transfer, teaching verb usage," the model's generated content will favor the use of language structures describing experimental activities in its syntactic structure and event design, reflecting a teaching orientation.

[0076] Embedding a pragmatic repetition distribution constraint mechanism during the generation process controls the frequency and positional distribution of the grammatical structures of target vocabulary in the generated content. This means that during language generation, preset frequency requirements for target vocabulary, expected contextual positions, and grammatical structure diversity indicators are used as constraint variables in real time to participate in the decoding decision. In each round of outputting lexical units, the language generation model needs to determine the cumulative occurrence of the current target vocabulary and its contextual grammatical state, comparing this state with the pragmatic repetition distribution constraint parameters. If the frequency is too low or the grammatical structure is repetitive, the generation probability is increased through a lexical unit probability adjustment mechanism, and it is forced to be output in a new context or structure, thereby achieving repetition control at the linguistic level. For example, if "absorb" appears in the initial paragraph as an active sentence, it will be forced to be embedded again in different paragraphs in the passive voice or the nominalized structure "absorption" to achieve semantic transfer and structural diversification.

[0077] When the model detects that the cumulative occurrence count or current position distribution of the target vocabulary does not meet the pragmatic repetition distribution constraint parameters, it adjusts the generation probability distribution and guides the target vocabulary to be embedded in the target English story content in a grammatical structure that has not yet appeared. This means that once the model determines that the current usage of the target vocabulary does not cover the preset teaching requirements, it initiates a generation probability reconstruction mechanism to increase the probability of the target vocabulary in the current lexical candidate space and guides it to be adopted in a novel grammatical form through structural instructions. For example, if "absorb" has appeared as a verb in the preceding text but has never appeared as a noun, the model will increase the probability of the lexical "absorption" in subsequent sentences from its original value. Upgraded to Simultaneously, it constrains the selection of generated sentence structures to be relative clauses or prepositional phrases, so that "absorption" can be reasonably embedded in the context with a new grammatical structure, such as "the process of absorption is influenced by temperature". This mechanism ensures that the target vocabulary is not only repeatedly used, but also deeply embedded in the generated content in diverse grammatical forms, achieving the pragmatic transfer and language internalization required by the teaching objectives.

[0078] This embodiment also discloses an intelligent English teaching story generation device, referring to... Figure 2 The device includes an acquisition module 201, a processing module 202, and an output module 203. It is used to execute any of the above-described intelligent English teaching story generation methods, wherein: The acquisition module 201 is used to acquire user input.

[0079] The processing module 202 is used to perform semantic parsing on the input content and identify the user's semantic intent vector and target learning vocabulary.

[0080] The processing module 202 is used to perform a query retrieval in a preset corpus based on semantic intent vectors and using retrieval enhancement generation technology to extract text paragraphs.

[0081] Processing module 202 is used to construct retrieval results containing the target learning vocabulary based on text paragraphs.

[0082] Processing module 202 is used to apply a pragmatic repetition distribution constraint mechanism, wherein the pragmatic repetition distribution constraint mechanism is used to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content.

[0083] Output module 203 is used to take the semantic intent vector and retrieval results as input and combine them with the pragmatic repetition distribution constraint mechanism to generate target English story content.

[0084] In one possible implementation, the processing module 202 is used to map the semantic intent vector to a target dimension vector space through vectorization technology to form a query vector that can be matched with text data in a preset corpus, the corpus containing structured text data.

[0085] The processing module 202 is used to calculate the similarity between the query vector and each text data in the preset corpus using a matching mechanism based on vector similarity calculation.

[0086] The processing module 202 is used to select candidate text data from multiple text data that meet a preset threshold for similarity with the query vector.

[0087] Processing module 202 is used to filter candidate text data through semantic correlation to obtain text paragraphs.

[0088] In one possible implementation, the processing module 202 is used to set pragmatic repetition distribution constraint parameters for the target learning vocabulary.

[0089] The processing module 202 is used to dynamically track the cumulative occurrence and current position distribution of the target learning vocabulary in the generated content.

[0090] The processing module 202 is used to compare the cumulative occurrence count and current position distribution with the pragmatic repetition distribution constraint parameters in real time. When the cumulative occurrence count or current position distribution does not meet the pragmatic repetition distribution constraint parameters, the generation probability distribution is adjusted to improve the priority of the target learning vocabulary in subsequent texts.

[0091] In one possible implementation, the processing module 202 is used to determine, based on the context dependency control mechanism, whether the target learning vocabulary has already completed semantic expression in the previous segment, and when the semantic expression has been closed, to suppress the continuous generation of the target learning vocabulary through a probability cooling strategy.

[0092] In one possible implementation, the processing module 202 is used to reweight the word probability distribution of the current decoding step when the cumulative number of occurrences does not reach a preset occurrence frequency threshold, or when the current position distribution does not conform to the context distribution pattern.

[0093] The processing module 202 is used to apply an adaptive boosting factor to the lexical items corresponding to the target learning words in the current word candidate list during the reweighting operation. The boosting factor is calculated based on the degree of usage absence of the target learning words, and the degree of absence is directly proportional to the boosting magnitude of the generation probability.

[0094] In one possible implementation, the processing module 202 is used to assign weights to text segments in the retrieval results based on an attention mechanism, and select text segments whose relevance to the current generated state meets preset requirements for language modeling.

[0095] Processing module 202 is used to guide the generation direction of generated content based on semantic intent vector.

[0096] The processing module 202 is used to embed a pragmatic repetition distribution constraint mechanism during the generation process to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content.

[0097] The output module 203 is used to adjust the generation probability distribution and guide the target learning vocabulary to be embedded in the target English story content in a grammatical structure form that has not appeared when it is found that the cumulative occurrence frequency or current position distribution of the target learning vocabulary does not meet the pragmatic repetition distribution constraint parameters.

[0098] In one possible implementation, the output module 203 is used to convert the input content into a feature vector.

[0099] The output module 203 is used to obtain semantic information of user needs by performing semantic analysis on the input content.

[0100] The output module 203 is used to identify and extract target learning words based on the semantic information of user needs, according to the frequency, contextual association and importance of different words in the input content and in the user's learning path.

[0101] Output module 203 is used to convert the target learning vocabulary into vocabulary intent vectors.

[0102] Output module 203 is used to construct semantic intent vectors based on feature vectors and lexical intent vectors.

[0103] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0104] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.

[0105] The communication bus 302 is used to enable communication between these components.

[0106] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0107] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0108] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications. The GPU is responsible for rendering and drawing the content required for display. The modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0109] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc. The data storage area may store data involved in the various method embodiments described above. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program for an intelligent generation method of English teaching stories.

[0110] exist Figure 3 In the illustrated electronic device, the user interface 303 is primarily used to provide an input interface for the user and to acquire user input data. The processor 301 can be used to call an application program stored in the memory 305 that represents an intelligent method for generating English teaching stories. When executed by one or more processors 301, the electronic device performs one or more methods as described in the above embodiments.

[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0112] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0113] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0114] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0115] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0116] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0117] This application also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 301, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0118] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for intelligently generating English teaching stories, characterized in that, The method includes: Get the user's input; The input content is semantically parsed to identify the user's semantic intent vector and target vocabulary. Based on the semantic intent vector, a search enhancement generation technique is used to perform a query search in a preset corpus and extract text paragraphs from it. Based on the text paragraph, construct search results containing the target vocabulary for learning; A pragmatic repetition distribution constraint mechanism is applied, wherein the pragmatic repetition distribution constraint mechanism is used to control the frequency and positional distribution of the grammatical structure of the target learning vocabulary in the generated content; The semantic intent vector and the retrieval results are used as input, and the target English story content is generated by combining the pragmatic repetition distribution constraint mechanism.

2. The method for intelligently generating English teaching stories according to claim 1, characterized in that, The step of extracting text paragraphs based on the semantic intent vector using retrieval enhancement generation technology specifically includes: The semantic intent vector is mapped to the target dimension vector space using vectorization technology to form a query vector that can match the text data in the preset corpus, wherein the corpus contains the structured text data; A matching mechanism based on vector similarity calculation is used to calculate the similarity between the query vector and each text data in the preset corpus; Select candidate text data from multiple text data that have a similarity to the query vector that meets a preset threshold; The candidate text data is filtered by semantic relevance to obtain the text paragraphs.

3. The method for intelligently generating English teaching stories according to claim 1, characterized in that, The applied pragmatic repetition distribution constraint mechanism is used to control the frequency and positional distribution of the grammatical structures of the target learning vocabulary in the generated content, specifically including: Set pragmatic repetition distribution constraint parameters for the target vocabulary; Dynamically track the cumulative occurrence count and current position distribution of the target vocabulary in the generated content; The cumulative occurrence count and the current position distribution are compared with the pragmatic repetition distribution constraint parameters in real time. When the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameters, the generation probability distribution is adjusted to improve the usage priority of the target learned vocabulary in subsequent texts.

4. The method for intelligently generating English teaching stories according to claim 1, characterized in that, Following the application of the repeated distribution constraint mechanism, the method further includes: Based on the context dependency control mechanism, it is determined whether the target vocabulary has already been semantically expressed in the previous segment. When the semantic expression has been closed, the continuous generation of the target vocabulary is suppressed through a probability cooling strategy.

5. The method for intelligently generating English teaching stories according to claim 3, characterized in that, When the cumulative occurrence count or the current position distribution does not meet the pragmatic repetition distribution constraint parameter, adjusting the generation probability distribution specifically includes: When the cumulative number of occurrences does not reach the preset occurrence frequency threshold, or when the current position distribution does not conform to the context distribution pattern, the word probability distribution of the current decoding step is reweighted. In the reweighting operation, an adaptive boosting factor is applied to the lexical items corresponding to the target learning vocabulary in the current vocabulary candidate list. The boosting factor is calculated based on the degree of usage absence of the target learning vocabulary, and the degree of absence is directly proportional to the increase in the generation probability.

6. The method for intelligently generating English teaching stories according to claim 1, characterized in that, The step of generating target English story content by taking the semantic intent vector and the retrieval results as input and combining them with the pragmatic repetition distribution constraint mechanism specifically includes: The text segments in the retrieval results are weighted based on the attention mechanism, and the text segments whose relevance to the current generated state meets the preset requirements are selected for language modeling. The semantic intent vector guides the direction of content generation; The pragmatic repetition distribution constraint mechanism is embedded in the generation process to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; When it is found that the cumulative occurrence count or current position distribution of the target learning vocabulary does not meet the pragmatic repetition distribution constraint parameter, the generation probability distribution is adjusted, and the target learning vocabulary is guided to be embedded in the target English story content in a grammatical structure that has not appeared before.

7. The method for intelligently generating English teaching stories according to claim 1, characterized in that, The step of performing semantic parsing on the input content to identify the user's semantic intent vector and target learning vocabulary specifically includes: Convert the input content into a feature vector; By performing semantic analysis on the input content, the semantic information of user needs can be obtained; Based on the semantic information of user needs, the target learning vocabulary is identified and extracted according to the frequency, contextual association and importance of different words in the input content and in the user's learning path. Convert the target vocabulary into vocabulary intent vectors; The semantic intent vector is constructed based on the feature vector and the lexical intent vector.

8. An intelligent story generation device for English teaching, characterized in that, The device is used to execute an intelligent English teaching story generation method as described in any one of claims 1-7, the device comprising an acquisition module (201), a processing module (202), and an output module (203), wherein: The acquisition module (201) is used to acquire the user's input content; The processing module (202) is used to perform semantic parsing on the input content and identify the user's semantic intent vector and target learning vocabulary; The processing module (202) is used to perform a query retrieval in a preset corpus based on the semantic intent vector and using retrieval enhancement generation technology to extract text paragraphs from it. The processing module (202) is used to construct retrieval results containing the target learning vocabulary based on the text paragraph; The processing module (202) is used to apply a pragmatic repetition distribution constraint mechanism, wherein the pragmatic repetition distribution constraint mechanism is used to control the frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; The output module (203) is used to take the semantic intent vector and the retrieval result as input, and combine them with the pragmatic repetition distribution constraint mechanism to generate target English story content.

9. An electronic device, characterized in that, The device includes a processor (301), a communication bus (302), a user interface (303), a network interface (304), and a memory (305). The memory (305) is used to store instructions. The user interface (303) and the network interface (304) are both used to communicate with other devices. The communication bus (302) is used to realize the connection and communication between the components within the electronic device. The processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device performs the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.

Citation Information

Cited By

  • Teaching resource recommendation method and system based on natural language processing

    CN121724811A