Intelligent oral English training system based on Prompt engineering

Through the Prompt project's intelligent spoken English training system, multi-role dialogue scenarios are dynamically generated and off-topic issues are detected in real time, solving the problems of rigid scenario generation and weak target control in existing systems, and achieving a personalized and flexible English spoken training experience.

CN120690182APending Publication Date: 2025-09-23CHUXIONG NORMAL UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510866326.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing English oral training systems are unable to instantly generate diverse topics and multi-role interaction scenarios based on users' simple natural language descriptions, and lack real-time detection and guidance of users' deviation from training goals, resulting in high barriers to personalized customization, insufficient flexibility, and poor training results.

Method used

An intelligent training system based on the Prompt project is adopted. Through the user input parsing module, Prompt generation module, large language model response module, ASR recognition module, off-topic detection and topic guidance module, multi-round dialogue context management module and student performance evaluation module, it can dynamically generate multi-role dialogue scenarios and detect and guide users to return to the topic in real time.

Benefits of technology

It enables users to generate multi-topic and multi-role dialogue scenarios with just one description, improves the personalization and flexibility of training, ensures the targetedness and effectiveness of training, and provides personalized feedback and natural interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690182A_ABST
    Figure CN120690182A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent oral English training system and method based on Prompt engineering, and relates to the technical field of oral English training. According to the method, a user is allowed to freely describe training requirements by using a natural language, the system automatically constructs a training Prompt, and personalized customization is realized; content deviating from a training target is accurately recognized through a deviation question detection module, a user is guided to return to a theme through real-time voice or text, and effectiveness and pertinence of training are guaranteed; a clear template definition ensures systematicness and reproducibility of a training task, and clear control and evaluation of a teaching target are facilitated; in combination with voice recognition, voice synthesis and text generation, natural interaction in a real context is realized, and the training experience is improved; on the basis of multi-dimensional automatic evaluation of a keyword hit rate, pronunciation quality and the like, personalized learning suggestions are supported; teachers are supported to generate training tasks in batches, students are supported to customize scenes, and different education environment requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of spoken English training, specifically an intelligent spoken English training system that dynamically generates multi-topic, multi-role dialogue scenarios based on natural language instructions. The core of this system is to utilize a large-scale language model to understand simple natural language descriptions from users and intelligently construct a training environment with diverse topics and interactive characters. Background Art

[0002] With the rapid development of artificial intelligence technology, language learning assistance tools based on large language models are increasingly used in the field of education. However, existing English speaking training systems have significant limitations:

[0003] (1) Rigid scenario generation: Mainstream systems typically rely on pre-set, fixed template dialogues or require complex parameter configuration to generate scenarios. This approach cannot meet the user's need to instantly generate diverse themes and multi-role interaction scenarios with a simple natural language description (for example, "Simulate an airport check-in conversation" or "Conduct a performance interview involving a manager and an employee"). This leads to a high threshold for personalized customization and a serious lack of flexibility.

[0004] (2) Weak target control: During the training process, existing systems generally lack effective real-time detection and guidance mechanisms for users’ deviation from the preset training targets, which affects the relevance and effectiveness of training.

[0005] (3) Insufficient dynamism and intelligence: Although Prompt Engineering is a key technology for guiding large-scale language models to generate high-quality output, its potential has not yet been fully tapped in existing spoken language training systems to dynamically build highly customized and structured training scenarios based on user intentions and to achieve intelligent guidance and control of the training process. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems mentioned in the background technology, such as the lack of personalized customization, insufficient scene flexibility, and difficulty in dynamically constructing training scenes in traditional systems.

[0007] The core of the invention provides an intelligent spoken English training system that can dynamically generate diverse theme scenes and multi-role dialogues based on the user's single natural language description.

[0008] Through innovative Prompt engineering technology, the system eliminates the tedious process of manually writing complex dialogue templates, greatly improving the flexibility and personalization of oral training.

[0009] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0010] An intelligent spoken English training system based on the Prompt project, comprising:

[0011] The user input parsing module is used to receive training requirements input by users or teachers in the form of natural language descriptions, and perform semantic analysis on the input to extract core information such as training scenarios, teaching objectives, and keywords;

[0012] A prompt generation module is used to construct a structured prompt template based on the output of the user input parsing module. The template includes scene background, user role, AI role, teaching objectives, keywords, and sensitive word fields;

[0013] A large language model response module is used to generate multi-round interactive dialogue content that conforms to the context based on the prompt template and role settings;

[0014] ASR recognition module, used to convert user voice input into text for system analysis;

[0015] The off-topic detection and topic guidance module is used to check the topic consistency of user input content. If it is detected that it deviates from the teaching objectives or keyword range set by the prompt, it will generate voice or text guidance content to prompt the user to return to the conversation topic;

[0016] The TTS synthesis module is used to convert AI-generated text into speech to achieve natural voice interaction;

[0017] The multi-round dialogue context management and memory retention module is used to record and manage historical dialogue information and key semantic content, ensuring information continuity between rounds. It also has a consistency check and context summary mechanism for generated content.

[0018] Dynamic role generation and assignment module, used to automatically generate multiple AI dialogue roles for users to choose from based on training topics, support user-defined roles, and dynamically adjust prompt templates and dialogue content based on user selections;

[0019] The student performance evaluation module is used to analyze the user's performance during the training process and generate feedback and scores based on keyword matching, voice fluency, and pronunciation accuracy.

[0020] Furthermore, the user input parsing module supports users to describe training objectives in just one sentence of natural language. The system automatically constructs the structured elements required for the corresponding training scenario through intent recognition, keyword extraction and teaching objective reasoning.

[0021] Furthermore, the Prompt generation module can dynamically generate a Prompt template based on the parsed teaching intention, and support adaptive adjustment of template fields, including but not limited to training objectives, dialogue keywords, banned words, and suggested expressions.

[0022] Furthermore, the workflow of the off-topic detection and topic guidance module includes the following steps:

[0023] Step B1: Speech transcription: The student's speech input is first transcribed into processable text content through the automatic speech recognition (ASR) module, which serves as the basis for subsequent semantic analysis;

[0024] Step B2: Keyword extraction and target topic modeling. The system extracts core keywords, phrases, and semantic tags from the current prompt or teaching objective, and constructs a semantic vector representation of the target topic based on the BERT and SimCSE models.

[0025] Step B3: Semantic analysis of user input: the student's transcribed text input is encoded using the same semantic vector model to obtain a semantic vector representation;

[0026] Step B4: Similarity calculation and topic consistency determination: Calculate the semantic similarity between the user input vector and the target topic vector. If the similarity is lower than the set threshold, it is determined to be off-topic input;

[0027] Step B5: Identify and categorize off-topic content. Categorize off-topic content for subsequent refined guidance strategies.

[0028] Step B6: generating guidance content, combining the current teaching stage and the type of the off-topic question to generate guidance prompt content or synthesized voice playback;

[0029] Step B7: User feedback and guidance loop, presenting guidance prompts to students to guide them back to the topic, and recording the frequency of off-topic deviations for subsequent teaching feedback and strategy optimization.

[0030] Furthermore, the student evaluation module provides students with personalized training feedback by integrating multiple AI models and scoring algorithms, including the following steps:

[0031] Step C1: Data collection during training. The system collects multimodal performance data of users during training in real time, including voice audio, transcribed text, interaction duration, and response delay information.

[0032] Step C2: Keyword matching analysis: perform keyword recognition and comparison on the user input text, and count the number and weight of the teaching keywords set by the prompt to evaluate the relevance of the user content;

[0033] Step C3: Speech fluency analysis: analyzing the user's speech data for speech rate, pauses, repetitions, and correction features, and using a fluency scoring algorithm to assess the naturalness and coherence of the expression;

[0034] Step C4: Pronunciation accuracy assessment: Use speech recognition and phoneme alignment technology to assess the accuracy of the user's pronunciation, compare it with the standard pronunciation model, and mark pronunciation problem words;

[0035] Step C5: Generate a multi-dimensional score, calculate a comprehensive score based on the above multiple dimensions, and provide feedback on the user's oral performance in the form of charts or text;

[0036] Step C6: Personalized feedback generation: the system combines the scoring results with historical performance to generate personalized learning suggestions.

[0037] Furthermore, the multi-round dialogue context management and memory retention module records the historical interaction content between the user and the AI ​​by constructing an information state map, including user intent, scenario settings, and keyword usage status, and uses this context information in subsequent responses to generate more consistent system responses, including the following steps:

[0038] Step D1: Conversation information collection and annotation. In each round of conversation, the system extracts and annotates the core information from user input and AI output in real time, including user intent, keywords, scenario settings, and emotional tendencies.

[0039] Step D2: Information state graph construction: The extracted information is used to construct an information state graph in a structured manner to represent the dynamic semantic state of the user-AI interaction process;

[0040] Step D3: Context information storage and update: The information state graph is used as the context storage structure and dynamically updated after each round of dialogue to ensure consistency and continuity between historical semantics and the latest user input.

[0041] Step D4: Contextual information retrieval and integration: When generating AI responses, key content from the current information state graph is retrieved to guide the language model to generate responses with contextual coherence.

[0042] Step D5: Consistency check and context summary: Consistency check is performed on the content generated by multiple rounds of dialogue to avoid information contradictions. At the same time, the long dialogue history is compressed through the context summary mechanism to improve the system response efficiency.

[0043] Furthermore, the dynamic role generation and assignment module supports multi-user collaborative training. Users can choose to participate in multiple rounds of role-playing dialogues among system-generated or customized AI roles. The system dynamically maintains the identity and responsibilities of each role and updates the prompt content to adapt to the current scenario, including the following steps:

[0044] Step E1: Multiple users enter the training scene. The system supports multiple users or users and multiple AI characters entering the same training task at the same time, starting the collaborative training mode.

[0045] Step E2: Role list generation and display. The system automatically generates multiple AI dialogue roles based on the current teaching scenario set by Prompt, and displays the role list and role responsibilities to the user.

[0046] Step E3: Role selection and confirmation. The user can select a role from the system recommended roles to join the training, or customize the role settings. After confirmation, enter the role-playing session;

[0047] Step E4: Dynamically update the prompt content. The system updates the prompt template in real time based on the user's selected role, including updating the participant role fields, conversation context, and assigned responsibilities to ensure that the generated multi-round conversations conform to the current division of labor settings.

[0048] Step E5: Collaborative conversation control and identity maintenance. During multiple rounds of conversation, the system dynamically maintains the identity tag and semantic memory of each character to ensure that the AI ​​or multiple users' responses are consistent with their respective character settings to avoid confusion.

[0049] Step E6: Collaborative training task guidance and feedback. The system generates feedback and prompts based on each role's task objectives and speech performance, guiding multiple parties to collaboratively complete training tasks, such as information exchange, problem solving, or process simulation.

[0050] An intelligent spoken English training method based on any of the above systems comprises the following steps:

[0051] Step F1: Training intention input: the user inputs the current training intention through natural language, such as "I want to practice English dialogue in airport scenes";

[0052] Step F2: Intent parsing and prompt construction: The user input parsing module identifies the scene, role, and task elements in the sentence, and sends them to the prompt generation module to build a structured prompt template;

[0053] Step F3: Dynamic generation and allocation of AI roles. The system calls the role generation and allocation module to generate or allocate corresponding AI dialogue roles based on the training scenario and update the role information in Prompt.

[0054] Step F4: Large language model content generation: The large language model response module generates multiple rounds of dialogue content based on prompts and role settings, guiding users to conduct scenario simulation training;

[0055] Step F5: Voice input and transcription: the user participates in oral dialogue training through voice, and the system transcribes the user's voice into text input through the automatic speech recognition module;

[0056] Step F6: Off-topic detection and topic guidance: The off-topic detection and topic guidance module analyzes the user input text in real time to determine whether it deviates from the training topic and provides text or voice guidance to return to the main line of the conversation when necessary;

[0057] Step F7: Conversation context and memory retention: The multi-round conversation context management and memory retention module continuously maintains the current conversation state and historical information to ensure the consistency and personalization of the AI ​​character's response;

[0058] Step F8: System output and speech synthesis: the system outputs the response content and synthesizes it into speech form through the text-to-speech module to achieve natural and smooth human-computer interaction;

[0059] Step F9: Student performance evaluation and feedback generation. The student performance evaluation module comprehensively analyzes the user's performance in this round of training and generates visual feedback.

[0060] The beneficial effects of the present invention are as follows:

[0061] 1. Single-sentence drive for multi-topic, multi-role dialogue scenario generation: Users only need to describe their training needs in natural language, such as "simulate a customer complaint handling" or "conduct a preoperative communication between a doctor and a patient," and the system will intelligently generate a complete dialogue scenario prompt containing diverse topics and multi-role interactions. This significantly lowers the threshold for customized training scenarios and enables unprecedented personalization and flexibility.

[0062] 2. Goal-oriented intelligent guidance: The off-topic detection module accurately identifies content that deviates from the training objectives, and guides users back to the topic through real-time voice or text, ensuring the effectiveness and pertinence of training.

[0063] 3. Standardized prompt structure: Clear template definition ensures the systematicness and reproducibility of training tasks, which is conducive to the clear control and evaluation of teaching objectives.

[0064] 4. Multimodal interaction: Combining speech recognition, speech synthesis, and text generation to achieve natural interaction in real contexts and enhance the training experience.

[0065] 5. Efficient feedback mechanism: Automatic evaluation based on multiple dimensions such as keyword hit rate and pronunciation quality, supporting personalized learning suggestions.

[0066] 6. Wide teaching applicability: Supports teachers to batch generate training tasks and students to independently customize scenarios to meet the needs of different educational environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a system structure diagram of the present invention;

[0068] Figure 2 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0069] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0070] See also Figure 1 - Figure 2 The present invention provides an intelligent spoken English training system based on the Prompt project, comprising:

[0071] The user input parsing module is used to receive training requirements input by users or teachers in the form of natural language descriptions, and perform semantic analysis on the input to extract core information such as training scenarios, teaching objectives, and keywords;

[0072] A prompt generation module is used to construct a structured prompt template based on the output of the user input parsing module. The template includes scene background, user role, AI role, teaching objectives, keywords, and sensitive word fields;

[0073] The prompt template uses structured JSON format and contains the following fields:

[0074] {"scene_name":"Scene name",

[0075] "background": "scene background description, providing conversation context",

[0076] "user_roles":["User Role 1","User Role 2 (optional)"],"ai_roles":["AI Role 1","AI Role 2 (optional)"],

[0077] "teaching_goals": ["The teaching goals of this training scenario"]

[0078] "topic_scope":["The scope of conversation topics allowed"],

[0079] "goal_keywords":["Keyword list for hit detection and evaluation"],

[0080] "per_turn_goals":[{"turn":1,"expected_keywords":["Keyword 1","Keyword 2"]} / / Can be expanded to multiple round goals],

[0081] "sensitive_words":["Sensitive topics, AI should avoid involving"],

[0082] "default_response": "Guidance words when students do not respond",

[0083] "post_response": "Summary or encouragement at the end of the conversation." The system parses this structured prompt to guide the large language model to generate conversation content that matches the scenario and teaching objectives. It also uses information such as keywords to match user input to the target and detect off-topic content, automatically guiding the user back to the topic.

[0084] / / User information file

[0085] "user_profile":{"name":"","gender":"","age":0,"profession":"","location":"","interests":[]},

[0086] / / Conversation context information

[0087] "dialogue_context":{"topic":"","history":[]},

[0088] / / Intent map

[0089] "intention_graph":{"nodes":[],"edges":[]},

[0090] / / Information status diagram

[0091] "info_state_graph":{"topic":"","known_slots":{"Information 1":null,"Information 2":null,"Information 3":null},"status":""},

[0092] / / Consistency detector

[0093] "consistency_checker":{"enabled":true,"last_checked":""},

[0094] / / Topic tracker

[0095] "topic_tracker":{"current_topic":"","topic_switch_log":[]}}

[0096] Prompt template structure example

[0097] {"scene_name":"Order Dialogue",

[0098] "background":"User simulates the ordering process, including cuisine selection, number of people, time and budget, etc."

[0099] "user_roles":["Customer A"],

[0100] "ai_roles":["waiter"],

[0101] "active_roles":{"user":"Customer A","ai":"Waiter"},

[0102] "user_defined_roles":{"Customer A":"A young woman who is health-conscious and on a budget"},

[0103] "teaching_goals": ["Master common expressions for ordering food", "Express personal needs fluently", "Have polite multi-round conversations"],

[0104] "topic_scope":["Restaurant Type","Food Preference","Time Arrangement","Budget Discussion"],

[0105] "topic_embedding":"[Topic vector of the ordering scene]",

[0106] "goal_keywords":["Order","Tonight","Number of People","Cuisine","Budget"],

[0107] "per_turn_goals":[{"turn":1,"expected_keywords":["Want to order food","tonight"]},

[0108] {"turn":2,"expected_keywords":["number of people","time"]},

[0109] {"turn":3,"expected_keywords":["like to eat","Sichuan cuisine"]},

[0110] {"turn":4,"expected_keywords":["budget","how much"]}],

[0111] "info_state_graph":{"topic":"Ordering Meals",

[0112] "known_slots":{"cuisine":null,"number of people":null,"budget":null,"time":null},

[0113] "status":{"Waiting for the user to add all order information"},

[0114] "consistency_checker":{"reminder_template":"Current task: Order food.\nKnown information:\n-Cuisine: {cuisine}\n-Number of people: {number of people}\n-Time: {time}\n-Budget: {budget}\nPlease continue to ask questions or complete the missing items. Do not complete the task early.",

[0115] "require_all_slots_before_closing":true},

[0116] "topic_tracker":{"topic_vector":"[Order Topic Vector]",

[0117] "threshold":{0.72,

[0118] "on_topic_fallback":"Let's get back to the topic of ordering food. What kind of dishes would you like to order?",

[0119] "off_topic_response":"We seem to have strayed from the topic of ordering. Please return to the topic of dishes or time."},

[0120] "sensitive_words":["politics","religion","violence"],

[0121] "default_response":["Please tell me your order requirements, such as time, number of people, etc."],

[0122] "post_response":"[Thanks for your practice! You've done a great job in fully expressing your ordering needs!"].

[0123] Dynamic generation process (pseudo code example):

[0124] The system first uses natural language understanding (NLU) technology to semantically analyze the training requirements entered by the user, extracting the core scenario names, teaching objectives, and keywords to form preliminary structured data. It then refines each field in the prompt by combining predefined rules and the knowledge base.

[0125] The pseudo code example is as follows:

[0126] defgenerate_prompt_from_description(description:str,user_custom_roles:dict=None)->dict:

[0127] scene_name=extract_scene_name(description),

[0128] background=generate_background(scene_name),

[0129] user_roles,ai_roles=assign_roles(scene_name),

[0130] ifuser_custom_roles:

[0131] user_roles=list(user_custom_roles.keys),

[0132] teaching_goals=extract_goals(description),

[0133] topic_scope=determine_topic_scope(scene_name),

[0134] goal_keywords=extract_goal_keywords(teaching_goals),

[0135] sensitive_words=load_sensitive_words,

[0136] return{"scene_name":scene_name,

[0137] "background":background,

[0138] "user_roles":user_roles,

[0139] "ai_roles":ai_roles,

[0140] "active_roles":{"user":user_roles[0],"ai":ai_roles[0]},

[0141] "user_defined_roles":{user_custom_rolesor},

[0142] "teaching_goals":teaching_goals,

[0143] "topic_scope":topic_scope,

[0144] "goal_keywords":goal_keywords,

[0145] "per_turn_goals":generate_turn_goals(goal_keywords),

[0146] "sensitive_words":sensitive_words,

[0147] "default_response":"Please try to express your thoughts.",

[0148] "post_response":"Training completed, good performance."}

[0149] The Prompt generation module workflow includes the following steps:

[0150] Step S1: Initialize the dialog context according to the prompt content;

[0151] Step S2: Call the large language model response module to generate AI responses that match the context and guide students to interact;

[0152] Step S3: The student inputs the speech, and the ASR module transcribes it into text;

[0153] Step S4: The real-time off-topic detection module monitors the content of the student's speech;

[0154] Step S5: If deviation from the topic is detected, the user is guided back to the target topic through a preset “default response”;

[0155] Step S6: Repeat the above steps until the dialogue reaches the teaching goal or the end condition

[0156] A large language model response module is used to generate multi-round interactive dialogue content that conforms to the context based on the prompt template and role settings;

[0157] Dialogue generation and interaction process:

[0158] Based on the generated prompt, the system calls a pre-trained large language model (such as the GPT series) to generate multiple rounds of dialogue content. The dialogue process includes:

[0159] 1. The user selects or customizes a role;

[0160] 2. The system generates a complete prompt based on the role information;

[0161] 3. AI initializes the scene and generates the opening dialogue;

[0162] 4. The user inputs voice, and the ASR module converts it into text;

[0163] 5. The system determines whether the question is off-topic based on the current conversation context;

[0164] 6. Repeat the above steps until the conversation reaches the teaching objective or ends.

[0165] Interaction pseudocode example:

[0166] defdialogue_interaction(prompt:dict):context=init_context(prompt),

[0167] forturninrange(MAX_TURNS):ai_text=LLM.generate_response(context),

[0168] tts_play(ai_text),

[0169] user_speech=get_user_speech(),

[0170] user_text=asr_recognize(user_speech);

[0171] ifdetect_off_topic(user_text,prompt):feedback=prompt["default_response"],

[0172] tts_play(feedback),

[0173] else:feedback="Very good, please continue."

[0174] update_context(context,ai_text,user_text);

[0175] give_feedback(feedback).

[0176] ASR recognition module, used to convert user voice input into text for system analysis;

[0177] The off-topic detection and topic guidance module is used to check the topic consistency of user input content. If it is detected that it deviates from the teaching objectives or keyword range set by the prompt, it will generate voice or text guidance content to prompt the user to return to the conversation topic;

[0178] In order to ensure that the system always revolves around the preset teaching objectives and topics during multiple rounds of interaction and effectively prevent user input from deviating from the topic, the present invention proposes a deviation detection and guidance mechanism based on topic tracking.

[0179] This mechanism uses a combination of explicit topic modeling and semantic similarity calculation to achieve dynamic monitoring and feedback of the semantic direction of user input. Its key components and technical points are as follows:

[0180] Topic vector modeling and semantic matching: The system builds a semantic representation model based on a predefined topic scope (specified by the topic_scope field in Prompt), encodes each round of user input into a vector, and calculates the similarity value between it and the current topic vector.

[0181] Off-topic judgment logic: Set a semantic similarity threshold. If the similarity between the current input and the topic vector is lower than the threshold, it is automatically judged as off-topic behavior and the off-topic event is recorded.

[0182] Topic tracking and switching management: Introduce the topic tracker topic_tracker to maintain the current valid topic tags (such as "ordering food", "job interview", "travel", etc.) and the topic switching log topic_switch_log in real time, to achieve continuous monitoring of the status of conversation topics and behavior auditing.

[0183] Guidance strategy and automatic remediation mechanism: When the system detects that the user has deviated from the topic, it automatically calls the default_response guidance module in Prompt to guide the user back to the intended topic scope through natural language, ensuring that the interaction direction does not deviate from the teaching objectives.

[0184] Multi-round dynamic topic consistency check: The system continuously updates the topic field dialogue_context.topic in the conversation context and verifies its consistency in each round of interaction to ensure the coherence of the topic.

[0185] This module achieves high-precision off-topic identification and flexible topic guidance by constructing an explicit topic semantic space and off-topic judgment logic, providing semantic closed-loop control for the language training system, and significantly improving the focus of the dialogue content, the goal achievement rate, and the teaching effectiveness.

[0186] The TTS synthesis module is used to convert AI-generated text into speech to achieve natural voice interaction;

[0187] The multi-round dialogue context management and memory retention module is used to record and manage historical dialogue information and key semantic content, ensuring information continuity between rounds. It also has a consistency check and context summary mechanism for generated content.

[0188] To improve the coherence and contextual consistency of multi-round dialogue interactions, the present invention provides a multi-round dialogue context management and memory retention module, which is mainly used to solve the problems of context forgetting and topic deviation existing in existing large language models during long dialogues, and to ensure semantic coherence and achievement of teaching objectives in language training tasks.

[0189] This module combines the Information State Graph with a multi-level memory mechanism to achieve efficient storage and dynamic access of historical information. Specific technical solutions include the following:

[0190] Information state graph construction mechanism: Build and dynamically maintain the information state graph info_state_graph, which includes a set of key information slots known_slots and the current conversation state identifier status. This graph is used to identify the valid information provided by the user, thereby avoiding repeated questions from the system and improving question-answering efficiency and user experience.

[0191] Multi-level memory structure design: Introducing a hierarchical memory system, including the short-term context cache dialogue_context.history, basic user information user_profile, and long-term preference memory area, to achieve hierarchical storage and retrieval of information at different time scales, supporting deep personalized dialogue generation.

[0192] Dynamic context fusion algorithm: By fusing the current input content with the historical context state, it generates comprehensive context input for the large language model in real time, ensuring the complete transmission of key semantic information and effectively avoiding the "mid-round semantic forgetting" phenomenon.

[0193] Consistency verification mechanism: A consistency checker module is designed to compare the semantic and factual consistency of the model output. If information conflicts or logical contradictions are detected, the system automatically triggers a corrective response strategy to ensure the quality and stability of the dialogue.

[0194] Incremental memory and summary mechanism: An incremental update strategy is used to dynamically maintain the information state map and conversation history, and a semantic compression algorithm is used to regularly generate conversation summaries. This not only reduces the burden of model calls, but also effectively prevents information forgetting and redundant accumulation.

[0195] This module uses the combined strategy of "information structured modeling + multi-layer memory fusion + consistency control" to achieve accurate management and effective inheritance of historical information in multiple rounds of language dialogues, fundamentally improving the system's semantic understanding ability and teaching interaction quality.

[0196] Dynamic role generation and assignment module, used to automatically generate multiple AI dialogue roles for users to choose from based on training topics, support user-defined roles, and dynamically adjust prompt templates and dialogue content based on user selections;

[0197] The student performance evaluation module is used to analyze the user's performance during the training process and generate feedback and scores based on keyword matching, voice fluency, and pronunciation accuracy.

[0198] The student performance evaluation module also includes a multi-dimensional automatic evaluation mechanism for student performance, which provides students with personalized training feedback by integrating multiple AI models and scoring algorithms.

[0199] (1) Evaluation dimensions and methods:

[0200] Description of the dimension's technical support;

[0201] Keyword hit rate: count the number of target keywords appearing in students’ answers and their proportion to the preset keywords;

[0202] Semantic coverage, using semantic embeddings (such as BERT vectors) to calculate the degree of overlap between student responses and the semantics of the teaching objectives;

[0203] Articulation clarity, using the signal-to-noise ratio or error rate (such as WER) during transcription by the ASR module as a reference indicator of speech clarity;

[0204] Fluency, which comprehensively assesses fluency by analyzing pause time, speaking speed and other indicators in the audio;

[0205] Grammatical accuracy: The syntax of user sentences is checked and scored by calling LLM (such as GPT).

[0206] (2) FastAPI backend interface example

[0207] @app.post(" / evaluate_response")

[0208] asyncdefevaluate_response(user_text:str,audio_path:str,prompt:dict):

[0209] keyword_score=calc_keyword_hit(user_text,prompt["goal_keywords"])

[0210] semantic_score=calc_semantic_similarity(user_text,prompt["teaching_goals"])

[0211] fluency_score=analyze_audio_fluency(audio_path)

[0212] pronunciation_score=asr_confidence_score(audio_path)

[0213] grammar_score=grammar_check_llm(user_text)

[0214] total_score=weighted_average([keyword_score,semantic_score,fluency_score,pronunciation_score,grammar_score])

[0215] return{"scores":{"Keyword hit rate":keyword_score,

[0216] "Semantic coverage": semantic_score,

[0217] "Pronunciation clarity":pronunciation_score,

[0218] "Fluency": fluency_score,

[0219] "Grammar Accuracy":grammar_score,

[0220] "Total score": total_score}

[0221] (3) Visual feedback example:

[0222]

[0223]

[0224] Technical architecture mapping relationship (system layer comparison)

[0225]

[0226] In this embodiment, preferably, the user input parsing module supports users to describe the training objectives in just one sentence of natural language, and the system automatically constructs the structured elements required for the corresponding training scenario through intention recognition, keyword extraction and teaching objective reasoning.

[0227] In this embodiment, preferably, the Prompt generation module can dynamically generate a Prompt template based on the parsed teaching intention, and support adaptive adjustment of template fields, including but not limited to training objectives, dialogue keywords, banned words, and suggested expressions.

[0228] In this embodiment, preferably, the workflow of the off-topic detection and topic guidance module includes the following steps:

[0229] Step B1: Speech transcription: The student's speech input is first transcribed into processable text content through the automatic speech recognition (ASR) module, which serves as the basis for subsequent semantic analysis;

[0230] Step B2: Keyword extraction and target topic modeling. The system extracts core keywords, phrases, and semantic tags from the current prompt or teaching objective, and constructs a semantic vector representation of the target topic based on the BERT and SimCSE models.

[0231] Step B3: Semantic analysis of user input: the student's transcribed text input is encoded using the same semantic vector model to obtain a semantic vector representation;

[0232] Step B4: Similarity calculation and topic consistency determination: Calculate the semantic similarity between the user input vector and the target topic vector. If the similarity is lower than the set threshold, it is determined to be off-topic input;

[0233] Step B5: Identify and categorize off-topic content. Categorize off-topic content for subsequent refined guidance strategies.

[0234] Step B6: generating guidance content, combining the current teaching stage and the type of the off-topic question to generate guidance prompt content or synthesized voice playback;

[0235] Step B7: User feedback and guidance loop, presenting guidance prompts to students to guide them back to the topic, and recording the frequency of off-topic deviations for subsequent teaching feedback and strategy optimization.

[0236] In this embodiment, preferably, the student evaluation module provides students with personalized training feedback by integrating multiple AI models and scoring algorithms, including the following steps:

[0237] Step C1: Data collection during training. The system collects multimodal performance data of users during training in real time, including voice audio, transcribed text, interaction duration, and response delay information.

[0238] Step C2: Keyword matching analysis: perform keyword recognition and comparison on the user input text, and count the number and weight of the teaching keywords set by the prompt to evaluate the relevance of the user content;

[0239] Step C3: Speech fluency analysis: analyzing the user's speech data for speech rate, pauses, repetitions, and correction features, and using a fluency scoring algorithm to assess the naturalness and coherence of the expression;

[0240] Step C4: Pronunciation accuracy assessment: Use speech recognition and phoneme alignment technology to assess the accuracy of the user's pronunciation, compare it with the standard pronunciation model, and mark pronunciation problem words;

[0241] Step C5: Generate a multi-dimensional score, calculate a comprehensive score based on the above multiple dimensions, and provide feedback on the user's oral performance in the form of charts or text;

[0242] Step C6: Personalized feedback generation: the system combines the scoring results with historical performance to generate personalized learning suggestions.

[0243] In this embodiment, preferably, the multi-round dialogue context management and memory retention module records the historical interaction content between the user and the AI ​​by constructing an information state map, including user intention, scene setting, and keyword usage status, and calls the context information in subsequent responses to generate a more consistent system reply, including the following steps:

[0244] Step D1: Conversation information collection and annotation. In each round of conversation, the system extracts and annotates the core information from user input and AI output in real time, including user intent, keywords, scenario settings, and emotional tendencies.

[0245] Step D2: Information state graph construction: The extracted information is used to construct an information state graph in a structured manner to represent the dynamic semantic state of the user-AI interaction process;

[0246] Step D3: Context information storage and update: The information state graph is used as the context storage structure and dynamically updated after each round of dialogue to ensure consistency and continuity between historical semantics and the latest user input.

[0247] Step D4: Contextual information retrieval and integration: When generating AI responses, key content from the current information state graph is retrieved to guide the language model to generate responses with contextual coherence.

[0248] Step D5: Consistency check and context summary: Consistency check is performed on the content generated by multiple rounds of dialogue to avoid information contradictions. At the same time, the long dialogue history is compressed through the context summary mechanism to improve the system response efficiency.

[0249] In this embodiment, preferably, the dynamic role generation and allocation module supports multi-user collaborative training. Users can choose to join multiple rounds of role-playing dialogues among system-generated or customized AI roles. The system dynamically maintains the identity and responsibilities of each role and updates the prompt content to adapt to the current scenario, including the following steps:

[0250] Step E1: Multiple users enter the training scene. The system supports multiple users or users and multiple AI characters entering the same training task at the same time, starting the collaborative training mode.

[0251] Step E2: Role list generation and display. The system automatically generates multiple AI dialogue roles based on the current teaching scenario set by Prompt, and displays the role list and role responsibilities to the user.

[0252] Step E3: Role selection and confirmation. The user can select a role from the system recommended roles to join the training, or customize the role settings. After confirmation, enter the role-playing session;

[0253] Step E4: Dynamically update the prompt content. The system updates the prompt template in real time based on the user's selected role, including updating the participant role fields, conversation context, and assigned responsibilities to ensure that the generated multi-round conversations conform to the current division of labor settings.

[0254] Step E5: Collaborative conversation control and identity maintenance. During multiple rounds of conversation, the system dynamically maintains the identity tag and semantic memory of each character to ensure that the AI ​​or multiple users' responses are consistent with their respective character settings to avoid confusion.

[0255] Step E6: Collaborative training task guidance and feedback. The system generates feedback and prompts based on each role's task objectives and speech performance, guiding multiple parties to collaboratively complete training tasks, such as information exchange, problem solving, or process simulation.

[0256] An intelligent spoken English training method based on any of the above systems comprises the following steps:

[0257] Step A1: Training intention input: The user inputs the current training intention through natural language, such as "I want to practice English conversation in airport scenes";

[0258] Step A2: Intent parsing and prompt construction: The user input parsing module identifies the scene, role, and task elements in the sentence, and then passes them to the prompt generation module to build a structured prompt template;

[0259] Step A3: Dynamically generate and assign AI roles. The system calls the role generation and assignment module to generate or assign corresponding AI dialogue roles based on the training scenario and update the role information in Prompt.

[0260] Step A4: Large language model content generation: The large language model response module generates multiple rounds of dialogue content based on prompts and role settings, guiding users to conduct scenario simulation training;

[0261] Step A5: Voice input and transcription: The user participates in oral dialogue training through voice, and the system transcribes the user's voice into text input through the automatic speech recognition module;

[0262] Step A6: Off-topic detection and topic guidance: The off-topic detection and topic guidance module analyzes the user input text in real time to determine whether it deviates from the training topic and provides text or voice guidance to return to the main line of the conversation when necessary.

[0263] Step A7: Conversation context and memory retention: The multi-round conversation context management and memory retention module continuously maintains the current conversation state and historical information to ensure the consistency and personalization of the AI ​​character's responses.

[0264] Step A8: System output and speech synthesis: the system outputs the response content and synthesizes it into speech form through the text-to-speech module to achieve natural and smooth human-computer interaction;

[0265] Step A9: Student performance evaluation and feedback generation,The student performance evaluation module comprehensively analyzes the user's performance in this round of training and generates visual feedback.

[0266] Example 1

[0267] The following is a flowchart of the overall implementation of the method of the present invention in the "student-side customized training scenario", which is used to describe the working mechanism and interaction logic between modules, and helps to more intuitively understand the system operation process and its key links:

[0268] User input parsing module

[0269] The system identifies the training need as spoken language related to "restaurant ordering" and the extraction intention is to learn common expressions in ordering scenarios. Keywords include "menu", "order", "I'd like", etc.

[0270] Prompt generation module

[0271] The system automatically generates a structured Prompt template based on the parsing results as follows:

[0272] {"scene_name":"Restaurant Order",

[0273] "background":"You are in a western restaurant preparing to order food.",

[0274] "user_roles":["customer"],

[0275] "ai_roles":["restaurant waiter"],

[0276] "active_roles":{"user":"Customer","ai":"Restaurant Waiter"},

[0277] "user_defined_roles":{"Customer":"A customer who comes to dine in and wishes to order a main course and a drink"},

[0278] "teaching_goals": ["Master common ordering terms in Western restaurants", "Correctly express dining needs"],

[0279] "topic_scope":["Order","Recommended Dishes","Checkout"],

[0280] "goal_keywords":["menu","order","I'd like","check please"],

[0281] "per_turn_goals":[{"turn":1,"expected_keywords":["I'd like","menu"]},

[0282] {"turn":2,"expected_keywords":["steak","drink"]},

[0283] {"turn":3,"expected_keywords":["check please"]}

[0284] "sensitive_words":["politics","religion"],

[0285] "default_response":"You can say: I'd like to order the steak.",

[0286] "post_response":"Very good, you have completed this ordering task,"

[0287] "user_profile":

[0288] {"name":"Student A","gender":"female","age":20,"profession":"College student","location":"Beijing","interests":["Travel","Food","English learning"]},

[0289] "dialogue_context":{"topic":"Ordering food at a restaurant","history"},

[0290] "intention_graph":{"nodes":["Order","Confirm Dishes","Request Bill"],

[0291] "edges":[{"from":"Order","to":"Confirm dishes"},{"from":"Confirm dishes","to":"Request bill"}],

[0292] "info_state_graph":{"topic":"Restaurant Order","known_slots":{"Main Course":null,"Drinks":null,"Dessert Required":null},

[0293] "status": "Initializing",

[0294] "consistency_checker":{"enabled":true,"last_checked":""},

[0295] "topic_tracker":{"current_topic":"Order","topic_switch_log"}

[0296] Dynamic role generation and assignment module

[0297] The system assigns the "customer" role to students by default and supports custom background descriptions; the AI ​​role is "restaurant waiter", which can be expanded to a collaborative dialogue between two AI waiters in the future.

[0298] Large language model response module

[0299] Based on the above prompt, the system generates greetings and guiding words in real context (such as: "Good evening! Here is the menu. What would you like to have today?"), and generates natural dialogue content based on the user's answer.

[0300] ASR speech recognition module

[0301] The student's voice input is converted into text in real time and fed into the semantic analysis and topic deviation detection module.

[0302] Off-topic detection and topic guidance module

[0303] It is used to identify whether user input has significant semantic deviations from the current teaching scenario during multiple rounds of human-computer interaction. For example, in the "restaurant ordering" training scenario, if a student inputs content such as "Do you support some political view?", the system will use the built-in semantic deviation detection mechanism to calculate the semantic relevance between the input and the teaching topic. When the relevance is lower than the set threshold, the system determines that it is inconsistent with the current task objective and triggers a preset guiding statement, such as: "Let's continue practicing the ordering part first, for example, what dishes would you like to order?" In this way, the system can effectively guide students back to the expected learning path and maintain the coherence and pertinence of the teaching task.

[0304] To improve judgment accuracy, the system optionally uses the sensitive_words field to flag keywords that may cause semantic mutations in specific teaching contexts (such as those related to society, faith, and violence). This field serves only as an auxiliary judgment basis. When input containing these words is detected, the system does not interrupt or block the conversation, but instead gently guides the topic back to the training objective. This mechanism is intended to enhance stability and contextual consistency during the learning process and does not provide content censorship, filtering, or restriction.

[0305] Multi-round dialogue context management and memory retention module

[0306] The system dynamically tracks and records the dishes students order, whether they order drinks, and whether they pay. If a student asks, "Give me another drink recommendation," the system can accurately make a recommendation based on the previous context.

[0307] TTS speech synthesis module

[0308] The AI ​​waiter's English response speech is synthesized and played back to provide an immersive voice interaction experience.

[0309] Student Performance Evaluation Module

[0310] After the training, the system will give visual feedback based on dimensions such as keyword hits, expression completeness, speech fluency, and speech clarity in the conversation, and save the training records for teachers or students to review and retrain in the future.

[0311] Example 2

[0312] The following is a flowchart of the implementation of the method of the present invention in "Teachers generate teaching tasks in batches", which shows the key steps and module interaction logic of the system automatically parsing and generating multiple training tasks after the teacher inputs the scenario:

[0313] User input parsing module

[0314] The system receives the training task requirements input by the teacher (such as: "Generate three oral training scenarios about airports, hotel check-in and shopping"), parses multiple scenario intentions, and extracts semantic information such as keywords, role preferences and teaching objectives as the basis for generating prompts in subsequent modules.

[0315] Prompt generation module

[0316] The system automatically builds a structured prompt template based on each identified training scenario. The template is as follows:

[0317] {"scene_name":"Airport Check-in","background":"The user simulates the check-in procedures with the counter staff at the international airport, including baggage check-in, flight confirmation, etc."

[0318] "user_roles":["Passenger A"],

[0319] "ai_roles":["Check-in staff"],

[0320] "active_roles":{"user":"Passenger A","ai":"Check-in clerk"},

[0321] "user_defined_roles":{"Passenger A":"A college freshman preparing to study abroad, traveling alone for the first time"},

[0322] "teaching_goals": ["Master common airport expressions", "Check in correctly", "Ask questions and confirm information politely"],

[0323] "topic_scope":["Baggage Check-in","Flight Confirmation","Get Boarding Pass","Gate Information"],

[0324] "topic_embedding":"[Airport boarding theme vector]",

[0325] "goal_keywords":["boarding pass","check-in","passport","baggage check-in"],

[0326] "per_turn_goals":[{"turn":1,"expected_keywords":["check-in","flight"]},

[0327] {"turn":2,"expected_keywords":["passport","boarding pass"]},

[0328] {"turn":3,"expected_keywords":["luggage","gate"]}],

[0329] "info_state_graph":{"topic":"Check-in","known_slots":{"flight number":null,"passport information":null,"number of checked bags":null},

[0330] "status":"Waiting for the user to complete all boarding information",

[0331] "consistency_checker":{"reminder_template":"Current task: Check-in at the airport. Known information: Flight number: {flight number} Passport: {passport information} Luggage: {number of checked luggage} Please continue to fill in the missing items","require_all_slots_before_closing":true},

[0332] "topic_tracker":{"topic_vector":"[Airport topic vector]","threshold":0.72,

[0333] "on_topic_fallback":"We continue with the boarding process. Do you have any checked baggage?",

[0334] "off_topic_response":"We seem to have strayed from the airport scene. Please return to the information exchange about boarding."},

[0335] "sensitive_words":["politics","religion","violence"],

[0336] "default_response":"Please provide me your flight details and passport information for check-in.",

[0337] "post_response":"You have successfully completed the check-in task. Well done!"}

[0338] Dynamic role generation and assignment module

[0339] The system automatically configures appropriate user roles and AI roles based on different scenarios (e.g., airport: passengers and ground staff; hotel: guests and front desk staff; shopping: customers and store clerks).

[0340] Large language model response module

[0341] After the teacher confirms the prompt, the system uses a large language model to generate AI-side guiding words, demonstration dialogue openings and example interactions (for example: "Welcome! Do you have servation?"), providing a language paradigm for students' subsequent training.

[0342] ASR speech recognition module

[0343] After a student starts a teaching task on the training side, his or her voice input is transcribed into text in real time by the system, which serves as the input basis for language understanding and dialogue guidance.

[0344] Off-topic detection and topic guidance module

[0345] The system will promptly correct students when they stray from the topic based on the topics and keyword ranges defined by the teacher (e.g., discussions on "political" content are prohibited in airport scenes), and will return them to the task topic through guiding words (e.g., "Let's continue to complete the dialogue exercise for check-in").

[0346] Multi-round dialogue context management and memory retention module

[0347] The system automatically tracks the information provided by students in each round of dialogue (such as whether their ID card has been presented or whether they have chosen breakfast) and dynamically updates the context in the AI ​​response, making the dialogue coherent and having a real sense of task advancement.

[0348] TTS speech synthesis module

[0349] Convert the English lines of the AI ​​character into voice broadcast, so that teachers can carry out immersive listening and speaking training for students in batch tasks, and improve the realism of interaction and the quality of language input.

[0350] Student Performance Evaluation Module

[0351] After each batch of generated training scenarios is completed, the system automatically scores based on the completion of the set goals for each round, keyword hit rate, speech clarity and coherence, generates a feedback report, and supports teachers to view the comprehensive performance records of all students in different tasks.

[0352] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent spoken English training system based on the Prompt project, characterized in that: include: The user input parsing module is used to receive training requirements input by users or teachers in the form of natural language descriptions, and perform semantic analysis on the input to extract core information such as training scenarios, teaching objectives, and keywords; A prompt generation module is used to construct a structured prompt template based on the output of the user input parsing module. The template includes scene background, user role, AI role, teaching objectives, keywords, and sensitive word fields; A large language model response module is used to generate multi-round interactive dialogue content that conforms to the context based on the prompt template and role settings; ASR recognition module, used to convert user voice input into text for system analysis; The off-topic detection and topic guidance module is used to check the topic consistency of user input content. If it is detected that it deviates from the teaching objectives or keyword range set by the prompt, it will generate voice or text guidance content to prompt the user to return to the conversation topic; The TTS synthesis module is used to convert AI-generated text into speech to achieve natural voice interaction; The multi-round dialogue context management and memory retention module is used to record and manage historical dialogue information and key semantic content, ensuring information continuity between rounds. It also has a consistency check and context summary mechanism for generated content. Dynamic role generation and assignment module, used to automatically generate multiple AI dialogue roles for users to choose from based on training topics, support user-defined roles, and dynamically adjust prompt templates and dialogue content based on user selections; The student performance evaluation module is used to analyze the user's performance during the training process and generate feedback and scores based on keyword matching, voice fluency, and pronunciation accuracy.

2. The English spoken language intelligent training system based on the Prompt project according to claim 1, characterized in that: The user input parsing module supports users to describe training objectives in just one sentence of natural language. The system automatically constructs the structured elements required for the corresponding training scenario through intent recognition, keyword extraction and teaching objective reasoning.

3. The system according to claim 1 is an intelligent spoken English training system based on the Prompt project, characterized in that: The Prompt generation module can dynamically generate a Prompt template based on the parsed teaching intention, and supports adaptive adjustment of template fields, including but not limited to training objectives, dialogue keywords, banned words, and suggested expressions.

4. According to the intelligent spoken English training system based on the Prompt project in claim 1, the workflow of the off-topic detection and topic guidance module comprises the following steps: Step B1: Speech transcription: The student's speech input is first transcribed into processable text content through the automatic speech recognition (ASR) module, which serves as the basis for subsequent semantic analysis; Step B2: Keyword extraction and target topic modeling. The system extracts core keywords, phrases, and semantic tags from the current prompt or teaching objective, and constructs a semantic vector representation of the target topic based on the BERT and SimCSE models. Step B3: Semantic analysis of user input: the student's transcribed text input is encoded using the same semantic vector model to obtain a semantic vector representation; Step B4: Similarity calculation and topic consistency determination: Calculate the semantic similarity between the user input vector and the target topic vector. If the similarity is lower than the set threshold, it is determined to be off-topic input; Step B5: Identify and categorize off-topic content. Categorize off-topic content for subsequent refined guidance strategies. Step B6: generating guidance content, combining the current teaching stage and the type of the off-topic question to generate guidance prompt content or synthesized voice playback; Step B7: User feedback and guidance loop, presenting guidance prompts to students to guide them back to the topic, and recording the frequency of off-topic deviations for subsequent teaching feedback and strategy optimization.

5. The English spoken language intelligent training system based on Prompt Project according to claim 1, characterized in that: The student evaluation module provides students with personalized training feedback by integrating multiple AI models and scoring algorithms. It includes the following steps: Step C1: Data collection during training. The system collects multimodal performance data of users during training in real time, including voice audio, transcribed text, interaction duration, and response delay information. Step C2: Keyword matching analysis: perform keyword recognition and comparison on the user input text, and count the number and weight of the teaching keywords set by the prompt to evaluate the relevance of the user content; Step C3: Speech fluency analysis: analyzing the user's speech data for speech rate, pauses, repetitions, and correction features, and using a fluency scoring algorithm to assess the naturalness and coherence of the expression; Step C4: Pronunciation accuracy assessment: Use speech recognition and phoneme alignment technology to assess the accuracy of the user's pronunciation, compare it with the standard pronunciation model, and mark pronunciation problem words; Step C5: Generate a multi-dimensional score, calculate a comprehensive score based on the above multiple dimensions, and provide feedback on the user's oral performance in the form of charts or text; Step C6: Personalized feedback generation: the system combines the scoring results with historical performance to generate personalized learning suggestions.

6. The English spoken language intelligent training system based on Prompt Project according to claim 1, characterized in that: The multi-round dialogue context management and memory retention module records the historical interaction content between the user and the AI ​​by building an information state map, including user intent, scenario settings, and keyword usage status. It then uses this context information in subsequent responses to generate more consistent system responses, including the following steps: Step D1: Conversation information collection and annotation. In each round of conversation, the system extracts and annotates the core information from user input and AI output in real time, including user intent, keywords, scenario settings, and emotional tendencies. Step D2: Information state graph construction: The extracted information is used to construct an information state graph in a structured manner to represent the dynamic semantic state of the user-AI interaction process; Step D3: Context information storage and update: The information state graph is used as the context storage structure and dynamically updated after each round of dialogue to ensure consistency and continuity between historical semantics and the latest user input. Step D4: Contextual information retrieval and integration: When generating AI responses, key content from the current information state graph is retrieved to guide the language model to generate responses with contextual coherence. Step D5: Consistency check and context summary: Consistency check is performed on the content generated by multiple rounds of dialogue to avoid information contradictions. At the same time, the long dialogue history is compressed through the context summary mechanism to improve the system response efficiency.

7. The English spoken language intelligent training system based on Prompt Project according to claim 1, characterized in that: The dynamic role generation and assignment module supports multi-user collaborative training. Users can choose to participate in multiple rounds of role-playing dialogues among system-generated or customized AI roles. The system dynamically maintains the identity and responsibilities of each role and updates the prompt content to adapt to the current scenario. The steps include the following: Step E1: Multiple users enter the training scene. The system supports multiple users or users and multiple AI characters entering the same training task at the same time, starting the collaborative training mode. Step E2: Role list generation and display. The system automatically generates multiple AI dialogue roles based on the current teaching scenario set by Prompt, and displays the role list and role responsibilities to the user. Step E3: Role selection and confirmation. The user can select a role from the system recommended roles to join the training, or customize the role settings. After confirmation, enter the role-playing session; Step E4: Dynamically update the prompt content. The system updates the prompt template in real time based on the user's selected role, including updating the participant role fields, conversation context, and assigned responsibilities to ensure that the generated multi-round conversations conform to the current division of labor settings. Step E5: Collaborative conversation control and identity maintenance. During multiple rounds of conversation, the system dynamically maintains the identity tag and semantic memory of each character to ensure that the AI ​​or multiple users' responses are consistent with their respective character settings to avoid confusion. Step E6: Collaborative training task guidance and feedback. The system generates feedback and prompts based on each role's task objectives and speech performance, guiding multiple parties to collaboratively complete training tasks, such as information exchange, problem solving, or process simulation.

8. An intelligent spoken English training method based on the system according to any one of claims 1 to 7, comprising the following steps: Step A1: Training intention input: The user inputs the current training intention through natural language, such as "I want to practice English conversation in airport scenes"; Step A2: Intent parsing and prompt construction: The user input parsing module identifies the scene, role, and task elements in the sentence, and the prompt generation module constructs a structured prompt template. Step A3: Dynamically generate and assign AI roles. The system calls the role generation and assignment module to generate or assign corresponding AI dialogue roles based on the training scenario and update the role information in Prompt. Step A4: Large language model content generation: The large language model response module generates multiple rounds of dialogue content based on prompts and role settings, guiding users to conduct scenario simulation training; Step A5: Voice input and transcription: The user participates in oral dialogue training through voice, and the system transcribes the user's voice into text input through the automatic speech recognition module; Step A6: Off-topic detection and topic guidance: The off-topic detection and topic guidance module analyzes the user input text in real time to determine whether it deviates from the training topic and provides text or voice guidance to return to the main line of the conversation when necessary. Step A7: Conversation context and memory retention: The multi-round conversation context management and memory retention module continuously maintains the current conversation state and historical information to ensure the consistency and personalization of the AI ​​character's responses. Step A8: System output and speech synthesis: the system outputs the response content and synthesizes it into speech form through the text-to-speech module to achieve natural and smooth human-computer interaction; Step A9: Student performance evaluation and feedback generation,The student performance evaluation module comprehensively analyzes the user's performance in this round of training and generates visual feedback.

Citation Information

Cited By

  • Conversational AI task execution method and system based on state diagram

    CN121255397A

  • English auxiliary teaching software multi-modal optimization method based on embedded system

    CN121456799A

  • English auxiliary teaching software multi-modal optimization method based on embedded system

    CN121456799B

  • Generative financial report writing method based on role division

    CN121997900A

  • A role-based generative financial report writing method

    CN121997900B