Coaching method, electronic equipment and computer readable storage medium

By obtaining the user's emotional state and adjusting the voice tone, we generate conversation content that fits the user's emotions, solving the problem of insufficient emotional expression in AI tutoring and improving the naturalness and realism of learning tutoring.

CN120687558APending Publication Date: 2025-09-23WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510639369.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing AI conversational learning tutoring technology lacks real emotional perception, resulting in insufficient emotional expression and mechanical responses that reduce students' learning enthusiasm and interactive experience.

Method used

By obtaining the user's emotional state, the dialogue content generation model is used to generate response content containing copy content and dialogue tone. Combined with the speech generation model and expression library, the tone, speed and pauses are adjusted to enhance emotional expression.

Benefits of technology

It improves the naturalness and authenticity of learning guidance, enhances students' emotional resonance and immersion, and improves their learning enthusiasm and interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687558A_ABST
    Figure CN120687558A_ABST
Patent Text Reader

Abstract

The invention discloses a tutoring method, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining an emotional state of a user under the condition that first dialogue content of the user is received, and the first dialogue content comprises a tutoring demand of the user; based on the first dialogue content and the emotional state, a dialogue content generation model is used to generate second dialogue content responding to the first dialogue content, and the second dialogue content comprises copywriting content and dialogue mood; and outputting the second dialogue content so as to tutorize the user. By means of the scheme, the emotion expression ability in AI dialogue tutoring can be enhanced, the mechanical feeling of AI dialogue tutoring is effectively improved, and the naturalness and reality sense of the interaction process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to the field of artificial intelligence technology. More specifically, the present application relates to a method, an electronic device, and a computer-readable storage medium for tutoring. Background Art

[0002] With the rapid development of artificial intelligence (AI), its application in education continues to expand. With its powerful natural language processing and data analysis capabilities, AI has brought a new model to learning tutoring for students of all ages. For children aged 6-12, who are in the critical early learning stages, AI conversational tutoring overcomes the time-consuming, costly, and resource-limited limitations of traditional manual instruction. Through real-time interaction, students can input their learning needs through natural conversation and quickly receive learning suggestions and guidance. Whether it's Chinese writing, math problem-solving, or English conversation practice, AI can efficiently absorb knowledge and improve skills. Currently, mainstream AI conversational tutoring solutions are mostly based on prompt debugging of the general LLM large language model. They assist students in completing learning tasks through role setting, goal planning, and step-by-step guidance. Even without prompt debugging and direct conversation, LLM can achieve basic interaction, but it often fails to achieve ideal learning results.

[0003] However, existing AI-powered conversational learning tutoring technology still has significant shortcomings. Due to AI's lack of authentic perception and personal experience with human emotions, it struggles to express contextually appropriate emotions based on students' emotional states and learning difficulties during tutoring conversations, hindering students from establishing effective emotional connections with AI. Furthermore, AI's inability to flexibly adjust language details like intonation, speed, and pauses when delivering content leads to mechanical responses that lack warmth in learning interactions, reducing students' enthusiasm and engagement, and severely impacting the effectiveness of tutoring and user experience.

[0004] In view of this, there is an urgent need to provide a solution for tutoring in order to enhance the emotional expression ability in AI dialogue tutoring, effectively improve the mechanical feel of AI dialogue tutoring, and enhance the naturalness and realism of the interaction process. Summary of the Invention

[0005] In order to at least solve one or more technical problems mentioned above, the present application proposes solutions for tutoring in the following aspects.

[0006] In a first aspect, the present application provides a method for tutoring, comprising: upon receiving a first conversation content of a user, obtaining the emotional state of the user, wherein the first conversation content includes the tutoring needs of the user; based on the first conversation content and the emotional state, using a conversation content generation model to generate a second conversation content in response to the first conversation content, wherein the second conversation content includes text content and conversational tone; and outputting the second conversation content to provide tutoring to the user.

[0007] In some embodiments, obtaining the emotional state of the user includes: determining the input order of the first conversation content in the current conversation; determining whether the input order is a preset order; if the input order is a preset order, using a sentiment analysis model to generate the emotional state of the user in the current preset order based on the first conversation content; if the input order is not a preset order, obtaining the emotional state of the user in the previous preset order.

[0008] In some embodiments, based on the first conversation content and the emotional state, using a conversation content generation model to generate a second conversation content in response to the first conversation content includes: based on the first conversation content, the emotional state and the user's coaching strategy, using a conversation content generation model and a preset knowledge base to generate a second conversation content in response to the first conversation content.

[0009] In some embodiments, the tutoring strategy is obtained by the following operations: obtaining the user's historical tutoring data, wherein the historical tutoring data includes historical conversation content, historical teaching methods, and historical teaching methods of related knowledge points; based on the historical tutoring data, using a preference analysis model to generate the user's preference characteristics, wherein the preference characteristics include learning level, learning style, and language adaptability; based on the preference characteristics, using a tutoring strategy generation model to generate the user's tutoring strategy, wherein the tutoring strategy includes tutoring difficulty, tutoring style, and tutoring language.

[0010] In some embodiments, outputting the second dialogue content includes: generating dialogue speech using a speech generation model based on the text content and the dialogue tone; and outputting the dialogue speech.

[0011] In some embodiments, the method further includes: selecting a preset expression in an expression library that matches the tone of the conversation according to an expression output rule; and outputting the text content and the preset expression.

[0012] In some embodiments, the first conversation content is in speech form; and based on the first conversation content, using the sentiment analysis model to generate the user's emotional state includes: performing noise reduction and standardization operations on the first conversation content to preprocess the first conversation content to obtain preprocessed first conversation content; based on the preprocessed first conversation content, using the sentiment analysis model to generate the user's emotional state.

[0013] In some embodiments, the dialogue content generation model is obtained by debugging using preset prompt words; and the preset prompt words include a first prompt word and a second prompt word, the first prompt word includes a coaching role, a coaching task and a coaching step, and the second prompt word includes an emotional state and / or a coaching strategy.

[0014] In a second aspect, the present application provides an electronic device comprising: a processor; and a memory storing program instructions for tutoring, wherein when the program instructions are executed by the processor, the method and multiple embodiments thereof described in the first aspect are implemented.

[0015] In a third aspect, the present application provides a computer-readable storage medium having stored thereon program instructions for tutoring, wherein when the program instructions are executed by a processor, the method described in the first aspect and its multiple embodiments are implemented.

[0016] The solution for tutoring provided above can obtain the user's emotional state when receiving the user's first dialogue content containing tutoring needs, and based on the emotional state and tutoring needs, use the dialogue content generation model to generate a second dialogue content containing both text content and dialogue tone. Compared with the mechanical interaction problem caused by the lack of emotional understanding of AI in the prior art, this solution incorporates the user's emotional state into the consideration of dialogue generation, so that the generated response content can fit the user's emotional state and learning confusion, significantly enhancing the emotional expression ability during the dialogue process; by incorporating adjustments to the dialogue tone (such as intonation, speaking speed, pauses and other details) into the output content, it effectively improves the mechanical feel of traditional AI responses and enhances the naturalness and realism of the interaction process. Using the solution of this application, a more effective emotional connection can be established, enhancing students' emotional resonance and immersion in the learning tutoring process, thereby improving learning enthusiasm and involvement, and effectively optimizing the actual effect and user experience of learning tutoring, solving the core problems of lack of emotional expression and poor interactive experience in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0018] Figure 1 is an exemplary flow chart illustrating a method for tutoring according to an embodiment of the present application;

[0019] Figure 2 is an exemplary schematic diagram illustrating output of a conversation content generation model according to an embodiment of the present application;

[0020] Figure 3 is a schematic diagram illustrating an exemplary interaction between a conversation content generation model and a RAG knowledge base according to an embodiment of the present application;

[0021] Figure 4 is an exemplary schematic diagram illustrating the use of a conversation content generation model and a RAG knowledge base to generate second conversation content according to an embodiment of the present application;

[0022] Figure 5 is an exemplary flow chart illustrating a process of acquiring an emotional state according to an embodiment of the present application;

[0023] Figure 6 is an exemplary schematic diagram illustrating obtaining an emotional state when the first conversation content is in voice form according to an embodiment of the present application;

[0024] Figure 7 is an exemplary flow chart illustrating a process of generating a tutoring strategy according to an embodiment of the present application;

[0025] Figure 8 is a schematic diagram illustrating a process of generating a tutoring strategy for a user according to an embodiment of the present application;

[0026] Figure 9 is an exemplary structural block diagram showing a system for tutoring according to an embodiment of the present application;

[0027] Figure 10 is an exemplary interaction diagram illustrating a method for tutoring implemented based on a two-end interaction system according to an embodiment of the present application;

[0028] Figure 11 FIG. 4 is a block diagram showing an exemplary structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0030] It should be understood that when the terms "first," "second," "third," and "fourth," etc., are used in the claims, specification, and drawings of this application, they are only used to distinguish different objects, rather than to describe a specific order. The terms "comprise" and "comprising" used in the specification and claims of this application indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0031] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this specification and claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this specification and claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0032] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0033] The specific implementation of the present application will be described in detail below with reference to the accompanying drawings. Figure 1 The following is an exemplary flow chart of a method 100 for tutoring according to an embodiment of the present application. It is understood that method 100 can be executed by any appropriate device with data processing capabilities, including but not limited to terminal devices, processors, and servers. Terminal devices include but are not limited to smartphones, smart learning machines, personal computers, laptops, tablet computers, and portable wearable devices.

[0034] like Figure 1As shown, at step S101, method 100 can obtain the user's emotional state upon receiving the user's conversation content (for distinction, this can be referred to as the first conversation content). Next, at step S102, method 100 can use the conversation content generation model to generate conversation content (for distinction, this can be referred to as the second conversation content) in response to the first conversation content based on the first conversation content and the emotional state. Here, the second conversation content includes the text content and the conversation tone. Finally, at step S103, method 100 can output the second conversation content to provide guidance to the user.

[0035] In step S101, the user can input the first conversation content in voice or text form through natural conversation, and the first conversation content includes the user's tutoring needs. The user's tutoring needs may include subject type, learning tasks, and personalized requirements. The subject type is used to identify the learning content, such as Chinese writing, math problem solving, or English conversation; the learning task is used to identify specific learning goals, which may include but is not limited to writing guidance, rhetorical device analysis, topic analysis requirements, language expression training, etc.; personalized requirements may include writing style preferences (such as the vividness of narrative essays and the logic of expository essays), the degree of detail of problem-solving steps, and the frequency of feedback.

[0036] The aforementioned feedback frequency refers to the time interval and frequency density at which users expect AI to respond or provide guidance during the learning tutoring process. In the learning tutoring scenario, different users have different needs for feedback. For example, some users hope that AI will give targeted evaluations and suggestions immediately after they input a small piece of content, complete a problem-solving step, or think for a few minutes. This demand corresponds to high-frequency feedback; while some users prefer to receive comprehensive feedback from AI all at once after completing an entire essay, solving a set of math problems, or after a relatively long period of study. This is low-frequency feedback. By allowing users to set the feedback frequency independently, the AI ​​tutoring rhythm can better adapt to personal learning habits and thinking patterns, thereby improving the adaptability of learning tutoring and user experience.

[0037] The aforementioned user's emotional state refers to the emotional characteristics displayed by the user during the conversation, which are obtained through natural language processing technology, speech recognition technology or multimodal interaction technology. Specifically, it may include: emotional type, emotional intensity and / or non-verbal emotional clues. Emotional types may include emotional tendencies such as happiness, confusion, anxiety, excitement, anger, etc. based on semantic content analysis; emotional intensity, that is, the intensity level of the above emotions, such as mild confusion, severe anxiety; non-verbal emotional clues, such as the use of punctuation marks in the conversation text (such as continuous exclamation marks reflecting excitement), input time intervals (long pauses reflecting difficulty in thinking), etc., which indirectly represent the emotional state. The process of obtaining the user's emotional state combines the text content input by the user, the voice and intonation features (if there is voice interaction) and the interactive behavior data to form a multi-dimensional characterization of the user's emotional state, and provide input parameters of the emotional dimension for the subsequent generation of conversation content. In order to facilitate understanding, how to obtain the user's emotional state will be combined later. Figure 5 The process 500 of acquiring the emotional state is described in detail.

[0038] In step S102 , the conversation content generation model can be constructed based on a general-purpose large language model (LLM). Specifically, it can adopt any publicly available or future general-purpose large language model architecture, such as GPT4o or DeepSeekV3. In practice, the second conversation content generated by the conversation content generation model can include at least one text content and at least one conversational tone corresponding to the text content.

[0039] Based on the first conversation content and the user's emotional state, the conversation content generation model can generate the second conversation content in a dual-track format of [content + tone], where the copy content provides a specific response to the user's tutoring needs; the conversation tone is generated based on the user's emotional state, and can include language details such as intonation, speaking speed, and pauses. For example, for users with anxious emotional states, a gentle and low tone can be used, the speaking speed can be slowed down by 20%-30%, and pauses can be added between sentences, thereby providing parameter basis for subsequent speech synthesis or expression display, and realizing immersive tutoring responses that fit the user's emotions and personalized needs.

[0040] In the aforementioned step S103, the method 100 may output the second conversation content in voice form, text form, or voice and text on the same screen form to provide guidance to the user.

[0041] When outputting in speech form, the second dialogue content can be output by performing the following operations: generating dialogue speech based on the text content and the conversational tone using a speech generation model; and outputting the dialogue speech. The speech generation model herein can be any currently available or future available speech generation model; as long as the speech generation model can convert the text content into speech and the conversational tone into specific speech parameters, it can be used to implement the tutoring solution of this application.

[0042] In the case of outputting in text form, the second dialogue content can be output by performing the following operations: selecting a preset expression in the expression library that matches the dialogue tone according to the expression output rules; and outputting the text content and the preset expression.

[0043] The expression library stores various preset expressions and their corresponding emotional attribute tags (such as "encouragement," "gentleness," "comfort," and "lively"). When selecting a matching preset expression from the expression library, the corresponding emotional attribute tag is retrieved from the expression library based on the tone of the conversation, and the preset expression associated with the corresponding emotional attribute tag is used as the preset expression that matches the tone of the conversation. It is understood that the various preset expressions in the expression library can be any form of visual data, such as images, text expressions, and animation files, and this application does not limit this.

[0044] Furthermore, in actual applications, preset emoticons can be output at a fixed location, such as the right side of the end of the text. This fixed location ensures the regularity of the preset emoticons, creating a predictable interaction pattern for users and avoiding visual clutter caused by random emoticon appearances.

[0045] In one example, if Figure 2 As shown, the second dialogue content output by the dialogue content generation model includes multiple textual content and multiple dialogue tones. On the one hand, the second dialogue content can be input into the speech generation model to generate corresponding dialogue speech, which is then output in speech form. On the other hand, based on the expression output rules, preset expressions matching each dialogue tone can be selected from the expression library, and each textual content and its matching preset expression can be output in text form.

[0046] Combination of the above Figure 1A method 100 for tutoring is described. The method 100 can obtain the user's emotional state when receiving the user's first dialogue content containing the tutoring needs, and based on the emotional state and the tutoring needs, generate a second dialogue content containing both the copy content and the dialogue tone using a dialogue content generation model. Compared with the mechanical interaction problem caused by the lack of emotional understanding of AI in the prior art, this solution incorporates the user's emotional state into the consideration of dialogue generation, so that the generated response content can fit the user's emotional state and learning confusion, significantly enhancing the emotional expression ability during the dialogue process; by incorporating the adjustment of the dialogue tone (such as intonation, speaking speed, pauses and other details) into the output content, the mechanical sense of the traditional AI response is effectively improved, and the naturalness and realism of the interaction process are enhanced. Using the solution of this application, a more effective emotional connection can be established, and the emotional resonance and immersion of students in the learning tutoring process can be enhanced, thereby improving learning enthusiasm and involvement, effectively optimizing the actual effect and user experience of learning tutoring, and solving the core problems of lack of emotional expression and poor interactive experience in the prior art.

[0047] In the process of obtaining the solution of this application, the inventors discovered that the knowledge of the large language model has certain limitations. Its own knowledge comes from the training data publicly available on the Internet, and it is unable to obtain real-time, non-public professional data. At the same time, the working principle of the large language model is based on text prediction. When no clear answer can be found, false information may be generated, which is the so-called "model hallucination." Although the knowledge limitations and hallucinations can be reduced to a certain extent by inputting specific professional knowledge and templates through Prompt debugging, the learning methods and cases are massive and dynamically developing, and the limited input of Prompt cannot completely solve the problem. Therefore, during the conversation, knowledge limitations and certain hallucinations may appear.

[0048] Furthermore, large language models can only maintain short-term memory and limited understanding during a conversation, lacking the long-term memory of humans. This makes it difficult to conduct in-depth personalized analysis based on a user's past conversations and behaviors, nor can it dynamically adjust tutoring plans based on a user's historical learning data. This limitation results in a poor user experience and poor learning outcomes, potentially leading to user churn over time.

[0049] Based on this, in order to generate a more personalized second dialogue content that better meets user needs, in some implementation scenarios, a dialogue content generation model and a preset knowledge base can be used to generate a second dialogue content in response to the first dialogue content based on the first dialogue content, emotional state, and the user's tutoring strategy. The preset knowledge base here serves as an external knowledge source, effectively supplementing professional and real-time data that is difficult for large language models to obtain; the user's tutoring strategy is based on an in-depth analysis of historical tutoring data, dynamically adjusting the tutoring direction and details. The two are linked with the dialogue content generation model to ensure that the generated second dialogue content has both knowledge accuracy and can accurately match the user's emotional state and personalized needs, fundamentally improving the interactive experience and learning outcomes, and avoiding the many drawbacks of traditional large language models when applied to tutoring scenarios.

[0050] In this implementation scenario, the dialogue content generation model can be debugged using preset prompts. These preset prompts can include a first prompt and a second prompt. The first prompt includes the tutoring role (e.g., "You are a teacher specializing in elementary school composition tutoring"), the tutoring task (e.g., "Guide students to complete the opening concept for a narrative essay"), and the tutoring steps (e.g., "First, ask students what scene they want to describe, then provide three or more opening examples"). The second prompt can include the emotional state and tutoring strategy.

[0051] The above-mentioned user's tutoring strategy may include tutoring difficulty, tutoring style and tutoring language. In order to facilitate understanding, how to generate a user's tutoring strategy will be combined later. Figure 7 The process 700 of generating a coaching strategy is described in detail.

[0052] The aforementioned tutoring difficulty can be divided into three levels: elementary, intermediate and advanced. It is understandable that different subject types have different knowledge objectives for the corresponding junior and senior high difficulty levels. In the case of Chinese writing, the knowledge objectives of elementary tutoring difficulty are basic language construction (such as mastering the basic structure of sentences), simple scene description (such as fragmented description of a single scene) and vocabulary accumulation (such as the use of common adjectives and onomatopoeia); the knowledge objectives of intermediate tutoring difficulty are paragraph coherence training (such as the use of logical cohesion words), basic rhetoric application (such as the simple application of metaphors and personification techniques) and theme clarification (such as describing around a single theme); the knowledge objectives of advanced tutoring difficulty are complex structure design (such as the "introduction, development, turn and conclusion" structure of narrative texts and the "general-specific-general" structure of expository texts), deep emotional expression (such as using actions and environment to set off mental activities) and literary enhancement (such as quoting poetry).

[0053] The aforementioned tutoring language may include lively, friendly or serious language styles. Tutoring styles may include case analysis and theoretical deduction. The case analysis type uses specific cases as the core carrier, and by disassembling examples, guides users to achieve learning goals through observation, imitation and comparison. The theoretical deduction type takes theory and logical deduction as the core, focusing on explaining principles, rules and thinking methods, and guiding users to achieve learning goals by understanding the underlying logic. In the case where the subject type is Chinese writing, the case analysis tutoring style can provide users with model essay decomposition (such as "Let's look at this sentence describing spring..."), and the understanding deduction tutoring style can provide users with writing framework deduction (such as "Narrative essays usually contain three elements: time, place, and events. We can first make an outline...").

[0054] In an embodiment of the present application, the aforementioned preset knowledge base may be a RAG knowledge base, which stores a large amount of subject-related professional knowledge data and information, and can facilitate querying by the dialogue content generation model when generating dialogue content.

[0055] In one example, if Figure 3 As shown in Figure 1, the interaction between the conversation content generation model and the RAG knowledge base is divided into three processes: data indexing, retrieval, and content generation. The purpose of data indexing can be understood as cataloging the professional knowledge data and information of the knowledge base to facilitate subsequent queries. Specifically, the first step is data extraction, which extracts data from the data sources represented by various application icons and converts it into binary code. Then, the embedding phase is entered to convert the data into vector representation, and then an index is created to organize these vectors through a specific structure for subsequent retrieval. In the retrieval phase, relevant information is searched from the index based on the user's first conversation content. The retrieved information will be automatically sorted and arranged in order according to criteria such as relevance. Finally, the conversation content generation model summarizes the sorted information to generate final content such as "Employees receive eight (8) weeks of paid maternity leave..."

[0056] Simply put, by integrating a large amount of professional knowledge data and information, accurate and comprehensive knowledge support is provided for the dialogue content generation model, ensuring the reliability and accuracy of the generated dialogue content. In addition, in the embodiments of the present application, RAG knowledge base update logic can also be added to ensure the validity of knowledge data and information.

[0057] In another example, Figure 4 FIG4 shows a process of generating the second conversation content using the conversation content generation model and the RAG knowledge base. Figure 4As shown in the figure, the first conversation content, emotional state and the user's counseling strategy are jointly input into the conversation content generation model, which queries the RAG knowledge base to obtain relevant information. Finally, the conversation content generation model combines these inputs with the query results and outputs the second conversation content, realizing the conversation content generation process that integrates the user's emotional state, counseling strategy and knowledge base data.

[0058] Next, combine Figure 5 The process 500 of obtaining the emotional state is described in detail. It can be understood that the following Figure 5 The description is a specific implementation of the above step S101. Figure 1 The features described can apply analogously here.

[0059] like Figure 5 As shown, at step S501, the input order of the first conversation content in the current conversation may be determined.

[0060] In practical applications, a user's input and an AI's response are called a conversation round. A complete AI conversation typically consists of multiple rounds of conversation between the user and the AI. The current conversation refers to the ongoing conversation between the user and the AI. In the embodiments of this application, the input order can indicate which sentence the user entered in the current conversation.

[0061] Next, at step S502 , it may be determined whether the input order is a preset order.

[0062] In actual applications, the sentiment analysis nodes and times follow the principle of business elastic adaptation. Those skilled in the art can select the specific value of the preset order according to actual needs, and this disclosure does not make specific limitations on this. Preferably, the preset order can be the third sentence or the eighth sentence. The third sentence is selected as the first analysis node based on the fact that users in dialogue scenarios usually initially show a stable emotional tendency after three rounds of interaction, and the eighth sentence is selected as the second analysis node to capture the progressive changes in user emotions after the dialogue deepens. The two analyses form an emotional evolution trajectory, providing two-dimensional data support for the adjustment of counseling strategies.

[0063] Next, at step S503, if the input order is the preset order, a sentiment analysis model can be used based on the first conversation content to generate the user's emotional state in the current preset order. The sentiment analysis model here can be any sentiment analysis model that is currently available or may be available in the future. As long as the sentiment analysis model can receive conversation content in voice or text form and output the user's emotional state, it can be used to implement the tutoring solution disclosed herein.

[0064] Furthermore, at step S504, if the input order is not the preset order, the user's emotional state at the previous preset order can be obtained. Additionally or alternatively, if the user's emotional state at the previous preset order cannot be obtained, the user's emotional state is determined to be null. The inability to obtain the user's emotional state at the previous preset order typically corresponds to the situation where the input order of the first conversation content in the current conversation is less than the minimum predicted order.

[0065] In one example, if Figure 6 As shown, in the case where the first conversation content is in the form of speech, in order to obtain a more accurate emotional state, noise reduction and standardization operations can be performed on the first conversation content to preprocess the first conversation content and obtain the preprocessed first conversation content. Next, the preprocessed first conversation content can be input into the sentiment analysis model so that the model generates and outputs the user's emotional state (such as Figure 5 happiness, patience, anger, impatience... shown in it).

[0066] The aforementioned noise reduction process uses a noise suppression algorithm to filter background noise in the first conversation in real time, ensuring the purity of the input speech and providing a clear, effective signal for subsequent analysis. Normalization processing unifies the physical parameters of the first conversation, including adjusting the volume to a preset dynamic range and standardizing the sampling rate to avoid signal distortion caused by device differences or recording environments, and ensuring consistency in the time and frequency domain characteristics of voice data from different sources.

[0067] Next, combine Figure 7 The process 700 of generating a coaching strategy according to the embodiment of the present disclosure is described in detail. It should be noted that the process of generating a coaching strategy is asynchronous with the execution process of the coaching method 100 described above, and is not generated in real time.

[0068] like Figure 7 As shown, at step S701, the user's historical tutoring data can be obtained. The historical tutoring data can be tutoring data within a preset analysis period, and can include historical conversation content, historical teaching methods, and relevant knowledge points. In actual applications, if the historical conversation content is in voice form, it can be converted into text information before input. The historical teaching method is the teaching method used when tutoring the user, and can include theoretical teaching and case teaching. Relevant knowledge points refer to knowledge points related to the historical conversation content.

[0069] The aforementioned preset analysis period can be one day, one week or one month, etc., and this application does not make specific restrictions on this. In an implementation scenario, the user's historical coaching data can be the coaching data of the last conversation the previous day, thereby achieving the generation of a user-specific coaching strategy based on the user's historical coaching data, and ensuring that the coaching content can closely fit the user's current learning status and preferences. In this implementation scenario, for any user, it is only necessary to execute the process 700 of generating a coaching strategy once a day to generate the user's coaching strategy. Next, in multiple conversations of the day, the coaching strategy is used. In short, in an embodiment of the present application, the user's coaching strategy can be updated regularly rather than in real time.

[0070] Next, at step S702, a preference analysis model can be used based on historical tutoring data to generate user preference profiles. Preference profiles may include learning level, learning style, and language adaptability. Learning level is an evaluation of the user's ability and can be classified into three levels: good, average, and poor. Learning style can include theoretical teaching and case-based teaching. Language adaptability can include, but is not limited to, serious language, lively language, and friendly language.

[0071] Finally, in step S703, a tutoring strategy generation model can be used to generate a tutoring strategy for the user based on the preference characteristics. The tutoring strategy includes tutoring difficulty, tutoring style, and tutoring language. It should be noted that the content of the tutoring strategy is consistent with the description in the previous embodiment and will not be repeated here.

[0072] The aforementioned preference analysis model and tutoring strategy generation model are large language models that have been debugged and optimized using prompt words. In practical applications, the preference analysis model and tutoring strategy generation model can be obtained by debugging the same general large language model with different prompt words, or by debugging different general large language models with different prompt words.

[0073] In one example, Figure 8 The process of generating a user's tutoring strategy is shown: historical tutoring data (historical conversation content, historical teaching methods, and related knowledge points) is used as input and enters the preference analysis model, which outputs preference features (learning level, learning style, and language adaptability) after processing; the preference features are then input into the tutoring strategy generation model, and finally the user's tutoring strategy (tutoring difficulty, tutoring style, and tutoring language) is generated, realizing a complete process from user information analysis to targeted strategy output.

[0074] Next, combine Figure 9 The following is an exemplary introduction to a system 900 for tutoring provided in an embodiment of the present application. Figure 9As shown, the system 900 for tutoring according to the embodiment of the present application may include an analysis module 901 , a generation module 902 and an output module 903 .

[0075] Furthermore, the analysis module 901 includes a preference analysis model, a counseling strategy generation model and a sentiment analysis model, which are used to analyze the user's preference characteristics, counseling strategies and emotional state; the generation module 902 includes a dialogue content generation model and a RAG knowledge base, which relies on the knowledge reserve of the RAG knowledge base and the language generation ability of the dialogue content generation model to generate the second dialogue content; the output module includes a speech generation model and an expression library, which is used to present the second dialogue content in a multimodal manner in the form of speech, text, and expressions, providing users with an intuitive and rich counseling interaction experience.

[0076] From the previous description, we can see that Figures 1 to 8 All steps of the described tutoring method are performed by a single device, which undoubtedly increases the computing pressure on the device. To balance the pressure on the device, in actual applications, some steps that directly interact with the user can be deployed on the client side, while others can be deployed on the server side, forming a dual-end interactive system.

[0077] For ease of understanding, let’s combine Figure 10 A detailed description of the tutoring method 1000 implemented based on a two-way interactive system is provided. It is understood that, to reduce the computational burden on the client, the aforementioned conversation content generation model, sentiment analysis model, preference analysis model, tutoring strategy generation model, speech generation model, and preset knowledge base can all be deployed on the server.

[0078] like Figure 10 As shown, at step S1001, the user can input the first conversation content through the client, and the client obtains the user's first conversation content and transmits it to the server, so that the server generates a second conversation content in response to the first conversation content based on the user's first conversation content.

[0079] Correspondingly, after receiving the first conversation content, the server can execute the following steps S1002 to S1004 to generate the second conversation content. Specifically, at step S1002, the server can obtain the user's emotional state; at step S1003, the server can obtain the user's counseling strategy; at step S1004, the server can use the conversation content generation model and the preset knowledge base to generate the second conversation content in response to the first conversation content based on the first conversation content, emotional state and the user's counseling strategy. It should be noted that the process of the server obtaining the user's emotional state and counseling strategy is combined with the above. Figure 5 and Figure 7 The contents described are the same and will not be repeated here.

[0080] Next, the server can execute step S1005 to send the second conversation content to the user's client. Correspondingly, the client can execute step S1006 to output the second conversation content in voice or text form to provide guidance to the user.

[0081] Combination of the above Figure 10 The present disclosure describes a method 1000 for tutoring implemented based on a two-terminal interactive system. It is understood that the description of various embodiments in this disclosure emphasizes the differences between the embodiments, and reference can be made to the similarities or correspondences between them. For the sake of brevity, this disclosure will not elaborate on each one.

[0082] Next, combine Figure 11 An electronic device 1100 provided in an embodiment of the present application is exemplarily introduced. Figure 11 As shown, the electronic device 1100 of the embodiment of the present application may include a processor 1101 , a memory 1102 and a communication bus 1103 .

[0083] In a specific embodiment, the processor 1101 may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor functions may also be other electronic devices, which is not specifically limited in this embodiment.

[0084] In the embodiment of the present application, the communication bus 1103 is used to realize the connection and communication between the processor 1101 and the memory 1102; the memory 1102 stores program instructions for tutoring; when the processor 1101 executes the program instructions stored in the memory 1102, the present application is realized. Figures 1 to 8 ,as well as Figure 10 The methods used for counseling are described.

[0085] Combination of the above Figure 11An electronic device for tutoring that can be used to perform the present application is described. It should be understood that the device structure or architecture here is merely exemplary, and the implementation method and implementation entity of the present application are not limited thereto, but can be changed without departing from the spirit of the present application. It is understood that the description of each embodiment in this disclosure emphasizes the differences between the various embodiments, and the same or corresponding parts can be referenced to each other. For the purpose of brevity, this disclosure will not go into details one by one.

[0086] According to the above description in conjunction with the accompanying drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by a software program. Therefore, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores program instructions for tutoring, and the program instructions can be used to implement the present application in conjunction with Figures 1 to 6 ,as well as Figure 10 The methods used for counseling are described.

[0087] It should be noted that although the operations of the present method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired results. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.

[0088] Although multiple embodiments of the present application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art can conceive of many changes, modifications, and alternatives without departing from the thought and spirit of the present application. It should be understood that in the process of practicing the present application, various alternatives to the embodiments of the present application described herein can be adopted. The accompanying claims are intended to define the scope of protection of the present application and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for tutoring, comprising: acquiring the user's emotional state upon receiving a first conversation content from the user, wherein the first conversation content includes the user's counseling needs; Based on the first conversation content and the emotional state, using a conversation content generation model to generate a second conversation content in response to the first conversation content, wherein the second conversation content includes text content and conversation tone; and The second conversation content is output to provide guidance to the user.

2. The method according to claim 1, wherein Acquiring the user's emotional state includes: Determining the input order of the first conversation content in the current conversation; determining whether the input order is a preset order; When the input order is a preset order, based on the first conversation content, using a sentiment analysis model to generate the user's emotional state in the current preset order; In a case where the input order is not a preset order, the emotional state of the user in the previous preset order is obtained.

3. The method according to claim 1, wherein, based on the first conversation content and the emotional state, using a conversation content generation model to generate a second conversation content in response to the first conversation content comprises: Based on the first dialogue content, the emotional state, and the user's coaching strategy, a dialogue content generation model and a preset knowledge base are used to generate second dialogue content in response to the first dialogue content.

4. The method according to claim 3, wherein: The coaching strategy is obtained by the following operations: Acquiring historical tutoring data of the user, wherein the historical tutoring data includes historical conversation content, historical teaching methods, and historical teaching methods of relevant knowledge points; generating a preference profile of the user using a preference analysis model based on the historical tutoring data, wherein the preference profile includes learning level, learning style, and language adaptability; Based on the preference features, a tutoring strategy generation model is used to generate a tutoring strategy for the user, wherein the tutoring strategy includes tutoring difficulty, tutoring style, and tutoring language.

5. The method according to claim 1, wherein Outputting the second conversation content includes: Based on the content of the copy and the tone of the conversation, using a speech generation model to generate conversational speech; The conversation voice is output.

6. The method according to any one of claims 1 or 5, further comprising: According to the expression output rules, a preset expression in the expression library is selected that matches the tone of the conversation; Output the text content and the preset expression.

7. The method according to claim 2, wherein: The first conversation content is in voice form; And based on the first conversation content, using a sentiment analysis model to generate the user's emotional state includes: performing noise reduction and standardization operations on the first conversation content to preprocess the first conversation content and obtain preprocessed first conversation content; Based on the preprocessed first conversation content, the sentiment analysis model is used to generate the user's emotional state.

8. The method according to claim 1, wherein The dialogue content generation model is obtained by debugging using preset prompt words; and the preset prompt words include a first prompt word and a second prompt word, the first prompt word includes a coaching role, a coaching task and a coaching step, and the second prompt word includes an emotional state and / or a coaching strategy.

9. An electronic device comprising: processor; and a memory storing program instructions for tutoring, wherein when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing program instructions for tutoring, wherein when the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.