Digital human student driving method based on large language model agent

Through the digital human-student-driven method based on large language model agents, the problem of insufficient teaching internship resources for traditional teacher students is solved, highly realistic simulated teaching is realized, teaching practice efficiency and interactivity are improved, and it is suitable for diverse teaching scenarios.

CN120375658APending Publication Date: 2025-07-25EAST CHINA NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510427194.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional teacher students’ teaching internships and simulated classrooms are limited by venue, time, resources and other conditions, and it is difficult to provide every teacher student with sufficient and diversified practical opportunities.

Method used

The digital human student-driven method based on the large language model agent is adopted to divide the teaching process into multiple stages, and the input and output of the stage agent is formatted and constrained, so that digital human students can simulate real classroom interactions and show different learning behaviors, emotional reactions and classroom performances.

Benefits of technology

Through a highly realistic, efficient and easy-to-use simulated teaching system, the teaching practice efficiency and interactivity of teacher students is improved, and diverse teaching needs are met, and it is suitable for teaching scenarios in different subjects and grades.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375658A_ABST
    Figure CN120375658A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human student driving method based on a large language model agent, and the method specifically comprises the steps: initializing digital human student configuration information, and loading a knowledge system and background setting; during teaching, voice of a teacher is converted into characters in real time, and accurate transmission of information is ensured; teaching content is accurately analyzed by means of an intention recognition module, and teaching intentions and stages are recognized, so that digital human students can respond in time; during question asking, the digital person student feeds back according to knowledge reserve and state, such as hand raising or doubt; through short-time memory updating and dynamic knowledge base adjustment, continuous learning adapts to the teaching progress, and diversified learning behaviors and emotional responses are shown. According to the digital human student driving method, the teaching process is divided into a plurality of stages, and input and output of stage intelligent agents are subjected to formatted constraint, so that digital human students can simulate real classroom interaction and show different learning behaviors, emotional responses and classroom performance, and the method is used for teaching simulation practice of normal teachers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a simulation classroom control technology based on a large language model, and in particular to a digital human student driving method based on a large language model intelligent agent. Background Art

[0002] With the transformation of the global education model and the rapid development of information technology, digital teaching tools and intelligent education systems have been widely used in education at all levels. Especially in the process of teacher training, the training of practical teaching ability is particularly important. Traditional teacher training teaching internships and simulation classrooms are usually limited by conditions such as venues, time, and resources, and it is difficult to provide each teacher training student with sufficient and diversified practical opportunities. Therefore, it is particularly important to build a simulation teaching system with digital human students as the core. The core research focus is to build a digital human student-driven method based on a large language model agent.

[0003] In recent years, with the rapid development of large language models (LLMs) and agent technology, the digital transformation of education has ushered in new opportunities. It is an effective technical solution to simulate student behavior and learning process in real classrooms through digital human student agents driven by large language models. Summary of the invention

[0004] The purpose of the present invention is to address the problem of insufficient practical teaching resources in the current normal student training, and propose a digital human student simulation driving method based on a large language model intelligent agent. This method divides the teaching process into multiple stages through the digital human student driving method, and formats the input and output of the stage intelligent agent, so that the digital human student can simulate real classroom interaction, show different learning behaviors, emotional reactions and classroom performance, and help normal students experience diverse teaching situations.

[0005] The specific technical solution for achieving the purpose of the present invention is:

[0006] A digital human student driving method based on a large language model agent comprises the following steps:

[0007] Step 1: Complete the initialization of the configuration information of the digital human student; configure the long-term memory module through the large language model system prompt, load the knowledge system, digital human student background and other information into the long-term memory module, and prepare the tools and expression action settings required for simulation in the tool library;

[0008] Step 2: The teacher starts teaching. The language information in the teaching process is converted into text information through voice recognition and enters the message queue in real time. The message queue caches and transmits the information in sequence.

[0009] Step 3: The message queue pushes the information to the intent recognition module. (This module is judged by the large language model combined with rules. Before the message is passed into the LLM, it first judges whether there are specific keyword matches for the intent in the message through rules, and then uses the large language model to judge the intent through the message semantics. Specifically, a Prompt containing a small number of labeled intent samples is constructed in a few-shot manner, enabling the LLM to better complete the teaching intent recognition task under the condition of a small number of samples.) This module uses the context learning ability of the large language model to analyze the teaching content and identify the teacher's teaching intent and teaching stage;

[0010] Step 4: If the intent recognition module determines that it enters the question-asking session, the question is input into each intelligent agent, namely the digital human student agent; the digital human student makes decision-making actions and expressions based on its own knowledge reserve, the set state of the digital human student, etc., including raising hands to answer and frowning to show confusion, etc., and passes the action and expression information generated by the large model according to the input into the downstream interface to drive the digital human student to make corresponding feedback;

[0011] Step 5: If the teacher calls on a certain digital human student, it enters the answering stage; every sentence of the teacher or the digital human student will enter the short-term memory, and the weight is changed according to the short-term memory decay function to reflect the timeliness and importance of different information;

[0012] Step 6: When the teacher asks the digital human student to sit down, a one-on-one conversation is completed; the historical conversation will be summarized and added to the short-term memory again, thereby updating the content of the short-term memory to make it more in line with the current teaching scenario and conversation process;

[0013] Step 7: In the learning process of the digital human student's combination of long-term and short-term memories, the knowledge base in the long-term memory module is dynamically added, deleted, modified, and queried in combination with the tools in the tool library, so as to generate diverse learning behaviors, emotional reactions, etc. of the digital human student as simulation learning data;

[0014] Step 8: Repeat Steps 3 to 7 until the teacher ends the teaching; where:

[0015] The learning process of the digital human student's combination of long-term and short-term memories includes:

[0016] a) Intent recognition. The digital human student uses intent recognition technology to analyze the teacher's speech text and judge whether the teacher's intent is to ask a question;

[0017] b) Knowledge point extraction and level determination. If it is determined to be a question, the digital human student uses the function call function of the model to extract the knowledge points in the text; the extracted knowledge points are judged at the level through a multi-classification model, and the knowledge points are compared with the knowledge boundary preset by the digital human student;

[0018] c) Knowledge boundary comparison. If the knowledge point level is higher than the knowledge boundary setting, the Retrieval-Augmented Generation (RAG) technology is used to retrieve the long-term memory knowledge base according to the content similarity; for short-term memory (class history), recall is performed based on the R score to ensure that recent information relevant to the current question is effectively called; if the knowledge point level does not exceed the knowledge boundary setting, the digital human student directly answers the question;

[0019] d) Knowledge base retrieval and reasoning. During the process of performing knowledge base retrieval operations, if relevant information is successfully retrieved, the digital human student will reason based on this information and then give an accurate answer to the question; if no relevant content is retrieved, the intelligent agent will truthfully feedback "don't know" to the teacher;

[0020] e) Learning and knowledge base update. After the teacher finishes explaining the question, the system will ask the digital human student about their mastery of the corresponding knowledge; if the digital human student confirms that they have learned, they will use the Function-call mechanism to update the knowledge base, aiming to incorporate new knowledge content or optimize the existing knowledge structure; if the intelligent agent indicates that they have not learned, no operation will be performed.

[0021] Furthermore, the short-term memory decay function is expressed as:

[0022] R time = W(t) = e -λt (1)

[0023] W(t) represents the weight of the conversation at time t, and λ is the decay constant, which determines the rate at which old conversations lose importance.

[0024] The recall based on the R score specifically includes: The memory correlation score is expressed as the cosine similarity of semantic vectors:

[0025]

[0026] R relevance represents the correlation or similarity between two vectors, calculated through cosine similarity; A and B represent the vectors of the text;

[0027] The recall of short-term memory during the question stage will take the top k according to the following scores from largest to smallest:

[0028] R = αR time + βR relevance (3)

[0029] R represents the final recall score, and α and β respectively represent the importance degree distribution of the time weight and similarity weight during short-term memory recall.

[0030] Furthermore, the knowledge system includes: subject knowledge and learning strategies and methods; the subject knowledge includes the basic content of different subjects, such as Chinese, mathematics, English, etc., so that the digital human students can make reasonable responses to the teaching of each subject; the learning strategies and methods include memory skills and problem-solving methods, etc., to help the digital human students better learn new knowledge.

[0031] Furthermore, the background of the digital human students includes: personal basic information, learning experience, hobbies and social relationships; the personal basic information includes name, age, gender, etc., to set an identity framework for them; the learning experience includes educational stage, grades and preferences, etc., which determine their classroom learning behaviors and responses; hobbies such as sports, music, etc., affect classroom participation and enthusiasm; social relationships cover relationships with classmates and teachers, which affect classroom interaction methods and emotional responses. These background information jointly shape the personalized characteristics of the digital human students.

[0032] Furthermore, the expression actions include: facial expressions, body movements and language expression actions; the facial expressions include smiling, frowning and surprise, etc., which reflect the emotional state and understanding of the teaching content; the body movements include raising hands, sitting down and nodding, etc., to enhance classroom expressiveness; the language expression actions involve intonation, speech rate and gestures when speaking, etc., to make the answer more natural.

[0033] The beneficial effects of the present invention are as follows:

[0034] The method of the present invention has a high degree of realism: through the combination of digital human technology and virtual reality, it can truly simulate various scenarios in the classroom, including digital human students of different ages, genders, learning abilities and learning styles. Each digital human student can make emotional responses and learning feedback according to the changes in the teaching content and environment. Normal school students need to dynamically adjust teaching strategies according to these feedbacks, thus greatly improving the interactivity of the system and the authenticity of the teaching scenario.

[0035] The method of the present invention has high efficiency: based on the classroom intention recognition module of the large language model and the intelligent agent-driven learning process simulation mechanism, it can quickly generate diverse learning behaviors and learning feedbacks, providing an instant and efficient teaching practice environment for normal school students, and significantly improving the efficiency of teaching practice.

[0036] The method of the present invention has ease of use: adopting a modular design, normal school students only need to perform simple operations to configure digital human students and teaching scenarios, without complex programming or manual intervention. Through computer-aided virtual teaching practice, it avoids the problems of insufficient teaching practice resources and inability to try out and iterate teaching strategies in traditional teaching internships, saving a large amount of human and time costs.

[0037] The method of the present invention has wide applicability: it supports simulations in different disciplines, different grades, and different teaching scenarios, and can meet the diverse teaching practice needs of normal students. At the same time, the behaviors and learning feedback of digital student can be dynamically adjusted according to teaching objectives and content, and are applicable to the cultivation of normal students at all levels of education. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of the present invention;

[0039] Figure 2 is an architecture diagram of the digital student of the present invention;

[0040] Figure 3 is a flowchart of the simulation learning process of the long-term and short-term memory combination technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In combination with the following specific embodiments and the accompanying drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no particularly restricted content.

[0042] Refer to the attached Figure 1 、 Figure 2 、 Figure 3 First, the present invention configures the persona information of the digital student, including knowledge background, learning style, and cognitive ability, and loads it into the long-term memory module. Then, the digital student Agent analyzes the teacher's teaching content in real time through the intention recognition ability of the large language model, identifies the teaching intention and teaching stage, and pushes the relevant information to the digital student intelligent agent. Finally, the digital student Agent combines the long-term and short-term memory mechanisms to dynamically manage the learning process: the short-term memory module captures the real-time conversation content and adjusts the information weight according to the memory decay function; the long-term memory module stores the knowledge system and historical learning trajectories, supporting the intelligent agent to generate diverse learning behaviors and emotional feedback. The specific operations are carried out according to the following steps:

[0043] 1) Dynamically set persona information

[0044] The present invention supports dynamic customization and configuration of the persona information of the digital student to meet the needs of diverse teaching scenarios. Teachers can flexibly set the personalized attributes of each digital student to ensure that their behaviors highly match those of real students.

[0045] Configure the long-term memory module through the system prompt of the large language model. The long-term memory module allows teachers to configure the Big Five personality traits of the digital human students, including openness, conscientiousness, extraversion, agreeableness, and neuroticism. For example, a student with high openness will show strong curiosity and creativity in class, while a student with high neuroticism may appear nervous and hesitant when answering questions. This personality configuration makes the behavior of digital human students more diverse and closer to real classroom situations.

[0046] In addition, teachers can dynamically adjust the learning ability and cognitive style of digital human students according to teaching needs. For example, set a student with strong memory but weak logical reasoning, so that they perform well in memorization tasks but encounter difficulties in mathematical reasoning. At the same time, the long-term memory module supports configuring the cognitive style of students, such as visual, auditory, or hands-on practical types, to simulate the learning behaviors of different types of students. In terms of basic information, the long-term memory module supports setting attributes such as the name, age, and gender of digital human students, and can generate diverse background stories according to needs, such as hobbies, family environment, learning experiences, etc. These information not only enhance the realism of digital human students but also provide richer teaching interaction materials for teachers.

[0047] To further enhance the realism of the classroom, the long-term memory module allows teachers to assign class seats to each digital human student and record their social relationships with surrounding classmates. For example, a student sitting in the back row may be more easily distracted or whisper to their neighbor, while a student sitting next to a friend may be more active in group discussions. This configuration of seat and social relationships makes the behavior of digital human students more in line with the actual classroom environment.

[0048] All configured persona information supports dynamic adjustment and real-time update. During the teaching process, teachers can modify their attributes in real time according to the performance of digital human students, such as adjusting learning ability or personality traits. The system will automatically update the relevant data and synchronize it to the long-term memory module to ensure that the behavior of digital human students is consistent with the latest configuration.

[0049] 2) Intention recognition and teaching stage control

[0050] The present invention realizes the precise analysis and stage control of the teacher's teaching content by combining the keyword fallback rule and the intention recognition method of the large language model. The intelligent agent can real-time recognize key intentions in the teaching process, such as "continue", "question", "answer", "sit down", etc., and dynamically adjust the behavior of digital human students according to different stages.

[0051] In the intention recognition module, the large language model performs in-depth semantic analysis on the teacher's speech or text input through its context learning ability to identify the teaching intention. For example, when the teacher asks "Who can answer this question?", the system will recognize that the current stage has entered the question-asking stage. At the same time, the keyword fallback rule, as an auxiliary mechanism, ensures that in cases where the model's recognition is uncertain or in complex contexts, the teaching stage can still be accurately judged through preset keywords (such as "question", "answer", "sit down").

[0052] In the question-asking stage, the digital human student will independently decide whether to raise their hand to answer the question based on their own learning ability, personality traits, and current knowledge mastery. For example, an extroverted and confident digital human student may quickly raise their hand, while an introverted or uncertain student may choose to remain silent. Through the agent decision-making mechanism, the diverse behaviors of real students are simulated.

[0053] When the teacher calls on a certain digital human student, the class enters the answering stage. The called student will stand up and generate an answer that matches their knowledge level and personality traits. During this process, the real-time conversation content will be incorporated into the short-term memory module for subsequent analysis and feedback. The answer is not only based on the knowledge reserve in long-term memory but also combines the dynamic information of the current teaching scenario to ensure the accuracy and situational relevance of the answer.

[0054] In the sitting-down stage, the teacher ends the conversation with the digital human student through an instruction (such as "Please sit down"). The class summary module (implemented by the large language model) will automatically summarize the conversation content and update it to the short-term memory module for reference in subsequent teaching. At the same time, the digital human student will adjust their learning state and emotional response according to the conversation result. For example, a student who answers correctly may show confidence, while a student who answers incorrectly may show confusion or a willingness to learn further.

[0055] By combining the keyword fallback rule and the large language model intention recognition, the present invention can accurately control the teaching stage and dynamically adjust the behavior of the digital human student, providing a highly realistic teaching practice environment for normal students.

[0056] 3) Agent Student Behavior Design

[0057] In order to highly restore the behaviors of students in a real classroom, the digital human student agent of the present invention is equipped with a rich library of expressions and actions, capable of dynamically responding to changes in the teaching scenario, enhancing the authenticity and engagement of teaching interactions. These expressions and actions not only enhance the expressiveness of the agent but also provide diverse teaching feedback for normal students, helping them better understand and respond to the behaviors of students in the classroom.

[0058] The expression library of the digital human student covers a variety of common classroom expressions to reflect their emotional states and cognitive responses. For example, a smile indicates happiness, agreement, or good interaction with the teacher; calmness represents concentration or a neutral state; confusion reflects a lack of understanding of the current knowledge point or learning difficulties; surprise indicates being unexpected about new information or having a strong interest; and dullness implies a distracted attention or a trance state. These expressions can vividly display the mental activities of the digital human student in the classroom and provide intuitive teaching feedback for normal school students.

[0059] The action library includes common body actions in the classroom, further enhancing the expressiveness of the digital human student. For example, standing up indicates answering a question or needing to leave the seat; sitting down represents finishing answering or refocusing; raising a hand indicates the hope to speak or answer a question; putting the hand down represents giving up speaking or waiting for an opportunity; nodding indicates agreement, understanding, or a positive response to the teacher's question; shaking the head indicates disagreement, lack of understanding, or the need for further explanation. In addition, yawning reflects fatigue or the lack of attractiveness of the classroom content; speaking simulates a conversation with the teacher or other students; sitting upright shows a good sitting posture and a focused state; dozing indicates tiredness or lack of interest in the classroom content; being absent-minded shows a lack of concentration or being attracted by the outside world; stretching reflects fatigue or the need for relaxation after maintaining a posture for a long time. These actions combined with expressions make the behavior of the digital human student more close to the real classroom situation.

[0060] The expressions and actions of the digital human student are not randomly triggered, but are dynamically generated based on the character settings and classroom situations. By analyzing the background settings of the intelligent agent and the real-time classroom situation through the large language model, the corresponding expressions and actions are intelligently triggered to ensure that the behavior is logical and close to reality. For example, for a lively and cheerful student, the intelligent agent is more inclined to show positive behaviors such as smiling and raising a hand; while for an introverted and quiet student, the system may more often trigger expressions such as confusion, calmness, or slight head shaking to show their introverted character. For a student with weaker learning ability, the system will trigger more expressions of confusion, dullness, or absent-mindedness to reflect their learning difficulties.

[0061] In the classroom situation, the behavior of the digital human student will be dynamically adjusted according to the teaching stage. For example, in the question-asking session, the digital human student will dynamically choose whether to raise a hand or show confusion according to their own knowledge mastery and personality characteristics; in the answering session, the student called on will stand up and generate an answer that matches their knowledge level, and at the same time combine expressions (such as a confident smile or a nervous frown) and actions (such as nodding or shaking the head) to enhance the expressiveness; during a long lecture, the digital human student may trigger actions such as yawning, stretching, or dozing, reflecting classroom fatigue or a decline in attention.

[0062] 4) Intelligent agent learning process based on long short-term memory

[0063] Reference Figure 3, the learning process of the Long Short-Term Memory (LSTM)-based agent can be described as the following steps:

[0064] Step 1: Intent recognition. The digital human student agent uses intent recognition technology to analyze the teacher's speech text and determine whether the teacher's intent is a question.

[0065] Step 2: Knowledge point extraction and hierarchy determination. If it is determined to be a question, the agent uses the function-call function of the model to extract the knowledge points in the text. The extracted knowledge points are hierarchically determined through a multi-classification model (this model is a large language model fine-tuned with knowledge hierarchy label data, with a piece of knowledge as the input and the level where the knowledge belongs as the output), and the knowledge points are compared with the knowledge boundaries preset by the agent.

[0066] Step 3: Knowledge boundary comparison. If the knowledge point level is higher than the set knowledge boundary, the Retrieval-Augmented Generation (RAG) technology is used to retrieve the long-term memory knowledge base according to the content similarity. For short-term memory (class history), recall is based on the R score to ensure that recent information relevant to the current question is effectively called; if the knowledge point level does not exceed the set knowledge boundary, the agent directly answers the question.

[0067] Step 4: Knowledge base retrieval and reasoning. During the knowledge base retrieval operation, if relevant information is successfully retrieved, the agent will reason based on this information and then give an accurate answer to the question; if no relevant content is retrieved, the agent will truthfully feedback "don't know" to the teacher.

[0068] Step 5: Learning and knowledge base update. After the teacher finishes explaining the question, the system will ask the agent about its mastery of the relevant knowledge. If the agent confirms that it has learned, it will use the Function-call mechanism to update the knowledge base to incorporate new knowledge content or optimize the existing knowledge structure; if the agent indicates that it has not learned, no operation will be performed.

[0069] The whole process closely revolves around the goals of intelligent interaction and efficient learning, fully demonstrating the intelligence and flexibility of the agent in the knowledge processing and learning process.

[0070] Among them, the weight decay function of short-term memory can be expressed as:

[0071] R time = W(t) = e λt (1)

[0072] W(t) represents the weight of the conversation at time t, λ is the decay constant, and its value range is from 0.1 to 0.5, with a default value of 0.3, which determines the rate at which old conversations lose importance.

[0073] The semantic vector uses an advanced text embedding model (such as OpenAI's text-embedding-ada-002), which is specially designed to generate high-quality text vector representations. It can effectively capture the semantic features of text and is suitable for tasks such as search, classification, and clustering. Through text embedding technology, the text data in the knowledge base can be converted into vector form. Subsequently, the knowledge retrieval module (which is based on the fasis vector database retrieval function) can semantically match and similarity search the embedding of the question text with the knowledge in the knowledge base, thereby improving retrieval efficiency and response accuracy. The memory relevance score can be expressed as the cosine similarity of the semantic vector:

[0074]

[0075] R relevance Represents the correlation or similarity between two vectors, calculated by cosine similarity. A and B represent the vectors of the text.

[0076] The final short-term memory recall in the questioning phase will be taken from the top k according to the following scores from largest to smallest:

[0077] R = αR time +βR relevance (3)

[0078] R represents the final recall score, α and β represent the importance distribution of time weight and similarity weight in short-term memory recall, and in the present invention, they are respectively taken as 0.5.

[0079] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A digital human student-driven method based on large language model agents, characterized in that, It includes the following steps: Step 1: Initialize the configuration information of the digital human student; configure the long-term memory module through the system prompt of the large language model. The long-term memory module loads the knowledge system and the background information of the digital human student, and the tool library prepares the tools required for simulation and sets the facial expressions and actions; Step 2: The teacher starts teaching. The language information during the teaching process is converted into text information in real time through speech recognition and enters the message queue. The message queue caches and transmits the information in order; Step 3: The message queue pushes the information to the intention recognition module. This module uses the context learning ability of the large language model to analyze the teaching content and identify the teacher's teaching intention and teaching stage; Step 4: If the intention recognition module determines that it enters the question-asking session, the question is input into each intelligent agent, that is, the digital human student; The digital human student makes decision-making actions and expressions based on its own knowledge reserve and the set state of the digital human student, including raising a hand to answer and frowning to show confusion, and transmits the action and expression information generated by the large language model according to the input to the downstream interface to drive the digital human student to make corresponding feedback; Step 6: If the teacher calls on a certain digital human student, it enters the answering stage; each sentence of the teacher or the digital human student will enter the short-term memory, and the weight is changed according to the short-term memory decay function to reflect the timeliness and importance of different information; Step 7: When the teacher asks the digital human student to sit down, a one-on-one conversation is completed; the historical conversation will be summarized and added to the short-term memory again, so as to update the content of the short-term memory and make it more in line with the current teaching scenario and conversation process; Step 8: During the learning process of the combination of the long-term and short-term memories of the digital human student, the knowledge base in the long-term memory module is dynamically added, deleted, modified, and queried in combination with the tools in the tool library, so as to generate diverse learning behaviors, emotional reaction simulation learning data of the digital human student; Step 9: Repeat steps 3 to 7 until the teacher ends the teaching; among them: The learning process of the combination of the long-term and short-term memories of the digital human student includes: a) Intention recognition. The digital human student uses intention recognition technology to analyze the teacher's speech text and judge whether the teacher's intention is to ask a question; b) Knowledge point extraction and hierarchy determination. If it is determined to be a question, the digital human student uses the function call (Function-call) function of the model to extract the knowledge points in the text; the extracted knowledge points are hierarchically determined through a multi-classification model, and the knowledge points are compared with the knowledge boundary preset by the digital human student; c) Knowledge boundary comparison. If the knowledge point level is higher than the knowledge boundary setting, the retrieval-enhanced generation technology is used to retrieve the knowledge base of the long-term memory according to the content similarity; for the short-term memory, that is, the classroom history, recall is performed according to the R score to ensure that the recent information related to the current question is effectively called; if the knowledge point level does not exceed the knowledge boundary setting, the digital human student directly answers the question; d) Knowledge base retrieval and reasoning. During the process of performing knowledge base retrieval operations, if relevant information is successfully retrieved, the digital human student will conduct reasoning based on this information and then give accurate answers to the questions; if no relevant content is retrieved, the digital human student will truthfully feedback "don't know" to the teacher. e) Learning and knowledge base update. After the teacher finishes explaining the question, the system will ask the digital human student about their mastery of the corresponding knowledge. If the digital human student confirms that they have learned, they will use the Function-call mechanism to update the knowledge base, aiming to incorporate new knowledge content or optimize the existing knowledge structure; if the agent indicates that they have not learned, no operation will be performed.

2. The digital human student driving method according to claim 1, wherein The short-term memory decay function is expressed as: R time = W(t) = e -λt (1) W(t) represents the weight of the conversation at time t, and λ is the decay constant, which determines the rate at which old conversations lose importance. The recall based on the R score specifically includes: The memory correlation score is expressed as the cosine similarity of semantic vectors. R relevance represents the correlation or similarity between two vectors, calculated by cosine similarity; A and B represent the vectors of the text; The recall in the short-term memory during the question stage will select the top k from largest to smallest according to the following scores: R = αR time + βR relevance (3) R represents the final recall score, and α and β respectively represent the importance degree distribution of the time weight and similarity weight during short-term memory recall.

3. The digital human student driving method according to claim 1, characterized in that, The knowledge system includes: subject knowledge and learning strategies and methods. Subject knowledge includes the basic content of different subjects so that the digital human student can make reasonable responses to various subject teachings; learning strategies and methods include memory skills and problem-solving methods to help the digital human student better learn new knowledge.

4. The digital human student driving method according to claim 1, characterized in that The background of the digital human student includes: personal basic information, learning experience, hobbies, and social relationships. Personal basic information includes name, age, and gender, which sets an identity framework for them; learning experience includes educational stage, grades, and preferences, etc., which determine their classroom learning behaviors and responses; hobbies such as sports, music, etc., affect classroom participation and enthusiasm; social relationships cover relationships with classmates and teachers, which affect classroom interaction methods and emotional responses. These background information together shape the personalized characteristics of the digital human student.

5. The digital human student driving method according to claim 1, wherein The expression actions include: facial expressions, body movements, and language expression actions. Facial expressions include smiling, frowning, and being surprised, which reflect the emotional state and understanding of the teaching content; body movements include raising hands, sitting down, and nodding, which enhance classroom expressiveness; language expression actions involve intonation, speech rate, and gestures during speaking, making the answers more natural.

Citation Information

Cited By

  • Intelligent analysis system for students to self-adjust learning strategies

    CN121503932A

  • Classroom simulation method and device, storage medium and electronic equipment

    CN122472082A