An individualized companion robot system based on dynamic student portrait and AI large model
The personalized learning robot system, which combines dynamic student profiles with AI big data models, achieves multimodal perception, deep cognitive understanding, and personalized decision-making. It solves the problems of dynamic understanding and adaptive learning paths in educational robot systems, thereby enhancing learners' participation and interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-16
AI Technical Summary
Existing educational robot systems lack a deep understanding of learners' dynamic changes, have a single perception dimension, and their teaching decisions and content generation are static. They cannot achieve personalized and adaptive learning path adjustments, and their human-computer interaction is rigid, making it difficult to establish a long-term learning companion relationship.
A personalized learning companion robot system based on dynamic student profiles and AI big data models is adopted. The system collects data in real time through a multimodal perception module, generates personalized teaching content by combining it with an AI brain, optimizes the learning path by using reinforcement learning, and conducts human-like interaction on the robot end to establish an emotional bond.
It achieves a deep understanding of learners' cognitive state and emotional changes, provides teaching content and interaction strategies that are highly tailored to individual characteristics, enhances immersion, trust and sustained participation in the learning process, while protecting user privacy and enriching the supply of teaching resources.
Smart Images

Figure CN122221901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of educational robots and artificial intelligence technology, and relates to a personalized learning companion robot system based on dynamic student profiles and AI large models. Background Technology
[0002] With the development of artificial intelligence and robotics, educational robots are evolving from teaching aids that execute fixed programs to auxiliary tools with certain interactive capabilities. Currently, educational robot technology mainly falls into several categories: 1) Programming educational robots, such as LEGO Mindstorms and Makeblock, primarily used for programming skills development, but lacking emotional interaction and personalized teaching capabilities. 2) Language learning robots, such as early childhood education robots, providing standardized language teaching content, but lacking a deep understanding of individual learner differences. 3) Intelligent educational assistants, mainly educational software based on tablets or computers, using simple algorithms to provide practice recommendations, lacking physical interaction and emotional connection. 4) Humanoid social robots, such as Pepper and Nao, possessing certain human-computer interaction capabilities, but insufficient deep application in the education field, lacking systematic teaching planning and personalized support.
[0003] Existing educational robots suffer from the following limitations: First, most systems rely on pre-programmed scripts or simple rules for interaction, lacking a deep understanding and modeling ability of learners' dynamically changing cognitive states (such as shifts in interest, skill development, and emotional fluctuations). Second, their perceptual dimensions are limited, often relying solely on voice or simple touch controls, failing to effectively integrate modalities rich in cognitive and emotional information, such as facial expressions and body language. Third, instructional decisions and content generation are relatively static, unable to make real-time, adaptive path adjustments and interventions based on a continuous understanding of learners. Finally, the human-computer interaction is rigid, making it difficult to establish a long-term, emotionally connected relationship, thus affecting learners' continued engagement and trust.
[0004] In recent years, large language models have made breakthroughs in understanding and generating natural language, multimodal perception and fusion technologies have matured, and reinforcement learning has shown great potential in sequence decision-making problems, providing new technological possibilities for humanoid robots to achieve deeper personalized and intelligent learning companionship. Furthermore, technologies such as computer vision, speech emotion recognition, and physiological signal detection are maturing, reinforcement learning has shown great potential in personalized path planning, and advancements in robot hardware technologies such as servo motors, sensors, and batteries have made low-cost, high-performance humanoid robots possible.
[0005] Therefore, there is an urgent need for a learning companion robot that can systematically solve the above problems and realize a closed loop from multi-dimensional perception and dynamic cognitive understanding to personalized intelligent intervention. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a personalized learning companion robot system based on dynamic student profiles and AI big data models, to establish an intelligent learning companion robot system that integrates "end, cloud, and human", to continuously track students' interests and abilities through dynamic student profiles, to generate personalized learning content by combining AI big data models, and to dynamically optimize the learning path by using reinforcement learning, so as to provide differentiated learning companion roles for students of different age groups, and to realize the transformation from "teaching implementer" to "growth partner".
[0007] To achieve the above objectives, the present invention provides the following technical solution: A personalized learning companion robot system based on dynamic student profiles and AI big data models. The system adopts an "end-cloud-human" collaborative architecture, including a robot end, a cloud end, and a user terminal. The three ends achieve data synchronization and collaborative interaction through data communication.
[0008] The robot terminal includes a multimodal perception module and a dynamic student modeling engine; the multimodal perception module is used to collect multimodal time-series data during student interaction in real time; the dynamic student modeling engine responds to the multimodal time-series data and updates and generates a structured student profile in real time. The cloud platform includes an AI brain and an adaptive curriculum planning module. The AI brain generates personalized teaching content based on real-time updated structured student profiles, combined with educational knowledge graphs and large language models. The adaptive curriculum planning module dynamically generates and adjusts personalized learning paths based on real-time updated structured student profiles using a proximal strategy optimization algorithm. The robot also includes an emotional interaction engine and an actuator; the emotional interaction engine plans the robot's anthropomorphic empathic interaction parameters based on the real-time updated structured student profiles and personalized teaching content; the actuator is used to execute the interaction parameters.
[0009] The beneficial effects of this invention are as follows: (1) By constructing a continuously updated dynamic student profile, the system can deeply understand the learner's cognitive state, interests and emotional changes, thereby providing teaching content and interaction strategies that are highly tailored to individual characteristics, overcoming the problems of rigid teaching content and lack of pertinence in traditional educational robots.
[0010] (2) By comprehensively utilizing multimodal perception technology and integrating multi-dimensional information such as speech, facial expressions, text and physiological signals, a comprehensive assessment of learners’ cognitive and emotional states can be achieved, breaking the limitations of traditional systems with single perception dimensions and one-sided understanding.
[0011] (3) By using advanced algorithms such as reinforcement learning, this invention can dynamically plan and adjust the learning path based on real-time feedback, making teaching intervention flexible and forward-looking, and solving the problem that existing educational robots make mechanical teaching decisions and cannot be dynamically optimized as learners grow.
[0012] (4) Through layered emotional interaction strategy and multimodal expression planning, the present invention can respond with empathy in an anthropomorphic way, establish a long-term learning relationship with stronger emotional ties, and significantly improve the sense of immersion, trust and continuous participation in the learning process.
[0013] (5) This invention uses multiple privacy protection technologies such as data localization processing, hierarchical encryption, and federated learning to achieve personalized services while ensuring the security of sensitive student data throughout the entire life cycle, protecting user privacy and complying with relevant regulations.
[0014] (6) Through the cloud ecosystem platform, teachers, developers and other parties are encouraged to participate in the creation and sharing of high-quality teaching resources, and the intelligent recommendation mechanism is used for precise distribution, which greatly enriches the supply of teaching resources and promotes the sharing and accumulation of educational wisdom.
[0015] (7) This invention realizes a closed loop of the whole process from multi-dimensional perception, deep cognitive understanding, personalized decision-making to human-like interaction, organically combining advanced artificial intelligence technology with robot hardware, and ultimately providing users with a truly intelligent, personalized and emotional long-term learning partner.
[0016] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of a personalized learning companion robot system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the AI brain architecture; Figure 3 This is a schematic diagram of the adaptive curriculum planning module architecture; Figure 4 This is a schematic diagram of the ecosystem platform architecture; Figure 5 A schematic diagram of the dynamic student modeling engine architecture; Figure 6This is a schematic diagram of the emotional interaction engine architecture; Figure 7 This is a schematic diagram of a data security architecture; Figure 8 A schematic diagram of the federated learning architecture; Figure 9 This is a specific application example of the personalized learning companion robot system proposed in this embodiment. Detailed Implementation
[0018] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0019] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0020] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0021] like Figure 1 As shown, one embodiment of the present invention provides a personalized learning companion robot system based on dynamic student profiles and AI large models. The system adopts a collaborative architecture of "end (perception, understanding) - cloud (optimization, decision-making) - human (interaction, supervision)", including robot end, cloud and user terminal. The three ends realize data synchronization and collaborative interaction through communication modules.
[0022] The cloud-based system includes an AI brain, an adaptive curriculum planning module, a safety and ethics control module, an ecosystem platform, and a safety and ethics control module. This enables the optimization of the learning support program, outputs curriculum planning and learning guidance, and monitors the ethical compliance of interactive content and behavior in real time through the safety and ethics control module.
[0023] The robot's end includes a dynamic student modeling engine, an emotion interaction engine, an actuator, a multimodal perception module, a local processor, and a communication module, enabling low-latency emotion computing and handling real-time interaction, lightweight perception, and local response.
[0024] The user terminals include terminals for parents and teachers, used for supervision, collaboration, and resource uploading.
[0025] 1. In the cloud, the AI brain, based on retrieval-enhanced generation technology and combined with educational knowledge graphs and large language models, achieves intent recognition, knowledge retrieval, content generation, educational filtering, and multimodal output planning, ensuring accurate responses and compliance with teaching principles. It supports interest customization and content generation, as well as learning games and intelligent problem-solving. Interest customization and content generation mainly include: constructing student interest graphs through multimodal data and updating interest intensity and relationships in real time; generating learning tasks, stories, projects, and dialogues highly relevant to students' interests based on interest and knowledge graphs; and supporting interdisciplinary project-based learning, such as "Mars Base Design" and "Ancient Civilization Exploration," to stimulate students' inquiry motivation.
[0026] like Figure 2 As shown, the AI brain includes an input layer, an intent recognition module, a knowledge retrieval module, a content generation module, an educational filtering and optimization module, and an output layer, which are connected in sequence. In addition, it also includes an API service layer.
[0027] (1) Input layer input content includes student queries Current situation , student portraits Among them, student queries are students' questions or requests, the current context is information such as the dialogue context and learning environment, and the student profile is a structured student profile obtained from the dynamic student modeling engine.
[0028] (2) The intent recognition module analyzes the query purpose through the intent classifier, identifies the intent as types such as knowledge question answering, concept explanation, task generation, and practice recommendation, and determines the path of subsequent knowledge retrieval.
[0029]
[0030] in, The result of the intent type determination. For intent classifiers.
[0031] (3) In the knowledge retrieval module, for input knowledge-based queries, relevant concepts and relationships are retrieved through a knowledge graph; for task-based intentions, content suitable for students' levels and interests is retrieved through a course database. The RAG (Retrieval Enhanced Generation) model is adopted to provide relevant knowledge fragments for LLM.
[0032]
[0033]
[0034]
[0035]
[0036] in, For the search results, For knowledge graph retrieval, It is a course database retriever.
[0037] The retrieval enhancement generative model is shown in the following equation:
[0038] in, For user input, To generate a response, For the relevant knowledge fragments retrieved, This is a collection of search results from the knowledge base.
[0039] (4) In the content generation module, the prompt word builder integrates the query, context, profile and search results to build the prompt; then the large language model generates the initial response content.
[0040]
[0041]
[0042] in, For prompt word constructor, This is the initial response content. For large language models.
[0043] (5) Educational filtering and optimization module, including educational filters, used to filter inappropriate or non-educational content; also includes a teaching optimizer, used to adjust teaching strategies based on student profiles, such as adjusting language complexity, adding or reducing scaffolding support, and incorporating personalized examples.
[0044]
[0045]
[0046] in, As an educational filter, This is filtered teaching content. As a teaching optimizer, The optimized teaching content.
[0047] (6) The multimodal output planning module includes: Facial Expression Planner: Plans the changes in facial expressions of the virtual teacher; Motion planner: plans body language and gestures; Speech parameter planner: Adjusts speech features such as intonation, speech rate, and pauses.
[0048] In the multimodal output planning module, if the output content requires facial expressions, the facial expression sequence is generated by the facial expression planner:
[0049] in, For facial expression planner, This is a sequence of facial expressions.
[0050] If the output requires actions, the action sequence is generated by the action planner:
[0051] in, For action planner, This is an action sequence.
[0052] (7) API service layer The main function of the API service layer is to provide standardized, callable interfaces for communication between modules within the system and between the system and external applications, thereby achieving functional decoupling, service reuse, and scalability. It is a component included in several core modules (such as the adaptive curriculum planning module, the dynamic student modeling engine, and the ecosystem platform).
[0053] Through the modules described above, the AI brain can possess the following characteristics: ① Personalized adaptation: All processing is based on student profiles to ensure that the content is suitable for the student's current level; ② Multimodal output: Not only generates text, but also plans facial expressions and actions to enhance teaching effectiveness; ③ Educational Principle: Built-in educational filters and teaching optimization ensure educational quality; ④ Scalable architecture: Modular design facilitates the replacement or upgrading of individual components; ⑤RAG Enhancement: Combines search results with large language models to improve accuracy and reliability.
[0054] 2. In the cloud, the adaptive curriculum planning module employs a near-end strategy optimization algorithm, defining a state space (student profile, learning history, current progress), an action space (selecting a topic, selecting a task type, adjusting difficulty, selecting an interaction method), and a reward function (task completion, emotional feedback, ability growth, parent satisfaction). This enables the dynamic generation and adjustment of personalized learning paths, such as... Figure 3 As shown.
[0055] (1) The near-end strategy optimization algorithm is as follows: Initialize policy network and value network ; for episode = 1 to M: Collection Track ; Calculate advantage estimation ; Update strategy: ; Update the value function:
[0056] PPO objective function:
[0057]
[0058] in, Â t This is the estimated value of the dominance function. This is the trimming parameter (usually 0.1~0.2). For a moment The student state vector, For a moment Teaching actions, For instant rewards, For policy network parameters, For value network parameters, Accumulate rewards for discounts. For the empirical expectation of the sampling trajectory, This is the clipping function, used to limit the range of values for the probability ratio.
[0059] (2) In the state space, the student profile is a cognitive profile encoding obtained from the student modeling engine. The learning history includes past learning records, grades, and interaction history. The current progress includes the current learning stage, completed tasks, and tasks to be completed. In the action space, the theme is a suitable learning theme selected from the knowledge graph. The task types include exercises, projects, experiments, discussions, etc. The difficulty refers to the complexity of the task (easy, medium, challenging). The interaction methods include explanation, Q&A, gamification, collaboration, etc.
[0060] (3) The reward function is expressed as:
[0061] in, The task completion rate is calculated based on task completion time, accuracy, and completion quality. This provides positive emotional feedback by analyzing emotional changes during the learning process, such as enthusiasm, participation, and interest. The ability growth rate is obtained by comparing the skill mastery rate before and after the test. Parent satisfaction was obtained through a feedback survey based on parents' subjective evaluations. , , , All are weights.
[0062] (4) In the adaptive curriculum planning module, the reinforcement learning engine is as follows: State encoder: Used to encode multidimensional state information into RL state vectors and perform feature normalization and standardization; Policy Network Actor network based on PPO algorithm, with input being the RL state vector processed by the state encoder. Output the probability distribution of each action in the action space. The objective function is:
[0063] Value Network As a Critic network, it evaluates the state value and outputs a state value estimate. Its loss function is:
[0064] Advantage estimator: based on value network output and actual reward sequence The advantage function is calculated using the generalized advantage estimation method. This is used to update the weights in the policy network. The formula is:
[0065] in, , γ As a discount factor, λ For GAE parameters.
[0066] PPO optimizer: Executes the proximal policy optimization algorithm and controls the magnitude of policy updates (parameters). ), balancing exploration and utilization.
[0067] (5) Learning path generator, including time planner and path converter.
[0068] Time planners include: Short-term planner: A learning plan of approximately 10 steps; Mid-term planner: A 30-step learning plan; Long-term planner: A 100-step learning plan; The path converter transforms RL action sequences into specific learning plans by taking into account resource availability and time constraints.
[0069] (6) Feedback processing module Real-time feedback reception: Receive feedback from multiple sources including students, teachers, and parents; Feedback reward calculation: converting qualitative feedback into quantitative reward signals; Path adjuster: Adjusts the generated learning path in real time based on feedback.
[0070] (7) External interface for connecting to student learning terminals, teacher monitoring panels and parent feedback systems.
[0071] The student learning terminal receives the learning plan and provides a learning experience; the teacher monitoring panel can view student progress and provide teaching feedback; and the parent feedback system allows parents to understand their child's progress and provide satisfaction feedback.
[0072] (8) Output layer: Output personalized learning paths, i.e. structured learning plans, including topic sequences (learning roadmaps), task arrangements (specific learning activities), difficulty levels (progressive challenges), and interaction plans (teaching interaction methods).
[0073] The workflow of the adaptive curriculum planning module is as follows: ① Initialization phase: State information → State encoder → RL state vector ② Planning Phase: RL State Vector → Policy Network → Action Selection → Learning Path Generation ③ Implementation Phase: Learning Path → Student Learning → Collection of Feedback Data ④ Evaluation Phase: Feedback Data → Reward Calculation → Advantage Estimation → PPO Update ⑤ Adjustment Phase: New Feedback → Feedback Processing → Path Adjustment → Replanning In addition, the adaptive curriculum planning module also integrates a curriculum database and an API service layer.
[0074] The course database includes: Theme bank: Learning themes organized by subject and grade level; Task library: various types of teaching tasks and activities; Difficulty gradient: Task variations of varying difficulty; Interactive Mode Library: Diverse teaching interaction methods.
[0075] The API service layer includes: generate_learning_path: Generates a personalized learning path based on student profiles and time ranges; update_policy: Updates the RL policy based on the interaction trajectory; adjust_path: Adjusts the learning path based on real-time feedback.
[0076] Through the modules described above, the adaptive curriculum planning module can have the following characteristics: Personalized adaptability: The learning path is dynamically adjusted based on the student's real-time status; Multi-objective optimization: Balancing multiple objectives such as task completion, emotional experience, and skill development; Security Exploration: The clip mechanism of the PPO algorithm prevents policy mutations; Real-time adjustment: Supports real-time adjustments to the learning plan based on feedback; Explainability: The RL decision-making process is traceable, making it easier for teachers to understand the system logic; Multi-source feedback integration: Integrating feedback from students, teachers, and parents.
[0077] 3. In the cloud, the ecosystem platform enables content uploading, review, metadata generation, and personalized recommendations (collaborative filtering + content filtering + context adaptation) through a lesson plan store. This supports collaborative participation in ecosystem development by teachers, developers, and parents, such as... Figure 4 As shown.
[0078] The ecosystem platform includes a lesson plan store API, a lesson plan upload layer, a lesson plan recommendation layer, a storage layer, an input layer, and external interfaces.
[0079] (1) The lesson plan store API includes: LessonPlaneAPI: The core interface class for the lesson plan store; upload_lesson_plan: Handles the function of uploading lesson plans; recommend_plans: Handles personalized lesson plan recommendation functionality.
[0080] (2) The lesson plan upload layer includes: Content review checks whether uploaded lesson plans comply with platform standards, filtering out inappropriate, illegal, or low-quality content. A ContentReviewError exception is thrown if the review fails.
[0081] Educational standards alignment check: This checks whether lesson plans are aligned with relevant educational standards, such as national curriculum standards and core competencies for specific subjects, and generates alignment results for use in metadata and subsequent recommendations.
[0082] Metadata generation: Based on the lesson plan content and review results, structured metadata is generated, including author information (the identity information of the uploading teacher), upload time (time stamp of the lesson plan upload), target age (the age range of students suitable for the lesson plan), subject classification (the subject or field to which the lesson plan belongs), estimated duration (the time required to complete the lesson plan), difficulty level (the difficulty level of the lesson plan), standard alignment (the alignment with educational standards), prerequisites (the knowledge required to learn this lesson plan), and learning objectives (the teaching objectives of the lesson plan).
[0083] The lesson plan storage function stores the lesson plan content and metadata in the database and returns a unique lesson plan ID (plan_id).
[0084] Returns the upload result, which includes the lesson plan ID, status, and metadata.
[0085] (3) The lesson plan recommendation layer includes: Collaborative filtering recommendation is based on the behavior of similar users. Given a student ID, it finds 10 similar students with similar interests and abilities. It analyzes the ratings and usage of lesson plans by similar students and outputs a recommendation list based on group preferences.
[0086] Content-based filtering recommendation is based on matching lesson plan content with student characteristics. It takes students' interest graphs and ability spectrums as input, matches student characteristics with lesson plan metadata, analyzes the subject, difficulty, and objectives of the lesson plan, and outputs a recommendation list based on content matching.
[0087] Hybrid recommendation combines the results of collaborative filtering and content-based filtering. The weight of the collaborative filtering result is set to 0.4, and the weight of the content-based filtering result is set to 0.6. The hybrid recommendation result is obtained by weighted averaging of the two results. The weights of the collaborative filtering and content-based filtering results can be adjusted based on actual performance.
[0088] Context-adapted filtering adjusts the recommendation results based on the current learning context, taking into account the learning time of day (morning / afternoon / evening), the student's current available learning time, and the learning location and equipment conditions, and outputs context-adapted personalized recommendations.
[0089] The Top 5 Recommendations selects the top 5 best lesson plans from the final recommendation list and returns them to the students' learning interface.
[0090] (4) Storage layer, used to store the specific content and metadata of the lesson plan, user rating data of the lesson plan, and user usage history and behavior patterns.
[0091] (5) The content received by the input layer includes: Teacher-uploaded lesson plan data: the original lesson plan content submitted by the teacher; Student profile data: Profile information obtained from the dynamic student modeling engine; Learning context data: Information about the current learning environment and context.
[0092] (6) External interfaces include Teacher upload interface: The interface for teachers to upload lesson plans; Student learning interface: The interface through which students view and use recommended lesson plans; Administrator review interface: The interface for administrators to review lesson plan content.
[0093] 4. In the robot's interface, the dynamic student modeling engine uses multimodal temporal data input. Through feature extraction, temporal modeling (LSTM / Transformer), and multi-task learning (interest graph update, ability spectrum tracking, emotion recognition, learning style classification, metacognitive assessment), it generates structured student profiles, supporting real-time updates and historical tracing, such as... Figure 5 As shown.
[0094] The dynamic student modeling engine includes an input layer, a feature extraction and fusion module, a temporal modeling module, a multi-task learning module, a multi-task learning branch, a profile integration layer, a storage layer, an API service layer, and external interfaces.
[0095] (1) The input layer receives four types of data sources, including speech features (tone, speech rate, emotion, etc.), facial expression features (attention, emotional response), text / interaction features (answer records, interactive behavior), and physiological state features (heart rate, skin conductance response, etc.).
[0096] (2) The core processing engine consists of the feature extraction and fusion module, the temporal modeling module, and the multi-task learning module. The feature extraction and fusion module uses the Embedding layer and FusionLayer to fuse multimodal features, the temporal modeling module uses LSTM / Transformer to process time series dependencies, and the multi-task learning module processes five cognitive dimensions in parallel.
[0097] In the dynamic student modeling engine, the temporal modeling module is the core intermediate layer connecting "feature extraction" and "multi-task learning," and its function is as follows: Capturing dynamic changes: Transforming discrete feature points into continuous state evolution trajectories, enabling the system to "understand" how students transition from "confusion" to "understanding".
[0098] Provide context: When performing emotion recognition, the multi-task learning module can not only rely on the current expression, but also combine the learning state of the previous few minutes (such as just making a mistake), thus improving the accuracy of judgment.
[0099] Support for prediction: hidden state It can be used to predict students' next behavior (such as whether they will give up on the task), thus enabling proactive intervention.
[0100] The specific processing content and output results of the time series modeling module are as follows: 1) Data processed The temporal modeling module receives a sequence of multimodal feature vectors after feature extraction and fusion as input. ,in, Indicates at time t The fused feature vector, In the formula, These are speech features (such as intonation, speech rate, and emotional tone). Facial expression features (such as emotion category, gaze direction, micro-expressions). For text / interaction features (such as answer accuracy, typing speed, and frequency of asking for help). These are physiological characteristics (such as heart rate, skin conductance response, and sitting pressure distribution).
[0101] The data processed by the temporal modeling module has the following characteristics: First, it is multimodal, integrating four types of information: auditory, visual, textual, and physiological; second, it is temporal, with the data arranged in chronological order (sampling frequency is usually on the order of seconds or minutes), reflecting the continuous changes in students' states; and third, it is heterogeneous, with different modal features having different dimensions and physical meanings.
[0102] 2) Handling method The temporal modeling module uses LSTM (Long Short-Term Memory) or Transformer to model the above multimodal feature vector sequences. Taking LSTM as an example, the specific calculation process is as follows: ① Input Gate : Determine the current input x t How much information flows into memory cells.
[0103] ②The Gate of Oblivion : Determines the memory unit of the previous moment How much information is retained?
[0104] ③ Output gate : Determines the current memory unit How much information is output as a hidden state? .
[0105] The final output is: , For a moment t The implicit state vector, typically with dimensions of 128 or 256, encodes all historical information from the beginning to the current moment.
[0106] If a Transformer is used, long-distance dependencies can be captured through a self-attention mechanism:
[0107] 3) Processing results The output of the time series modeling module is a hidden state sequence containing time-dependent information:
[0108] This output is passed to the subsequent multi-task learning module for parallel computation of the real-time states of the five cognitive dimensions, as shown in Table 1: Table 1
[0109] The interest graph is stored in a graph structure, with nodes representing interest topics, weights representing interest intensity, and edges representing the strength of association between topics, used for cross-disciplinary content generation. The ability spectrum is represented by multi-dimensional vectors, with each subject or ability dimension corresponding to a normalized score (0-1) reflecting the current level of mastery. Emotional state includes dominant emotion, probability distribution, and confidence level, used for decision-making in the emotional interaction engine. Learning style includes main preferences and scores for each category, facilitating the selection of appropriate interaction methods (e.g., providing charts for visual learners). Metacognitive assessment includes sub-dimension scores for self-monitoring, planning, and evaluation, used to determine whether students need guidance on learning strategies.
[0110] (3) The profile integration layer adopts a key-value data structure, merging the outputs of the five dimensions as independent fields into a single object, and attaching necessary metadata. The integration process includes: ① Dimension Mapping: The output of each dimension is directly mapped to a subfield of the portrait object, with the field name corresponding one-to-one with the dimension.
[0111] ② Format consistency: If the output of certain dimensions is a tensor or graph structure, it needs to be converted into a serializable format (such as a JSON-compatible list or dictionary).
[0112] ③ Metadata addition: Automatically generate fields such as timestamp (Unix timestamp or ISO format), student ID, and profile version number to ensure the timeliness and traceability of the profile.
[0113] ④ Historical version control: Each time an update is performed, the old image is stored in the historical database, and the new image overwrites the current version. Historical images can be queried by time range.
[0114] The profile integration layer integrates the output from five dimensions to form a structured student profile, while also including a timestamp in the profile to support historical version tracing.
[0115] (4) The storage layer serves as a student profile database, supporting real-time updates and historical queries.
[0116] (5) The API service layer includes: update_profile: Receives new data and updates the profile; get_profile: Query current or historical profiles; predict_next_actions: Predicts actions based on profiles and context.
[0117] (6) External interfaces include: Front-end application system: Displays student status in real time; Recommendation system: Personalized content recommendation; Teacher panel: Monitors the learning status of the class.
[0118] The workflow of the dynamic student modeling engine is as follows: ① Input multimodal timing data stream .
[0119] ②Feature extraction and fusion:
[0120] in, The fusion layer is a multimodal feature fusion function that combines embedding vectors from different modalities. The embedding function is a trainable mapping function that transforms the original input features from the original space into a low-dimensional, dense vector space, generating the corresponding embedding vector.
[0121] ③ Timing modeling is achieved through LSTM or Transformer encoders.
[0122] ④ Conduct multi-task learning, including: Interest graph update: ;
[0123] Capability spectrum update: ;
[0124] Emotional state recognition: ; Learning style classification: ; Metacognitive assessment: .
[0125] 5. In the robot's interface, the emotional interaction engine, based on a hierarchical strategy library (preschooler-playmate, child-mentor, teenager-advisor), identifies students' emotions through a multimodal emotion computing model and plans the robot's facial expressions, actions, and voice parameters to achieve human-like empathetic interaction, such as... Figure 6 As shown in Table 2: The age stratification strategy is as follows: Table 2
[0126] The emotion interaction engine includes an input layer, an emotion detection module, an interaction planning module, a knowledge base, an emotion detection output layer, and an interaction planning output layer.
[0127] (1) The input layer receives input including: Audio data: Students' voice input, including features such as intonation, speaking speed, and volume; Video data: Videos of students' facial expressions and body language; Content to be expressed: The text of the teaching content that the virtual teacher needs to express; Student profile: includes information such as age, learning style, and emotional history; Interaction context: current dialogue state, learning environment, previous emotional state, etc.
[0128] (2) The emotion detection module includes a feature extraction network, a multimodal feature fusion unit, an emotion classifier, and an emotion analyzer.
[0129] Feature extraction networks include: Audio feature extraction network ( Extracting emotion-related features from speech data; Visual feature extraction network ( ): Extracting facial expression features from facial images; Text feature extraction network ( ): Extracting emotional semantic features from dialogue text.
[0130] A multimodal feature fusion unit is used to fuse features from three modalities:
[0131] in, The weight matrix for sentiment classification. For bias terms; The symbol ] represents a vector concatenation operation; the final output is the sentiment probability distribution via the Softmax function. .
[0132] The sentiment classifier classifies sentiments based on fused features and outputs a sentiment probability distribution.
[0133] The sentiment analyzer is used to determine the dominant sentiment category, calculate the classification confidence score, and add a timestamp record. The dominant sentiment category also serves as input to the interaction context to form a feedback loop.
[0134] (3) Interaction planning module ① The strategy selector considers factors such as student age (child, teenager, adult), learning style (visual, auditory, kinesthetic), and current emotional state (happy, confused, frustrated), and selects interaction strategies based on student profile and current context.
[0135] In the emotion-based interaction engine, the selection of interaction strategies is a dynamic decision based on the student's profile (age, cognitive level, learning style, emotional history) and the current context (task difficulty, learning stage, and immediate emotional state). Below are some typical examples of interaction strategies, each corresponding to specific triggering conditions and the robot's expression: Strategy 1: Encouragement Strategy The triggering conditions are as follows: Student profile: Low score (<0.5) in a certain subject in the ability spectrum, or low score in "self-monitoring" in metacognitive assessment.
[0136] Current context: The student has just completed a challenging task (even if not entirely correct), or their facial expression shows confusion / frustration.
[0137] The strategic goal is to enhance students' self-confidence and maintain their learning motivation.
[0138] Example of robot expression: Facial expressions: smiling, nodding, and expectant eyes.
[0139] Action: Raise a thumb and gently pat the shoulder (if the robot has tactile feedback).
[0140] Language: "Wow, this problem is indeed difficult, but you've persisted in thinking about it for so long, that's amazing!" or "It's okay, let's look at the steps again, you'll definitely understand!" Strategy 2: Challenge Strategy The triggering conditions are as follows: Student profile: High score (>0.8) in a certain subject on the ability spectrum, and strong interest in the topic shown on the interest graph.
[0141] Current context: The student completes the current task quickly with an accuracy rate of >90%, or actively requests "something more difficult".
[0142] The strategic objective is to provide appropriate challenges, prevent boredom, and promote further development of abilities.
[0143] Example of robot expression: Expression: Raised eyebrows and a slightly mysterious smile.
[0144] Actions: Spread your hands out to present the new task, and lean forward to show focus.
[0145] Language: "You did it too fast! Try this advanced question and see if you can push your limits?" or "I have a secret mission that only a little expert like you can complete. Want to give it a try?" Strategy 3: Guiding Strategy The triggering conditions are as follows: Student profile: Learning style is classified as "visual" or "auditory", and "planning" ability is low in metacognitive assessment.
[0146] Current context: Students are faced with complex tasks and don't know where to start, or they repeatedly try the same incorrect steps.
[0147] The strategy aims to provide structured prompts to help students break down the problem and gradually approach the answer.
[0148] Example of robot expression: Expression: Focused, slightly thoughtful.
[0149] Actions: Point to key information on the screen and use gestures to simulate step-by-step breakdowns.
[0150] Language: "This problem can be broken down into three steps. Let's look at the first step: What do you think we need to know first?", or "Remember the chart we used last time? We can draw one to help this time too." Strategy 4: Empathy Strategy The triggering conditions are as follows: Student profile: Emotional history shows frequent recent negative emotions (such as frustration and anxiety).
[0151] Current context: Students are showing signs of fatigue (yawning, rubbing their eyes) due to prolonged studying, or show signs of nervousness as exams approach.
[0152] The strategic goals are: to provide emotional support, alleviate negative emotions, and build trust.
[0153] Example of robot expression: Expression: A gentle gaze and a slight upturn of the corners of the mouth indicate understanding.
[0154] Action: Nod slowly and make a "take a break" gesture.
[0155] Language: "You seem a little tired today. How about taking a 5-minute break and listening to a joke before continuing?", or "I understand you're a little nervous right now. Actually, everyone feels like that when facing challenges. Take a deep breath, and let's take it slow together." Strategy 5: Gamification Strategy The triggering conditions are as follows: Student profile: younger (3-6 years old), or high score in "kinesthetic" learning style.
[0156] Current context: The task is highly repetitive, and students' attention is starting to wander (eyes wandering, fidgeting more).
[0157] The strategic goal is to increase participation through gamification elements, transforming tedious practice into fun interaction.
[0158] Example of robot expression: Facial expressions: exaggerated surprise, blinking.
[0159] Actions: Imitating game character actions, clapping and cheering.
[0160] Language: "Next up, challenge mode! Answer three questions correctly to collect a star, and collect five stars to exchange for a secret Easter egg!" or "Let's have a competition! Let's see if you can figure out the answer first, or if I can think of the next question first?" Strategy 6: Metacognitive Strategy The triggering conditions are as follows: Student profile: The scores for "self-monitoring" or "evaluation" in the metacognitive assessment are low, but the ability spectrum is not low.
[0161] Current context: The student immediately asked for the next task after completing the previous one, without reflecting on the task.
[0162] The strategic goal is to cultivate metacognitive abilities and guide students to reflect on their learning process.
[0163] Example of robot expression: Expression: Serious, with a slight questioning tone.
[0164] Action: Place both hands on the table (simulating a teacher's posture) and slowly nod to indicate that you are thinking.
[0165] Language: "You solved that problem very quickly. Can you tell me how you came up with that solution?", or "If we encounter a similar problem next time, which step do you think is most likely to result in a mistake? How can we avoid it?" The logical flow for selecting the above different strategies is as follows: Input: Student profile (age, ability, style, emotional history), current context (task, immediate emotion, environment).
[0166] Strategy selector: Matches the most suitable strategy using a rule engine or lightweight classification model.
[0167] Output: The selected strategy type, passed to the subsequent "target sentiment determiner" and "expression selector".
[0168] Through the aforementioned strategy library, the robot can flexibly switch interaction methods in different situations, just like an experienced teacher, truly achieving a combination of "personalized instruction" and "emotional companionship".
[0169] ② Content sentiment analyzer: The content sentiment analyzer analyzes the “teaching content text that the virtual teacher needs to express”, and analyzes the emotional tendency of the content that the virtual teacher wants to express, such as: encouraging content, corrective content, explanatory content, etc., as shown in Table 3.
[0170] Table 3
[0171] ③ The target emotion determiner combines content emotion, students' current emotions, and interaction strategies to determine the target emotion that the virtual teacher should convey. For example, when a student is frustrated, the target emotion might be "encouragement" or "comfort."
[0172] ④ Expression selector: Expression selector: Selects a suitable sequence of facial expressions from the expression library; Gesture selector: Selects a suitable sequence of body movements from a gesture library; Voice parameter adjuster: Adjust voice parameters such as tone, speech rate, and volume.
[0173] ⑤ Timing planner coordinates the timing of facial expressions, gestures, and speech to ensure the synchronicity and naturalness of multimodal expression.
[0174] (4) Knowledge base, including: Facial Expression Library: Predefined virtual teacher facial expression templates; Gesture library: Predefined virtual teacher body movement templates; Strategy rule base: Interaction strategy selection rules and heuristic knowledge.
[0175] (5) The emotion detection output layer calculates the emotion probability distribution, i.e. the probability value of each emotion category; determines the dominant emotion, i.e. the most likely emotion category; calculates the confidence level as a measure of the reliability of the classification result; and adds a timestamp as a record of the detection time.
[0176] (6) Interactive planning output layer, the output content includes: Interaction strategy: The chosen interaction method (e.g., encouraging, guiding, corrective). Facial expression sequence: A sequence of facial expression changes ordered over time; Gesture sequence: a sequence of limb movements ordered by time; Speech parameters: Adjustment parameters for speech synthesis; Timing: A time coordination plan for multimodal expression.
[0177] The workflow of the emotion interaction engine is as follows: 1) Emotion detection process: Audio / video data / text → Feature extraction → Feature fusion → Sentiment classification → Sentiment analysis → Detection results 2) Interaction planning process: The robot's intended message + student profile + interaction context → interaction strategy selection → sentiment analysis → target sentiment determination → expression selection → time-series planning → interaction plan 6. In the robot end, the actuators include multi-degree-of-freedom bionic joints, high-resolution facial expression display units, speakers, etc., to realize limb movements, display facial expressions, and emit voice.
[0178] The multimodal perception module includes a distributed tactile sensor array and a multimodal perception head composed of a stereo vision camera and microphone array. In the personalized learning companion robot system, the distributed tactile sensor array is a crucial component of the multimodal perception module. Its main function is to endow the robot with tactile perception capabilities, enabling it to understand the student's physical interactions and adjust its interaction strategies and emotional expressions accordingly. This module transforms the student's tactile behavior into quantifiable emotional and intentional signals through physical contact perception, allowing the robot to understand and respond to non-verbal interactions like a human companion, while ensuring the safety and human-like nature of the interaction process. This design allows the learning companion robot to break through the limitations of traditional educational robots that rely solely on voice and vision, truly achieving multimodal, comprehensive human-machine emotional connection.
[0179] 7. Security and Privacy Protection Mechanisms Design a systematic approach that spans both the edge and cloud ends, namely a security and privacy mechanism. This mechanism employs multiple mechanisms such as hierarchical data encryption, local feature extraction, anonymization, federated learning, and privacy policy checks to provide basic data security and privacy protection capabilities for all other modules, ensuring the security of student data throughout its entire lifecycle and complying with regulations such as COPPA.
[0180] (1) Data security architecture, such as Figure 7 As shown, it includes: 1) Data Security API Layer SecurityPrivacyAPI: The main entry point for security and privacy protection; process_sensitive_data: Handles different types of sensitive data; enforce_privacy_policies: Perform privacy policy checks.
[0181] 2) Core Security Service Layer EncryptionService: Provides data encryption / decryption functionality; AccessControl: Manages data access permissions and policies. 3) Sensitive data processing layer, sensitive data includes biometric data, interaction logs and evaluation results.
[0182] The biometric data processing method is as follows: raw biometric data from sensors such as cameras and microphones are processed locally on the user's device to avoid uploading the raw data; representative feature vectors are extracted from the raw data, and the raw data is deleted immediately after feature extraction; the feature vectors are encrypted, and then the encrypted data is stored / transmitted.
[0183] The interaction logs are processed as follows: For the original interaction logs consisting of user and system interaction records, they are first anonymized to remove or obfuscate personal identification information; then the anonymized logs are encrypted and stored.
[0184] The evaluation results are processed as follows: the original learning evaluation and test results are directly encrypted without any additional processing.
[0185] 3) Privacy check layer: Receives data access requests and performs privacy checks. The process includes: Upon receiving a request from an internal system component to access sensitive data, the request is examined using three parallel inspection strategies. These strategies include: ①COPPA Compliance Check: Checks whether the data complies with the Children's Online Privacy Protection Act, ensuring special protection for the data of children under the age of 13; ② Parental consent verification: Verify whether parental consent has been obtained for data processing, and check whether the consent form is within its validity period; ③ Data minimization principle: Ensure that only the minimum data necessary for access is provided, and filter out unnecessary data fields according to the purpose of access.
[0186] Inspection result processing: If the inspection passes, the minimal data is returned for the visitor to use; if the inspection fails, an exception is thrown, indicating the corresponding failed inspection strategy, and data access is blocked.
[0187] 4) Data storage layer: All sensitive data is stored in encrypted form through an encrypted database.
[0188] 5) External data input layer, which receives incoming external data. External data includes biometric data provided by sensors / cameras, user interaction logs recording user interactions with the system, and learning evaluation results generated by the evaluation system.
[0189] In an encrypted database, data users include: AI brain engine: requires data for analysis and modeling; Teacher Management System: Teachers can view students' learning progress; Parental control panel: Parents can understand their child's learning progress.
[0190] (2) Federated learning architecture, such as Figure 8 As shown, the federated learning architecture includes a server-side component, a client-side component, a privacy protection module, a model update module, and an output layer.
[0191] 1) On the server, maintain the global model and manage the set of clients participating in federated learning through the client manager.
[0192] 2) The client consists of various devices or users participating in the training. Each client independently trains the model using local data, and the original data is always kept on the client device and is not uploaded to the server.
[0193] 3) Privacy protection module, differential privacy protection: Add differential privacy protection to the model update for each client, which can protect privacy by adding noise; the client model update after privacy protection ensures that even if the server receives the model update, it cannot infer the original data.
[0194] 4) The model update module first securely aggregates all client-side privacy updates through encryption or secure multi-party computation. All client model updates are aggregated using a weighted average and applied to the global model to update its parameters. The updated global model is then distributed back to each client for the next round of training.
[0195] 5) Output layer: Outputs the global model improved by federated training, evaluation metrics during the training process, and statistical information on training progress and participation.
[0196] The process of federated learning is as follows: Ⅰ. Global Model Distribution: The server sends the current global model to the clients participating in the training. II. Local Training: Each client trains the received model using local data; III. Model Update Calculation: Calculate the model update (gradient difference) generated during local training; IV. Differential Privacy Processing: Adding noise to the client's model updates to protect privacy; V. Update Upload: The client uploads the privacy update to the server; VI. Secure Aggregation: The server securely aggregates updates from all clients; VII. Global Model Update: Update the global model using the aggregation results; VIII. Model Evaluation: Evaluate the performance of the updated global model; IX. Preparations for the next round: Prepare to begin the next round of federal training.
[0197] The federated learning architecture enables the core functionality of training machine learning models while protecting user data privacy, making it particularly suitable for handling sensitive student data in educational settings.
[0198] In this embodiment, a 10-year-old student participating in a "Mars Base" project is used as an example to illustrate the data flow, emotion recognition, path adjustment, and feedback reinforcement processes of the personalized learning companion robot system proposed in this invention during multi-stage interactions. Figure 9 As shown.
[0199] (1) Initial interaction phase (0-5 minutes) Student to robot: "I want to design a Mars base!" (excited tone) Robot end -> Edge node: Send voice data + facial expression data; Edge node -> Cloud: Forwarding data + preliminary sentiment analysis (excitement); Cloud platform -> AI brain: Generate response strategies; AI Brain -> Dynamic Student Modeling Engine: Update Interest Graph (Space Interest+); Dynamic Student Modeling Engine -> Adaptive Curriculum Planning Module: Recommend the "Mars Base" project; Cloud to Edge Node: Send Interaction Plan; Edge node -> Robot end: Execution plan; Robot side -> Student: "Great! Let's design a Mars base together!" (accompanied by a happy emoji).
[0200] (2) Task execution phase (5-30 minutes) Student -> Robot: Start designing, asking "How much oxygen is needed?"; Robot client -> Edge node: Forwarding issues + Current design status; Edge node -> Cloud: Problem + Contextual Information; Cloud-based AI Brain: Retrieving knowledge about the Martian atmosphere; AI Brain -> Adaptive Curriculum Planning Module: Recognizing Mathematical Computer Learning; Adaptive Curriculum Planning Module -> AI Brain: It is recommended to incorporate proportion calculation; Cloud to edge node: Send solutions + math challenge; Edge node -> Robot end: Execute teaching; Robot on the robot -> Student: "Oxygen only accounts for 0.13% of the Martian atmosphere. Let's calculate how much we need to bring..." (accompanied by a thinking expression).
[0201] (3) Difficulty identification and adjustment phase (30-35 minutes) Student -> Robot: "This calculation is so difficult..." (frowning, voice low); Robot client -> Edge node: Frustration detected; Edge node -> Cloud: Emotional alert + Current task status; Cloud-based dynamic student modeling engine: updating emotional state (frustration); Dynamic Student Modeling Engine -> Adaptive Curriculum Planning Module: It is recommended to reduce the difficulty. Adaptive Curriculum Planning Module -> Emotional Interaction Engine: Requesting Encouragement Strategies; Cloud to Edge Node: Send Adjustment Plan; Edge node -> Robot end: Perform adjustments; Robot side -> Student: "It's okay, we'll take it one step at a time. Look at this chart..." (accompanied by encouraging emojis and visual aids).
[0202] (4) Task completion and feedback phase (35-45 minutes) Student -> Robot: "I've calculated it! We need to bring XXX liters of oxygen!" Robot end -> Edge node: Send success signal; Edge node -> Cloud: Task completion data + positive sentiment detection; Cloud-based -> Dynamic Student Modeling Engine: Updating Ability Spectrum (Mathematical Confidence+); Dynamic Student Modeling Engine -> AI Brain: Further learning is recommended; Cloud to Edge Node: Send celebration messages and suggestions for expansion; Edge node -> Robot: Execute celebration; Robot side -> Student: "Amazing! You've solved the key problem of survival on Mars!" (Celebratory dance)
[0203] (5) Long-term tracking phase (follow-up) Cloud-based dynamic student modeling engine: records learning trajectories; Dynamic student modeling engine: Updates long-term profiles; Adaptive Curriculum Planning Module: Planning the next advanced mission for the Mars project.
[0204] In summary, this invention provides a personalized learning companion robot system based on dynamic student profiles and large-scale AI models. Through multimodal time-series data fusion, it can deeply and continuously model the learner's dynamic cognitive state. Based on dynamic cognitive state and structured knowledge, it can generate highly personalized teaching content and dialogues. Utilizing advanced learning algorithms, it can autonomously plan and adjust long-term learning paths in real time to match the learner's growth pace. According to the learner's cognitive development stage and immediate emotional state, it can drive the robot to perform natural, friendly, and emotionally supportive anthropomorphic interactions. Furthermore, throughout the learning process, federated learning ensures data security and privacy protection.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A personalized learning companion robot system based on dynamic student profiles and large-scale AI models, characterized in that, The system adopts an "end-cloud-human" collaborative architecture, including a robot end, a cloud end, and a user terminal. The three ends achieve data synchronization and collaborative interaction through data communication. The robot terminal includes a multimodal perception module and a dynamic student modeling engine; the multimodal perception module is used to collect multimodal time-series data during student interaction in real time; the dynamic student modeling engine responds to the multimodal time-series data and updates and generates a structured student profile in real time. The cloud platform includes an AI brain and an adaptive curriculum planning module. The AI brain generates personalized teaching content based on real-time updated structured student profiles, combined with educational knowledge graphs and large language models. The adaptive curriculum planning module dynamically generates and adjusts personalized learning paths based on real-time updated structured student profiles using a proximal strategy optimization algorithm. The robot also includes an emotional interaction engine and an actuator; the emotional interaction engine plans the robot's anthropomorphic empathic interaction parameters based on the real-time updated structured student profiles and personalized teaching content; the actuator is used to execute the interaction parameters.
2. The personalized learning companion robot system according to claim 1, characterized in that, The AI brain includes an input layer, an intent recognition module, a knowledge retrieval module, a content generation module, an educational filtering and optimization module, and an output layer. Input layer input content includes student queries Current situation , student portraits Student queries are student questions or requests, the current context is the dialogue context and learning environment information, and the student profile is a structured student profile obtained from the dynamic student modeling engine. The intent recognition module analyzes the query purpose through an intent classifier and identifies the intent as knowledge question answering, concept explanation, task generation, or practice recommendation type; Based on the intent recognition results, the knowledge retrieval module uses a retrieval-enhanced generative model to provide retrieval results to the content generation module. Specifically, for knowledge-based queries, relevant concepts and relationships are retrieved through a knowledge graph; for task-based intents, content suitable for students' levels and interests is retrieved from a course database. The content generation module uses a prompt word builder to construct prompt words for a large language model by integrating student queries, current context, student profiles, and search results. Then, the large language model generates preliminary response content. The educational filtering and optimization module includes an educational filter and a teaching optimizer; the educational filter is used to filter out inappropriate or non-educational content; the teaching optimizer is used to adjust the filtered teaching content based on a structured student profile. The multimodal output planning module plans the robot's facial expression changes, body language and gestures, and voice parameters based on the optimized teaching content.
3. The personalized learning companion robot system according to claim 1, characterized in that, The adaptive curriculum planning module uses a proximal policy optimization algorithm to define the state space, action space, and reward function, enabling the dynamic generation and adjustment of personalized learning paths. The state space includes the student's profile, learning history, and current progress. The student profile is a structured student profile output by the dynamic student modeling engine. The learning history includes the student's past learning records, grades, and interaction history. The current progress includes the current learning stage, completed tasks, and tasks to be completed. The action space includes selecting a theme, selecting a task type, adjusting the task difficulty, and selecting an interaction method; the reward function is obtained by weighted integration of task completion, emotional feedback, ability growth, and parental satisfaction.
4. The personalized learning companion robot system according to claim 1, characterized in that, The ecosystem platform includes an input layer, a lesson plan upload layer, a lesson plan recommendation layer, and a storage layer; The input layer includes lesson plan data uploaded by the teacher's terminal, student profile data, and learning context data representing the student's current learning environment and context information; The lesson plan upload layer performs content review and educational standard alignment checks on the uploaded lesson plan data; it generates structured metadata based on the lesson plan content and review results, including author information, upload time, suitable student age range, subject classification, time required to complete the lesson plan, difficulty level, standard alignment, knowledge required to learn this lesson plan, and teaching objectives; it transmits the lesson plan content and structured metadata to the storage layer for storage and returns a unique lesson plan ID; The lesson plan recommendation layer employs a hybrid recommendation strategy that includes collaborative filtering and content-based filtering to recommend suitable lesson plans to students. The hybrid recommendation strategy involves taking a weighted average of the results of collaborative filtering and content-based filtering to obtain a hybrid recommendation result. Then, the recommendation result is adjusted according to the current learning context, taking into account the learning time of day, the student's current available learning time, and the learning location and equipment conditions. Personalized recommendations adapted to the context are then output. Finally, the top 5 best lesson plans are selected from the final recommendation list and returned to the student's learning interface.
5. The personalized learning companion robot system according to claim 1, characterized in that, The dynamic student modeling engine includes an input layer, a feature extraction and fusion module, a temporal modeling module, a multi-task learning module, a multi-task learning branch, a profile integration layer, and a storage layer. The input layer receives four types of data sources, including speech features, facial expression features, text / interaction features, and physiological state features; The feature extraction and fusion module extracts features from the four types of input feature data and fuses the extracted features. The time series modeling module uses LSTM / Transformer to handle time series dependencies; the multi-task learning module processes five cognitive dimensions in parallel; the profile integration layer integrates the outputs of the five dimensions to form a structured student profile; and the storage layer serves as a student profile database, supporting real-time updates and historical queries.
6. The personalized learning companion robot system according to claim 1, characterized in that, The emotional interaction engine includes an input layer, an emotion detection module, an interaction planning module, an emotion detection output layer, and an interaction planning output layer. The input layer receives input including students' audio data, video data, structured student profiles, and the content and interaction context that the virtual teacher wants to express. The emotion detection module includes a feature extraction network, a multimodal feature fusion unit, an emotion classifier, and an emotion analyzer. The feature extraction network is used to extract emotion-related features, facial expression features, and emotional semantic features from audio data, video data, and interactive context, respectively. The multimodal feature fusion unit is used to fuse features from the three modalities. The sentiment classifier classifies sentiments based on fused features and outputs a sentiment probability distribution; the sentiment analyzer determines the dominant sentiment category and calculates the classification confidence score. The interaction planning module includes a strategy selector, a content sentiment analyzer, a target sentiment determiner, an expression selector, and a timing planner. The strategy selector considers students' age, learning style, and current emotional state, and selects interaction strategies based on structured student profiles and the current context; the content sentiment analyzer analyzes the emotional tendency of the content that the virtual teacher wants to express. The target emotion determiner combines the content's emotional tendency, the student's current emotion, and interaction strategies to determine the target emotion that the virtual teacher should convey; the expression selector selects facial expression sequences, body movement sequences, and voice parameters based on the target emotion; and the timing planner coordinates the timing of facial expressions, gestures, and voice to ensure the synchronicity and naturalness of multimodal expression. The sentiment detection output layer calculates the sentiment probability distribution to determine the dominant sentiment. And calculate the confidence score as a measure of the reliability of the classification results; The interaction planning output layer outputs interaction strategies, facial expression sequences, gesture sequences, voice parameters, and timing arrangements.
7. The personalized learning companion robot system according to claim 1, characterized in that, The actuator includes a multi-degree-of-freedom bionic joint, a high-resolution facial expression display unit, and a speaker, used to realize limb movements, display facial expressions, and emit voice.
8. The personalized learning companion robot system according to claim 1, characterized in that, The multimodal sensing module includes a distributed tactile sensor array and a multimodal sensing head composed of a stereo vision camera and a microphone array.