Learning auxiliary robot and system thereof
Through multimodal knowledge presentation and remote collaboration modules, the learning assisted robot dynamically adjusts the teaching content and difficulty according to students' needs, solving the problem of personalized learning in the existing technology, and achieving an efficient and interesting learning experience and improvement of team collaboration capabilities.
Patent Information
- Application Number
- CN202510651048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing learning assisted robots cannot dynamically adjust teaching strategies and content based on students' learning progress, interest preferences and abilities, resulting in difficulty in meeting students' personalized learning needs, the learning process is boring, and even aversion to learning.
The multimodal knowledge presentation module combines media forms of text, images, audio and video, combined with natural language processing and computer vision technology, generates personalized learning paths, and supports students to communicate and cooperate with others through the remote collaboration module, and dynamically adjusts the learning content and difficulty.
It has realized dynamic adjustment of personalized learning paths, stimulated students' interest in learning, improved learning efficiency and effectiveness, enhanced team collaboration capabilities, and provided a rich and diverse learning experience.
Smart Images

Figure CN120260358A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence teaching, and in particular relates to a learning assistance robot and a system thereof. Background Art
[0002] With the rapid development of artificial intelligence technology, the field of education has also ushered in an intelligent transformation. In recent years, learning assistance technology has become a research hotspot in the field of education, aiming to provide students with a more personalized and efficient learning experience by combining advanced technical means. Among them, learning assistance robots and their systems, as emerging educational tools, are gradually showing their great potential in the field of education.
[0003] At present, there are some learning-assistive robot products on the market. These products usually have basic functions such as knowledge presentation and problem answering. For example, some robots can answer students' questions through voice interaction, or display relevant knowledge content through a display screen. However, existing learning-assistive robots often lack in-depth understanding and flexible response to students' personalized learning needs. They can usually only impart knowledge according to preset procedures or processes, but cannot dynamically adjust teaching strategies and content according to students' learning progress, interest preferences and ability differences. This "one-size-fits-all" teaching method is not only difficult to meet students' personalized learning needs, but may also cause students to feel boring and tedious during the learning process, and even develop an aversion to learning. Therefore, staff need to improve it. Summary of the invention
[0004] The purpose of the present invention is to provide a learning-assisting robot and a system thereof to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A learning assistance system, comprising:
[0007] Multimodal knowledge presentation module, which is used to combine text, image, audio and video media forms to flexibly select and present the most appropriate knowledge transfer media according to the nature of the learning content and the needs of students;
[0008] The natural language processing module is used to receive and understand students' text and voice input questions and convert them into instructions and information that can be recognized by the system;
[0009] Computer vision module, used to recognize students’ gestures, expressions and learning materials through images and videos to help understand students’ questions and needs;
[0010] The multimodal answer generation module is used to combine the results of natural language processing and computer vision technologies to generate and present multimodal answers, including but not limited to text, image, audio, and video forms;
[0011] The learning content management module is used to store, organize, and manage various forms of learning content so that the multimodal knowledge presentation module can select and present according to needs;
[0012] The user interface module is used to provide a user-friendly interaction interface, enabling students to conveniently interact with the learning assistance system, including selecting learning content, asking questions, and viewing answers.
[0013] Preferably, it further includes:
[0014] The personalized learning path generation module is used to dynamically generate a personalized learning path based on the student's historical learning data, ability assessment, and real-time feedback, and guide the multimodal knowledge presentation module to present knowledge according to this path. Its degree of personalization P l is evaluated by the following formula:
[0015]
[0016] where S is the score of the student's historical learning data, A is the score of the student's ability assessment, F is the score of the student's real-time feedback, α, β, and λ are the weights of each score, and θ is the threshold parameter.
[0017] Preferably, the multimodal answer generation module further includes:
[0018] The answer selection unit is used to select the most appropriate answer from a preset answer library;
[0019] The answer generation unit is used to automatically generate a new answer when there is no appropriate answer in the answer library. Its ability Q g is evaluated by the following formula:
[0020]
[0021] where T is the answer generation time, N is the number of factors affecting the answer generation quality, w i is the weight of the i-th factor, f i (x i ,t) is the function of the i-th factor changing with time t, x i is the specific value of the i-th factor, and λ is the attenuation coefficient;
[0022] The answer optimization unit is used to continuously optimize the quality and accuracy of the answer according to the student's learning feedback and assessment results. Its optimization effect O p is represented by the following formula:
[0023]
[0024] Among them, M is the number of answers before and after optimization, and R j and are the scores of the j-th answer after and before optimization respectively, and E k and are the actual value and the expected value of the k-th evaluation index respectively, and E max,k and E min,k are the maximum value and the minimum value of the k-th evaluation index respectively, and L is the number of evaluation indexes.
[0025] Preferably, it further includes:
[0026] A remote collaboration module for supporting students to communicate and cooperate with other learners and teachers in real time, including sharing learning experiences, discussing problems, and carrying out group cooperation.
[0027] A learning assistance robot, comprising:
[0028] A robot body;
[0029] The top of the surface of the robot body is fixedly connected with a camera module, and the surface of the robot body is fixedly connected with a sensor module at the bottom of the camera module;
[0030] The surface of the robot body is fixedly connected with a microphone module at the bottom of the sensor module, the surface of the robot body is fixedly connected with a display module, the top of the display module is fixedly connected with a control module, and the bottom of the robot body is fixedly connected with a moving module.
[0031] Preferably, a path planning unit is fixedly connected to the inner bottom wall of the robot body, and an interaction strategy generation unit is fixedly connected to the top of the path planning unit.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] (1) By collecting students' historical learning data, conducting ability assessments, and combining with real-time feedback, a personalized learning path is dynamically generated. Through historical learning data and ability assessments, the system can accurately identify students' strengths and weaknesses in learning, so as to customize a learning path for students. The real-time feedback mechanism allows the system to dynamically adjust the difficulty and rhythm of the learning path according to students' learning progress and reactions, ensuring that learning is both efficient and not too strenuous for students. The personalized learning path can guide students to learn at the most suitable rhythm and way for themselves, thereby improving learning efficiency and reducing ineffective learning time. By introducing learning content related to students' interests, the personalized learning path can stimulate students' learning interests and make learning more active and interesting.
[0034] (2) Through the multi-modal knowledge presentation module, by combining various forms such as text, images, audio, and video, the learning content is presented to students in an intuitive and vivid way. Through the combination of multiple media forms, the system can provide students with a richer and more diverse learning experience, making learning more vivid and interesting. Different students have different preferences and acceptance degrees for different media forms. Multi-modal knowledge presentation can meet these differences, help students understand the learning content faster. Forms such as images, audio, and video can provide students with more intuitive learning materials, which helps to deeply understand and master knowledge points. The stimulation of multiple senses can enhance the memory effect of students, making the learning content more unforgettable.
[0035] (3) Through the remote collaboration module, it supports students to communicate and cooperate with other learners and teachers in real time, including sharing learning experiences, discussing problems, and carrying out group cooperation, etc. Through communication with other learners, students can understand different learning methods and ideas, thus broadening their learning horizons. Real-time discussion and cooperation help students solve problems in learning faster and improve their problem-solving abilities. Learning and cooperating with other learners can enhance students' sense of belonging and achievement, thereby stimulating learning motivation. The remote collaboration module can let students exercise their teamwork abilities in practice, which is of great significance for future career development. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the system flow chart of the present invention;
[0037] Figure 2 is the three-dimensional view of the present invention;
[0038] Figure 3 is the three-dimensional view of the path planning unit of the present invention;
[0039] In the figure: 1, robot body; 2, camera module; 3, sensor module; 4, microphone module; 5, display module; 6, control module; 7, mobile module; 8, path planning unit; 9, interaction strategy generation unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1:
[0042] Please refer to Figure 1As shown in the figure, a learning assistance system includes:
[0043] A multimodal knowledge presentation module, which is used to combine media forms such as text, images, audio, and video, and flexibly select and present the most suitable knowledge dissemination media according to the nature of the learning content and the needs of students;
[0044] A natural language processing module, which is used to receive and understand the text and voice input questions of students, and convert them into instructions and information recognizable by the system;
[0045] A computer vision module, which is used to recognize the gestures, expressions, and learning materials of students through images and videos to assist in understanding the questions and needs of students;
[0046] A multimodal answer generation module, which is used to combine the results of natural language processing and computer vision technologies to generate and present multimodal answers, including but not limited to text, image, audio, and video forms;
[0047] A learning content management module, which is used to store, organize, and manage various forms of learning content so that the multimodal knowledge presentation module can select and present according to needs;
[0048] A user interface module, which is used to provide a user-friendly interaction interface to enable students to conveniently interact with the learning assistance system, including selecting learning content, asking questions, and viewing answers.
[0049] It also includes:
[0050] A personalized learning path generation module, which is used to dynamically generate a personalized learning path according to the historical learning data, ability assessment, and real-time feedback of students, and guide the multimodal knowledge presentation module to present knowledge according to this path. Its degree of personalization P l Is evaluated by the following formula:
[0051]
[0052] Where S is the score of the student's historical learning data, A is the score of the student's ability assessment, F is the score of the student's real-time feedback, α, β, λ are the weights of each score, and θ is the threshold parameter.
[0053] The multimodal answer generation module further includes:
[0054] An answer selection unit, which is used to select the most suitable answer from a preset answer library;
[0055] An answer generation unit, which is used to automatically generate a new answer when there is no suitable answer in the answer library. Its ability Q to generate a new answer g Is evaluated by the following formula:
[0056]
[0057] Among them, T is the answer generation time, N is the number of factors affecting the answer generation quality, w i is the weight of the i-th factor, f i (x i , t) is the function of the i-th factor changing with time t, x i is the specific value of the i-th factor, and λ is the attenuation coefficient;
[0058] The answer optimization unit is used to continuously optimize the quality and accuracy of the answer according to the learning feedback and evaluation results of the students. Its optimization effect O p is represented by the following formula:
[0059]
[0060] Among them, M is the number of answers before and after optimization, R j and are the scores of the j-th answer after and before optimization respectively, E k and are the actual value and the expected value of the k-th evaluation index respectively, E max,k and E min,k are the maximum value and the minimum value of the k-th evaluation index respectively, and L is the number of evaluation indexes.
[0061] T (T is the time parameter) is the answer generation time;
[0062] N (N is the quantity parameter) is the number of factors affecting the answer generation quality;
[0063] w i is the weight of the i-th factor;
[0064] f i (x i , t) (f i is the function parameter) is the function of the i-th factor changing with time t;
[0065] x i is the specific value of the i-th factor;
[0066] λ is the attenuation coefficient;
[0067] M is the number of answers before and after optimization;
[0068] R j and (R j and are the scoring parameters) are the scores of the j-th answer after and before optimization respectively;
[0069] E k and (E k and is the evaluation index parameter) are the actual value and the expected value of the k-th evaluation index respectively;
[0070] E max,k and E min,k (E max,k and E min,k is the extreme value parameter) are the maximum value and the minimum value of the k-th evaluation index respectively;
[0071] L is the number of evaluation indexes;
[0072] S is the score of the student's historical learning data;
[0073] A is the score of the student's ability evaluation;
[0074] F is the score of the student's real-time feedback;
[0075] α, β, γ (α, β, γ are weight parameters) are the weights of each score;
[0076] θ is the threshold parameter.
[0077] Range interpretation:
[0078] Q g The value range of is [0, +∞), and the larger the value, the stronger the ability of the answer generation unit;
[0079] O p The value range of is [-1, 1]. The closer the value is to 1, the better the optimization effect, and the closer the value is to -1, the worse the optimization effect;
[0080] P l The value range of is (0, 1). The larger the value, the higher the degree of personalization.
[0081] It also includes:
[0082] A remote collaboration module, which is used to support real-time communication and cooperation between students and other learners and teachers, including sharing learning experiences, discussing problems, and carrying out group cooperation.
[0083] Embodiment 2:
[0084] Please refer to Figures 2 to 3 As shown in, a learning assistance robot includes:
[0085] The robot body 1;
[0086] A camera module 2 is fixedly connected to the top of the surface of the robot body 1, and a sensor module 3 is fixedly connected to the bottom of the surface of the robot body 1 where the camera module 2 is located;
[0087] The surface of the robot body 1 is fixedly connected with a microphone module 4 at the bottom of the sensor module 3, the surface of the robot body 1 is fixedly connected with a display module 5, the top of the display module 5 is fixedly connected with a control module 6, and the bottom of the robot body 1 is fixedly connected with a moving module 7.
[0088] The inner bottom wall of the robot body 1 is fixedly connected with a path planning unit 8, and the top of the path planning unit 8 is fixedly connected with an interaction strategy generation unit 9.
[0089] The robot body 1 has the ability to move and a display module 5 and an audio output device for interacting with users;
[0090] The control module 6 is used to control the operation of the moving module 7, the display module 5 and the audio output device of the robot;
[0091] The sensor module 3 is used to sense the position, actions and / or environmental information of the student to assist the control unit in making decisions;
[0092] The camera module 2 is used to capture the image or video information of the student and send it to the learning assistance system for computer vision processing;
[0093] The microphone module 4 is used to receive the voice input questions of the student and send them to the learning assistance system for natural language processing.
[0094] The path planning unit 8 is used to plan the moving path of the robot according to the learning position and environmental information of the student;
[0095] The interaction strategy generation unit 9 is used to generate appropriate interaction strategies according to the learning status and needs of the student, such as adjusting the content of the display interface, emitting a prompt sound, etc.
[0096] Embodiment III:
[0097] Please refer to Figures 1 to 3 As shown, with the rapid development of artificial intelligence technology, the education field has also ushered in an intelligent transformation. This application proposes a learning assistance robot and its system, aiming to provide personalized learning experiences for students by combining various media forms such as text, image, audio and video. The following will detail the actual application scenarios and effects of this learning assistance robot and its system.
[0098] Personalized learning path
[0099] Student situation: A junior high school student who has a strong interest in mathematics and physics but encounters difficulties in English learning.
[0100] Application process: The learning assistance system collected the student's historical learning data, ability assessment, and real-time feedback, and used the personalized learning path generation module to customize a learning path for him. This path focused on the learning of basic English knowledge in the initial stage, and gradually introduced more English materials related to mathematics and physics as the student's English ability improved to meet his learning interests.
[0101] Effect: With the help of the learning assistance robot, the student's English score has improved significantly, and at the same time, the learning of mathematics and physics has become more interesting and efficient.
[0102] Multimodal knowledge presentation
[0103] Student situation: A primary school student who is particularly interested in scientific experiments.
[0104] Application process: The learning assistance system provided the student with experimental tutorials combining text, images, audio, and video through the multimodal knowledge presentation module. The student watched videos to understand the experimental steps, listened to audio explanations of the experimental principles, and read text and viewed images at the same time to deepen understanding.
[0105] Effect: The student showed extremely high enthusiasm and creativity during the experiment, not only mastering scientific knowledge but also cultivating practical skills and problem-solving abilities.
[0106] Natural language processing and computer vision
[0107] Student situation: A high school student who is preparing for the college entrance examination.
[0108] Application process: The student input questions to the learning assistance robot by voice. The robot used the natural language processing module to understand the questions and convert them into instructions recognizable by the system. At the same time, the computer vision module judged the student's learning status and needs by recognizing the student's expressions and gestures. The system combined this information to generate and present multimodal answers in the form of text, images, and audio.
[0109] Effect: With the help of the learning assistance robot, the student can obtain answers to questions more quickly and at the same time feel a more natural and user-friendly interaction experience.
[0110] Remote collaboration
[0111] Student situation: A college student who is participating in a multinational team cooperation project.
[0112] Application process: The student used the remote collaboration module of the learning assistance system to communicate and cooperate with other team members in real time. They shared learning experiences, discussed questions, carried out group cooperation, and jointly completed project tasks.
[0113] Effect: During the process of remote collaboration, students not only improve their communication and teamwork skills, but also expand their international perspectives and cross-cultural communication abilities.
[0114] Working principle: The learning assistance system, through the multi-modal knowledge presentation module, combines various media forms such as text, images, audio, and video, and flexibly selects and presents the most suitable knowledge dissemination media according to the nature of the learning content and the needs of students. Students can conveniently interact with the learning assistance system through the user interface module to select learning content, ask questions, and view answers. The natural language processing module is responsible for receiving and understanding the text and speech input questions of students and converting them into instructions and information recognizable by the system. At the same time, the computer vision module recognizes the gestures, expressions, and learning materials of students through images and videos to assist in understanding the questions and needs of students. The multi-modal answer generation module combines the results of natural language processing and computer vision technologies to generate and present multi-modal answers, including but not limited to text, image, audio, and video forms. This module further includes an answer selection unit, an answer generation unit, and an answer optimization unit. The answer selection unit selects the most suitable answer from the preset answer library; if there is no suitable answer in the answer library, the answer generation unit automatically generates a new answer; the answer optimization unit continuously optimizes the quality and accuracy of the answer according to the learning feedback and evaluation results of students. The personalized learning path generation module dynamically generates a personalized learning path based on the historical learning data, ability assessment, and real-time feedback of students, and guides the multi-modal knowledge presentation module to present knowledge according to this path. In addition, the remote collaboration module supports students to communicate and cooperate with other learners and teachers in real time, including sharing learning experiences, discussing questions, and carrying out group cooperation.
[0115] The robot body captures the image or video information of students through the camera module, the sensor module senses the position, movement, and environmental information of students, and the microphone module receives the voice input questions of students. These information are sent to the learning assistance system for corresponding processing. The control module is responsible for controlling the operation of the mobile module, display module, and audio output device of the robot. The path planning unit conducts path planning according to the learning position and environmental information of students, and the interaction strategy generation unit generates strategies for interacting with students.
[0116] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A learning assistance system, characterized in that, Including: A multi-modal knowledge presentation module, which is used to combine media forms of text, images, audio, and video, and flexibly select and present the most suitable knowledge dissemination media according to the nature of the learning content and the needs of students; A natural language processing module, which is used to receive and understand the text and speech input questions of students and convert them into instructions and information recognizable by the system; A computer vision module, which is used to identify the gestures, expressions, and learning materials of students through images and videos to assist in understanding the questions and needs of students; A multi-modal answer generation module, which is used to combine the results of natural language processing and computer vision technologies to generate and present multi-modal answers, including but not limited to text, image, audio, and video forms; A learning content management module, which is used to store, organize, and manage various forms of learning content so that the multi-modal knowledge presentation module can select and present according to needs; A user interface module, which is used to provide a user-friendly interaction interface to enable students to conveniently interact with the learning assistance system, including selecting learning content, asking questions, and viewing answers.
2. The learning assistance system according to claim 1, wherein It also includes: A personalized learning path generation module, which is used to dynamically generate a personalized learning path based on the student's historical learning data, ability assessment, and real-time feedback, and guide the multi-modal knowledge presentation module to present knowledge according to this path, with its personalization degree P l Evaluated by the following formula: Among them, S is the historical learning data score of the student, A is the ability assessment score of the student, F is the real-time feedback score of the student, α, β, and λ are the weights of each score, and θ is the threshold parameter.
3. A learning assistance system according to claim 1, characterized in that: The multi-modal answer generation module further includes: An answer selection unit, which is used to select the most suitable answer from a preset answer library; An answer generation unit, which is used to automatically generate a new answer when there is no appropriate answer in the answer library, and its ability Q to generate a new answer g is evaluated by the following formula: Among them, T is the response generation time, N is the number of factors affecting the response generation quality, w i is the weight of the i-th factor, f i (x i , t) is the function of the i-th factor changing with time t, x i is the specific value of the i-th factor, and λ is the attenuation coefficient; Answer optimization unit, which is used to continuously optimize the quality and accuracy of answers according to the learning feedback and assessment results of students, and its optimization effect O p is represented by the following formula: Among them, M is the number of answers before and after optimization, R j and are the scores of the j-th answer after and before optimization respectively, E k and are the actual value and the expected value of the k-th evaluation index respectively, E max,k and E min,k are the maximum value and the minimum value of the k-th evaluation index respectively, and L is the number of evaluation indexes.
4. A learning assistance system according to claim 1, characterized in that, It also includes: A remote collaboration module, which is used to support students to communicate and cooperate with other learners and teachers in real time, including sharing learning experiences, discussing questions, and conducting group cooperation.
5. A learning assistance robot, applicable to a learning assistance system according to claims 1 to 4, characterized in that, Including: A robot body (1); A camera module (2) is fixedly connected to the top of the surface of the robot body (1), and a sensor module (3) is fixedly connected to the bottom of the surface of the robot body (1) where the camera module (2) is located; A microphone module (4) is fixedly connected to the bottom of the surface of the robot body (1) where the sensor module (3) is located, a display module (5) is fixedly connected to the surface of the robot body (1), a control module (6) is fixedly connected to the top of the display module (5), and a mobile module (7) is fixedly connected to the bottom of the robot body (1).
6. The learning assistance robot and system according to claim 1, characterized in that: A path planning unit (8) is fixedly connected to the inner bottom wall of the robot body (1), and an interaction strategy generation unit (9) is fixedly connected to the top of the path planning unit (8).
Citation Information
Cited By
Personalized learning path dynamic generation method based on large model
CN121190276A