AI digital human family education method and device based on large language model

By building an educational resource database through large language models and web crawler technology, combined with Edge-TTS and optimized video generation models, the problems of insufficient accuracy, flexibility and real-time performance of AI digital tutoring methods in existing technologies are solved, and an AI digital tutoring method with high accuracy, high flexibility, high real-time performance and high interactivity is realized.

CN120707344APending Publication Date: 2025-09-26HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711125.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies lack AI digital tutoring methods with high accuracy, high flexibility, high real-time performance and high interactivity. They cannot make real-time dynamic adjustments based on the needs and knowledge levels of different users. They lack interactivity and cannot answer user questions in real time. The content update cycle is long and the adaptability is insufficient, making it difficult to meet diverse and rapidly changing learning needs.

Method used

A large language model combined with web crawler technology is used to crawl educational resource data, build an educational resource database, obtain student information and questions, use the large language model to generate personalized interaction styles and answer questions, use the Edge-TTS module for voice generation, and combine with the optimized video generation model to generate teaching videos and animations to achieve efficient and personalized interaction.

Benefits of technology

It achieves highly real-time and highly interactive educational tutoring, provides personalized and highly targeted tutoring based on user characteristics, is not restricted by time and region, and utilizes a wide database of educational resources to achieve highly accurate tutoring, thereby improving learning experience and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707344A_ABST
    Figure CN120707344A_ABST
Patent Text Reader

Abstract

The invention provides an AI digital human family education method and device based on a large language model, and relates to the technical field of artificial intelligence. The method comprises the following steps: according to student information, performing adaptive style generation by using a large language model to obtain a personalized interaction style; based on an educational resource database, according to the personalized interaction style, the student information and the student questions, a big language model is used for question answering, and a question answering text is obtained; based on a personalized interaction style, according to the question answering text, performing voice generation by using an Edge-TTS module to obtain a teaching audio; performing video generation by using an optimized video generation model according to the teaching audio and the digital human image picture to obtain a silent teaching video; and performing audio synthesis according to the silent teaching video and the teaching audio to obtain a teaching animation. The AI digital family education method based on artificial intelligence is high in accuracy, high in flexibility, high in real-time performance and high in interactivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI digital tutoring method and device based on a large language model. Background Art

[0002] With the continuous development of artificial intelligence, natural language processing, and multimodal technologies, virtual intelligent tutors have shown great potential in improving educational effectiveness and reach. Virtual tutors can replace teachers, significantly reducing the workload of educational video production and saving teachers time and energy. The integration of virtual intelligent tutors with large language models and professional knowledge bases can generate interactive virtual teachers with real-time conversations, making it easier for students to access professional educational content. This highly interactive learning experience can enhance student understanding and engagement, further strengthening educational effectiveness.

[0003] In the existing field of teaching assistance that relies on artificial intelligence, pre-recorded teaching videos and teaching tutoring management systems are used to provide personalized tutoring based on students' learning preferences and teachers' teaching characteristics. They have good professionalism and rich teaching content, but lack interactivity in the tutoring process, cannot answer users' questions in real time, and cannot provide timely and targeted solutions. In addition, learning tutoring software and teacher tutoring videos have problems such as long content update cycles and insufficient adaptability, making it difficult to meet diverse and rapidly changing learning needs. Based on student-side learning data for manual feedback, an intelligent tutoring system is used to achieve instant matching and sending of learning materials to achieve targeted tutoring. This method realizes the combination of teachers and intelligent systems for online real-time tutoring, and has a certain degree of data interactivity, but relies on teachers' human input and energy consumption, is costly and has poor scalability. Long-term online real-time tutoring may further affect teachers' work efficiency and judgment accuracy, and cannot provide learning support anytime and anywhere.

[0004] In the existing technology, there is a lack of an AI digital tutoring method based on artificial intelligence with high accuracy, high flexibility, high real-time performance and high interactivity. Summary of the Invention

[0005] To address the technical issues of existing technologies, such as a lack of personalization and flexibility, an inability to dynamically adjust to the needs and knowledge levels of different users, a lack of interactivity, an inability to answer users' questions in real time, a long content update cycle, insufficient adaptability, and difficulty meeting diverse and rapidly changing learning needs, the present invention provides an AI digital tutoring method and device based on a large language model. The technical solution is as follows:

[0006] On the one hand, an AI digital tutoring method based on a large language model is provided, the method being implemented by an AI digital tutoring device, the method comprising:

[0007] Use web crawler technology to crawl learning materials and obtain educational resource data; based on the educational resource data, use the large language model to build an educational resource database;

[0008] Obtain student information and student questions; based on the preset educational information prompt template, use the large language model to generate an adaptive style according to the student information to obtain a personalized interaction style;

[0009] Based on the educational resource database, a large language model is used to answer questions according to personalized interaction styles, student information, and student questions, and obtain the answer text;

[0010] Based on the personalized interaction style and the question-answer text, the Edge-TTS module is used for speech generation to obtain teaching audio;

[0011] Obtain a digital human image; use the optimized video generation model to perform inference based on the teaching audio and the digital human image to obtain a silent teaching video;

[0012] Audio synthesis is performed based on silent teaching videos and teaching audios to obtain teaching animations.

[0013] On the other hand, an AI digital tutoring device based on a large language model is provided, which is applied to an AI digital tutoring method based on a large language model, and the device includes:

[0014] The database construction module is used to crawl learning materials using web crawler technology to obtain educational resource data; based on the educational resource data, a large language model is used to build an educational resource database;

[0015] The interaction style generation module is used to obtain student information and student questions. Based on the preset educational information prompt template, it uses a large language model to generate an adaptive style according to the student information to obtain a personalized interaction style.

[0016] The question answering module is used to answer questions based on the educational resource database, personalized interaction style, student information and student questions using a large language model to obtain the answer text;

[0017] The teaching audio generation module is used to generate speech based on the personalized interaction style and the question-answer text using the Edge-TTS module to obtain teaching audio;

[0018] The teaching video generation module is used to obtain the image of the digital human. Based on the teaching audio and the image of the digital human, the optimized video generation model is used for inference to obtain a silent teaching video.

[0019] The teaching animation synthesis module is used to synthesize audio based on silent teaching videos and teaching audio to obtain teaching animations.

[0020] On the other hand, an AI digital tutoring device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned AI digital tutoring methods based on a large language model is implemented.

[0021] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned AI digital tutoring methods based on a large language model.

[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0023] The present invention proposes an AI digital tutoring method based on a large language model. By collecting user voice information, the user's voice is converted into text and sent to the large language model for questioning based on the user's age and personalized needs. The method also connects to an external educational resource database, and through vector similarity matching based on the questions, relevant knowledge is retrieved to provide a text answer to the question. The generated text answer is converted into speech, and a corresponding lip shape change video is generated based on the speech content. The generated speech and lip shape change video are synthesized to generate the final output.

[0024] The AI ​​digital virtual intelligent tutoring assistance system based on artificial intelligence (AI) of the present invention utilizes AI technology and a large language model to achieve highly real-time and highly interactive tutoring. It also provides personalized and highly targeted tutoring based on user characteristics. It allows for flexible tutoring anytime and anywhere, regardless of time or location, and utilizes a broad database of educational resources to achieve highly accurate tutoring. This invention provides an AI digital tutoring method based on AI that is highly accurate, flexible, real-time, and interactive. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 This is a flow chart of an AI digital tutoring method based on a large language model provided by an embodiment of the present invention;

[0027] Figure 2This is a block diagram of an AI digital tutoring device based on a large language model provided by an embodiment of the present invention;

[0028] Figure 3 This is a structural diagram of an AI digital tutoring device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0030] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0031] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.

[0032] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0033] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0034] The embodiment of the present invention provides an AI digital tutoring method based on a large language model, which can be implemented by an AI digital tutoring device, which can be a terminal or a server. Figure 1 The flowchart of the AI ​​digital tutoring method based on a large language model is shown. The processing flow of the method may include the following steps:

[0035] S1. Use web crawler technology to crawl learning materials and obtain educational resource data; based on the educational resource data, use the large language model to build an educational resource database.

[0036] Optionally, based on the educational resource data, a large language model is used to construct an educational resource database, including:

[0037] Perform text segmentation on educational resource data to obtain a collection of resource sentence fragments;

[0038] Based on the nomic-embed-text model, text embedding is performed based on the combination of resource sentence fragments to obtain the resource text embedding set;

[0039] Based on the resource text embedding set, a large model is used for vectorized mapping to obtain a resource text vector set;

[0040] Construct an educational resource database based on the resource text vector set.

[0041] In one feasible implementation, to address the problem of general large-scale model hallucinations and efficiently utilize limited computing resources to achieve excellent performance, the present invention utilizes Retrieval-Augmented Generation (RAG) technology. This invention constructs a specialized educational resource database. This database crawls a rich and diverse range of text content, including but not limited to educational materials, popular science knowledge, literary works, common sense, and other knowledge fragments from various fields. These knowledge fragments are rigorously screened and organized, with high accuracy and authority, providing a comprehensive and reliable data source for subsequent information retrieval.

[0042] The crawled learning materials' text content is segmented and broken down into semantically distinct text segments that meet the required token counts for the embedding model. During the segmentation process, the logical structure and semantic coherence of the text are fully considered, with the text segmented at the sentence level to ensure that each segment independently conveys complete and unambiguous information. The nomic-embed-text model is used for embedding, converting the segmented text segments into high-dimensional vectors and storing them in the ChromaDB vector library. A deep neural network model is used to map the semantic information of the text segments into a high-dimensional vector space, ensuring that text segments with similar semantics are closely spaced in the vector space.

[0043] During the subsequent search process, the system can quickly and efficiently retrieve the most relevant text snippets from the vector library based on the user's input question vector, providing strong support for generating accurate answers. For example, when a user asks, "What are Newton's main contributions?" the system can accurately match text snippets stored in the database on Newtonian mechanics, the law of universal gravitation, and other related knowledge using a cosine similarity algorithm.

[0044] S2. Obtain student information and student questions; based on the preset educational information prompt template and student information, use the large language model to generate an adaptive style to obtain a personalized interaction style.

[0045] In one feasible implementation, a personalized interaction style is adapted based on the user's educational background, personal preferences, and educational scenarios provided by the students. For example, a relaxed and encouraging tone is mainly used for lower grade users, and a formal and serious tone is mainly used for higher grade users. A slightly faster speaking speed is mainly used for simple questions that are easy to understand, and a medium or slow speaking speed is mainly used for complex and difficult to understand questions.

[0046] Among them, student information includes age, gender, voice preference and learning progress.

[0047] In a feasible implementation, key information such as the user's age and educational background is collected through the front-end interactive interface, and used for accurate classification of user categories, so as to provide personalized support for the large language model module and ensure that the generated output content is highly matched with the user's characteristics. For example, when the user is 10 years old and in elementary school, the user information collection module adds a prompt to the large language model module by calling langchain: "I have an elementary school education, please answer the following questions in a relaxed and encouraging tone:..." In addition, the present invention fully considers the needs of user privacy protection, and can minimize the risk of user privacy information leakage by adopting technical means such as encrypted transmission and permission control, thereby ensuring data security and system reliability.

[0048] S3. Based on the educational resource database, a large language model is used to answer questions according to personalized interaction style, student information, and student questions to obtain question answer text.

[0049] Optionally, based on the educational resource database, and according to the personalized interaction style, student information, and student questions, a large language model is used to answer questions and obtain question answer text, including:

[0050] Based on the educational resource database, vector similarity matching is performed according to student questions to obtain background knowledge;

[0051] Based on the preset teaching text template, construct the teaching prompt text according to background knowledge and student information;

[0052] Based on the personalized interaction style, according to the student's questions and teaching prompt text, a large language model is used to perform question response reasoning to obtain the question response text.

[0053] In one feasible implementation, the closest segment vector is found in a vector database based on cosine similarity and identified as background knowledge. This personalized teaching prompt is then generated. This teaching prompt combines the user's student information with the retrieved embedding vectors of the background knowledge related to the question, providing the DeepSeek R1:7B large language model with clear and targeted instructions and sufficient background information. For example, if a 10-year-old elementary school student asks, "What are Newton's main contributions?" the generated prompt is: "I have an elementary school education. Please answer the following question in a relaxed and encouraging tone, using basic concepts such as Newtonian mechanics and the law of universal gravitation, in simple and understandable language: What are Newton's main contributions? Relevant information includes [a brief description of the knowledge content corresponding to the embedding vector of the relevant text segment]." This guides the large language model to generate responses that fully consider the user's comprehension and acceptance level while accurately aligning the output with the relevant knowledge content, ensuring that the generated answers meet the user's personalized needs while being highly accurate and relevant.

[0054] S4. Based on the personalized interaction style and the question-answer text, the Edge-TTS module is used to generate speech and obtain teaching audio.

[0055] Optionally, based on the personalized interaction style and the question-answer text, the Edge-TTS module is used for speech generation to obtain teaching audio, including:

[0056] Extract features from the question answer text to obtain semantic information features;

[0057] Based on the semantic information features, audio signals are converted through deep neural networks to obtain preliminary teaching audio;

[0058] Based on the personalized interaction style, personalized audio is generated according to the preliminary teaching audio to obtain the teaching audio.

[0059] In one feasible implementation, the present invention employs the Edge-TTS module to implement speech generation and optimization. The Edge-TTS module enables efficient, low-latency speech synthesis on the local device without relying on an external server, improving response speed and enhancing the real-time interactive experience.

[0060] The Edge-TTS module receives text input, which can be from question-and-answer content, personalized learning suggestions, reminders, etc. It analyzes the input text using natural language processing (NLP) technology to extract semantic information for subsequent speech synthesis.

[0061] After processing the input text, a series of speech generation algorithms convert it into audio signals. This process involves multiple steps, including phoneme recognition, intonation analysis, rhythm control, and emotional expression. Using models such as deep neural networks (DNNs) or convolutional neural networks (CNNs), Edge-TTS can generate audio output with natural and fluent speech characteristics.

[0062] During the audio output process, various optimization techniques are employed, such as acoustic model optimization, noise suppression, and speech smoothing. These technologies ensure that the generated speech not only has high-quality sound but also accurately expresses specific tone, emotion, and intonation, making the speech output more vivid and authentic, and enhancing user interactivity. The generated speech signal is ultimately played back through the user's device (such as a smart tutoring device, mobile phone, or tablet), providing clear voice feedback. During this process, the Edge-TTS module not only provides clear and coherent speech content but also adjusts parameters such as speech rate, pitch, and volume to suit different user needs and environments.

[0063] In order to achieve personalized and targeted question-and-answer functions, the present invention also generates audio based on personalized interaction styles. The Edge-TTS module can adjust the style of speech synthesis according to the personalized needs of different users. For example, for students of different age groups, the system can generate voice feedback with a slower speaking speed and a gentler tone; while for adult users, speech synthesis may pay more attention to clarity and conciseness. In this way, the smart tutoring system can provide more personalized voice interaction and enhance the user experience. Through this voice interaction design based on Edge-TTS, the present invention can not only provide users with high-quality voice feedback, but also perform customized optimization according to the user's personal characteristics, thereby improving the targetedness and personalization capabilities of the smart tutoring system.

[0064] S5. Obtain a digital human image; use the optimized video generation model to perform inference based on the teaching audio and the digital human image to obtain a silent teaching video.

[0065] Optionally, based on the teaching audio and the digital human image, an optimized video generation model is used for inference to obtain a silent teaching video, including:

[0066] Based on the digital human image, the JoyHallo model is used to generate digital human videos to obtain high-performance training videos.

[0067] Use the training video to optimize the ER-NeRF model and obtain the optimized ER-NeRF model;

[0068] Based on the digital human image and teaching audio, the optimized ER-NeRF model is used for efficient teaching video inference to generate silent teaching videos.

[0069] In a feasible implementation, the present invention uses the Efficient Region-Aware Neural Radiance Fields (ER-NeRF) model as a basic model and integrates the JoyHallo model as its prior setting.

[0070] The JoyHallo model generates a video of the target virtual human by adjusting the corresponding digital mouth shape based on the input audio signal, combining facial expressions and torso movements. This model's key advantage lies in its ability to accurately generate corresponding lip shapes based on Chinese audio input, making it particularly suitable for tutoring tasks and providing highly accurate language synchronization. However, due to the long inference time required by the JoyHallo model, it performs poorly in real-time interactions, making it difficult to quickly respond to user questions and impacting the user experience.

[0071] To overcome this problem, the present invention introduces the ER-NeRF model based on the JoyHallo model, aiming to improve real-time performance and response efficiency. By training on a five-minute speech video of a virtual character, the ER-NeRF model can fully learn the character's voice patterns, facial expressions, and torso movements, generating more natural and smooth virtual human videos. In practice, it is very difficult to obtain a sufficiently long speech video of a specified digital human image, which limits the optimization training of the ER-NeRF model for a specified digital human image. In particular, in the absence of sufficient training data, the model's performance is difficult to achieve as expected.

[0072] Unlike the ER-NeRF model, the JoyHallo model has low data requirements. Therefore, this paper innovatively uses the JoyHallo model to generate avatar videos and uses them as training data for the ER-NeRF model, thus eliminating the ER-NeRF model's need for training video data. This approach not only effectively addresses the problem of insufficient training data but also enables efficient, real-time video generation during the inference phase.

[0073] The fused lip shape video generation model combines the JoyHallo model's high accuracy in generating virtual digital humans with the ER-NeRF model's high efficiency during the inference phase, ensuring the real-time and personalized responsiveness of the virtual intelligent tutoring system. This innovative fusion approach not only achieves industry-leading accuracy in generating virtual digital humans but also significantly improves inference speed, significantly optimizing the user experience of the intelligent tutoring system.

[0074] S6. Perform audio synthesis based on the silent teaching video and the teaching audio to obtain a teaching animation.

[0075] In a feasible implementation, animation synthesis is performed based on silent teaching videos and teaching audio to ensure the coherence and temporal consistency of the generated animation, so that the changes in each frame are synchronized with the audio-driven actions, the time dimension of the video is optimized, and the smooth transition between frames is enhanced.

[0076] In one feasible implementation, the present invention also designs and develops a Streamlit UI interface using the present invention's method, dividing the interface into an input / output interface and a virtual intelligent tutor display interface to enhance user experience and interaction. In the input / output interface, users can choose to interact directly with the digital human through voice input or manually enter text information. When the digital human is dynamically displayed, the input / output interface will display the corresponding text content to help users better understand. The digital human display interface uses dynamic visual and auditory dual presentations and preset backgrounds to achieve efficient, accurate, and realistic interaction with users, enhancing the user's immersive experience and service experience.

[0077] The method proposed in the present invention can also be connected to the interface provided by the video interaction platform to jump into the function page for online one-to-one communication and tutoring with a real teacher, so as to increase the practicality and functionality of the virtual intelligent tutor.

[0078] In a feasible implementation manner, the method proposed in the present invention can also realize the corresponding user management function, where users can send graphic and text questions, online video teaching applications, etc. to online teachers through the module; the classroom management function can obtain real-person teacher information online, including teaching courses, academic qualifications, professor ratings, etc. In addition, it also has the functions of receiving user management module information, sending graphic and text answers to users, and receiving online video teaching applications; by setting the system parameters of the background server of the teacher online tutoring subsystem, the privacy of users and teachers is protected.

[0079] The present invention proposes an AI digital tutoring method based on a large language model. By collecting user voice information, the user's voice is converted into text and sent to the large language model for questioning based on the user's age and personalized needs. The method also connects to an external educational resource database, and through vector similarity matching based on the questions, relevant knowledge is retrieved to provide a text answer to the question. The generated text answer is converted into speech, and a corresponding lip shape change video is generated based on the speech content. The generated speech and lip shape change video are synthesized to generate the final output.

[0080] The AI ​​digital virtual intelligent tutoring assistance system based on artificial intelligence (AI) of the present invention utilizes AI technology and a large language model to achieve highly real-time and highly interactive tutoring. It also provides personalized and highly targeted tutoring based on user characteristics. It allows for flexible tutoring anytime and anywhere, regardless of time or location, and utilizes a broad database of educational resources to achieve highly accurate tutoring. This invention provides an AI digital tutoring method based on AI that is highly accurate, flexible, real-time, and interactive.

[0081] Figure 2 This is a block diagram of an AI digital tutoring device based on a large language model according to an exemplary embodiment. The device is used in an AI digital tutoring method based on a large language model. Figure 2 The device includes a database construction module 210, an interactive style generation module 220, a teaching audio generation module 240, a teaching video generation module 250, and a teaching animation synthesis module 260.

[0082] The database construction module 210 is used to crawl learning materials using web crawler technology to obtain educational resource data; based on the educational resource data, the educational resource database is constructed using a large language model;

[0083] The interaction style generation module 220 is used to obtain student information and student questions; based on the preset educational information prompt template, the large language model is used to generate an adaptive style according to the student information to obtain a personalized interaction style;

[0084] The question answering module 230 is used to answer questions using a large language model based on the educational resource database and the personalized interaction style, student information, and student questions to obtain question answer text;

[0085] The teaching audio generation module 240 is used to generate speech based on the personalized interaction style and the question answer text using the Edge-TTS module to obtain the teaching audio;

[0086] The teaching video generation module 250 is used to obtain a digital human image; based on the teaching audio and the digital human image, the optimized video generation model is used for inference to obtain a silent teaching video;

[0087] The teaching animation synthesis module 260 is used to perform audio synthesis based on the silent teaching video and the teaching audio to obtain the teaching animation.

[0088] Among them, student information includes age, gender, voice preference and learning progress.

[0089] Optionally, the database construction module 210 is further configured to:

[0090] Perform text segmentation on educational resource data to obtain a collection of resource sentence fragments;

[0091] Based on the nomic-embed-text model, text embedding is performed based on the combination of resource sentence fragments to obtain the resource text embedding set;

[0092] Based on the resource text embedding set, a large model is used for vectorized mapping to obtain a resource text vector set;

[0093] Construct an educational resource database based on the resource text vector set.

[0094] Optionally, the question answering module 230 is further configured to:

[0095] Based on the educational resource database, vector similarity matching is performed according to student questions to obtain background knowledge;

[0096] Based on the preset teaching text template, construct the teaching prompt text according to background knowledge and student information;

[0097] Based on the personalized interaction style, according to the student's questions and teaching prompt text, a large language model is used to perform question response reasoning to obtain the question response text.

[0098] Optionally, the teaching audio generation module 240 is further configured to:

[0099] Extract features from the question answer text to obtain semantic information features;

[0100] Based on the semantic information features, audio signals are converted through deep neural networks to obtain preliminary teaching audio;

[0101] Based on the personalized interaction style, personalized audio is generated according to the preliminary teaching audio to obtain the teaching audio.

[0102] Optionally, the teaching video generating module 250 is further configured to:

[0103] Based on the digital human image, the JoyHallo model is used to generate digital human videos to obtain high-performance training videos.

[0104] Use the training video to optimize the ER-NeRF model and obtain the optimized ER-NeRF model;

[0105] Based on the digital human image and teaching audio, the optimized ER-NeRF model is used for efficient teaching video inference to generate silent teaching videos.

[0106] The present invention proposes an AI digital tutoring method based on a large language model. By collecting user voice information, the user's voice is converted into text and sent to the large language model for questioning based on the user's age and personalized needs. The method also connects to an external educational resource database, and through vector similarity matching based on the questions, relevant knowledge is retrieved to provide a text answer to the question. The generated text answer is converted into speech, and a corresponding lip shape change video is generated based on the speech content. The generated speech and lip shape change video are synthesized to generate the final output.

[0107] The AI ​​digital virtual intelligent tutoring assistance system based on artificial intelligence (AI) of the present invention utilizes AI technology and a large language model to achieve highly real-time and highly interactive tutoring. It also provides personalized and highly targeted tutoring based on user characteristics. It allows for flexible tutoring anytime and anywhere, regardless of time or location, and utilizes a broad database of educational resources to achieve highly accurate tutoring. This invention provides an AI digital tutoring method based on AI that is highly accurate, flexible, real-time, and interactive.

[0108] Figure 3 This is a structural diagram of an AI digital tutoring device provided by an embodiment of the present invention. Figure 3 As shown, the AI ​​digital tutoring device may include the above Figure 2 The AI ​​digital tutoring device based on the large language model shown. Optionally, the AI ​​digital tutoring device 310 may include a first processor 2001.

[0109] Optionally, the AI ​​digital tutoring device 310 may further include a memory 2002 and a transceiver 2003 .

[0110] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0111] The following combination Figure 3 The following is a detailed introduction to the various components of the AI ​​digital tutoring device 310:

[0112] The first processor 2001 is the control center of the AI ​​digital tutoring device 310 and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0113] Optionally, the first processor 2001 can perform various functions of the AI ​​digital tutoring device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0114] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 are shown in FIG.

[0115] In a specific implementation, as an embodiment, the AI ​​digital tutoring device 310 may also include multiple processors, such as Figure 3 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0116] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0117] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and accessed through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0118] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0119] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0120] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0121] It should be noted that Figure 3 The structure of the AI ​​digital tutor device 310 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0122] In addition, the technical effects of the AI ​​digital tutoring device 310 can refer to the technical effects of the AI ​​digital tutoring method based on the large language model described in the above method embodiment, and will not be repeated here.

[0123] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0124] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0125] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0126] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0127] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0128] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0129] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0130] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0131] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0132] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0133] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0134] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0135] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. An AI digital tutoring method based on a large language model, characterized in that: The method comprises: Use web crawler technology to crawl learning materials and obtain educational resource data; based on the educational resource data, use the large language model to build an educational resource database; Obtain student information and student questions; based on the preset educational information prompt template, use the large language model to generate an adaptive style according to the student information to obtain a personalized interaction style; Based on the educational resource database, a large language model is used to answer questions according to personalized interaction styles, student information, and student questions, and obtain the answer text; Based on the personalized interaction style and the question-answer text, the Edge-TTS module is used for speech generation to obtain teaching audio; Obtain a digital human image; use the optimized video generation model to perform inference based on the teaching audio and the digital human image to obtain a silent teaching video; Audio synthesis is performed based on silent teaching videos and teaching audios to obtain teaching animations.

2. The AI ​​digital tutoring method based on a large language model according to claim 1 is characterized in that: The student information includes age, gender, voice preference and learning progress.

3. The AI ​​digital tutoring method based on a large language model according to claim 1 is characterized in that: The method of constructing an educational resource database using a large language model based on educational resource data includes: Perform text segmentation on educational resource data to obtain a collection of resource sentence fragments; Based on the nomic-embed-text model, text embedding is performed based on the combination of resource sentence fragments to obtain the resource text embedding set; Based on the resource text embedding set, a large model is used for vectorized mapping to obtain a resource text vector set; Construct an educational resource database based on the resource text vector set.

4. The AI ​​digital tutoring method based on a large language model according to claim 1 is characterized in that: The method uses a large language model to answer questions based on the educational resource database, personalized interaction style, student information, and student questions, and obtains question answer text, including: Based on the educational resource database, vector similarity matching is performed according to student questions to obtain background knowledge; Based on the preset teaching text template, construct the teaching prompt text according to background knowledge and student information; Based on the personalized interaction style, according to the student's questions and teaching prompt text, a large language model is used to perform question response reasoning to obtain the question response text.

5. The AI ​​digital tutoring method based on a large language model according to claim 1 is characterized in that: The method uses the Edge-TTS module to generate speech based on the personalized interaction style and the question-answer text to obtain the teaching audio, including: Extract features from the question answer text to obtain semantic information features; Based on the semantic information features, audio signals are converted through deep neural networks to obtain preliminary teaching audio; Based on the personalized interaction style, personalized audio is generated according to the preliminary teaching audio to obtain the teaching audio.

6. The AI ​​digital tutoring method based on a large language model according to claim 1 is characterized in that: The method uses the optimized video generation model to perform reasoning based on the teaching audio and the digital human image to obtain the silent teaching video, including: Based on the digital human image, the JoyHallo model is used to generate digital human videos to obtain high-performance training videos. Use the training video to optimize the ER-NeRF model and obtain the optimized ER-NeRF model; Based on the digital human image and teaching audio, the optimized ER-NeRF model is used for efficient teaching video inference to generate silent teaching videos.

7. An AI digital tutoring device based on a large language model, wherein the AI ​​digital tutoring device based on a large language model is used to implement the AI ​​digital tutoring method based on a large language model as described in any one of claims 1 to 6, characterized in that: The device comprises: The database construction module is used to crawl learning materials using web crawler technology to obtain educational resource data; based on the educational resource data, a large language model is used to build an educational resource database; The interaction style generation module is used to obtain student information and student questions. Based on the preset educational information prompt template, it uses a large language model to generate an adaptive style according to the student information to obtain a personalized interaction style. The question answering module is used to answer questions based on the educational resource database, personalized interaction style, student information and student questions using a large language model to obtain the answer text; The teaching audio generation module is used to generate speech based on the personalized interaction style and the question-answer text using the Edge-TTS module to obtain teaching audio; The teaching video generation module is used to obtain the image of the digital human. Based on the teaching audio and the image of the digital human, the optimized video generation model is used for inference to obtain a silent teaching video. The teaching animation synthesis module is used to synthesize audio based on silent teaching videos and teaching audio to obtain teaching animations.

8. The AI ​​digital tutoring device based on a large language model according to claim 1, characterized in that: The database construction module is further used to: Perform text segmentation on educational resource data to obtain a collection of resource sentence fragments; Based on the nomic-embed-text model, text embedding is performed based on the combination of resource sentence fragments to obtain the resource text embedding set; Based on the resource text embedding set, a large model is used for vectorized mapping to obtain a resource text vector set; Construct an educational resource database based on the resource text vector set.

9. An AI digital tutoring device, characterized in that: The AI ​​digital tutoring device includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 6.