Natural language understanding method and AI teaching assistant system based on deep learning

Through the natural language understanding method based on deep learning, a knowledge and problem database is built, AI large language model is used for learning and understanding, and various forms of reply are generated, which solves the problems of large amount of data and low efficiency of existing intelligent teaching assistant systems, realizes multimedia free dialogue and immersive interaction, and improves user interactivity and information communication efficiency.

CN117252259BActive Publication Date: 2025-08-22SHANGHAI ZHIZHIQUAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310978221.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-08-22
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

The existing intelligent teaching assistant system requires a large amount of data and insufficient efficiency in the training algorithm, cannot support users' free dialogue, lacks interaction and immersion, and has a single reply form.

Method used

We adopt natural language understanding methods based on deep learning to build a knowledge and problem database, pre-process it through various forms of learning materials and user input information, use AI large language models to learn and understand, generate various forms of replies, and use cloud computing optimization algorithms to support interactive and immersive replies in multiple media.

Benefits of technology

It has achieved the realization of saving computing resources, improving efficiency, supporting free dialogue and interaction in multiple media, enhancing user interaction and information communication efficiency, and is suitable for artificial intelligence teaching assistant system in the VR metaverse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117252259B_ABST
    Figure CN117252259B_ABST
Patent Text Reader

Abstract

The present invention discloses a natural language understanding method and intelligent teaching assistant system based on deep learning. First, a knowledge database and a question database are constructed, and learning material documents are stored in the knowledge database. Preprocessed natural language information is stored in the question database. Then, the natural language information in the question database is learned and understood. Based on the understood content, relevant knowledge points are searched in the knowledge database. The learning material corresponding to the best matching knowledge point is selected as a sample to reply to the natural language information. A record including the question, reply, and evaluation is generated and saved in the knowledge database. Finally, multiple forms of replies are generated and output according to corresponding requirements. This achieves the effect of saving computing resources and improving efficiency, and multiple users can use it in real time without interfering with each other. It also improves the practicality of user interaction and the efficiency of information transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence teaching assistant systems, and specifically relates to a natural language understanding method based on deep learning and an AI teaching assistant system. Background Art

[0002] Currently, many intelligent teaching assistant systems only support answering pre-programmed questions and do not support free conversations with users. Some intelligent teaching assistants are developed based on natural language processing algorithms such as long short-term memory networks and convolutional networks. However, these algorithms require large amounts of new data. The training time and computing resources consumed by these systems limit their expressiveness and potential. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a natural language understanding method based on deep learning, which solves the problem in the prior art that the training algorithm requires a large amount of data and is inefficient.

[0004] The present invention adopts the following technical solutions to solve the above technical problems:

[0005] The natural language understanding method based on deep learning includes the following steps:

[0006] Step 1: Build a knowledge database. First, obtain various forms of learning materials that are pre-stored or uploaded by users, clean and pre-process them, and then summarize and organize the cleaned and pre-processed learning materials into documents and save them into the knowledge database.

[0007] Step 2: Build a question database, pre-process the various forms of natural language information input by the user, and then save the pre-processed natural language information into the question database;

[0008] Step 3: Learn and understand the natural language information in the question database. Based on the understanding, search for relevant knowledge points in the knowledge database. Then, based on one or more scoring or matching algorithms, select the learning materials corresponding to the best matching knowledge points as samples to respond to the natural language information.

[0009] Step 4: Generate a record including questions, replies and comments and save it to the knowledge database;

[0010] Step 5: Generate responses in various forms and output them according to corresponding requirements.

[0011] The learning materials, natural language information input by the user, and the responses include but are not limited to at least one of text, voice, video, and image. The cleaning and pre-processing of the learning materials in step 1 include but are not limited to the following:

[0012] Filtering effective information: Identifying and eliminating invalid, redundant, or irrelevant information from learning materials, retaining only information that contributes to understanding the text content;

[0013] Record and understand knowledge points: record and understand several knowledge points in the learning materials and their internal relationships;

[0014] Marking knowledge categories: Conduct content analysis on learning materials to identify and mark the knowledge categories they cover;

[0015] Preprocessing of video and audio learning materials: For video and audio learning materials, first generate subtitles and perform preprocessing on the subtitles, including but not limited to semantic segmentation, timestamp tagging, and speaker tagging;

[0016] Preprocessing of image-based learning materials: For image-based learning materials, identify and extract information from images, including but not limited to text, features of objects in images, visual elements in images, and various image parameter information, convert this information into text descriptions, and understand and process the text descriptions;

[0017] Standardization: Standardize the learning materials to reduce data noise;

[0018] Noise removal: Identify and remove noise in learning materials, including but not limited to grammatical errors, typos, and irrelevant vocabulary.

[0019] The pre-processing of the natural language information in step 2 includes but is not limited to the following:

[0020] Marking categories of natural language information: Performing content analysis on natural language information input by users to identify and mark the knowledge categories it covers;

[0021] Preprocessing of natural language information from video and speech: First, generate subtitles for natural language information from video and speech, and then preprocess the subtitles, including but not limited to semantic segmentation, timestamp tagging, and speaker tagging.

[0022] Preprocessing of natural language information from images: Identify and extract information from natural language information from images, including but not limited to text, features of objects in the images, visual elements in the images, and various parameter information of the images, convert this information into text descriptions, and understand and process the text descriptions;

[0023] Standardization processing: Standardize text information generated based on natural language information to reduce data noise;

[0024] Noise removal: Identify and remove noise information in natural language information, including but not limited to grammatical errors, typos, and irrelevant vocabulary.

[0025] In step 3, the learning and understanding of the natural language information in the question database includes but is not limited to the following:

[0026] Extract key points: Use AI large language models to learn and extract key points from natural language information;

[0027] Understand key points: Use natural language processing models to understand and record key points.

[0028] In step 3, the learning materials corresponding to the best matching knowledge points are selected as samples to reply to the natural language information, including but not limited to the following parts:

[0029] Query and search for relevant learning materials: Compare the key points in the natural language information with the knowledge points in the knowledge database, and find several knowledge points that are closest to the key points in the vector space;

[0030] Select the best matching learning materials: Compare the key points of the found learning materials with the natural language information and select the best matching learning materials;

[0031] Reply using learning materials: Reply to natural language messages based on the selected best-matching learning materials and the trained AI language model.

[0032] The step 4 also includes a self-learning scoring process, which is as follows:

[0033] First, collect and record user feedback on the responses, including both positive and negative feedback;

[0034] Then, based on the collected feedback, a certain scoring rule is used to score the responses;

[0035] Finally, the scoring results are fed back to optimize the response, including but not limited to adjusting parameter weights, re-understanding the key points of the instruction questions, and regenerating more detailed and accurate responses.

[0036] The responses in step 5 include but are not limited to the following:

[0037] If it is text, it is output directly to the terminal;

[0038] If it is voice, it will be converted through the text-to-speech function and output synchronously in the form of audio;

[0039] If it is learning materials, including but not limited to knowledge graphs and slides, the embedded image generator will be used to generate relevant material output according to needs;

[0040] If it is a video, output the video link or use a small window to play the video.

[0041] In order to further solve the problems of the existing intelligent teaching assistant system, which lacks a sense of interaction and immersion with the intelligent teaching assistant and has a single response form, the present invention also provides an AI teaching assistant system. The specific technical solution is as follows:

[0042] The AI ​​teaching assistant system includes a cloud-based backend and a user terminal; wherein the user terminal collects various command questions and evaluation information of the reply input by the user, and transmits them to the cloud-based backend; the cloud-based backend applies the natural language understanding method based on deep learning to process the command questions and feeds back the reply information to the user terminal; the user chooses whether to evaluate the reply and the content of the evaluation through the terminal based on the received reply information.

[0043] The cloud backend includes a knowledge base storage module, a question input module, a backend learning module, a self-learning scoring module, and a knowledge output module; wherein,

[0044] Knowledge base storage module: used to store knowledge database;

[0045] Question input module: used to receive natural language information sent by the user terminal, including but not limited to text, voice, video, and images, and pre-process the received natural language information;

[0046] Backend learning module: used to learn and understand the natural language information input by users and generate responses in various forms;

[0047] Self-learning scoring module: used to assign weights to backend learning modules and optimize them and their responses;

[0048] Knowledge output module: used to output the responses generated by the back-end learning module to the user terminal.

[0049] The user terminal is any hardware carrier having a user interaction interface, and the hardware carrier is equipped with reply output modules in various forms.

[0050] The user terminal supports users to upload learning materials in various forms including but not limited to text, voice, video, and images, and the cloud backend stores the learning materials uploaded by users into the knowledge database through the knowledge base storage module.

[0051] A computer storage medium, characterized in that: the computer storage medium stores a plurality of computer instructions, and when the computer instructions are called, they are used to execute all or part of the steps of the natural language understanding method based on deep learning.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1. By adopting multiple deep learning natural language processing algorithms, the problem of large amounts of data required for training algorithms and insufficient efficiency is solved, achieving the effect of saving computing resources and improving efficiency.

[0054] 2. By placing algorithm learning and calculation in the cloud, the problem of excessive server load and untimely response caused by multiple processes is solved, and the effect of real-time use by multiple users without interfering with each other is achieved.

[0055] 3. By allowing questions to be asked in multiple media, generating responses using natural language processing, and responding in multiple media, it solves the problems of high costs, insufficient manpower, and low efficiency of human teaching assistants. It also solves the problems of traditional intelligent teaching assistants having a single response format, only understanding pre-programmed questions, and only being able to communicate through text. It achieves the effect that users can freely choose the question method (text, voice questions, screenshot questions) and get easy-to-understand and rich answers.

[0056] 4. By adding Web3.0 technology, the artificial intelligence teaching assistant system is allowed to access the VR metaverse, solving the problem of lack of interaction and immersion with the intelligent teaching assistant, and achieving the effect of improving the practicality of user interaction and the efficiency of information transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a schematic diagram of the functional modules of the teaching assistant system of the present invention.

[0058] Figure 2 A flow chart of a sample acquisition method of the present invention is shown.

[0059] Figure 3 This is a flow chart of the self-learning module of the method of the present invention. DETAILED DESCRIPTION

[0060] The structure and working process of the present invention will be further described below with reference to the accompanying drawings.

[0061] The natural language understanding method based on deep learning includes the following steps:

[0062] Step 1: Build a knowledge database. First, obtain various forms of learning materials that are pre-stored or uploaded by users, clean and pre-process them, and then summarize and organize the cleaned and pre-processed learning materials into documents and save them into the knowledge database.

[0063] Step 2: Build a question database, pre-process the various forms of natural language information input by the user, and then save the pre-processed natural language information into the question database;

[0064] Step 3: Learn and understand the natural language information in the question database. Based on the understanding, search for relevant knowledge points in the knowledge database. Then, based on one or more scoring or matching algorithms, select the learning materials corresponding to the best matching knowledge points as samples to respond to the natural language information.

[0065] Step 4: Generate a record including questions, replies and comments and save it to the knowledge database;

[0066] Step 5: Generate responses in various forms and output them according to corresponding requirements.

[0067] Query and search for relevant learning materials use text embedding technology, which maps text content into vector space and calculates the similarity between text vectors.

[0068] Specific embodiments, such as Figures 1 to 3 As shown,

[0069] The natural language understanding method based on deep learning includes the following steps:

[0070] Step 1: Build a knowledge database and obtain various forms of learning materials that are pre-stored or uploaded by users, including but not limited to text, voice, video, images, etc. Clean and pre-process these learning materials. The specific implementation includes but is not limited to the following steps:

[0071] 1. Filtering effective information: Identify and eliminate invalid, redundant, or irrelevant information from learning materials, including but not limited to modal particles, repeated content, and words and phrases that contribute little to understanding the text, to ensure that only information that contributes to understanding the text is retained;

[0072] Specifically, the contributing information is meaningful information, which is determined by information contribution. The specific determination method includes but is not limited to using word embedding technology to map the text content into a vector space, and determine the information that is closer to the overall content of the text in the vector space, that is, more relevant, that is, to filter out the effective information and only retain the information that contributes to understanding the text content.

[0073] 2. Record and Understand Knowledge Points: Use one or more natural language processing techniques, including but not limited to word embedding, to record and understand several knowledge points and their inherent relationships within the learning materials. Word embedding maps words or phrases from a vocabulary into a vector space. By capturing the semantic and grammatical relationships between words, semantically similar words are placed closer together in the vector space.

[0074] 3. Label knowledge categories: Conduct content analysis on learning materials to identify and label the knowledge categories they cover;

[0075] 4. Preprocessing video and audio learning materials: For video and audio learning materials, generate subtitles using one or more speech recognition technologies and preprocess the subtitle content. Preprocessing includes but is not limited to semantic segmentation, timestamp tagging, speaker tagging, etc.

[0076] 5. Preprocessing image-based learning materials: For image-based learning materials, use one or more image recognition technologies to identify and extract important information from the image, including but not limited to text information, object features, visual elements, and various image parameters. This information is then converted into text descriptions, which are then further understood and processed.

[0077] 6. Standardization: Standardize the learning materials, including but not limited to changing English characters to lowercase, changing Chinese characters to simplified characters, and removing special symbols, to reduce data noise and complexity.

[0078] 7. Remove noise information: Identify and remove other noise information in the learning materials, including but not limited to grammatical errors, typos, irrelevant vocabulary, etc.

[0079] These cleaned and pre-processed learning materials are then further summarized and organized into documents, which are then saved in the knowledge database.

[0080] Step 2: Preprocess the natural language information in various forms input by the user, including but not limited to text, voice, video, and images. The specific implementation includes but is not limited to the following steps:

[0081] 1. Mark the categories of natural language information: Perform content analysis on the natural language information input by the user to identify and mark the knowledge categories it covers;

[0082] 2. Preprocessing of natural language information from video and speech: Generate subtitles using one or more speech recognition technologies and preprocess the subtitle content. Preprocessing includes but is not limited to semantic segmentation, timestamp tagging, speaker tagging, etc.

[0083] 3. Preprocessing natural language information from images: For natural language information from images, use one or more image recognition technologies to identify and extract important information from the image, including but not limited to text information, features of objects in the image, visual elements in the image, and various image parameters. This information is then converted into text descriptions, which are then further understood and processed.

[0084] 4. Standardization: Standardize text information generated based on natural language information, including but not limited to lowercase English characters, simplified Chinese characters, and removal of special symbols, to reduce data noise and complexity.

[0085] 5. Remove noise information: Identify and remove other noise information in natural language information, including but not limited to grammatical errors, typos, irrelevant words, etc.

[0086] The preprocessed natural language information is then saved into the question database.

[0087] Step 3: Learn and understand the natural language information in the question database. The specific implementation includes but is not limited to the following steps:

[0088] 1. Extract key points: Use one or more AI language models, including but not limited to GPT, to learn and extract key points from natural language information.

[0089] 2. Understand key points: Use one or more natural language processing technologies, including but not limited to word embedding technology, to understand and record each key point.

[0090] Based on the content of understanding, relevant knowledge points are searched in the knowledge database. Then, based on one or more scoring or matching algorithms, the learning materials corresponding to the best matching knowledge points are selected as samples to respond to the natural language message. The specific implementation of selecting the learning materials corresponding to the best matching knowledge points as samples to respond to the natural language message includes but is not limited to the following steps:

[0091] 1. Query and search for relevant learning materials: Use one or more algorithms, including but not limited to the cosine similarity algorithm, to compare the key points in the natural language information with the knowledge points in the knowledge database, and find several knowledge points that are closest to the key points in the vector space. The cosine similarity algorithm is a measurement method for calculating the cosine value of the angle between two vectors, and is usually used to evaluate the similarity between two vectors. The common form of its mathematical formula is: cos(θ) = (A·B) / (||A||||B||). Among them, θ is the angle between the two vectors, A·B is the dot product of vector A and vector B, ||A|| and ||B|| are the moduli of vector A and vector B;

[0092] 2. Select the best matching learning materials: Use one or more algorithms, including but not limited to AI large language models, to compare the key points in the found learning materials with the natural language information to select the best matching learning materials;

[0093] 3. Reply using learning data: Using one or more algorithms, including but not limited to few-shot learning, to respond to natural language messages based on selected, best-matched learning data and a pre-trained AI language model. Few-shot learning refers to machine learning techniques that can make accurate predictions even with a small number of training samples, by combining them with a pre-trained AI language model.

[0094] Step 4: Generate a record including the question, answer, and evaluation, and save it to the knowledge database; this step also includes the self-learning scoring process, as follows:

[0095] 1. Collect feedback: Collect and record user feedback on responses, including both positive and negative feedback;

[0096] 2. Rating calculation: Based on the collected feedback, a certain scoring algorithm is used to score the responses;

[0097] 3. Feedback: The scoring results will be used to optimize the response, including but not limited to adjusting parameter weights, re-understanding the key points of the instruction questions, and regenerating more detailed and accurate responses.

[0098] Step 5: Generate responses in various forms, including but not limited to text, voice, video, and images, and output them according to the corresponding requirements.

[0099] 1. If it is text, it will be output directly to the terminal;

[0100] Second, if it is voice, it will be converted through the text-to-speech function and output synchronously in the form of audio;

[0101] 3. If it is learning materials, including but not limited to knowledge graphs, slides, etc., the embedded image generator will be used to generate relevant material output according to needs;

[0102] 4. If it is a video, output the video link or use a small window to play the video.

[0103] The AI ​​teaching assistant system includes a cloud-based backend and a user terminal; wherein the user terminal collects various command questions and evaluation information of the reply input by the user, and transmits them to the cloud-based backend; the cloud-based backend applies the natural language understanding method based on deep learning to process the command questions and feeds back the reply information to the user terminal; the user chooses whether to evaluate the reply and the content of the evaluation through the terminal based on the received reply information.

[0104] The cloud backend includes a knowledge base storage module, a question input module, a backend learning module, a self-learning scoring module, and a knowledge output module; wherein,

[0105] Knowledge base storage module: used to store learning materials or school course content on online learning platforms. These learning materials can be in the form of text, images, videos, etc., especially the materials on online learning platforms are mainly recorded videos. These materials will go through a series of cleaning and pre-processing steps. These include but are not limited to filtering out valid information, recording knowledge point keywords, marking the knowledge categories covered by images and videos, generating subtitles and marking timestamps for videos, extracting text from images, etc. The next step is to summarize and organize all data into documents and save them in the database. Data can be stored according to knowledge categories, but a more practical method is to store them according to the uploading institution and school. This ensures that the answers users receive are relevant to their affiliated institution or school.

[0106] The question input module receives natural language information sent by user terminals, including but not limited to text, voice, video, and images, and pre-processes the received natural language information. For example, it collects and records questions posed by students to teaching assistants or instructions given by teachers to teaching assistants. Instructions can take various forms, including text, voice, video, and images. Text instructions go directly to the pre-processing step, while voice instructions are converted to text using speech recognition. The teaching assistant recognizes the uploaded image, extracts the text, and then sends it to the pre-processing stage. The video is split into audio and image information and processed separately. During the pre-processing stage, the teaching assistant identifies valid instructions and other noise. Valid instructions are transmitted back to the back-end learning module.

[0107] Back-end learning module: This module is used to learn and understand the natural language input from users and generate responses in various forms. For example, it is used to learn and understand the aforementioned word sequence and search the knowledge database for relevant knowledge points based on the understood meaning. The back-end learning module includes two established deep learning natural language understanding algorithms: GPT-3 and BERT. The BERT algorithm can learn instructions bidirectionally and capture keywords. It can also process multiple lines of instructions simultaneously, learning and understanding the key points and functions of each line and selecting the most appropriate response. GPT-3 uses a small sample learning method, which consumes fewer resources to learn user instructions. GPT-3 also assigns a weight to each possible response and transmits it back to the self-learning scoring module.

[0108] The self-learning scoring module is used to assign weights to the backend learning module, optimize the backend learning module and its responses, and assign weights to the responses output by the backend learning module. The module uses ensemble learning to integrate the results of two language processing algorithms into several classifiers. Each classifier fits the command and response using a linear model and returns the result with the smallest mean square error. Ensemble learning summarizes the results of all classifiers and selects the most likely result to pass to the text and voice output modules. The frontend also includes a scoring system, allowing users to rate each response sent by the teaching assistant. Whether positive or negative, the results are fed back to the self-learning scoring module for self-optimization. Optimization includes adjusting parameter weights, re-understanding the key points of the user's question, and regenerating more detailed responses. A record containing the user's question, the teaching assistant's response, and the user's evaluation is generated and saved in the storage module.

[0109] Knowledge output module: used to output the responses generated by the back-end learning module to the user terminal. For example, it is used to convey the content generated by the teaching assistant to the user. The responses that the teaching assistant can generate can also be in various forms, including text, audio, images, and videos. The text part will be directly transmitted back to the interactive interface. The voice part will be synchronously output in the form of sound through the text-to-speech service. In some cases, such as when the user requires the generation of a knowledge tree or a knowledge graph, the embedded image generator can be used to generate relevant images based on the user's needs. The teaching assistant can also return a hyperlink or play in a small window to display a video clip according to the situation. The content of this video clip can solve the questions raised by the user and is easier to understand than pure text. This video can be a recorded course video or a course video from an online learning platform.

[0110] The user terminal is any hardware carrier having a user interaction interface, and the hardware carrier is equipped with reply output modules in various forms.

[0111] The user terminal supports users to upload learning materials in various forms including but not limited to text, voice, video, and images, and the cloud backend stores the learning materials uploaded by users into the knowledge database through the knowledge base storage module.

[0112] To further illustrate the solution, the following is a detailed description of the specific implementation process using a specific example:

[0113] Step 1: Clean and pre-process the pre-collected learning materials, which specifically includes (taking video learning materials as an example):

[0114] Step 1.1: Use one or more speech recognition technologies to generate subtitles for the video learning materials, record the start and end timestamps of each word, and mark the sound source;

[0115] Step 1.2: Use the pre-trained AI language model to segment the subtitles into sentences and remove noise information (such as grammatical errors, typos, and irrelevant words such as modal particles).

[0116] Step 1.3: Use natural language processing technology to standardize each complete sentence (including but not limited to converting English characters to lowercase, converting Chinese characters to simplified characters, removing special symbols, etc.) and filter out valid information (including but not limited to removing modal particles, removing repeated sentences, removing sentences with too many repeated words, etc.);

[0117] Step 1.4: Based on the records from step 1.1, add start and end timestamps and a sound source tag to each cleaned complete sentence. Also include all relevant information about the knowledge category, including but not limited to the video title, author, creation date, video URL, video cover, and the corresponding course. Combine all subtitles of the video learning material into a complete text and summarize its content using an AI large language model.

[0118] Step 1.5: Use one or more word embedding models to map each complete sentence into a different vector space and record all the vector information of each complete sentence.

[0119] Step 1.6: The cleaned and pre-processed learning materials are further summarized and organized according to the information related to the knowledge category to form documents, and these documents are saved in the knowledge database;

[0120] Step 2: The user selects course information (for example, the user selects a knowledge category contained in a knowledge database, such as English teaching);

[0121] Step 3: The user inputs natural language information (Example: What are the test points for Expression and Majority? Please summarize the content of the video "Day 3". Reply in the form of a list.);

[0122] Step 4: Use the question input module to pre-process the "natural language information input by the user". The specific steps include:

[0123] Step 4.1: Use natural language processing technology to standardize the "natural language information input by the user" (including but not limited to changing English characters to lowercase, changing Chinese characters to simplified characters, and removing special symbols) and remove noise information (including but not limited to grammatical errors, typos, irrelevant words, etc.)

[0124] Step 4.2: Match, identify, and mark the knowledge category corresponding to the "natural language information input by the user" based on the course information selected by the user in step 2.

[0125] Step 5: Use the backend learning module to learn, understand, and analyze the "natural language information input by the user" and ultimately generate responses in various forms. The specific steps include:

[0126] Step 5.1: Use the pre-trained AI large language model to identify, separate, and store the "question" and "request for a response" in the "natural language information input by the user." (Example: "Question": What are the test points for Expression and Majority? Please summarize the content of the video "Day 3." "Request for a response": Reply in the form of a list.)

[0127] Step 5.2: Use the pre-trained AI language model to learn, understand, and analyze the "question" and then break it down into several "independent questions." (Example: "Independent questions": 1. What are the key points for "Expression"? 2. What are the key points for "Majority"? 3. What was discussed in the video "Day 3"?)

[0128] Step 5.3: Use the pre-trained AI language model to learn, understand, and analyze each "independent question" to determine its type of question, including but not limited to "questions about a specific knowledge point" or "questions summarizing a particular learning resource." Based on this determination, the best matching learning resource type is selected. (Example: "Questions about a specific knowledge point": 1. What is the key point of Expression? 2. What is the key point of Majority?; "Questions summarizing a particular learning resource": 1. What is covered in the video "Day 3"?)

[0129] Step 5.4: For each "independent question", if the "independent question" is a "question about a specific knowledge point", use the following steps to generate a response to the "independent question":

[0130] Step 5.4.1: Use one or more word embedding models to map the “independent question” to the corresponding vector space; (Example: Use several word embedding models in the Python framework of Sentence Transformers)

[0131] Step 5.4.2: Use the cosine similarity algorithm to compare the "independent question" with each knowledge point in the knowledge database. Use one or more evaluation criteria to select the knowledge points that are closest to the "independent question" in each evaluation criterion and record them; (Example: Calculate the distance between each sentence in the learning material and the "independent question" using the cosine similarity algorithm, select the first few sentences closest to the "independent question" in each model, and record their positions in the learning material. Calculate the distance between each sentence in the learning material and the "independent question" using the cosine similarity algorithm, select the first few sentences closest to the average value of the distance between a sentence and the next N consecutive sentences in each model and the "independent question", and record their positions in the learning material.)

[0132] Step 5.4.3: Using the pre-trained AI large language model, compare all the knowledge points recorded in step 3.3.2 and the portions of the learning materials in which they are located with the "independent question," and select one or more best-matching learning materials. (Example: Based on the locations of all the knowledge points recorded in step 5.4.2, combine several sentences of learning materials near their locations as a sample paragraph, and use the AI ​​large language model to compare all the sample paragraphs with the "independent question" to select one or more best-matching sample paragraphs.)

[0133] Step 5.4.4: Use the small sample learning method and the pre-trained AI large prediction model, utilizing the best matching learning data combined with the "requirements for responses" in step 5.1 to generate a response to the "independent question." (Example: Use the best matching sample paragraph selected in step 5.4.3 as a sample to train the AI ​​large prediction model, and then generate a response to the "independent question" based on this sample paragraph.)

[0134] Step 5.5: If the "independent question" is a "summary and conclusion question about a particular learning material," use the following steps to generate a response to the "independent question":

[0135] Step 5.5.1: Use one or more word embedding models to map the “independent question” to the corresponding vector space; (Example: Use several word embedding models in the Python framework of Sentence Transformers)

[0136] Step 5.5.2: Use the cosine similarity algorithm to compare the "independent question" with the learning materials in the knowledge database. Using one or more evaluation criteria, select the learning materials that are closest to the "independent question" for each evaluation criterion and record them. (Example: Using the cosine similarity algorithm, calculate the distance between the knowledge category-related information in the learning materials, including but not limited to the video name, author, creation date, video URL, video cover, corresponding course, etc., and the "independent question". Select the first few learning materials closest to the "independent question" in each model and record them.)

[0137] Step 5.5.3: Use the pre-trained AI large language model to compare all the learning materials and their knowledge category-related information recorded in step 3.4.2 with the "independent question" and select one or more best-matching learning materials. (Example: Based on the knowledge category-related information recorded in step 5.5.2, and using the AI ​​large language model to compare the knowledge category-related information with the "independent question," select one or more best-matching learning materials. "For a summary question of a particular learning material": 1. What does the video "Day 3" cover? The learning material with the video title "Day 3" is selected as the best-matching learning material.)

[0138] Step 5.5.4: Use the small sample learning method and the pre-trained AI large prediction model to generate a response to the "independent question" using the best matching learning material and its pre-summarized and summarized content, combined with the "Response Requirements" in Step 2.1. (Example: Use the pre-summarized and summarized content of the best matching learning material selected in Step 5.5.3 as a sample to train the AI ​​large prediction model, and then generate a response to the "independent question" based on this sample paragraph.)

[0139] Step 5.6: Integrate all generated responses and generate a complete response. The specific steps include:

[0140] Step 5.6.1. Combine the best matching learning material knowledge and relevant information of each "independent question" corresponding to the paragraph, including but not limited to the video timestamp, video link, video author, creation date, etc., with the response generated in step 5.5 to form a complete response; (Example: Expression test points:

[0141] 1. The meaning expressed;

[0142] 2. The meaning of the expression;

[0143] 3. The specific embodiment of key elements of values;

[0144] 4. Specific manifestation, translated as of;

[0145] 5. The adjective "expressive" means expressive;

[0146] 6. The adjective expressible means able to express clearly;

[0147] 7. The negative adjective is inexpressible, which means something that cannot be expressed in words.

[0148] This answer is based on the video "Day 3".

[0149] The test points of Majority are as follows:

[0150] 1. Majority: Most, often used to describe the proportion of quantity or number of people.

[0151] 2.minority: minority, opposite to majority.

[0152] 3.a majority of: most, often used to modify nouns.

[0153] 4. take something seriously: take something seriously, often used to express attitude or behavior.

[0154] 5.it is obvious that: obviously, often used to introduce an obvious point of view or conclusion.

[0155] 6.System / Systematic: System / systematic, often used to describe a system or method.

[0156] 7.Warning System: Warning system, often used to describe a certain warning mechanism.

[0157] 8. Red Alert: Red alert, often used to indicate the highest level of danger.

[0158] 9.Endangered Species: Endangered species, often used to describe the protection of biodiversity.

[0159] 10.Systematic Survey Methods: Systematic survey methods, often used to describe scientific research methods.

[0160] 11. Systematic drug abuse: Systematic drug abuse, often used to describe the scale and impact of drug abuse.

[0161] This answer is based on the video "Day 2".

[0162] -Introduced vocabulary related to value, such as value, evaluate, valuable, etc.

[0163] -Discussed the best way of education and the concept of school district housing.

[0164] -It talks about common prefixes in English and some words related to value, work, production, etc.

[0165] - Explains the concepts of devaluation and undervaluation and their differences, as well as the subjectivity of value and how to quantify it.

[0166] -Discussed the dedication and personal choices of scientific research, as well as related vocabulary such as overvalue and toxic assets. -Introduced vocabulary and collocations related to the head, as well as vocabulary related to titles, and emphasized the simple rules of verbs modifying nouns.

[0167] -Explained several economic and financial related words, such as clickbait, headlong, overhead, finance and financial.

[0168] -Introduces some vocabulary related to finance and technology, such as financial, fiscal, technology, high technology and technological.

[0169] -Discusses English words and expressions related to consumption and shopping, such as "consumed sparingly", "consuming", "style conscious consumers" and "consumerism".

[0170] - Mentioned terms related to online shopping, such as "e-commerce platform" and "shop around," and touched on privacy issues in modern society. - Explained the meaning and usage of the word "assume," as well as related vocabulary and expressions.

[0171] -Introduces English vocabulary related to the courier service industry.

[0172] - The course discusses the characteristics of poetry in expressing emotions and the meaning and usage of some vocabulary, including expressible, inexpressible, explicable, inexplicable, call, call out, issue, recall, etc. - The course introduces vocabulary related to insurance and investment and their meanings.

[0173] -Choose your own lifestyle test point.

[0174] -Insurance in English is insurance, assurance means self-confirmation, investment in English is invest, and fund in English is fund.

[0175] -The English word for choice is choose. To choose one’s own lifestyle can be expressed as choose their own way of life.

[0176] -Making friends is important. Policy refers to policies and guidelines. Address means solving problems. Thorny questions refer to truly difficult problems.

[0177] -The lecture also emphasized the importance of independent thinking and not following the crowd, reminding young people not to be influenced by peer pressure and to choose the path that suits them.

[0178] -Introduces the meaning and usage of some English words and phrases, including headhunter, explore, fine, search, seek, find, job seeker, official, gain, benefit, formation flight, eraser, eradicate, ready, radiate, etc.

[0179] - The proverb "No pain, no gain" is quoted. The correct usage is "No pain, no gain", not "No pains, no gains". The content of the video "Day 3" is summarized as follows:

[0180] -When using English, be sure to use common expressions rather than unconventional ones.

[0181] -Introduce the concept of happiness in the movie "The Pursuit of Happiness", that is, happiness is a process of acquisition, called "The Pursuit of Happiness".

[0182] -The letter I in the word Happiness in the movie also expresses the meaning of happiness, that is, the answer to happiness lies in oneself, and one needs to find it by oneself, rather than looking for reasons.

[0183] -The clip ends with a line from the movie: The answer lies within yourself, It is an eye in happiness.

[0184] This answer is based on the video "Day 3".

[0185] Step 6: Use the knowledge output module to output the responses generated by the backend learning module to the user terminal;

[0186] Step 7: Users can rate the generated responses. The self-learning scoring module will record the scores and provide feedback to the back-end learning module to optimize responses to similar questions.

[0187] Step 8: Users can request related learning materials based on their responses. The specific steps are as follows:

[0188] Step 8.1. Select the required learning material category, such as knowledge graph, slides, etc.

[0189] Step 8.2: Integrate the subtitles of one or more of the best learning materials used in the reply into a complete text. Use the AI ​​language model to generate content corresponding to the learning material category.

[0190] Step 8.3: Import the content generated by the AI ​​Big Prediction Model into the generator of the corresponding learning material category. And export the corresponding learning material;

[0191] Step 8.4: Send the learning materials to the user.

[0192] This solution also provides an AI teaching assistant system that integrates multiple deep learning natural language understanding algorithms and Web3.0 technologies. The AI ​​teaching assistant system has the following components: a question input module, a backend learning module, a knowledge base storage module, a self-learning scoring module, and a knowledge output module.

[0193] The question input module collects and records questions posed by students to teaching assistants or instructions given by teachers to them. Instructions can take various forms, including text, voice, video, and images. Text instructions go directly to the preprocessing step, while voice instructions are converted to text using speech recognition. The teaching assistant recognizes uploaded images, extracts the text, and then sends them to the preprocessing stage. Videos are split into audio and image components and processed separately. During the preprocessing stage, the teaching assistant identifies valid instructions and other noise. Valid instructions are then transmitted back to the backend learning module.

[0194] The back-end learning module is responsible for learning and understanding the aforementioned word sequence and searching the knowledge database for relevant knowledge points based on the understood meaning. This module includes two improved deep learning natural language understanding algorithms: an improved self-attention-based GPT model and an improved BERT model. The BERT algorithm can bidirectionally learn instructions and extract key words. It can also process multiple lines of instructions simultaneously, learning and understanding the key points and purpose of each line and selecting the most appropriate response. GPT-3 uses a small-sample learning approach, consuming fewer resources to learn user instructions. GPT-3 also assigns a weight to each possible response and transmits it back to the self-learning scoring module.

[0195] The knowledge base storage module is used to store learning materials from online learning platforms or school course content. These learning materials can be in the form of text, images, videos, etc., especially the materials on online learning platforms are mainly recorded videos. These materials will go through a series of cleaning and pre-processing steps. These include but are not limited to filtering out valid information, recording knowledge point keywords, marking the knowledge categories covered by images and videos, generating subtitles and marking timestamps for videos, extracting text from images, etc. The next step is to summarize and organize all data into documents and save them in the database. Data can be stored according to knowledge categories, but a more practical method is to store it according to the uploading institution and school. This ensures that the answers users receive are relevant to their affiliated institution or school.

[0196] The self-learning scoring module assigns weights to the responses output by the back-end learning module. The module uses ensemble learning to place the results of the two language processing algorithms into several classifiers. Each classifier uses a linear model to fit the instructions and responses, and then returns the result with the smallest mean square error. Ensemble learning summarizes the results of all classifiers and then selects the most likely result to pass to the text and voice output modules. At the same time, the front-end is equipped with a scoring system, allowing users to evaluate each response sent by the teaching assistant. Whether positive or negative, the results will be sent back to the self-learning scoring module for self-optimization. Optimization includes adjusting parameter weights, re-understanding the key points of the user's question, and regenerating more detailed responses. A record containing the user's question, the response generated by the teaching assistant, and the user's evaluation will be generated and saved in the storage module.

[0197] The knowledge output module is used to convey the content generated by the teaching assistant to the user. The responses that the teaching assistant can generate can also be in various forms, including text, audio, images, and videos. The text part will be directly transmitted back to the interactive interface. The voice part will be output synchronously in the form of sound through the text-to-speech service. In some cases, such as when the user requires the generation of a knowledge tree or knowledge graph, the embedded image generator can be used to generate relevant images based on the user's needs. The teaching assistant can also return a hyperlink or play a small window to display a video clip according to the situation. The content of this video clip can solve the questions raised by the user and is easier to understand than pure text. This video can be a recorded course video or a course video from an online learning platform.

[0198] In this embodiment,

[0199] Command inputs can be received via voice, text, images, and videos. After preprocessing, this information is sent to the cloud backend. Multiple deep learning natural language understanding algorithms connected to the cloud process the information into a sequence of strings and begin learning. The algorithms use an attention mechanism to identify key words in the command and calculate the likelihood of multiple possible responses stored in a module. These responses are also learned alongside the command, using a feedback mechanism to help the algorithm identify which response is most appropriate.

[0200] In addition to conventional text and voice replies, the present invention also generates a knowledge graph to help users systematically understand knowledge concepts.

[0201] The raw data, which can be from course websites or recorded videos, is scraped and initially organized using a crawler. Text data is tagged, and subtitles are recognized for video data. The subtitles are saved as documents, with each sentence timestamped. All of this data is saved as CSV files and then cleaned. Data cleaning uses methods such as information entropy or cosine distance to remove meaningless text from the documents, retaining only the core knowledge and relevant information. These organized documents are then fed into the backend deep learning model.

[0202] After receiving feedback from the intelligent teaching assistant, users can rate their satisfaction. This rating, along with the user's question and the machine's feedback, is integrated into a JSON file and sent to the backend. Based on the question and rating, the learning module adjusts the weighting of various parameters in the generated response and relearns to understand the user's question. It then generates a new response based on the adjusted parameters and presents it to the user.

[0203] This invention integrates multiple deep learning language understanding models and uses small-sample learning methods to reduce resource consumption, improve operational efficiency, and generate more natural sentences. Users can freely communicate with the teaching assistant of this invention and ask questions in any tone and manner. The intelligent teaching assistant uses natural language models to understand which parts of the message are questions and which parts are free conversations, and respond accordingly.

[0204] The present invention also expands the information carriers that can be processed, and can accept not only text and voice, but also images and videos, thereby improving the practicality of the intelligent teaching assistant system, the diversity of processing tasks and the portability of user use.

[0205] In addition, the intelligent teaching assistant of the present invention can also use these information carriers as replies.

[0206] Compared with traditional intelligent teaching assistants that can generally only reply with text, and some support playing sound replies, the present invention can use images and videos to provide a richer user experience.

[0207] The present invention also allows for embedding into more usage scenarios. Not only can it be embedded in online learning platforms as a small assistant or used independently on a web page, but it can also be applied to the virtual environment of the metaverse, interacting with users through modeling images and providing an immersive experience.

[0208] The following is a detailed description of the solution using the example of using it in a Metaverse classroom.

[0209] First, students can access the intelligent conversational robot through interactive web pages within the Metaverse and send commands through various channels. In this virtual immersive environment, students can interact with the robot by typing text, speaking, drawing, sending pictures, and more. These commands are transmitted to the cloud backend for processing. Finally, the robot responds to students synchronously with voice and text, answering their questions or providing requested information.

[0210] For example, if a student wants to learn more about a course, they can first enter the course code and then say, "I want to know more about this course." The robot will learn and understand this command based on these two pieces of information. Once the robot understands the command, it will search the data storage module for relevant course data. This data can include information provided by the school, feedback from students who have taken the course, and so on. The robot will then select the most likely response and output it to the student. Finally, the robot will display the response on the interactive screen and read it aloud to the student. The student can then rate the response and provide feedback to the robot on their satisfaction level.

[0211] A computer storage medium storing a plurality of computer instructions, wherein the computer instructions are used to execute all or part of the steps of the deep learning-based natural language understanding method when called.

[0212] This invention integrates multiple deep learning language understanding models and uses small-sample learning methods to reduce resource consumption, improve operational efficiency, and generate more natural sentences. Users can freely communicate with the teaching assistant of this invention, freely asking questions in any tone and manner. The intelligent teaching assistant uses natural language models to understand which part of the information is a question and which is a free conversation, and responds accordingly. This invention also expands the information carriers that can be processed, accepting not only text and voice but also images and videos, improving the practicality, task diversity, and portability of the intelligent teaching assistant system. Furthermore, the intelligent teaching assistant of this invention can also use these information carriers as replies. Compared to traditional intelligent teaching assistants, which generally only reply with text and some support audio playback, this invention can use images and videos to provide a richer user experience. This invention also allows for integration into a wider range of use cases. It can not only be embedded in online learning platforms as a small assistant or used independently on a web page, but can also be applied to the virtual environment of the metaverse, interacting with users through modeled characters and providing an immersive experience.

Claims

1. A natural language understanding method based on deep learning, characterized by: The steps include: Step 1: Build a knowledge database. First, obtain various forms of learning materials that are pre-stored or uploaded by users, and clean and pre-process these learning materials. Then, summarize and organize the cleaned and pre-processed learning materials to form documents and save them into the knowledge database. The learning materials, natural language information input by users, and responses all include but are not limited to at least one of text, voice, video, and image formats. The cleaning and pre-processing of learning materials in step 1 include but are not limited to the following: Filtering effective information: Identifying and eliminating invalid, redundant, or irrelevant information from learning materials, retaining only information that contributes to understanding the text content; Record and understand knowledge points: record and understand several knowledge points in the learning materials and their internal relationships; Marking knowledge categories: Conduct content analysis on learning materials to identify and mark the knowledge categories they cover; Preprocessing of video and audio learning materials: For video and audio learning materials, first generate subtitles and perform preprocessing on the subtitles, including but not limited to semantic segmentation, timestamp tagging, and speaker tagging; Preprocessing of image-based learning materials: For image-based learning materials, identify and extract information from images, including but not limited to text, features of objects in images, visual elements in images, and various image parameter information, convert this information into text descriptions, and understand and process the text descriptions; Standardization: Standardize the learning materials to reduce data noise; Noise removal: Identify and remove noise in learning materials, including but not limited to grammatical errors, typos, and irrelevant vocabulary; Step 2: Build a question database, pre-process the various forms of natural language information input by the user, and then save the pre-processed natural language information into the question database; Step 3: Learn and understand the natural language information in the question database. Based on the understanding, search for related knowledge points in the knowledge database. Then, based on one or more scoring or matching algorithms, select the learning materials corresponding to the best matching knowledge points as samples to respond to the natural language information. The learning materials corresponding to the best matching knowledge points as samples to respond to the natural language information include but are not limited to the following: Query and search for relevant learning materials: Compare the key points in the natural language information with the knowledge points in the knowledge database, and find several knowledge points that are closest to the key points in the vector space; Select the best matching learning materials: Compare the key points of the found learning materials with the natural language information and select the best matching learning materials; Reply using learning materials: Reply to natural language messages based on the selected most suitable learning materials and the trained AI language model; Step 4: Generate a record including questions, replies and comments and save it to the knowledge database; Step 5: Generate responses in various forms and output them according to corresponding requirements.

2. The natural language understanding method based on deep learning according to claim 1, characterized in that: In step 3, the learning and understanding of the natural language information in the question database includes but is not limited to the following: Extract key points: Use AI large language models to learn and extract key points from natural language information; Understand key points: Use natural language processing models to understand and record key points.

3. The natural language understanding method based on deep learning according to claim 1, characterized in that: The step 4 also includes a self-learning scoring process, which is as follows: First, collect and record user feedback on the responses, including both positive and negative feedback; Then, based on the collected feedback, a certain scoring rule is used to score the responses; Finally, the scoring results are fed back to optimize the response, including but not limited to adjusting parameter weights, re-understanding the key points of the instruction questions, and regenerating more detailed and accurate responses.

4. The natural language understanding method based on deep learning according to claim 1, characterized in that: The responses in step 5 include but are not limited to the following: If it is text, it is output directly to the terminal; If it is voice, it will be converted through the text-to-speech function and output synchronously in the form of audio; If it is learning materials, including but not limited to knowledge graphs and slides, the embedded image generator will be used to generate relevant material output according to needs; If it is a video, output the video link or use a small window to play the video.

5. AI teaching assistant system, characterized by: It includes a cloud backend and a user terminal; wherein the user terminal collects various command questions input by the user and evaluation information of the reply, and transmits them to the cloud backend; the cloud backend applies the natural language understanding method based on deep learning as described in any one of claims 1 to 4 to process the command questions, and feeds back the reply information to the user terminal; the user chooses whether to evaluate the reply and the content of the evaluation through the terminal based on the received reply information.

6. The AI ​​teaching assistant system according to claim 5, characterized in that: The cloud backend includes a knowledge base storage module, a question input module, a backend learning module, a self-learning scoring module, and a knowledge output module; wherein, Knowledge base storage module: used to store knowledge database; Question input module: used to receive natural language information sent by the user terminal, including but not limited to text, voice, video, and images, and pre-process the received natural language information; Backend learning module: used to learn and understand the natural language information input by users and generate responses in various forms; Self-learning scoring module: used to assign weights to backend learning modules and optimize them and their responses; Knowledge output module: used to output the responses generated by the back-end learning module to the user terminal.

7. The AI ​​teaching assistant system according to claim 5, characterized in that: The user terminal is any hardware carrier having a user interaction interface, and the hardware carrier is equipped with reply output modules in various forms.

8. The AI ​​teaching assistant system according to claim 5, characterized in that: The user terminal supports users to upload learning materials in various forms including but not limited to text, voice, video, and images, and the cloud backend stores the learning materials uploaded by users into the knowledge database through the knowledge base storage module.

9. A computer storage medium, characterized in that: The computer storage medium stores a plurality of computer instructions, which, when called, are used to execute all or part of the steps of the deep learning-based natural language understanding method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method for constructing domain intelligent question-answering system based on zero sample of knowledge graph

    CN116383352A

  • Chat robot and method integrating natural language understanding algorithm and Web3.0 technology

    CN116521831A