Teaching aid robot intelligent dialogue system and method
By using large models and artificial intelligence technology to perform voice recognition and semantic embedding coding on students' questions, combined with students' basic information, the problem that traditional systems cannot deeply understand students' intentions is solved, more comprehensive and accurate answers are generated, and learning efficiency is improved.
Patent Information
- Application Number
- CN202510805824.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional educational assistance systems find it difficult to deeply analyze the true intentions behind students' questions and are unable to effectively expand and improve the content of questions, resulting in answers that lack depth and comprehensiveness, affecting learning efficiency.
Using text analysis technology based on large models and artificial intelligence, through speech recognition, word segmentation processing and word semantic embedding coding, combined with students' basic information, it optimizes the representation of learning question content, generates a more comprehensive and accurate understanding, and ensures that the content is consistent with students' intentions through the intention confirmation module.
It achieves an in-depth understanding and content expansion of students' questions, generates more complete and clear answers, provides sufficient basis for subsequent replies, and improves learning efficiency and experience.
Smart Images

Figure CN120706439A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent dialogue, and more specifically, to an intelligent dialogue system and method for a teaching robot. Background Art
[0002] In today's educational environment, students' learning needs are becoming increasingly diverse and complex. As knowledge systems continue to evolve and expand, students' questions are no longer limited to basic knowledge, but often involve interdisciplinary and in-depth exploration. However, traditional educational assistance systems struggle to keep pace with this change, exposing significant shortcomings when responding to student questions.
[0003] Specifically, traditional systems typically only provide a superficial understanding of students' questions, failing to deeply analyze the true intent and implicit information behind the questions, nor effectively expand and refine the content of the questions. As a result, the generated answers often lack depth and comprehensiveness. Furthermore, traditional systems lack mechanisms for interacting with students to confirm the intent of the questions, often providing answers based on the system's own understanding of the questions. This can deviate from students' actual needs, forcing them to repeatedly correct or re-ask the questions, impacting their learning efficiency and experience.
[0004] Therefore, an optimized intelligent dialogue solution for teaching robots is desired. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a teaching robot intelligent dialogue system and method, which performs word segmentation processing and word semantic embedding coding of the speech recognition results of the learning question content by using text analysis and processing technology based on large models and artificial intelligence. At the same time, the basic information of the student object is scheduled and semantic embedding coding is performed on it, so as to achieve the expansion of the learning question content based on the core clue semantic optimization representation between the semantic embedding features of the student object basic information and the semantic embedding features of each learning question content word. In this way, the system can identify and utilize the complex relationship between the question content and the basic information of the student, generate a more comprehensive and accurate understanding, and make the questions more complete and clear, so that the content expansion is more comprehensive, providing sufficient basis for subsequent replies.
[0006] According to one aspect of the present application, a teaching aid robot intelligent dialogue system is provided, which includes: a voice signal input module for acquiring a learning question voice signal input by a student object; a voice recognition module for performing voice recognition on the learning question voice signal to obtain a learning question content voice recognition result; a content expansion module for inputting the learning question content voice recognition result into a question improvement and expansion module based on a large language model to obtain expanded learning question content; an intention confirmation module for displaying the expanded learning question content, and for the student object to confirm whether the expanded learning question content conforms to its question intention; a reply module for, in response to receiving a confirmation instruction from the student object, inputting the learning question expanded content into an intelligent dialogue engine based on a Deepseek model to obtain a reply text; and a voice broadcast module for converting the reply text into a voice signal and then performing voice broadcast.
[0007] According to another aspect of the present application, a teaching robot intelligent dialogue method is provided, which includes:
[0008] Acquire a learning question voice signal recorded by a student object;
[0009] Performing speech recognition on the learning question voice signal to obtain a learning question content speech recognition result;
[0010] Inputting the speech recognition result of the learning question content into the question improvement and expansion module based on the large language model to obtain the learning question expansion content;
[0011] Displaying the expanded content of the learning question, and having the student object confirm whether the expanded content of the learning question meets the intention of the student object;
[0012] In response to receiving a confirmation instruction from the student object, inputting the learning question expansion content into an intelligent dialogue engine based on a Deepseek model to obtain a reply text;
[0013] The reply text is converted into a voice signal and then voice broadcasted.
[0014] Compared with the existing technology, the present application provides a teaching robot intelligent dialogue system and method, which uses text analysis and processing technology based on large models and artificial intelligence to perform word segmentation processing and word semantic embedding coding of the speech recognition results of the learning question content. At the same time, it dispatches the basic information of the student object and performs semantic embedding coding on it. In this way, the core clue semantic optimization representation between the semantic embedding features of the student object basic information and the semantic embedding features of each learning question content word is used to achieve the expansion of the learning question content. In this way, the system can identify and utilize the complex relationship between the question content and the student's basic information, generate a more comprehensive and accurate understanding, and make the questions more complete and clear, so that the content expansion is more comprehensive, providing sufficient basis for subsequent replies. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 is a block diagram of an intelligent dialogue system for a teaching robot according to an embodiment of the present application;
[0017] Figure 2 Schematic diagram of data flow of the teaching robot intelligent dialogue system according to an embodiment of the present application;
[0018] Figure 3 1 is a block diagram of a content expansion module in an intelligent dialogue system of a teaching robot according to an embodiment of the present application;
[0019] Figure 4 is a block diagram of an identity constraint semantic optimization unit in an intelligent dialogue system of a teaching robot according to an embodiment of the present application;
[0020] Figure 5 Flowchart of the intelligent dialogue method of the teaching robot according to the embodiment of the present application. DETAILED DESCRIPTION
[0021] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0022] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0023] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0024] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0025] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0026] In the technical solution of this application, an intelligent dialogue system for a teaching robot is proposed. Figure 1 A block diagram of a teaching robot intelligent dialogue system according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the teaching robot intelligent dialogue system according to the embodiment of the present application. Figure 1 and Figure 2As shown, the teaching aid robot intelligent dialogue system 300 according to the embodiment of the present application includes: a voice signal input module 310, which is used to obtain a learning question voice signal input by a student object; a voice recognition module 320, which is used to perform voice recognition on the learning question voice signal to obtain a learning question content voice recognition result; a content expansion module 330, which is used to input the learning question content voice recognition result into a question improvement and expansion module based on a large language model to obtain a learning question expanded content; an intention confirmation module 340, which is used to display the learning question expanded content, and the student object confirms whether the learning question expanded content is consistent with its question intention; a reply module 350, which is used to input the learning question expanded content into an intelligent dialogue engine based on a Deepseek model in response to receiving a confirmation instruction from the student object to obtain a reply text; and a voice broadcast module 360, which is used to convert the reply text into a voice signal and then perform voice broadcast.
[0027] In particular, the voice signal recording module 310 is used to obtain the learning question voice signal recorded by the student object. It should be understood that by obtaining the learning question voice signal recorded by the student object, the system can better understand and analyze the actual intention behind the question based on real voice input, and then generate more accurate answers. In addition, through voice signal recording, the system can also identify non-verbal information such as tone and intonation, thereby further optimizing the dialogue experience and making it closer to the effect of face-to-face communication. In one example, a voice signal can be obtained by a highly sensitive microphone or other audio acquisition device to clearly capture the student's voice input. In this way, the system can not only accurately obtain the content of the student's questions, but also lay a solid foundation for further content analysis, semantic understanding and even personalized replies.
[0028] In particular, the speech recognition module 320 is used to perform speech recognition on the learning question speech signal to obtain a speech recognition result of the learning question content. Accordingly, considering that the learning question speech signal contains a large amount of students' question intentions and question contents, and the speech signal is an analog signal, the computer cannot directly understand and process the information contained therein. Therefore, through speech recognition, the speech signal can be converted into a text-based learning question content speech recognition result, thereby converting the speech information into a data format that the system can analyze and operate, laying the foundation for subsequent intelligent processing. For example, after the speech "how to prove the Pythagorean theorem" is converted into text, the system can further analyze the mathematical concepts and question intentions therein, and accurately understand the problem, so as to improve the accuracy of subsequent replies. In one example, the learning question voice signal can be subjected to speech recognition by adopting advanced automatic speech recognition (ASR) technology. First, the acquired learning question voice signal is preprocessed. Then, key features, such as Mel-frequency cepstral coefficients (MFCCs), are extracted from the preprocessed audio signal to effectively extract important information from the original audio signal. Then, the extracted features are input into a pre-trained deep learning model, such as a model based on the Transformer architecture. These models are trained with a large amount of speech data and can recognize different pronunciations, accents, and background noise, thereby converting the input speech features into corresponding text output.
[0029] In particular, the content expansion module 330 is used to input the speech recognition results of the learning question content into the question improvement and expansion module based on the large language model to obtain the expanded content of the learning question. It should be understood that when students ask learning questions, they may be brief, vague or lack information due to limitations in their language expression ability, current thinking state or questioning scenario. If they answer directly, key information may be omitted. Therefore, in order to better understand the students' intentions, the questions are made more specific and clear to ensure that more comprehensive answers are provided. In particular, in a specific example of the present application, if Figure 3As shown, the content expansion module 330 includes: a question content word segmentation encoding unit 331, which is used to perform word segmentation processing and word semantic embedding encoding on the speech recognition result of the learning question content to obtain a sequence distribution of the word semantic embedding features of the learning question content; a basic information encoding unit 332, which is used to schedule the basic information of the student object and perform semantic embedding encoding on the basic information of the student object to obtain the student object basic information semantic embedding features; an identity constraint semantic optimization unit 333, which is used to perform basic information constraint question content core clue semantic optimization encoding on the sequence distribution of the student object basic information semantic embedding features and the learning question content word semantic embedding features to obtain the basic information constraint question content semantic optimization features; a learning question expansion content generation unit 334, which is used to obtain the learning question expansion content based on the basic information constraint question content semantic optimization features.
[0030] Specifically, the question content word segmentation encoding unit 331 is used to perform word segmentation processing and word semantic embedding encoding on the learning question content speech recognition result to obtain a sequence distribution of the learning question content word semantic embedding features. That is, in the technical solution of the present application, first, the learning question content speech recognition result is word segmented to obtain a sequence distribution of the learning question content words; considering that the learning question content speech recognition result contains a large number of words or phrases, some of which contain important semantic information, and in order to be able to understand and analyze the information contained in each question content word more carefully and clearly, in the technical solution of the present application, the learning question content speech recognition result is word segmented to decompose a continuous text into meaningful units (i.e., words) to obtain a sequence distribution of the learning question content words. Further, each learning question content word in the sequence distribution of the learning question content words is input into a word semantic embedding encoder based on the Bert model to obtain a sequence distribution of the learning question content word semantic embedding encoding vector as the sequence distribution of the learning question content word semantic embedding features. Here, in order to further capture and extract the semantic information and meaning in each learning question content word, the present application inputs each learning question content word in the sequence distribution of the learning question content word into the word semantic embedding encoder based on the Bert model to obtain a sequence distribution of the learning question content word semantic embedding encoding vector. It can be understood that the Bert model is constructed based on the encoder in the Transformer architecture and consists of multiple stacked Transformer encoder layers. Each encoder layer contains components such as a multi-head attention mechanism and a feedforward neural network. It can process text sequences in parallel and efficiently capture long-term dependencies in the text. For a given word, the Bert model will consider the words on its left and right at the same time, so as to better understand the specific meaning of the learning question content word and generate a more accurate word vector representation.
[0031] Specifically, the basic information encoding unit 332 is used to schedule the basic information of the student object and perform semantic embedding encoding on the basic information of the student object to obtain the semantic embedding features of the student object's basic information. That is, in the technical solution of the present application, first, the basic information of the student object is scheduled, and the basic information includes identity information and learning status; further, the basic information of the student object is input into a semantic encoder based on the Bert model to obtain a semantic embedding encoding vector of the student object's basic information as the semantic embedding features of the student object's basic information. Considering that the basic information of the student object contains potential semantic information and key semantic features, in order to better understand and analyze the content expressed in the basic information, in the technical solution of the present application, the basic information of the student object is semantically embedded to obtain a semantic embedding encoding vector of the student object's basic information. In particular, in a specific embodiment of the present application, the basic information of the student object can be input into a semantic encoder based on the Bert model to adopt a bidirectional Transformer architecture to fully consider the mutual relationship between the various parts of the input text, explore the hidden contextual semantic relationship therein, and obtain the semantic embedding encoding vector of the student object's basic information. For example, by analyzing the contextual connection between a student's grade level and their learning progress, we can more accurately understand where the student is in the learning process. This allows us to better understand the student's specific needs and the context of the question, thereby generating more precise responses.
[0032] Specifically, the identity constraint semantic optimization unit 333 is used to perform basic information constraint question content core clue semantic optimization encoding on the sequence distribution of the student object basic information semantic embedding feature and the learning question content word semantic embedding feature to obtain basic information constraint question content semantic optimization features. Considering that the student object basic information semantic embedding feature represents the individual characteristics of the student, such as identity, learning situation, etc., and the sequence distribution of the learning question content word semantic embedding feature reflects the semantics of the question itself, this information is crucial to understanding the learning question content. Students of different grades have different knowledge reserves and cognitive levels, and the focus and depth of their questions will also be different. Therefore, in order to better characterize the question content semantics based on the student's basic information as a constraint condition, in the technical solution of the present application, the sequence distribution of the student object basic information semantic embedding feature and the learning question content word semantic embedding feature is performed basic information constraint question content core clue semantic optimization encoding to obtain basic information constraint question content semantic optimization features. Specifically, this method uses attention mechanism, heterogeneous transformer and template constraint learning to deeply integrate students' basic information with the semantic features of the question content, explore the correlation and complementarity between the two, and thus generate highly consistent and optimized features that fit the students' actual situation, significantly improving the understanding of the question intention and the accuracy of the response. In particular, in a specific example of this application, such as Figure 4 As shown, the identity constraint semantic optimization unit 333 includes: a core clue extraction subunit 3331, which is used to extract core clues from the sequence distribution of the student object basic information semantic embedding coding vector and the learning question content word semantic embedding coding vector to obtain a basic information core clue coding vector and a learning question content word semantic core clue coding vector; a core clue weaving template construction subunit 3332, which is used to construct a basic information-learning question content core clue weaving template matrix between the basic information core clue coding vector and the learning question content word semantic core clue coding vector; a basic information-learning question content encoding subunit 3333, which is used to perform cross-modal key clue guided encoding on the sequence distribution of the student object basic information semantic embedding coding vector and the learning question content word semantic embedding coding vector based on the basic information-learning question content core clue weaving template matrix to obtain a basic information constraint question content semantic optimization coding vector as the basic information constraint question content semantic optimization feature.
[0033] More specifically, the core clue extraction subunit 3331 is configured to extract core clues from the sequence distribution of the student object's basic information semantic embedding encoding vector and the learning question content word semantic embedding encoding vector, respectively, to obtain a basic information core clue encoding vector and a learning question content word semantic core clue encoding vector. In an embodiment of the present application, the student object's basic information semantic embedding encoding vector is first subjected to core clue extraction based on point convolution coding to obtain the basic information core clue encoding vector. It should be understood that the student object's basic information semantic embedding encoding vector contains rich but potentially redundant information, such as the student's identity, learning status, and other semantic features. Point convolution coding, by performing a convolution operation on each dimension of the vector, can focus on local information and extract key features that are likely closely related to the content of the learning question, namely, core clues. For example, in a vector representing information such as a student's grade, subject performance, and learning style, point convolution coding can highlight the information most helpful for understanding the question's intent, such as grade, which may affect the expected difficulty of the question, and subject performance, which may reflect the degree of mastery of the knowledge point. This allows subsequent processing to more efficiently utilize key elements of the student's basic information, providing core support for understanding the context of the question. Furthermore, a core clue extractor based on convolutional coding and statistical features is used to extract the semantic core clue encoding vector of the learning question content word from the sequence distribution of the semantic embedding encoding vector of the learning question content word. It should be understood that the sequence of semantic embedding encoding vectors of the learning question content word contains the semantic information of the question, but it is relatively scattered. Through convolutional coding, local semantic patterns in the sequence can be captured, and semantic fragments expressed by closely related word combinations can be identified. At the same time, combined with statistical features (such as word frequency, word importance measurement, etc.), the core semantic clues of the question content can be grasped as a whole. The semantic core clue encoding vector of the learning question content word accurately summarizes the core semantics of the question and highlights the key information, which helps to combine with the core clues of the student's basic information in the subsequent step to accurately understand the core appeal of the student's question. In a specific example, the core clues are extracted from the sequence distribution of the student object basic information semantic embedding encoding vector and the learning question content word semantic embedding encoding vector to obtain the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector respectively according to the following formula; wherein, the formula is:
[0034] v c1 =Leaky ReLU{Conv 1×1 (V1)+b1}
[0035]
[0036] Among them, V1 is the semantic embedding encoding vector of the basic information of the student object, Conv 1×1is the point convolution code, b1 is the bias vector, Leaky ReLU is the activation function, v c1 is the basic information core clue encoding vector, V2 is the sequence distribution of the semantic embedding encoding vector of the learning question content word, v 21 , v 22 , v 2i and v 2n are the first, second, i-th and n-th learning question content word semantic embedding encoding vectors in the sequence distribution of the learning question content word semantic embedding encoding vector, v 2i ' is the i-th learning question content word semantic embedding enhanced encoding vector in the sequence distribution of the learning question content word semantic embedding enhanced encoding vector, sigmoid is the sigmoid function, max(·) and min(·) are the maximum and minimum values respectively, α, β and γ are modulation parameters, μ(v 2i ') and σ(v 2i ') 2 v 2i ' The mean and variance of D i is the i-th learning question content word semantic embedding local significant coefficient in the sequence distribution of the learning question content word semantic embedding local significant coefficient, exp(·) represents the exponential function value with the natural constant e as the base, n is the number of vectors in the sequence distribution of the learning question content word semantic embedding encoding vector, v c2 It is to learn the semantic core clue encoding vector of the question content words.
[0037] More specifically, the core clue weaving template construction subunit 3332 is used to construct a basic information-learning question content core clue weaving template matrix between the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector. By constructing the basic information-learning question content core clue weaving template matrix, the elements in the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector can be organized in a specific way to clarify the correspondence and mutual influence between different clues. For example, the elements in the matrix can quantitatively represent the relationship between the student's grade and the difficulty requirement of the knowledge point in the question, or the relationship between the student's subject performance and the depth of the question. In this way, a structured basis is provided for the subsequent fusion of the two core clues, so that the system can analyze how the student's basic information affects the understanding of the question content based on this matrix, and provide strong support for generating responses that are more in line with student needs. In a specific example, the basic information-learning question content core clue weaving template matrix between the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector is constructed using the following formula; wherein, the formula is:
[0038]
[0039] Where T is the transpose operation, and φ(·) is a feature mapping function, such as a linear mapping or a nonlinear kernel function, M tem It is basic information - learning question content core clue weaving template matrix.
[0040] More specifically, the basic information-learning question content encoding subunit 3333 is used to perform cross-modal key clue guided encoding on the sequence distribution of the student object basic information semantic embedding encoding vector and the learning question content word semantic embedding encoding vector based on the basic information-learning question content core clue weaving template matrix to obtain the basic information constraint question content semantic optimization encoding vector as the basic information constraint question content semantic optimization feature. In an embodiment of the present application, first, the basic information-learning question content core clue weaving template matrix is subjected to explicit semantic space adaptive compensation to obtain the optimized basic information-learning question content core clue weaving template matrix; it should be understood that the original basic information-learning question content core clue weaving template matrix may have some limitations, which may be due to the inherent differences between different modalities (such as text and student background information). This difference will lead to "missing" information in the preliminary structured semantic alignment process, that is, some fine-grained significant difference information is not fully captured. For example, in an educational scenario, students' questions may involve interdisciplinary knowledge points, and their background information contains multi-dimensional information such as personal interests and past learning experiences. If this information cannot be well integrated, the generated answer may lack pertinence and depth, and cannot meet the students' in-depth learning needs. Therefore, it is preferred to further target the feature vector v used to represent the core clues of different modal features. c1 and v c2Through its residual mechanism based on heteromodal mapping, a preliminary structured semantic alignment of common explicit associations is performed within a shared representation space to obtain preliminary structured semantic alignment vectors for basic information and learning question content, and preliminary structured semantic alignment vectors for learning question content and basic information. Then, based on the global semantic spatialization representation of the explicitly modeled graph structure under mapping preferences, the residual mechanism is further utilized to capture "missing" information during the alignment process under heteromodal mapping. Adaptive compensation is performed through explicit associations to further capture fine-grained inter-modal saliency information during the preliminary alignment process. Specifically, by performing feature mapping and residual processing on the encoding vectors of the core clues of the basic information and the semantic core clues of the learning question content words, the fine-grained heteromodal saliency features between the two are identified. Based on this, a global semantic optimization of the basic information-learning question content core clue braided template matrix is performed. This reduces the illusion of global semantic constraints between modalities when fine-grained relationships between local salient regions are misaligned, allowing the resulting optimized template matrix to more accurately reflect the complex relationship between question content and student basic information. More specifically, in a specific example of the present application, the specific steps of performing explicit semantic space adaptive compensation on the basic information-learning question content core clue weaving template matrix are as follows: first, feature mapping and residual processing are performed on the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector to obtain the basic information-learning question content preliminary structured semantic alignment vector and the learning question content-basic information preliminary structured semantic alignment vector; then, the different modal fine-grained significant difference features between the basic information-learning question content preliminary structured semantic alignment vector and the learning question content-basic information preliminary structured semantic alignment vector are captured to obtain the basic information-learning question content display matrix; and then, based on the basic information-learning question content display matrix, the basic information-learning question content core clue weaving template matrix is globally semantically optimized to obtain the optimized basic information-learning question content core clue weaving template matrix.
[0041] Preferably, in this example, the basic information-learning question content core clue weaving template matrix is subjected to explicit semantic space adaptive compensation using the following compensation formula to obtain an optimized basic information-learning question content core clue weaving template matrix; wherein, the formula is:
[0042]
[0043]
[0044] in, It is subtracted by position point, v m and v m' are respectively the basic information-learning question content preliminary structured semantic alignment vector and the learning question content-basic information preliminary structured semantic alignment vector, To calculate the square of the vector norm, log2 is the logarithmic function value with base 2. is the matrix multiplication, M cor It is the basic information-learning question content display matrix, M tem 'It is to optimize basic information - learn the core clues of question content and weave the template matrix.
[0045] Next, the sequence distribution of the semantic embedding coding vector of the content words of the learning question is reshaped to obtain the semantic embedding coding matrix of the content words of the learning question; it should be understood that in an educational scenario, each question raised by a student is unique and contains specific semantic information and background associations. After preliminary word segmentation processing and semantic embedding coding, this information forms a sequence distribution, but this sequence distribution often lacks a structured organizational form and is difficult to effectively integrate and compare with other basic information of the students (such as interests, hobbies, learning history, etc.). Through reshaping, this disordered sequence distribution can be converted into an ordered matrix form to more efficiently capture and utilize the core clues in the question content and the relationship between it and the basic information of the students. In a specific example, the sequence distribution of the semantic embedding coding vector of the content words of the learning question is reshaped to obtain the semantic embedding coding matrix of the content words of the learning question; wherein, the formula is:
[0046] M=reshape{v 21 ,v 22 ,...,v 2i ,...,v 2n}
[0047] Among them, reshape is the shape reshaping operation, and M is the semantic embedding encoding matrix of the learned question content words.
[0048] Subsequently, the heterogeneous transformer performs cross-modal constraint encoding on the student object's basic information semantic embedding encoding vector as the query vector, the learning question content word semantic embedding encoding matrix as the key matrix, and the optimized basic information-learning question content core clue weaving template matrix as the prior information constraint matrix to obtain the basic information-constrained question content semantic optimized encoding vector. The heterogeneous transformer's cross-modal constraint encoding utilizes these three matrices to deeply fuse the student's basic information with the semantic information of the learning question content. Using the student object's basic information semantic embedding encoding vector as the query vector, relevant information can be searched within the learning question content word semantic embedding encoding matrix. By optimizing the constraints of the basic information-learning question content core clue weaving template matrix, the fusion process ensures that the core clue semantic relationships between the two are closely integrated. For example, when generating the encoding vector, the importance of each semantic element in the question content is adjusted based on basic information such as the student's grade and academic performance, highlighting the parts relevant to the student's actual situation. Specifically, the basic information-constrained question content semantic optimized encoding vector highly integrates the semantics of the student's basic information and the learning question content, comprehensively and accurately reflecting the student's true intention when asking the question under the specific basic information constraints. This provides a key basis for the intelligent dialogue system to generate highly personalized responses that precisely meet students' needs, greatly improving the system's intelligent interaction capabilities. In a specific example, the following cross-modal constraint encoding formula is used to perform cross-modal constraint encoding on the heterogeneous converter to obtain the semantically optimized encoding vector of the basic information constraint question content; wherein, the cross-modal constraint encoding formula is:
[0049]
[0050] Among them, L is the scale of M, that is, the width of the matrix multiplied by the height of the matrix, softmax is the normalization function, v ti It is the semantic optimization encoding vector of the question content constrained by the basic information.
[0051] Specifically, the learning question expansion content generation unit 334 is used to obtain the learning question expansion content based on the basic information constraint question content semantic optimization feature. In the technical solution of the present application, the basic information constraint question content semantic optimization coding vector is input into the question improvement and expansion module based on the large language model to obtain the learning question expansion content. That is, the basic information constraint question content semantic optimization feature obtained by constraint optimization using the sequence distribution of the student object basic information semantic embedding coding vector and the learning question content word semantic embedding coding vector is generated and processed to utilize the powerful language understanding and generation capabilities of the large language model to process complex semantic information. Based on its pre-trained knowledge system and powerful reasoning ability, this information is deeply mined and analyzed, thereby better completing the question improvement and expansion task. This can make it easier for the subsequent intelligent dialogue engine based on the Deepseek model to understand the student's true intentions, thereby providing more accurate and targeted responses, and improving the interactive effect and practicality of the intelligent dialogue system.
[0052] In the technical solution of the present application, the sequence distribution of the student object basic information semantic embedding coding vector and the learning question content word semantic embedding coding vector respectively represent the text sequence distribution of the student object basic information semantic coding features and the learning question content word granularity semantic coding features. When performing basic information constrained question content semantic optimization coding, due to the inconsistent granularity of text semantic features, there will be a fine-grained offset in the alignment of multi-granularity text semantic distributions during cross-domain semantic optimization, resulting in the obtained basic information constrained question content semantic optimization coding vector having weak feature expression capabilities at the fine-grained micro-semantic interaction level, affecting the text quality of the learning question expansion content obtained by inputting the question improvement and expansion module based on the large language model.
[0053] In a preferred example, in the process of inputting the semantically optimized coding vector of the basic information constraint question content into the large language model-based question improvement and expansion module to obtain the learning question expansion content, the semantically optimized coding vector of the basic information constraint question content is subjected to semantic coding optimization, and the process includes:
[0054] Based on the global feature distribution characteristics of the semantic optimization coding vector of the basic information constraint question content, a first basic information constraint question content semantic optimization tuning coefficient and a second basic information constraint question content semantic optimization tuning coefficient are constructed, which are expressed as:
[0055]
[0056] v i ∈V1∈R n
[0057] Among them, V represents the semantic optimization encoding vector of the basic information constraint question content, μ and σ 2 represent the mean and variance of the feature set composed of all eigenvalues of the semantic optimization encoding vector of the basic information constraint question content, respectively, represents positional subtraction, ε represents a predetermined hyperparameter, which can be set based on experience or based on common hyperparameter optimization tools, such as 0.9. Of course, the above is only an example and does not constitute a limitation. V1 represents the basic information constraint question content semantic optimization coding tuning vector, R represents a real number set, n represents the total number of eigenvalues of the basic information constraint question content semantic optimization coding tuning vector, v i represents the eigenvalue of the i-th position of the basic information constrained question content semantic optimization coding tuning vector, θ1 represents the first basic information constrained question content semantic optimization tuning coefficient, and θ2 represents the second basic information constrained question content semantic optimization tuning coefficient.
[0058] Based on the first basic information constraint question content semantic optimization tuning coefficient, the basic information constraint question content semantic optimization coding vector is pseudo-phase shifted to obtain a first basic information constraint question content semantic optimization pseudo-phase shift vector, which is expressed as:
[0059]
[0060] Where V2 represents the pseudo-phase transfer vector of the semantic optimization of the question content constrained by the first basic information, ⊙ represents the point multiplication by position, max represents the maximum value function, [·] ⊙-1 Indicates that the inverse of the eigenvalue at each position of the vector is calculated.
[0061] Based on the second basic information constraint question content semantic optimization tuning coefficient, the basic information constraint question content semantic optimization coding vector is subjected to heterogeneous pseudo phase migration to obtain a second basic information constraint question content semantic optimization pseudo phase migration vector, which is expressed as:
[0062]
[0063] Among them, V3 represents the pseudo-phase transfer vector of the semantic optimization of the question content constrained by the second basic information;
[0064] Based on the first basic information constraint question content semantic optimization pseudo phase transition vector and the second basic information constraint question content semantic optimization pseudo phase transition vector, the basic information constraint question content semantic optimization encoding vector is fine-tuned to obtain an optimized basic information constraint question content semantic optimization encoding vector, which is expressed as:
[0065]
[0066] Among them, Sigmoid represents the activation function, Indicates addition by position, α represents the first weight hyperparameter, β represents the second weight hyperparameter. Similarly, it can be set based on experience or based on common hyperparameter optimization tools, for example, α = 0.4, β = 0.6. Of course, the above is only an example and does not constitute a limitation. V' represents the optimized basic information constraint question content semantic optimization encoding vector.
[0067] In the preferred embodiment, the gradient difference of the eigenvalue of the semantic optimization coding vector of the question content constrained by basic information relative to the heterogeneous tuning representation of its overall feature set is calculated and mapped to the semantic evolution energy level parameter. Then, the energy level modulation pseudo-phase migration of the coordinate constraint is realized by adopting the heterogeneous modulation representation paradigm, and the staggered stacking spatial displacement transformation is performed under the condition of set dimension normalization. In this way, the fusion and enhancement of the semantic evolution phase perception can be expanded along the feature fusion trajectory to expand the axial perception domain, thereby improving the ability of the semantic optimization coding vector of the question content constrained by basic information to capture micro-semantic changes, and finally optimizing the expression effect of the semantic optimization coding vector of the question content constrained by basic information. In this way, the text quality of the learning question expansion content obtained by inputting the question improvement and expansion module based on the large language model is improved.
[0068] In particular, the intention confirmation module 340 is used to display the expanded content of the learning question, and the student object confirms whether the expanded content of the learning question is consistent with its question intention. It should be understood that although the system processes and expands the learning questions through a series of complex technologies, due to the complexity and diversity of natural language and the differences in individual students' thinking patterns, the system's understanding of the students' question intentions may still be biased. Even if the student's basic information and semantic features are comprehensively considered for optimized coding, it cannot be completely guaranteed that the expanded content is completely consistent with what the students have in mind. Therefore, students need to confirm in person to correct possible misunderstandings. Through student confirmation, it can be verified whether the expanded content of the learning question accurately reflects its true intention. If the student confirms that it is consistent, it means that the system's understanding and expansion of the question is correct, and the subsequent process can continue to provide students with accurate responses; if it is not consistent, the student can provide timely feedback, and the system will adjust based on the feedback to ensure that the final response provided is what the student really needs, thereby improving the accuracy of the entire intelligent dialogue system. In one example, after the content expansion module completes a series of operations on the original speech signal, including recognition, word segmentation, and semantic embedding encoding, the system generates a more comprehensive and accurate learning question expansion based on the large language model. This expansion is intended to supplement and improve the student's initial question, making it more specific and clear, and facilitating subsequent intelligent responses. The optimized and expanded question content is then presented to the student. Typically, the system presents the expanded question content through the system's user interface (UI) in concise and clear text or chart format to ensure clarity. The system then invites the student to confirm whether the expanded content accurately reflects the student's question intent. This can be achieved through an interactive interface, such as providing a "Confirm" or "Re-enter" button option, allowing the student to make a selection directly on the interface. If the student believes the displayed content does meet their question intent, they can click the "Confirm" button to inform the system to proceed to the next step—submitting the expanded question to the Deepseek model to generate the final response text. Conversely, if students find that the displayed content doesn't match their actual question, they can select "Re-enter," allowing them to re-enter or modify their question until they are satisfied. This interactive confirmation mechanism not only helps improve the system's accuracy in understanding questions and reduces incorrect answers due to misunderstandings, but also enhances user engagement and satisfaction.
[0069] In particular, the reply module 350 is configured to, in response to receiving a confirmation instruction from the student object, input the expanded learning question content into an intelligent dialogue engine based on the Deepseek model to obtain a reply text. In particular, considering that the Deepseek model has been trained on a large amount of text data and possesses excellent language comprehension capabilities, it is able to deeply analyze the semantics, logic, and key knowledge points in the expanded learning question content. At the same time, the Deepseek model also possesses powerful language generation capabilities and can generate high-quality, logical reply text based on its understanding of the question. For example, for complex academic questions, the Deepseek model can organize clear and accurate language to provide answers.
[0070] Specifically, the voice broadcast module 360 is configured to convert the reply text into a voice signal and then broadcast it. It should be understood that voice is a natural and intuitive means of communication. Through voice broadcast, students can quickly receive responses from the system and interact with the system in a more natural way, increasing the realism and intimacy of the interaction. In one example, a suitable text-to-speech (TTS) engine, such as Google Text-to-Speech, Amazon Polly, or Microsoft Azure Text-to-Speech, is first selected to generate speech output that is close to human voice. These engines are trained with extensive data and are capable of mimicking very natural human speech characteristics and adjusting factors such as intonation, rhythm, and stress based on different contexts. Next, the system passes the reply text generated by the Deepseek model as input to the engine. The TTS engine then analyzes the input text, determines the correct pronunciation of each word, and adjusts factors such as intonation, rhythm, and stress based on the context, ultimately synthesizing a coherent speech signal. It is worth noting that the TTS engine uses a deep neural network model to simulate the vocal characteristics of human speech, making the generated speech sound more natural. In addition, to better meet personalized needs, the TTS engine also allows users to choose different voice styles (such as male, female, children's voices, etc.), and even adjust parameters such as speaking speed and pitch to generate the voice output that best suits the user; further, after the voice signal is successfully generated, it will be broadcast voicely. This design enhances the user's sense of participation and satisfaction, and meets the needs encountered by users in the learning process.
[0071] As described above, the teaching aid robot intelligent dialogue system 300 according to the embodiment of the present application can be implemented in various wireless terminals, such as a server equipped with a teaching aid robot intelligent dialogue algorithm. In one possible implementation, the teaching aid robot intelligent dialogue system 300 according to the embodiment of the present application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the teaching aid robot intelligent dialogue system 300 can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the teaching aid robot intelligent dialogue system 300 can also be one of the many hardware modules of the wireless terminal.
[0072] Alternatively, in another example, the teaching robot intelligent dialogue system 300 and the wireless terminal may also be separate devices, and the teaching robot intelligent dialogue system 300 may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0073] Furthermore, a teaching robot intelligent dialogue method is also provided.
[0074] Figure 5 Flowchart of the teaching robot intelligent dialogue method according to the embodiment of the present application. Figure 5 As shown, the intelligent dialogue method of the teaching aid robot according to the embodiment of the present application includes the following steps: S1, obtaining a learning question voice signal recorded by a student object; S2, performing voice recognition on the learning question voice signal to obtain a learning question content voice recognition result; S3, inputting the learning question content voice recognition result into a question improvement and expansion module based on a large language model to obtain learning question expanded content; S4, displaying the learning question expanded content, and confirming by the student object whether the learning question expanded content is consistent with its question intention; S5, in response to receiving a confirmation instruction from the student object, inputting the learning question expanded content into an intelligent dialogue engine based on a Deepseek model to obtain a reply text; S6, converting the reply text into a voice signal and then performing voice broadcast.
[0075] In summary, the teaching aid robot intelligent dialogue method according to the embodiment of the present application is explained, which performs word segmentation processing and word semantic embedding coding of the speech recognition results of the learning question content by using text analysis and processing technology based on large models and artificial intelligence. At the same time, the basic information of the student object is scheduled and semantic embedding coding is performed on it, so as to achieve the expansion of the learning question content based on the core clue semantic optimization representation between the semantic embedding features of the student object basic information and the semantic embedding features of the words of each learning question content. In this way, the system can identify and utilize the complex relationship between the question content and the basic information of the student, generate a more comprehensive and accurate understanding, and make the questions more complete and clear, so that the content expansion is more comprehensive, providing sufficient basis for subsequent replies.
[0076] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A teaching robot intelligent dialogue system, characterized in that: include: The voice signal input module is used to obtain the learning question voice signal input by the student object; a speech recognition module for performing speech recognition on the learning question speech signal to obtain a speech recognition result of the learning question content; a content expansion module for inputting the speech recognition result of the learning question content into a question improvement and expansion module based on a large language model to obtain expanded content of the learning question; An intention confirmation module, configured to display the expanded content of the learning question and allow the student subject to confirm whether the expanded content of the learning question is consistent with the student's intention to ask the question; A reply module, configured to input the expanded content of the learning question into an intelligent dialogue engine based on a Deepseek model to obtain a reply text in response to receiving a confirmation instruction from the student object; A voice broadcast module, used to convert the reply text into a voice signal and then broadcast it; The content expansion module includes: A question content word segmentation encoding unit, configured to perform word segmentation processing and word semantic embedding encoding on the speech recognition result of the learning question content to obtain a sequence distribution of word semantic embedding features of the learning question content; A basic information encoding unit, configured to schedule the basic information of the student object and perform semantic embedding encoding on the basic information of the student object to obtain a semantic embedding feature of the basic information of the student object; An identity constraint semantic optimization unit, configured to perform semantic optimization encoding of the basic information constraint question content core clues on the sequence distribution of the student object basic information semantic embedding features and the learning question content word semantic embedding features to obtain a basic information constraint question content semantic optimization feature; The learning question expansion content generating unit is used to obtain the learning question expansion content based on the semantic optimization features of the question content constrained by the basic information.
2. The teaching robot intelligent dialogue system according to claim 1, characterized in that: The question content word segmentation encoding unit includes: A word segmentation subunit is used to perform word segmentation processing on the speech recognition result of the learning question content to obtain a sequence distribution of the learning question content words; The word embedding encoding subunit is used to input each learning question content word in the sequence distribution of the learning question content words into a word semantic embedding encoder based on the Bert model to obtain a sequence distribution of the learning question content word semantic embedding encoding vector as the sequence distribution of the learning question content word semantic embedding features.
3. The teaching robot intelligent dialogue system according to claim 2, characterized in that: The basic information encoding unit includes: A basic information scheduling subunit is used to schedule the basic information of the student object, the basic information including identity information and learning status; The student object basic information semantic encoding subunit is used to input the basic information of the student object into a semantic encoder based on the Bert model to obtain a student object basic information semantic embedding encoding vector as the student object basic information semantic embedding feature.
4. The teaching robot intelligent dialogue system according to claim 3, characterized in that: The identity constraint semantic optimization unit includes: A core clue extraction subunit is used to extract core clues from the sequence distribution of the student object basic information semantic embedding coding vector and the learning question content word semantic embedding coding vector to obtain a basic information core clue coding vector and a learning question content word semantic core clue coding vector; A core clue weaving template construction subunit is used to construct a basic information-learning question content core clue weaving template matrix between the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector; The basic information-learning question content encoding sub-unit is used to weave a template matrix based on the core clues of the basic information-learning question content, and perform cross-modal key clue-guided encoding on the sequence distribution of the student object basic information semantic embedding encoding vector and the learning question content word semantic embedding encoding vector to obtain the basic information constraint question content semantic optimization encoding vector as the basic information constraint question content semantic optimization feature.
5. The teaching robot intelligent dialogue system according to claim 4, characterized in that: The core clue extraction subunit is used to: Performing point convolution coding-based core clue extraction on the semantic embedding coding vector of the basic information of the student object to obtain the basic information core clue coding vector; The learning question content word semantic core clue encoding vector is extracted from the sequence distribution of the learning question content word semantic embedding encoding vector using a core clue extractor based on convolutional coding and statistical features.
6. The teaching robot intelligent dialogue system according to claim 5, characterized in that: The basic information-learning question content encoding sub-unit includes: An explicit compensation secondary sub-unit, configured to perform explicit semantic space adaptive compensation on the basic information-learning question content core clue weaving template matrix to obtain an optimized basic information-learning question content core clue weaving template matrix; a reshape secondary subunit, configured to reshape the sequence distribution of the semantic embedding encoding vectors of the learning question content words to obtain a semantic embedding encoding matrix of the learning question content words; The cross-modal constraint encoding secondary sub-unit is used to use the student object basic information semantic embedding encoding vector as the query vector, the learning question content word semantic embedding encoding matrix as the key matrix and the optimized basic information-learning question content core clue weaving template matrix as the prior information constraint matrix, and perform cross-modal constraint encoding of the heterogeneous converter on them to obtain the basic information constraint question content semantic optimization encoding vector.
7. The teaching robot intelligent dialogue system according to claim 6, characterized in that: The explicit compensation secondary subunit is used for: Performing feature mapping and residual processing on the basic information core clue encoding vector and the learning question content word semantic core clue encoding vector to obtain a basic information-learning question content preliminary structured semantic alignment vector and a learning question content-basic information preliminary structured semantic alignment vector; Capturing the different modal fine-grained significant difference features between the basic information-learning question content preliminary structured semantic alignment vector and the learning question content-basic information preliminary structured semantic alignment vector to obtain a basic information-learning question content display matrix; Based on the basic information-learning question content display matrix, the basic information-learning question content core clue weaving template matrix is globally semantically optimized to obtain the optimized basic information-learning question content core clue weaving template matrix.
8. The teaching robot intelligent dialogue system according to claim 7, characterized in that: The learning question expansion content generating unit is used to input the semantic optimization coding vector of the basic information constraint question content into the question improvement and expansion module based on the large language model to obtain the learning question expansion content.
9. A teaching robot intelligent dialogue method, characterized in that: include: Acquire a learning question voice signal recorded by a student object; Performing speech recognition on the learning question voice signal to obtain a learning question content speech recognition result; Inputting the speech recognition result of the learning question content into the question improvement and expansion module based on the large language model to obtain the learning question expansion content; Displaying the expanded content of the learning question, and having the student object confirm whether the expanded content of the learning question meets the intention of the student object; In response to receiving a confirmation instruction from the student object, inputting the learning question expansion content into an intelligent dialogue engine based on a Deepseek model to obtain a reply text; The reply text is converted into a voice signal and then voice broadcasted.