Touch and talk pen intelligent children education method based on emotion recognition
By integrating multimodal data acquisition and generation adversarial networks and Transformer architecture emotion recognition technology in the dot reading pen, the problem that the dot reading pen cannot perceive children's emotional state in real time is solved, and a personalized and intelligent learning experience is achieved, which improves learning efficiency and interest.
Patent Information
- Application Number
- CN202510083553.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The existing dot reading pen lacks real-time perception of children's emotional state and is unable to adjust teaching strategies according to children's emotional changes, resulting in a single learning experience and it is difficult to maintain children's learning interest and concentration.
The built-in camera, microphone and touch sensor of the point-reading pen collects children's facial expression data, voice tone data and touch behavior data. The generative adversarial network and Transformer architecture are used to enhance multimodal emotional data and extract deep features to generate personalized recommended content that matches the current emotional state of children.
Real-time and high-precision emotional recognition, dynamically adjust learning content, significantly improve learning efficiency, reduce error rates, enhance learning interest, and provide a personalized and intelligent learning experience.
Smart Images

Figure CN119992624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point reading pens, and in particular to an intelligent children's education method using a point reading pen based on emotion recognition. Background Art
[0002] With the rapid development of artificial intelligence and intelligent hardware technology, intelligent teaching tools such as reading pens have gradually emerged in the field of education. Reading pens assist children's learning through voice, images and text, bringing convenience to children's education.
[0003] At present, the reading pens on the market are mainly based on a preset teaching content library, and use touch sensing to achieve voice playback or simple interactive functions of teaching materials. Although they can improve children's interest in learning, they still have the following limitations: On the one hand, traditional reading pens lack the ability to perceive children's emotional state in real time, and are unable to adjust teaching strategies according to children's emotional changes. When children are tired or frustrated in learning, the reading pens cannot provide timely encouragement or adjust the teaching content, resulting in a monotonous learning experience and difficulty in maintaining children's learning interest and concentration. On the other hand, the teaching content logic of existing reading pens is fixed, and it is impossible to dynamically recommend more suitable teaching content based on children's individual differences and immediate emotional state, and the personalized education experience is seriously insufficient.
[0004] To sum up, the existing technology has significant deficiencies in emotion perception ability, multimodal data fusion, dynamic adjustment of educational content and personalized recommendation, making it difficult to provide children with a comprehensive, intelligent and emotional learning experience. The defects of existing technology make it impossible for existing reading pen products to fully realize the potential of intelligent educational tools. A new method is urgently needed to solve the above problems. Summary of the invention
[0005] One purpose of the present invention is to propose an intelligent children's education method based on emotion recognition using a reading pen. The present invention solves the fuzzy classification problem of emotional states in the emotion recognition scenario of the reading pen, and also significantly improves the adaptability of the model to complex emotional states.
[0006] According to an embodiment of the present invention, a method for intelligent children's education using a reading pen based on emotion recognition comprises the following steps:
[0007] S1. Collect children’s facial expression data, voice intonation data and touch behavior data through the built-in camera, microphone and touch sensor of the reading pen;
[0008] S2. Use a generative adversarial network to enhance the collected facial expression data, voice intonation data, and touch behavior data to generate multimodal high-quality emotional data that matches the actual emotional state;
[0009] S3. Extract features from the enhanced multimodal high-quality emotion data through an emotion recognition model based on the Transformer architecture to generate deep emotion features related to facial expressions, voice intonation, and touch behavior;
[0010] S4. using the emotion state classification unit to classify the deep emotion features and generate an emotion state label corresponding to the current emotion state of the child;
[0011] S5. Generate voice feedback content, image animation content, and interactive game content that are appropriate to the current emotional state using a content generation model based on the Transformer architecture according to the emotional state label;
[0012] S6. Generate a personalized list of recommended content that matches the child's emotional state based on the matching results of the voice feedback content, image animation content, and interactive game content with the educational content library of the reading pen;
[0013] S7. Present educational content in a personalized recommended content list through the display screen, voice output unit, and touch interaction unit of the reading pen, output teaching content including voice feedback, image animation, or interactive games, and record relevant operation data in real time when the child interacts with the reading pen.
[0014] Optionally, the S1 specifically includes:
[0015] S11. Collect the child's facial expression data through the built-in camera of the reading pen, the facial expression data includes key point coordinate data P f (x,y), where x and y represent the horizontal and vertical coordinates of the facial landmarks, respectively;
[0016] S12. The voice and intonation data of the child is collected through the built-in microphone of the reading pen, wherein the voice and intonation data includes the audio signal amplitude A(t) and the frequency distribution F(ω), where t represents the time variable and ω represents the frequency variable;
[0017] S13. The touch sensor built into the reading pen is used to collect the child's touch behavior data, wherein the touch behavior data includes the touch intensity S(t) and the touch position P t (x,y).
[0018] Optionally, the S2 specifically includes:
[0019] S21. Decompose the collected facial expression data, voice intonation data, and touch behavior data into time and space, and use a multi-level feature encoder to hierarchically represent the data of different modalities:
[0020] Facial expression data f (x,y,t) is converted into a dynamic expression feature vector Extract local time features through the time window Δt;
[0021] Speech intonation data A(t), F(ω) is converted into spectral time series feature tensor Capture intonation and emotional patterns;
[0022] Touch behavior data S(t),P t (x,y) is converted into a touch event flow matrix Mapping the temporal changes and intensity distribution of touch behavior;
[0023] S22. Build a generative adversarial network model, including the module design of the generator and the discriminator:
[0024] G(z,X)=Attention(W g ·[z,X]+b g );
[0025] Among them, G(z,X) represents the generated pseudo multimodal sentiment data, W g and b g are the weight and bias of the generator, and Attention is the multimodal fusion attention mechanism;
[0026] D(X)=σ(W d ·X+b d );
[0027] Among them, D(X) represents the true and false probability output of the discriminator, W d and b d is the weight and bias of the discriminator, σ is the activation function;
[0028] S23. A joint optimization mechanism is used to train the generative adversarial network, and the generalization ability of the generative adversarial network is optimized by balancing the overall data authenticity loss and the modality consistency loss:
[0029]
[0030] in, Represents resistance to loss, represents the modal consistency loss, M i (X) is the i-th mode projection of the original data, M i (G(z,X)) is the i-th mode projection of the generated data;
[0031] S24. The generator combines sentiment and educational domain knowledge to adjust the strategy:
[0032]
[0033] Among them, G enh(z, X) is the enhanced data, K(X, z) represents the domain-specific sentiment dynamic correlation function, and λ is the adjustment coefficient;
[0034] S25. The multimodal pseudo emotion data G generated by the generator G enh (z,X) is fused with the original data, and the subset G closest to the actual emotional state is extracted through the embedded optimization algorithm opt (X), and finally output enhanced multimodal high-quality sentiment data:
[0035]
[0036] Among them, G opt (X) is the final output high-quality multimodal sentiment data, R(G enh ) is a regularization term used to balance data complexity and consistency, and α is a weight parameter.
[0037] Optionally, the S3 specifically includes:
[0038] S31. Multimodal high-quality sentiment data G opt (X) Perform modal decomposition to obtain facial expression feature data Speech and intonation feature data and touch behavior characteristic data The feature data of each modality is processed in time series to generate a sequence representation S f (t), S a (t) and S s (t), where each S i (t) represents the characteristic sequence of mode i in the time dimension;
[0039] S32. Construct an emotion recognition model T based on the Transformer architecture, wherein the emotion recognition model includes:
[0040] Embedding layer, which transforms the multimodal feature sequence S f (t), S a (t), S s (t) is mapped to a high-dimensional feature space to generate an embedding vector E f (t), E a (t), E s (t):
[0041]
[0042] in, and Represent the embedding weight and bias of modality i respectively;
[0043] Attention mechanism, which calculates the correlation within and between modalities through a multi-head self-attention mechanism to obtain the attention weight
[0044]
[0045] Among them, Q k ,K k are the query matrix and key matrix of the kth head respectively, and d is the dimension of the embedding vector;
[0046] Position encoding: adding position encoding P(t) to time series data to generate position-aware features E pos (t):
[0047] E pos (t) = concat(E f (t),E a (t),E s (t))+P(t);
[0048] S33. Use the encoder structure of the Transformer model to deeply fuse multimodal features and extract the global sentiment feature vector F global .
[0049] Optionally, the S4 specifically includes:
[0050] S41. The global sentiment feature vector F global Input the emotional state classification unit and use the emotional state mapping function M c Perform nonlinear mapping on the global sentiment feature vector to generate the classification probability distribution P c :
[0051] P c =softmax(W c ·F global +b c );
[0052] Among them, W c is the classification weight matrix, b c is the bias vector, P c Represents the probability distribution of each emotional state category;
[0053] S42. Generate the classification probability distribution P c Optimize and use the cross entropy loss function to calculate the true sentiment label Y and the predicted probability P c The differences:
[0054]
[0055] Among them, K represents the number of emotional state categories, Y krepresents the annotation value of the true sentiment label in the kth category, P c,k represents the predicted probability of the k-th emotional state;
[0056] S43. According to the optimized classification probability distribution P c Select the emotional state with the largest probability as the current child’s emotional state label T label :
[0057] T label = argmax k (P c,k );
[0058] Among them, T label Indicates the emotional state label corresponding to the classification result, which is used to identify the current emotional state of the child.
[0059] Optionally, the S5 specifically includes:
[0060] S51. Using emotional state label T label And the enhanced multimodal high-quality sentiment data G opt (X), Generate the emotion-driven content vector V by jointly embedding multimodal emotion features c :
[0061]
[0062] in, High-quality emotional features representing facial expressions, voice intonation, and touch behavior, respectively. is the modality feature embedding weight, W t ,W l is the weight of the joint embedding of the emotional state label and the modality feature, b g ,b t is the bias term;
[0063] S52.Content generation model construction based on multi-stage generator:
[0064] Stage 1 is the global content template T template Generation of:
[0065]
[0066] Among them, α i is the attention weight, Q i ,K i are the query matrix and the key matrix, representing the global correlation between the emotion-driven content and the content library. Generate weights for content, P(i) is the position encoding function;
[0067] Phase 2 is the modal content refinement module, which generates voice feedback, image animation, and interactive game content:
[0068]
[0069] in, is the weight matrix of the modality decoder, b speech ,b image ,b game is the bias term;
[0070] Phase 3 is the emotion enhancement module, where the content is generated after emotion enhancement optimization:
[0071]
[0072] in, is the final optimized modal content, m∈{speech,image,game}, represents the matching loss between modality content and sentiment label, λ 1 To adjust the coefficient and control the intensity of emotional optimization.
[0073] Optionally, the S6 specifically includes:
[0074] S61. Optimized voice feedback content Image animation content and interactive gaming content Perform modal feature extraction and generate matching feature vector F match :
[0075]
[0076] Among them, W s ,W i ,W g is the modal feature mapping weight matrix, which acts on speech, image and interactive content respectively, b m is the matching bias term, F match Represents the integrated multimodal matching features, which are used for matching calculation with the educational content library;
[0077] S62. The content in the reading pen education content library is embedded in the function E lib Transformed into embedded feature vector F lib :
[0078] F lib =E lib (C lib );
[0079] Among them, C lib Represents all available content in the educational content library, E libis a feature embedding function that generates a high-dimensional feature vector F based on content semantics and features. lib ;
[0080] S63. Using content matching function M sim Matching feature vector F match and the embedding feature vector F lib Calculate the similarity score S sim :
[0081]
[0082] Among them, M sim Indicates the matching degree calculation based on cosine similarity, S sim Indicates the matching score between the optimized content and the education library content;
[0083] S64. Similarity score S sim Sort and select the k contents with the highest scores as recommendation candidates:
[0084] C rec =Top k (S sim );
[0085] Among them, C rec Represents the final recommended content set, Top k is a function for selecting the top k contents according to similarity scores;
[0086] S65. Recommend content set C rec Perform personalized fine-tuning based on the child's emotional state label T label and historical preference features to generate the final personalized recommendation content list C final :
[0087]
[0088] in, represents the deviation loss between the recommended content and the emotional state label, λ 2 It is the adjustment coefficient used to control the fine-tuning intensity.
[0089] The beneficial effects of the present invention are:
[0090] (1) The present invention combines the generative adversarial network and the Transformer architecture to achieve real-time, high-precision emotion recognition by enhancing multimodal emotion data and extracting deep features. A multi-layer attention mechanism is used in the generative adversarial network to reconstruct features and optimize emotion consistency of the collected multimodal data, so that the model can generate high-quality pseudo-multimodal data from limited training data, thereby enhancing the generalization ability of the model. Through the multi-head attention mechanism of the Transformer architecture, the temporal and semantic correlation of multimodal data is fully explored, and the fuzzy classification problem of emotional state is solved in the emotion recognition scenario of the reading pen, and the adaptability of the model to complex emotional states is significantly improved.
[0091] (2) The present invention generates multimodal educational content that matches the current emotional state of children by designing a multi-stage content generation model, including three stages: global planning, modal refinement, and emotional reinforcement. The global planning module generates content templates using emotion-driven embedding vectors to ensure that content generation has clear teaching objectives and logical frameworks. The modal refinement module optimizes voice feedback, image animation, and interactive game content based on different emotional states. The emotional reinforcement module further adjusts the correlation between the generated content and the emotional state through a gradient optimization matching function. The multi-stage design makes up for the lack of personalization in traditional content generation technology, making the generated content more accurate and in line with children's immediate learning needs.
[0092] (3) The present invention introduces a dynamic mapping and recommendation optimization mechanism in the matching of generated content with the educational content library of the reading pen, and utilizes cosine similarity calculation and emotion-driven fine-tuning to generate a personalized recommended content list that meets children's needs. In the similarity calculation, efficient screening of content matching is achieved by combining multimodal emotional features and semantic embedding of educational content. In the recommendation optimization link, the recommended content is dynamically adjusted in combination with children's historical learning data and immediate emotional state, ensuring that the content recommendation not only meets personalized needs but also has diversity and breadth. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0094] Figure 1 The present invention provides a flowchart of an intelligent children's education method based on emotion recognition using a reading pen. DETAILED DESCRIPTION
[0095] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0096] refer to Figure 1 , an intelligent children's education method based on emotion recognition using a reading pen, comprising the following steps:
[0097] S1. Collect children’s facial expression data, voice intonation data and touch behavior data through the built-in camera, microphone and touch sensor of the reading pen;
[0098] S2. Use a generative adversarial network to enhance the collected facial expression data, voice intonation data, and touch behavior data to generate multimodal high-quality emotional data that matches the actual emotional state;
[0099] S3. Extract features from the enhanced multimodal high-quality emotion data through an emotion recognition model based on the Transformer architecture to generate deep emotion features related to facial expressions, voice intonation, and touch behavior;
[0100] S4. using the emotion state classification unit to classify the deep emotion features and generate an emotion state label corresponding to the current emotion state of the child;
[0101] S5. Generate voice feedback content, image animation content, and interactive game content that are appropriate to the current emotional state using a content generation model based on the Transformer architecture according to the emotional state label;
[0102] S6. Generate a personalized list of recommended content that matches the child's emotional state based on the matching results of the voice feedback content, image animation content, and interactive game content with the educational content library of the reading pen;
[0103] S7. Present educational content in a personalized recommended content list through the display screen, voice output unit, and touch interaction unit of the reading pen, output teaching content including voice feedback, image animation, or interactive games, and record relevant operation data in real time when the child interacts with the reading pen.
[0104] In this implementation, S1 specifically includes:
[0105] S11. Collect the child's facial expression data through the built-in camera of the reading pen, and the facial expression data includes the key point coordinate data P f (x,y), where x and y represent the horizontal and vertical coordinates of the facial landmarks, respectively;
[0106] S12. The voice and intonation data of the child are collected through the built-in microphone of the reading pen. The voice and intonation data include the audio signal amplitude A(t) and the frequency distribution F(ω), where t represents the time variable and ω represents the frequency variable;
[0107] S13. The touch sensor built into the reading pen is used to collect the child's touch behavior data, which includes touch intensity S(t) and touch position P t (x,y).
[0108] In this implementation, S2 specifically includes:
[0109] S21. Decompose the collected facial expression data, voice intonation data and touch behavior data in terms of time and space characteristics, and use a multi-level feature encoder to hierarchically represent data of different modalities:
[0110] Facial expression data f (x,y,t) is converted into a dynamic expression feature vector Extract local time features through the time window Δt;
[0111] Speech intonation data A(t), F(ω) is converted into spectral time series feature tensor Capture intonation and emotional patterns;
[0112] Touch behavior data S(t),P t (x,y) is converted into a touch event flow matrix Mapping the temporal changes and intensity distribution of touch behavior;
[0113] S22. Build a generative adversarial network model, including the module design of the generator and the discriminator:
[0114] G(z,X)=Attention(W g ·[z,X]+b g );
[0115] Among them, G(z,X) represents the generated pseudo multimodal sentiment data, W g and b g are the weight and bias of the generator, and Attention is the multimodal fusion attention mechanism;
[0116] D(X)=σ(W d ·X+b d );
[0117] Among them, D(X) represents the true and false probability output of the discriminator, W d and b d is the weight and bias of the discriminator, σ is the activation function;
[0118] S23. A joint optimization mechanism is used to train the generative adversarial network, and the generalization ability of the generative adversarial network is optimized by balancing the overall data authenticity loss and the modality consistency loss:
[0119]
[0120] in, Represents resistance to loss, represents the modal consistency loss, M i (X) is the i-th mode projection of the original data, M i (G(z,X)) is the i-th mode projection of the generated data;
[0121] S24. The generator combines sentiment and educational domain knowledge to adjust the strategy:
[0122]
[0123] Among them, G enh (z, X) is the enhanced data, K(X, z) represents the domain-specific sentiment dynamic correlation function, and λ is the adjustment coefficient;
[0124] S25. The multimodal pseudo emotion data G generated by the generator G enh (z,X) is fused with the original data, and the subset G closest to the actual emotional state is extracted through the embedded optimization algorithm opt (X), and finally output enhanced multimodal high-quality sentiment data:
[0125]
[0126] Among them, G opt (X) is the final output high-quality multimodal sentiment data, R(G enh ) is a regularization term used to balance data complexity and consistency, and α is a weight parameter.
[0127] In this implementation, S3 specifically includes:
[0128] S31. Multimodal high-quality sentiment data G opt (X) Perform modal decomposition to obtain facial expression feature data Speech and intonation feature data and touch behavior characteristic data The feature data of each modality is processed in time series to generate a sequence representation S f (t), S a (t) and S s (t), where each S i (t) represents the characteristic sequence of mode i in the time dimension;
[0129] S32. Construct an emotion recognition model T based on the Transformer architecture. The emotion recognition model includes:
[0130] Embedding layer, which transforms the multimodal feature sequence S f (t), S a(t), S s (t) is mapped to a high-dimensional feature space to generate an embedding vector E f (t), E a (t), E s (t):
[0131]
[0132] in, and Represent the embedding weight and bias of modality i respectively;
[0133] Attention mechanism, which calculates the correlation within and between modalities through a multi-head self-attention mechanism to obtain the attention weight
[0134]
[0135] Among them, Q k ,K k are the query matrix and key matrix of the kth head respectively, and d is the dimension of the embedding vector;
[0136] Position encoding: adding position encoding P(t) to time series data to generate position-aware features E pos (t):
[0137] E pos (t) = concat(E f (t),E a (t),E s (t))+P(t);
[0138] S33. Use the encoder structure of the Transformer model to deeply fuse multimodal features and extract the global sentiment feature vector F global .
[0139] In this implementation, S4 specifically includes:
[0140] S41. The global sentiment feature vector F global Input the emotional state classification unit and use the emotional state mapping function M c Perform nonlinear mapping on the global sentiment feature vector to generate the classification probability distribution P c :
[0141] P c =softmax(W c ·F global +b c );
[0142] Among them, W c is the classification weight matrix, bc is the bias vector, P c Represents the probability distribution of each emotional state category;
[0143] S42. Generate the classification probability distribution P c Optimize and use the cross entropy loss function to calculate the true sentiment label Y and the predicted probability P c The differences:
[0144]
[0145] Among them, K represents the number of emotional state categories, Y k represents the annotation value of the true sentiment label in the kth category, P c,k represents the predicted probability of the k-th emotional state;
[0146] S43. According to the optimized classification probability distribution P c Select the emotional state with the largest probability as the current child’s emotional state label T label :
[0147] T label = argmax k (P c,k );
[0148] Among them, T label Indicates the emotional state label corresponding to the classification result, which is used to identify the current emotional state of the child.
[0149] In this implementation, S5 specifically includes:
[0150] S51. Using emotional state label T label And the enhanced multimodal high-quality sentiment data G opt (X), Generate the emotion-driven content vector V by jointly embedding multimodal emotion features c :
[0151]
[0152] in, High-quality emotional features representing facial expressions, voice intonation, and touch behavior, respectively. is the modality feature embedding weight, W t ,W l is the weight of the joint embedding of the emotional state label and the modality feature, b g ,b t is the bias term;
[0153] S52.Content generation model construction based on multi-stage generator:
[0154] Stage 1 is the global content template Ttemplate Generation of:
[0155]
[0156] Among them, α i is the attention weight, Q i ,K i are the query matrix and the key matrix, respectively, representing the global correlation between the emotion-driven content and the content library, W i e Generate weights for content, P(i) is the position encoding function;
[0157] Phase 2 is the modal content refinement module, which generates voice feedback, image animation, and interactive game content:
[0158]
[0159]
[0160] in, is the weight matrix of the modality decoder, b speech ,b image ,b game is the bias term;
[0161] Phase 3 is the emotion enhancement module, where the content is generated after emotion enhancement optimization:
[0162]
[0163] in, is the final optimized modal content, m∈{speech,image,game}, represents the matching loss between modality content and sentiment label, λ 1 To adjust the coefficient and control the intensity of emotional optimization.
[0164] In this implementation, S6 specifically includes:
[0165] S61. Optimized voice feedback content Image animation content and interactive gaming content Perform modal feature extraction and generate matching feature vector F match :
[0166]
[0167] Among them, W s ,W i ,W g is the modal feature mapping weight matrix, which acts on speech, image and interactive content respectively, b m is the matching bias term, Fmatch Represents the integrated multimodal matching features, which are used for matching calculation with the educational content library;
[0168] S62. The content in the reading pen education content library is embedded in the function E lib Transformed into embedded feature vector F lib :
[0169] F lib =E lib (C lib );
[0170] Among them, C lib Represents all available content in the educational content library, E lib is a feature embedding function that generates a high-dimensional feature vector F based on content semantics and features. lib ;
[0171] S63. Using content matching function M sim Matching feature vector F match and the embedding feature vector F lib Calculate the similarity score S sim :
[0172]
[0173] Among them, M sim Indicates the matching degree calculation based on cosine similarity, S sim Indicates the matching score between the optimized content and the education library content;
[0174] S64. Similarity score S sim Sort and select the k contents with the highest scores as recommendation candidates:
[0175] C rec =Top k (S sim );
[0176] Among them, C rec Represents the final recommended content set, Top k is a function for selecting the top k contents according to similarity scores;
[0177] S65. Recommend content set C rec Perform personalized fine-tuning based on the child's emotional state label T label and historical preference features to generate the final personalized recommendation content list C final :
[0178]
[0179] in, represents the deviation loss between the recommended content and the emotional state label, λ 2 It is the adjustment coefficient used to control the fine-tuning intensity.
[0180] Embodiment 1:
[0181] Embodiment At 3 pm on December 5, 2024, at a children's education and training center in City A, Xiao Ming (pseudonym), an 8-year-old student, was using a smart reading pen based on emotion recognition to learn scientific knowledge. The reading pen is equipped with multimodal emotion recognition and content generation technology. Its supporting educational content library includes 5,000 voice feedback units, 1,200 image animation materials and 300 interactive game modules. The education and training center hopes to improve Xiao Ming's learning attention and interest through this device.
[0182] In the first 15 minutes of learning, Xiao Ming began to learn basic science knowledge modules. The reading pen collected Xiao Ming's facial expression data through the camera, and analyzed in real time that his eyebrows were naturally stretched, the corners of his mouth were slightly raised, and his voice was steady. The touch sensing data recorded that the force of each touch was about 1.2N, and the interval was between 1.5 seconds and 2 seconds, indicating that Xiao Ming was in a "focused" state (the emotional state label was "focused"). The system generated a learning feedback log:
[0183] Time: 15:03:45
[0184] Facial features: Normal concentration, no obvious frowning or squinting;
[0185] Voice characteristics: steady intonation, speaking speed about 120 words per minute;
[0186] Touch data: Touch force is uniform, 1.2±0.2N;
[0187] Emotional state: Focused;
[0188] Feedback content: Voice encouragement "You are doing great, let's keep going!"
[0189] The reading pen then pushed a popular science animation, which described the principles of plant photosynthesis and combined it with interesting interactive games to help Xiao Ming consolidate the knowledge points. During the interaction, the system recorded that Xiao Ming's error click rate was 5% and the completion time was 3 minutes and 20 seconds, which was 20% shorter than expected.
[0190] At 15:20, Xiao Ming started to work on the relatively difficult task - the periodic table learning module. The system captured Xiao Ming's frown and downward mouth through the camera, and his facial expression showed that he was in a "frustrated" state (emotional state label was "frustrated"). Voice analysis found that Xiao Ming's speech speed dropped to 80 words per minute, accompanied by multiple pauses. The touch sensor recorded that his touch force dropped from 1.2N to 0.8N, and the touch interval lengthened to more than 4 seconds. The system automatically generated an emotional state analysis report:
[0191] Time: 15:21:12
[0192] Facial features: frowning brows, corners of mouth depressed;
[0193] Voice characteristics: The speaking speed decreases and pauses occur;
[0194] Touch data: The touch force dropped to 0.8±0.1N, and the interval increased to 4.3 seconds;
[0195] Emotional state: frustrated;
[0196] System behavior: Switch the teaching content to "encouragement mode".
[0197] The reading pen played the voice message "Don't worry, we can do it step by step!" and pushed a cartoon animation about a scientific explorer who overcame difficulties to learn the periodic table. It also provided a step-by-step interactive game to help Xiao Ming gradually become familiar with the basic classification of elements. After completing the interactive game, the system recorded that his emotional state changed from "frustrated" to "positive".
[0198] At 15:40, Xiao Ming completed the periodic table learning module, and his emotional state label was displayed as "positive". Voice analysis recorded that his speech speed returned to 130 words per minute, and touch data recorded that the touch force returned to 1.3N, and the touch interval was shortened to 1.8 seconds. The system pushed the advanced task "Basic Operations of Chemical Experiments" and generated a feedback log:
[0199] Time: 15:41:00
[0200] Facial features: eyebrows stretched, eyes bright;
[0201] Voice characteristics: speech speed recovered, tone rose;
[0202] Touch data: Stable force, 1.3±0.1N;
[0203] Emotional state: positive;
[0204] System behavior: Recommend advanced learning tasks.
[0205] The system pushed a lab operation animation and combined it with real-time voice guidance to help Xiao Ming complete the lab operation simulation. The lab was completed in 10 minutes, 15% shorter than expected.
[0206] It can be seen from the embodiments that the reading pen of the present invention can accurately perceive the emotional state of children, and significantly improve learning efficiency and reduce error rate by dynamically adjusting learning content, while enhancing learning interest, truly realizing the personalization and intelligence of children's education.
[0207] The present invention combines generative adversarial networks and Transformer architectures, and realizes real-time, high-precision emotion recognition by enhancing multimodal emotion data and extracting deep features. A multi-layer attention mechanism is used in the generative adversarial network to reconstruct features and optimize emotion consistency of the collected multimodal data, so that the model can generate high-quality pseudo-multimodal data from limited training data and enhance the generalization ability of the model. Through the multi-head attention mechanism of the Transformer architecture, the temporal and semantic correlation of multimodal data is fully explored, which solves the fuzzy classification problem of emotional states in the emotion recognition scenario of the reading pen and significantly improves the adaptability of the model to complex emotional states.
[0208] The present invention generates multimodal educational content that matches the current emotional state of children by designing a multi-stage content generation model, including three stages: global planning, modal refinement and emotional reinforcement. The global planning module generates content templates using emotion-driven embedding vectors to ensure that content generation has clear teaching objectives and logical frameworks. The modal refinement module optimizes voice feedback, image animation and interactive game content based on different emotional states. The emotional reinforcement module further adjusts the correlation between the generated content and the emotional state through a gradient optimization matching function. The multi-stage design makes up for the lack of personalization in traditional content generation technology, making the generated content more accurate and in line with children's immediate learning needs.
[0209] The present invention introduces a dynamic mapping and recommendation optimization mechanism in the matching of generated content with the educational content library of the reading pen, and uses cosine similarity calculation and emotion-driven fine-tuning to generate a personalized recommended content list that meets children's needs. In the similarity calculation, efficient screening of content matching is achieved by combining multimodal emotional features and semantic embedding of educational content. In the recommendation optimization link, the recommended content is dynamically adjusted in combination with children's historical learning data and immediate emotional state, ensuring that the content recommendation not only meets personalized needs but also has diversity and breadth.
[0210] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. An intelligent children's education method based on emotion recognition using a reading pen, characterized in that: The steps include: S1. Collect children’s facial expression data, voice intonation data and touch behavior data through the built-in camera, microphone and touch sensor of the reading pen; S2. Use a generative adversarial network to enhance the collected facial expression data, voice intonation data, and touch behavior data to generate multimodal high-quality emotional data that matches the actual emotional state; S3. Extract features from the enhanced multimodal high-quality emotion data through an emotion recognition model based on the Transformer architecture to generate deep emotion features related to facial expressions, voice intonation, and touch behavior; S4. using the emotion state classification unit to classify the deep emotion features and generate an emotion state label corresponding to the current emotion state of the child; S5. Generate voice feedback content, image animation content, and interactive game content that are appropriate to the current emotional state using a content generation model based on the Transformer architecture according to the emotional state label; S6. Generate a personalized list of recommended content that matches the child's emotional state based on the matching results of the voice feedback content, image animation content, and interactive game content with the educational content library of the reading pen; S7. Present educational content in a personalized recommended content list through the display screen, voice output unit, and touch interaction unit of the reading pen, output teaching content including voice feedback, image animation, or interactive games, and record relevant operation data in real time when the child interacts with the reading pen.
2. According to claim 1, the method for intelligent children's education based on emotion recognition using a point reading pen is characterized in that: The S1 specifically includes: S11. Collect the child's facial expression data through the built-in camera of the reading pen, the facial expression data includes key point coordinate data P f (x,y), where x and y represent the horizontal and vertical coordinates of the facial landmarks, respectively; S12. The voice and intonation data of the child is collected through the built-in microphone of the reading pen, wherein the voice and intonation data includes the audio signal amplitude A(t) and the frequency distribution F(ω), where t represents the time variable and ω represents the frequency variable; S13. The touch sensor built into the reading pen is used to collect the child's touch behavior data, wherein the touch behavior data includes the touch intensity S(t) and the touch position P t (x,y).
3. The method for intelligent children's education based on emotion recognition using a point reading pen according to claim 1, characterized in that: The S2 specifically includes: S21. Decompose the collected facial expression data, voice intonation data and touch behavior data in terms of time and space characteristics, and use a multi-level feature encoder to hierarchically represent data of different modalities: Facial expression data f (x,y,t) is converted into a dynamic expression feature vector Extract local time features through the time window Δt; Speech intonation data A(t), F(ω) is converted into spectral time series feature tensor Capture intonation and emotional patterns; Touch behavior data S(t),P t (x,y) is converted into a touch event flow matrix Mapping the temporal changes and intensity distribution of touch behavior; S22. Build a generative adversarial network model, including the module design of the generator and the discriminator: G(z,X)=Attention(W g ·[z,X]+b g ); Among them, G(z,X) represents the generated pseudo multimodal sentiment data, W g and b g are the weight and bias of the generator, and Attention is the multimodal fusion attention mechanism; D(X)=σ(W d ·X+b d ); Among them, D(X) represents the true and false probability output of the discriminator, W d and b d is the weight and bias of the discriminator, σ is the activation function; S23. A joint optimization mechanism is used to train the generative adversarial network, and the generalization ability of the generative adversarial network is optimized by balancing the overall data authenticity loss and the modality consistency loss: in, Represents resistance to loss, represents the modal consistency loss, M i (X) is the i-th mode projection of the original data, M i (G(z,X)) is the i-th mode projection of the generated data; S24. The generator combines sentiment and educational domain knowledge to adjust the strategy: Among them, G enh (z, X) is the enhanced data, K(X, z) represents the domain-specific sentiment dynamic correlation function, and λ is the adjustment coefficient; S25. The multimodal pseudo emotion data G generated by the generator G enh (z,X) is fused with the original data, and the subset G closest to the actual emotional state is extracted through the embedded optimization algorithm opt (X), and finally output enhanced multimodal high-quality sentiment data: Among them, G opt (X) is the final output high-quality multimodal sentiment data, R(G enh ) is a regularization term used to balance data complexity and consistency, and α is a weight parameter.
4. The method for intelligent children's education based on emotion recognition using a point reading pen according to claim 1, characterized in that: The S3 specifically includes: S31. Multimodal high-quality sentiment data G opt (X) Perform modal decomposition to obtain facial expression feature data Speech and intonation feature data and touch behavior characteristic data The feature data of each modality is processed in time series to generate a sequence representation S f (t), S a (t) and S s (t), where each S i (t) represents the characteristic sequence of mode i in the time dimension; S32. Construct an emotion recognition model T based on the Transformer architecture, wherein the emotion recognition model includes: Embedding layer, which transforms the multimodal feature sequence S f (t), S a (t), S s (t) is mapped to a high-dimensional feature space to generate an embedding vector E f (t), E a (t), E s (t): in, and Represent the embedding weight and bias of modality i respectively; Attention mechanism, which calculates the correlation within and between modalities through a multi-head self-attention mechanism to obtain the attention weight Among them, Q k ,K k are the query matrix and key matrix of the kth head respectively, and d is the dimension of the embedding vector; Position encoding: adding position encoding P(t) to time series data to generate position-aware features E pos (t): E pos (t)=concat(E f (t),E a (t),E s (t))+P(t); S33. Use the encoder structure of the Transformer model to deeply fuse multimodal features and extract the global sentiment feature vector F global .
5. The method for intelligent children's education based on emotion recognition using a point reading pen according to claim 1, characterized in that: The S4 specifically includes: S41. The global sentiment feature vector F global Input the emotional state classification unit and use the emotional state mapping function M c Perform nonlinear mapping on the global sentiment feature vector to generate the classification probability distribution P c : P c =softmax(W c ·F global +b c ); Among them, W c is the classification weight matrix, b c is the bias vector, P c Represents the probability distribution of each emotional state category; S42. Generate the classification probability distribution P c Optimize and use the cross entropy loss function to calculate the true sentiment label Y and the predicted probability P c The differences: Among them, K represents the number of emotional state categories, Y k represents the annotation value of the true sentiment label in the kth category, P c,k represents the predicted probability of the k-th emotional state; S43. According to the optimized classification probability distribution P c Select the emotional state with the largest probability as the current child’s emotional state label T label : T label =argmax k (P c,k ); Among them, T label Indicates the emotional state label corresponding to the classification result, which is used to identify the current emotional state of the child.
6. The method for intelligent children's education based on emotion recognition using a point reading pen according to claim 1, characterized in that: The S5 specifically includes: S51. Using emotional state label T label And the enhanced multimodal high-quality sentiment data G opt (X), Generate the emotion-driven content vector V by jointly embedding multimodal emotion features c : in, High-quality emotional features representing facial expressions, voice intonation, and touch behavior, respectively. is the modality feature embedding weight, W t ,W l is the weight of the joint embedding of the emotional state label and the modality feature, b g ,b t is the bias term; S52.Content generation model construction based on multi-stage generator: Stage 1 is the global content template T template Generation of: Among them, α i is the attention weight, Q i ,K i are the query matrix and the key matrix, representing the global correlation between the emotion-driven content and the content library. Generate weights for content, P(i) is the position encoding function; Phase 2 is the modal content refinement module, which generates voice feedback, image animation, and interactive game content: in, is the weight matrix of the modality decoder, b speech ,b image ,b game is the bias term; Phase 3 is the emotion enhancement module, where the content is generated after emotion enhancement optimization: in, is the final optimized modal content, m∈{speech,image,game}, It represents the matching loss between modal content and sentiment label, and λ1 is the adjustment coefficient to control the intensity of sentiment optimization.
7. The method for intelligent children's education based on emotion recognition using a point reading pen according to claim 1, characterized in that: The S6 specifically includes: S61. Optimized voice feedback content Image animation content and interactive game content V o ga pt me Perform modal feature extraction and generate matching feature vector F match : Among them, W s ,W i ,W g is the modal feature mapping weight matrix, which acts on speech, image and interactive content respectively, b m is the matching bias term, F match Represents the integrated multimodal matching features, which are used for matching calculation with the educational content library; S62. The content in the reading pen education content library is embedded in the function E lib Transformed into embedded feature vector F lib : F lib =E lib (C lib ); Among them, C lib Represents all available content in the educational content library, E lib is a feature embedding function that generates a high-dimensional feature vector F based on content semantics and features. lib ; S63. Using content matching function M sim Matching feature vector F match and the embedding feature vector F lib Calculate the similarity score S sim : Among them, M sim Indicates the matching degree calculation based on cosine similarity, S sim Indicates the matching score between the optimized content and the education library content; S64. Similarity score S sim Sort and select the k contents with the highest scores as recommendation candidates: C rec =Top k (S sim ); Among them, C rec Represents the final recommended content set, Top k is a function for selecting the top k contents according to similarity scores; S65. Recommend content set C rec Perform personalized fine-tuning based on the child's emotional state label T label and historical preference features to generate the final personalized recommendation content list C final : in, represents the deviation loss between the recommended content and the emotional state label, and λ2 is the adjustment coefficient used to control the fine-tuning strength.
Citation Information
Cited By
Multi-mode self-adaptive child voice interaction method for education service robot
CN120164478A
Intelligent management method for interactive data of touch and talk pen based on AI analysis
CN120234357A
An intelligent management method for interactive data of reading pens based on AI analysis
CN120234357B
Child multi-modal data output method and system
CN120296704A
AI-based learning question personalized recommendation method
CN120318042A