Education robot with camera and information interaction identification processing system and method
By integrating cameras and multiple recognition modules in educational robots, and adjusting teaching strategies based on user emotions, the problem of existing educational robots lacking flexibility and adaptability is solved, and the teaching effect and user experience are improved.
Patent Information
- Application Number
- CN202510392374.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Existing educational robots lack flexibility and adaptability and cannot adjust teaching strategies based on users' emotional state.
An educational robot with a camera is adopted to integrate the user's facial expressions and voice through the image acquisition module, voice acquisition module, expression recognition module, voice recognition module and emotion recognition module to judge the user's emotional state, and the central processing module selects appropriate teaching strategies, and outputs teaching content through the speaker and display module.
It realizes that educational robots adjust teaching strategies based on user emotions, improves teaching effect and user experience, and provides a personalized learning experience.
Smart Images

Figure CN120276599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent devices, and particularly to an educational robot with a camera, an information interaction recognition and processing system, and a method. Background Art
[0002] With the rapid development of information technology, various intelligent devices and technologies have gradually been introduced into the education field to improve teaching quality and learning efficiency. These technologies include, but are not limited to, online learning platforms, virtual reality, augmented reality, and educational robots. Among many emerging teaching tools, educational robots have received extensive attention because they can provide personalized and highly interactive learning experiences.
[0003] Educational robots can provide users with personalized learning experiences through their intelligent designs and rich interaction functions. They can adjust teaching content according to users' learning progress and interests, thereby better meeting the needs of different users. Strong interactivity is another major feature of educational robots. They can communicate with users in various ways such as voice and vision, making the learning process more vivid and interesting.
[0004] Although educational robots have many advantages, most current educational robots still adopt fixed preset teaching strategies. This means that regardless of the actual emotional state of the user, the teaching content and methods provided by the robot are pre-set, lacking flexibility and adaptability. Summary of the Invention
[0005] The purpose of the present invention is to provide an educational robot with a camera, an information interaction recognition and processing system, and a method, which can adjust teaching strategies according to users' emotions, greatly improving the teaching effect and user experience of educational robots.
[0006] To achieve the above object, in a first aspect, the present invention provides an information interaction recognition and processing system for an educational robot with a camera, including an image acquisition module, a voice acquisition module, an expression recognition module, a voice recognition module, an emotion recognition module, a central processing module, a speaker module, and a display module;
[0007] The expression recognition module is connected to the image acquisition module; the voice recognition module is connected to the voice acquisition module; the emotion recognition module is respectively connected to the expression recognition module and the voice recognition module; the central processing module is respectively connected to the emotion recognition module and the voice recognition module; the speaker module is connected to the central processing module; the display module is connected to the central processing module;
[0008] The image acquisition module is used to capture the user's facial image from multiple angles;
[0009] The voice collection module is used to collect the user's voice;
[0010] The facial expression recognition module is used to recognize the user's facial expression according to the user's facial image,
[0011] The voice recognition module is used to recognize the user's voice;
[0012] The emotion recognition module is used to comprehensively judge the user's emotional state based on the user's facial expression and voice;
[0013] The central processing module is used to select appropriate teaching strategies according to the user's emotional state;
[0014] The loudspeaker module is used to output teaching content in the form of voice;
[0015] The display module is used to output teaching content in the form of video.
[0016] Among them, the image collection module includes a camera unit and an image preprocessing unit; the image preprocessing unit is connected to the camera unit;
[0017] The camera unit is used to capture the user's facial image from multiple angles;
[0018] The image preprocessing unit is used to preprocess the user's facial image to facilitate subsequent processing.
[0019] Among them, the voice collection module includes a microphone array unit and a voice preprocessing unit; the voice preprocessing unit is connected to the microphone array unit;
[0020] The microphone array unit is used to accurately capture the user's voice input;
[0021] The voice preprocessing unit is used to preprocess the voice input by the user to facilitate subsequent processing.
[0022] Among them, the facial expression recognition module includes an expression feature extraction unit and an expression classification unit; the expression classification unit is connected to the expression feature extraction unit;
[0023] The expression feature extraction unit is used to extract high-level features according to the user's facial image;
[0024] The expression classification unit is used to output the probability distribution of various expressions according to the extracted high-level features.
[0025] Among them, the speech recognition module includes a speech feature extraction unit, a speech recognition unit, a semantic understanding unit, and an intonation analysis unit; the speech recognition unit is connected to the speech feature extraction unit; the semantic understanding unit is connected to the speech recognition unit; the intonation analysis unit is connected to the speech feature extraction unit;
[0026] The speech feature extraction unit is configured to extract high-level speech features according to the user's speech input;
[0027] The speech recognition unit is configured to decode the user's speech input into text output according to the extracted high-level features;
[0028] The semantic understanding unit is configured to parse the semantics of the user's speech input according to the text output;
[0029] The intonation analysis unit is configured to analyze the user's intonation according to the extracted high-level features.
[0030] Among them, the emotion recognition module includes an information integration unit and an emotion judgment unit; the emotion judgment unit is connected to the information integration unit;
[0031] The information integration unit is configured to integrate the probability distribution of various expressions output by the expression classification unit and the user's intonation analyzed by the intonation analysis unit;
[0032] The emotion judgment unit is configured to judge the user's emotional state according to the information integrated by the information integration unit and feedback it to the central processing module.
[0033] Among them, the central processing module includes a storage unit, a strategy adjustment unit, and a communication unit; the strategy adjustment unit is connected to the storage unit; the communication unit is connected to the storage unit;
[0034] The storage unit is configured to store a teaching strategy database;
[0035] The strategy adjustment unit is configured to select the most suitable teaching strategy from the teaching strategy database according to the user's emotional state and the semantics of the user's speech input, and output teaching content through the display module and the speaker module;
[0036] The communication unit is configured to connect the storage unit to an external server to update the teaching strategy database in real time.
[0037] In a second aspect, the present invention also provides an information interaction recognition and processing method for an educational robot with a camera, including
[0038] The image acquisition module captures the user's facial images from multiple angles, and the voice acquisition module collects the user's voice;
[0039] The facial expression recognition module recognizes the user's facial expression based on the user's facial image, and the voice recognition module recognizes the user's voice;
[0040] The emotion recognition module comprehensively judges the user's emotional state based on the user's facial expression and voice;
[0041] The central processing module selects an appropriate teaching strategy according to the user's emotional state, and outputs teaching content through the speaker module and the display module.
[0042] Thirdly, the present invention also provides an educational robot with a camera, including a robot body, two eye cameras, two ear cameras, a mouth microphone array, a chest display, and two shoulder speakers;
[0043] The two eye cameras are respectively arranged on the robot body; the two ear cameras are respectively arranged on the robot body; the mouth microphone array is arranged on the robot body; the chest display is arranged on the robot body; the two shoulder speakers are respectively arranged on the robot body.
[0044] For an educational robot, an information interaction recognition and processing system and method of the present invention, the image acquisition module accurately captures the user's facial image from multiple angles, the voice acquisition module is used to collect the user's voice, the facial expression recognition module judges the user's expression through the facial image, the voice recognition module recognizes the user's voice, mainly recognizing semantics and intonation, the emotion recognition module can judge the user's emotional state according to the recognized user's expression and intonation, the central processing module understands the user's semantics, selects the corresponding teaching content, and selects an appropriate teaching method according to the user's emotional state, and then outputs the teaching content through the speaker module and the display module; for example, if the user is happy, provide teaching content with strong interesting interactivity, if the user is in a low mood, slow down the teaching content rhythm, decompose complex tasks into smaller and easier-to-complete tasks, guide step by step, and stop teaching in due time to play some soothing music. Thus, the teaching strategy can be adjusted according to the user's emotion, greatly improving the teaching effect and user experience of the educational robot. Brief Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art.
[0046] Figure 1 It is a schematic structural diagram of the first embodiment of the present invention.
[0047] Figure 2 It is a schematic structural diagram of the image acquisition module according to the first embodiment of the present invention.
[0048] Figure 3 It is a schematic structural diagram of the voice acquisition module according to the first embodiment of the present invention.
[0049] Figure 4 It is a schematic structural diagram of the facial expression recognition module according to the first embodiment of the present invention.
[0050] Figure 5 It is a schematic structural diagram of the speech recognition module according to the first embodiment of the present invention.
[0051] Figure 6 It is a schematic structural diagram of the emotion recognition module according to the first embodiment of the present invention.
[0052] Figure 7 It is a schematic structural diagram of the central processing module according to the first embodiment of the present invention.
[0053] Figure 8 It is a schematic flow diagram of the second embodiment of the present invention.
[0054] Figure 9 It is a schematic structural diagram of the third embodiment of the present invention.
[0055] 1 - Image acquisition module, 2 - Voice acquisition module, 3 - Facial expression recognition module, 4 - Speech recognition module, 5 - Emotion recognition module, 6 - Central processing module, 7 - Speaker module, 8 - Display module, 11 - Camera unit, 12 - Image preprocessing unit, 21 - Microphone array unit, 22 - Voice preprocessing unit, 31 - Facial expression feature extraction unit, 32 - Facial expression classification unit, 41 - Speech feature extraction unit, 42 - Speech recognition unit, 43 - Semantic understanding unit, 44 - Intonation analysis unit, 51 - Information integration unit, 52 - Emotion judgment unit, 61 - Storage unit, 62 - Strategy adjustment unit, 63 - Communication unit, 100 - Robot body, 200 - Eye camera, 300 - Ear camera, 400 - Mouth microphone array, 500 - Chest display, 600 - Shoulder speaker. Detailed implementation manners
[0056] The first embodiment of this application is as follows:
[0057] Please refer to Figures 1-7 , wherein, Figure 1 It is a schematic structural diagram of the first embodiment of the present invention. Figure 2 It is a schematic structural diagram of the image acquisition module according to the first embodiment of the present invention. Figure 3 It is a schematic structural diagram of the voice acquisition module according to the first embodiment of the present invention. Figure 4It is a schematic structural diagram of the facial expression recognition module according to the first embodiment of the present invention. Figure 5 It is a schematic structural diagram of the speech recognition module according to the first embodiment of the present invention. Figure 6 It is a schematic structural diagram of the emotion recognition module according to the first embodiment of the present invention. Figure 7 It is a schematic structural diagram of the central processing module according to the first embodiment of the present invention.
[0058] The present invention provides an information interaction recognition processing system for an educational robot with a camera, including an image acquisition module 1, a voice acquisition module 2, a facial expression recognition module 3, a speech recognition module 4, an emotion recognition module 5, a central processing module 6, a speaker module 7, and a display module 8. The image acquisition module 1 includes a camera unit 11 and an image preprocessing unit 12. The voice acquisition module 2 includes a microphone array unit 21 and a voice preprocessing unit 22. The facial expression recognition module 3 includes a facial expression feature extraction unit 31 and a facial expression classification unit 32. The speech recognition module 4 includes a speech feature extraction unit 41, a speech recognition unit 42, a semantic understanding unit 43, and an intonation analysis unit 44. The emotion recognition module 5 includes an information integration unit 51 and an emotion judgment unit 52. The central processing module 6 includes a storage unit 61, a strategy adjustment unit 62, and a communication unit 63. Through the foregoing solution, the teaching strategy can be adjusted according to the user's emotion, greatly improving the teaching effect and user experience of the educational robot.
[0059] Further, the facial expression recognition module 3 is connected to the image acquisition module 1. The speech recognition module 4 is connected to the voice acquisition module 2. The emotion recognition module 5 is respectively connected to the facial expression recognition module 3 and the speech recognition module 4. The central processing module 6 is respectively connected to the emotion recognition module 5 and the speech recognition module 4. The speaker module 7 is connected to the central processing module 6. The display module 8 is connected to the central processing module 6.
[0060] The image acquisition module 1 is used to capture the user's facial image from multiple angles.
[0061] The voice acquisition module 2 is used to collect the user's voice.
[0062] The facial expression recognition module 3 is used to recognize the user's facial expression according to the user's facial image.
[0063] The speech recognition module 4 is used to recognize the user's speech.
[0064] The emotion recognition module 5 is used to comprehensively judge the user's emotion state based on the user's facial expression and voice.
[0065] The central processing module 6 is used to select a suitable teaching strategy according to the user's emotion state.
[0066] The loudspeaker module 7 is used to output teaching content in the form of voice;
[0067] The display module 8 is used to output teaching content in the form of video.
[0068] In this embodiment, the image acquisition module 1 accurately captures the user's facial image from multiple angles, the voice acquisition module 2 is used to acquire the user's voice, the expression recognition module 3 judges the user's expression through the facial image, the voice recognition module 4 recognizes the user's voice, mainly recognizing semantics and intonation, and the emotion recognition module 5 can judge the user's emotional state according to the recognized user's expression and intonation. The central processing module 6 understands the user's semantics, selects the corresponding teaching content, and selects a suitable teaching method according to the user's emotional state, and then outputs the teaching content through the loudspeaker module 7 and the display module 8; for example, if the user is happy, provide teaching content with strong interesting interactivity. If the user is in a low mood, slow down the rhythm of the teaching content, decompose complex tasks into smaller and easier-to-complete tasks, guide step by step, and pause the teaching in due course to play some soothing music. Thus, the teaching strategy can be adjusted according to the user's emotions, greatly improving the teaching effect and user experience of the educational robot.
[0069] Further, the image acquisition module 1 includes a camera unit 11 and an image preprocessing unit 12; the image preprocessing unit 12 is connected to the camera unit 11;
[0070] The camera unit 11 is used to capture the user's facial image from multiple angles;
[0071] The image preprocessing unit 12 is used to preprocess the user's facial image for convenient subsequent processing.
[0072] In this embodiment, the camera unit 11 uses high-definition cameras at multiple positions to capture the user's facial image from different angles, ensuring that facial details can be accurately captured even when the user is moving. The graphic preprocessing is used to convert the color image into a grayscale image, reducing the computational complexity and increasing the processing speed; adjusting the image size and brightness to meet the requirements of subsequent processing; using a filtering algorithm to remove noise in the image and improve the image quality.
[0073] Further, the voice acquisition module 2 includes a microphone array unit 21 and a voice preprocessing unit 22; the voice preprocessing unit 22 is connected to the microphone array unit 21;
[0074] The microphone array unit 21 is used to accurately capture the user's voice input;
[0075] The voice preprocessing unit 22 is used to preprocess the voice input by the user to facilitate subsequent processing.
[0076] In this embodiment, the microphone array unit 21 uses a highly sensitive microphone array to capture the voice input of the user to ensure high-quality sound collection. The voice preprocessing unit 22 applies a noise reduction algorithm to remove background noise, improve voice clarity, and detect the start and end points of the voice, remove the silent part, and extract the effective voice segment to facilitate subsequent processing.
[0077] Furthermore, the facial expression recognition module 3 includes a facial expression feature extraction unit 31 and a facial expression classification unit 32; the facial expression classification unit 32 is connected to the facial expression feature extraction unit 31;
[0078] The facial expression feature extraction unit 31 is used to extract high-level features according to the facial image of the user;
[0079] The facial expression classification unit 32 is used to output the probability distribution of various facial expressions according to the extracted high-level features.
[0080] In this embodiment, the facial expression feature extraction unit 31 inputs the processed facial image into a pre-trained ResNet model, automatically extracts high-level features through multiple convolutional layers, pooling layers, and fully connected layers, and uses a Softmax classifier to output the probability distribution of each facial expression category, such as the probability of a happy expression, the probability of a confused expression, and the probability of a low-mood expression.
[0081] Furthermore, the voice recognition module 4 includes a voice feature extraction unit 41, a voice recognition unit 42, a semantic understanding unit 43, and an intonation analysis unit 44; the voice recognition unit 42 is connected to the voice feature extraction unit 41; the semantic understanding unit 43 is connected to the voice recognition unit 42; the intonation analysis unit 44 is connected to the voice feature extraction unit 41;
[0082] The voice feature extraction unit 41 is used to extract high-level voice features according to the voice input of the user;
[0083] The voice recognition unit 42 is used to decode the voice input of the user into text output according to the extracted high-level features;
[0084] The semantic understanding unit 43 is used to parse the semantics of the voice input of the user according to the text output;
[0085] The intonation analysis unit 44 is used to analyze the intonation of the user according to the extracted high-level features.
[0086] In this embodiment, the speech feature extraction unit 41 extracts MFCC features, fundamental frequency features, and prosody features, and uses a convolutional neural network to extract high-level speech features. The speech recognition unit 42 inputs the extracted features into an acoustic model, recognizes the corresponding phoneme sequence, and decodes it in combination with a language model to generate a text output. The semantic understanding unit 43 uses natural language processing to understand the user's semantics. The intonation analysis unit 44 judges the user's emotion based on the fundamental frequency features and prosody features.
[0087] Furthermore, the emotion recognition module 5 includes an information integration unit 51 and an emotion judgment unit 52. The emotion judgment unit 52 is connected to the information integration unit 51.
[0088] The information integration unit 51 is used to integrate the probability distribution of various expressions output by the expression classification unit 32 and the intonation of the user analyzed by the intonation analysis unit 44.
[0089] The emotion judgment unit 52 is used to judge the user's emotional state based on the information integrated by the information integration unit 51 and feedback it to the central processing module 6.
[0090] In this embodiment, the information integration unit 51 integrates the probability distribution of various expressions of the user and the intonation, and provides it to the emotion judgment unit 52. The emotion judgment unit 52 uses a deep learning model to judge the user's emotional state and feedback it to the central processing module 6.
[0091] Furthermore, the central processing module 6 includes a storage unit 61, a strategy adjustment unit 62, and a communication unit 63. The strategy adjustment unit 62 is connected to the storage unit 61. The communication unit 63 is connected to the storage unit 61.
[0092] The storage unit 61 is used to store a teaching strategy database.
[0093] The strategy adjustment unit 62 is used to select the most appropriate teaching strategy from the teaching strategy database according to the user's emotional state and the semantics of the user's speech input, and output teaching content through the display module 8 and the speaker module 7.
[0094] The communication unit 63 is used to connect the storage unit 61 to an external server to update the teaching strategy database in real time.
[0095] In this embodiment, there is a teaching strategy database in the storage unit 61, which is used to store different teaching strategies in various emotional states and is connected to an external server for real-time update through the communication unit 63. The strategy adjustment unit 62 understands the semantics of the user, combines the current emotional state of the user, selects the correct and appropriate teaching strategy, and outputs the teaching content through the display module 8 and the speaker module 7.
[0096] An information interaction recognition and processing system of an educational robot with a camera described in this embodiment. In this embodiment, the image acquisition module 1 accurately captures the facial image of the user from multiple angles. The voice acquisition module 2 is used to collect the voice of the user. The expression recognition module 3 judges the expression of the user through the facial image. The voice recognition module 4 recognizes the voice of the user, mainly recognizing semantics and intonation. The emotion recognition module 5 can judge the emotional state of the user according to the recognized expression and intonation of the user. The central processing module 6 understands the semantics of the user, selects the corresponding teaching content, and selects an appropriate teaching method according to the emotional state of the user, and then outputs the teaching content through the speaker module 7 and the display module 8. For example, if the user is happy, provide teaching content with strong interesting interactivity. If the user is in a low mood, slow down the rhythm of the teaching content, decompose complex tasks into smaller and easier-to-complete tasks, guide step by step, and pause the teaching in due time to play some soothing music. Thus, the teaching strategy can be adjusted according to the user's emotion, greatly improving the teaching effect and user experience of the educational robot.
[0097] The second embodiment of this application is:
[0098] On the basis of the first embodiment, please refer to Figure 8 , where Figure 8 is the flowchart of the second embodiment of the present invention.
[0099] An information interaction recognition and processing method of an educational robot with a camera provided by the present invention includes:
[0100] S1 The image acquisition module 1 captures the facial image of the user from multiple angles, and the voice acquisition module 2 collects the voice of the user;
[0101] S2 The expression recognition module 3 recognizes the facial expression of the user according to the facial image of the user, and the voice recognition module 4 recognizes the voice of the user;
[0102] S3 The emotion recognition module 5 comprehensively judges the emotional state of the user based on the facial expression and voice of the user;
[0103] S4 The central processing module 6 selects an appropriate teaching strategy according to the emotional state of the user, and outputs the teaching content through the speaker module 7 and the display module 8.
[0104] The third embodiment of this application is as follows:
[0105] Based on the first embodiment, please refer to Figure 9 , wherein, Figure 9 is the structural schematic diagram of the third embodiment of the present invention.
[0106] An educational robot with a camera provided by the present invention includes a robot body 100, two eye cameras 200, two ear cameras 300, a mouth microphone array 400, a chest display 500, and two shoulder speakers 600;
[0107] The two eye cameras 200 are respectively arranged on the robot body 100; the two ear cameras 300 are respectively arranged on the robot body 100; the mouth microphone array 400 is arranged on the robot body 100; the chest display 500 is arranged on the robot body 100; the two shoulder speakers 600 are respectively arranged on the robot body 100.
[0108] For the educational robot with a camera in this embodiment, an expression recognition module 3, a voice recognition module 4, an emotion recognition module 5, and a central processing module 6 are arranged in the robot body 100. The two eye cameras 200 and the two ear cameras 300 are used to capture user images from multiple angles. The mouth microphone array 400 is used to collect the voice input of the user. The chest display 500 and the two shoulder speakers 600 cooperate to output teaching interaction content.
[0109] The above-disclosed are only one or more preferred embodiments of this application. It cannot be used to limit the scope of rights of this application. Those of ordinary skill in the art can understand the whole or part of the processes of realizing the above embodiments, and the equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. An information interaction recognition processing system for an educational robot with a camera, characterized in that it includes an image acquisition module, a voice acquisition module, a facial expression recognition module, a speech recognition module, an emotion recognition module, a central processing module, a speaker module, and a display module; the facial expression recognition module is connected to the image acquisition module; the speech recognition module is connected to the voice acquisition module; the emotion recognition module is respectively connected to the facial expression recognition module and the speech recognition module; the central processing module is respectively connected to the emotion recognition module and the speech recognition module; the speaker module is connected to the central processing module; the display module is connected to the central processing module; the image acquisition module is used to capture the user's facial image from multiple angles; the voice acquisition module is used to acquire the user's voice; the facial expression recognition module is used to recognize the user's facial expression according to the user's facial image, the speech recognition module is used to recognize the user's voice; the emotion recognition module is used to comprehensively judge the user's emotional state based on the user's facial expression and voice; the central processing module is used to select an appropriate teaching strategy according to the user's emotional state; the speaker module is used to output teaching content in the form of voice; the display module is used to output teaching content in the form of video.
2. The information interaction recognition processing system for an educational robot with a camera according to claim 1, characterized in that the image acquisition module includes a camera unit and an image preprocessing unit; the image preprocessing unit is connected to the camera unit; the camera unit is used to capture the user's facial image from multiple angles; the image preprocessing unit is used to preprocess the user's facial image to facilitate subsequent processing.
3. The information interaction recognition processing system for an educational robot with a camera according to claim 2, characterized in that the voice acquisition module includes a microphone array unit and a voice preprocessing unit; the voice preprocessing unit is connected to the microphone array unit; the microphone array unit is used to accurately capture the user's voice input; the voice preprocessing unit is used to preprocess the user's input voice to facilitate subsequent processing.
4. The information interaction recognition processing system for an educational robot with a camera according to claim 3, characterized in that the facial expression recognition module includes a facial expression feature extraction unit and a facial expression classification unit; the facial expression classification unit is connected to the facial expression feature extraction unit; the facial expression feature extraction unit is used to extract high-level features according to the user's facial image; the facial expression classification unit is used to output the probability distribution of various facial expressions according to the extracted high-level features.
5. The information interaction recognition processing system for an educational robot with a camera according to claim 4, characterized in that The speech recognition module includes a speech feature extraction unit, a speech recognition unit, a semantic understanding unit, and an intonation analysis unit; the speech recognition unit is connected to the speech feature extraction unit; the semantic understanding unit is connected to the speech recognition unit; the intonation analysis unit is connected to the speech feature extraction unit; The speech feature extraction unit is configured to extract high-level speech features based on the user's speech input. The speech recognition unit is configured to decode the user's speech input into text output according to the extracted high-level features. The semantic understanding unit is configured to parse the semantics of the user's speech input according to the text output. The intonation analysis unit is configured to analyze the user's intonation according to the extracted high-level features.
6. The information interaction recognition processing system of an educational robot with a camera as claimed in claim 5, wherein The emotion recognition module includes an information integration unit and an emotion judgment unit; the emotion judgment unit is connected to the information integration unit; The information integration unit is configured to integrate the probability distribution of various expressions output by the expression classification unit and the intonation of the user analyzed by the intonation analysis unit. The emotion judgment unit is configured to judge the user's emotional state according to the information integrated by the information integration unit and feedback it to the central processing module.
7. The information interaction recognition processing system of an educational robot with a camera as claimed in claim 6, wherein The central processing module includes a storage unit, a strategy adjustment unit, and a communication unit; the strategy adjustment unit is connected to the storage unit; the communication unit is connected to the storage unit; The storage unit is configured to store a teaching strategy database. The strategy adjustment unit is configured to select the most appropriate teaching strategy from the teaching strategy database according to the user's emotional state and the semantics of the user's speech input, and output teaching content through the display module and the speaker module. The communication unit is configured to connect the storage unit to an external server to update the teaching strategy database in real time.
8. A method for information interaction recognition and processing of an educational robot with a camera, applied to an information interaction recognition and processing system of an educational robot with a camera as described in any one of claims 1 to 7, characterized in that, Comprising: The image acquisition module captures the user's facial image from multiple angles, and the voice acquisition module acquires the user's voice. The expression recognition module recognizes the user's facial expression according to the user's facial image, and the speech recognition module recognizes the user's speech. The emotion recognition module comprehensively judges the user's emotional state based on the user's facial expression and speech. The central processing module selects an appropriate teaching strategy according to the user's emotional state and outputs teaching content through the speaker module and the display module.
9. An educational robot with a camera, adopting the information interaction recognition processing system of an educational robot with a camera as claimed in any one of claims 1 to 7, wherein It includes a robot body, two eye cameras, two ear cameras, a mouth microphone array, a chest display, and two shoulder speakers. Two of the eye cameras are respectively arranged on the robot body; two of the ear cameras are respectively arranged on the robot body; the mouth microphone array is arranged on the robot body; the chest display is arranged on the robot body; two of the shoulder speakers are respectively arranged on the robot body.