Educational evaluation type desktop pet capable of replacing family education and teaching through lively activities and control method
Through computer vision and speech recognition technology, combined with artificial intelligence, real-time monitoring and analysis of children's behavior, the problem of existing electronic pets being unable to detect and interact is solved, and intelligent educational evaluation and personalized companionship are realized.
Patent Information
- Application Number
- CN202510254966.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-11
Smart Images

Figure CN120295457A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent electronic devices, and particularly relates to an educational evaluation type desktop pet for teaching in entertainment to replace a tutor and a control method thereof. Background Art
[0002] With the development of technology, electronic pets have gradually become an important tool for children's education and entertainment at home. With the rise of artificial intelligence, people pay more attention to personalized, intelligent and better interactive electronic pets. Especially parents hope that the electronic pets are more suitable for their children's personalities and habits, so as to better accompany their children. Traditional electronic pets mainly provide simple interaction and entertainment functions. Users interact with the electronic pets through buttons or touch screens to complete simple tasks such as feeding, bathing, and playing. They lack the ability to detect and guide children's behavior habits. In addition, existing electronic devices usually cannot monitor children's learning status, sitting postures, attention and other behaviors in real time, and it is difficult to help parents effectively supervise their children's behavior habits. Therefore, companion desktop electronic pets suitable for children are expected to be applied in families, and the prospect is broad.
[0003] At present, the development of artificial intelligence (AI) technology in the fields of behavior detection, emotion recognition, speech interaction, etc. has been relatively mature, and there are many pre-trained open-source models, but less research focuses on children's education and companionship. Modern family parents are busy with work and it is difficult to pay attention to their children's behaviors and learning status all day long. At the same time, children need emotional support and interaction, especially in single-child families or when parents are busy. Popularization of AI technology: AI technologies such as computer vision and natural language processing have matured and can be applied to children's education and behavior supervision. Popularization of intelligent devices: The cost of intelligent hardware (such as cameras, microphones, sensors) has decreased, providing a basis for the development of intelligent electronic pets.
[0004] Among them, how to perform behavior detection to better interact and provide more personalized companionship is the key point for the development of desktop electronic pets. Summary of the Invention
[0005] The purpose of the present invention is to provide an educational evaluation type desktop pet for teaching in entertainment to replace a tutor and a control method thereof to solve the above technical problems.
[0006] To solve the above technical problems, the present invention uses computer vision, speech recognition and artificial intelligence technologies to monitor the behavior of children in real time and provide feedback and interaction functions. The specific technical solutions of an educational evaluation type desktop pet for teaching in entertainment to replace a tutor and a control method thereof of the present invention are as follows:
[0007] An educational evaluation desktop pet that combines education with entertainment to replace private tutoring, including a behavior detection module, a speech recognition module, an action control module, a data recording and analysis module, and a controller module;
[0008] The behavior detection module is used to monitor the child's sitting posture, attention, and distance from the screen in real time. OpenCV and MediaPipe are used to detect the child's sitting posture and attention, and the YOLO model is used to detect interfering objects;
[0009] The speech interaction module is used to recognize the child's speech;
[0010] The action control module is used to control the electronic pet to generate actions according to the child's speech or emotions;
[0011] The data recording and analysis module stores the behavior data locally or in the cloud and uses data analysis tools to generate result reports;
[0012] The controller module is used to process the data of the behavior detection module and implement communication with each module. Further, the behavior detection module includes a camera and a sensor module. The sensor module includes a distance sensor, an accelerometer, and a gyroscope. The camera uses computer vision technology to detect the child's posture and interfering objects. The camera uses a built-in high-definition camera that supports 720P or 1080P resolution and is used to capture the child's facial expressions, body movements, and learning environment images. It also real-time monitors the movement and posture changes of the pet itself. The distance sensor is used to detect the distance between the child and the screen, and the accelerometer and gyroscope are used to detect the child's actions.
[0013] Further, the speech interaction module, which is used to recognize the child's speech, includes a microphone, a speaker, and a display screen. The microphone is used for speech interaction, the speaker is used for speech feedback and playing music, and the display screen is used to display the pet's expressions and animations and for interaction; the speech interaction module first uses Google Speech-to-Text to convert the child's speech input into text, then uses the GPT model for natural language processing to understand the user's intention and generate an appropriate response, and finally converts the text into speech output.
[0014] Further, the action control module, according to the child's speech commands and behavioral actions, program-controls the desktop pet to make corresponding reactions. At the same time, the program intelligently adjusts the interaction content and methods according to the child's interest preferences and usage habits to provide personalized companion services, uses speech recognition technology to understand the child's speech input, and speech synthesis technology to generate the pet's speech feedback.
[0015] Furthermore, the data recording and analysis module stores the behavior data locally or in the cloud, uses data analysis tools to generate result reports, records the child's behavior data and interaction data; then statistically analyzes the recorded data, uses machine learning models to classify or predict behaviors, and finally visualizes the data: generates reports for parents to view.
[0016] Furthermore, the controller module includes a core processor, a power module, and a communication module. The core processor selects a high-performance and low-power ARM architecture processor. The power module is used to provide power for the desktop pet; the communication module includes a Wifi module and a Bluetooth module, and is connected to the home network or school network through Wifi or Bluetooth to achieve data synchronization, obtain online learning resources, and remote communication with the parent mobile APP.
[0017] The present invention also discloses a control method for an educational and entertaining educational evaluation desktop pet that replaces a tutor, including the following steps:
[0018] Step 1: Behavior detection: The behavior detection module analyzes the child's behavior through pose detection and object detection;
[0019] Step 2: Voice interaction: Recognize the child's voice and make corresponding feedback according to the voice, output voice reminders;
[0020] Step 3: Action control: Generate the actions of the desktop electronic pet according to the child's input or emotional state; Step 4: Data analysis and visualization: Record the child's behavior and interaction data with the electronic pet and analyze. Furthermore, the above Step 1 includes the following specific steps:
[0021] Step 1.1: Pose detection uses a deep learning model to predict the positions of human key points in the image to judge the user's pose;
[0022] Input: Image I
[0023] Output: Key point coordinates (x i , y i ), where i represents the i-th key point;
[0024] Step 1.2: Judge the pose according to the relative positions of the key points;
[0025] Step 1.3: Use a large model to detect specific objects, and transform the object detection problem into a regression problem, directly predicting the bounding box and class probability;
[0026] Input: Image I
[0027] Output: Bounding box (x, y, w, h) and class probability p
[0028] Loss function:
[0029]
[0030] Among them, S 2 is the number of grids, B is the number of bounding boxes for each grid, indicates whether the j-th bounding box of the i-th grid contains a target
[0031] Step 1.4: Estimate the distance between the user and the screen through the camera and the size of the known object;
[0032] Distance formula:
[0033]
[0034] Among them, D is the distance, f is the focal length of the camera, W is the actual width of the object, and p is the pixel width of the object in the image.
[0035] Furthermore, the said Step 2 includes the following specific steps:
[0036] Step 2.1: Speech recognition converts the speech signal into text, uses the Mel Frequency Cepstral Coefficient (MFCC) extraction method to extract the feature vector m, and then uses the deep learning model Transformer for the conversion from speech to text T;
[0037] Input: Speech signal s(t);
[0038] Converted to: MFCC feature vector m;
[0039] Output: Text T;
[0040] Step 2.2: Understand the child's intention and generate a response, use the Transformer model for natural language processing, the core of which is the self-attention mechanism, and the self-attention formula:
[0041]
[0042] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the vector,
[0043] Step 2.3: Convert the text into a speech signal for speech synthesis.
[0044] Furthermore, the said Step 3 and Step 4 include the following specific steps:
[0045] Step 3.1: Predesigned the actions of the desktop electronic pet and stored them in the action library;
[0046] Step 3.2: Generate actions according to the child's instructions or the environmental state through rules;
[0047] Step 3.3: Control the desktop electronic pet to perform actions through hardware.
[0048] Step 4: Data analysis and visualization: After recording the child's behavior and the interaction data with the electronic pet, the following steps are carried out in sequence:
[0049] Step 4.1: Statistically calculate the mean, maximum, minimum, and standard deviation statistics of the learning duration, and analyze the trend of the behavior data;
[0050] Step 4.2: Use a classification model to classify the behavior, and use a regression model to predict the future behavior trend of the child;
[0051] Decision tree classification selects the best splitting point through information gain:
[0052]
[0053] Among them, H(D) is the entropy of the dataset D;
[0054] Linear regression minimizes the loss function:
[0055]
[0056] Step 4.3: Conduct data visualization, and use the visualization tool Matplotlib to generate a bar chart of the learning duration, a line chart of the sitting posture angle, and a pie chart of the behavior classification.
[0057] An educational evaluation desktop pet that combines education with entertainment and a control method instead of a tutor according to the present invention has the following advantages:
[0058] 1. The desktop electronic pet of the present invention can interact with children;
[0059] 2. The desktop electronic pet of the present invention can detect and evaluate the behavior of children, and can timely remind or encourage children to help them develop good living and learning habits;
[0060] 3. Compared with traditional electronic pets, the present invention is more intelligent, can record the interaction between children and pets, detect and analyze the actions and emotions of children, etc., is more personalized, and can help parents better understand their children;
[0061] 4. Based on the present invention, many in-depth expansion studies can be carried out, such as not only targeting the after-school study and life of children, but also developing to children's classroom learning, after-school interest cultivation, etc., and has strong function expandability. Description of the Drawings
[0062] Figure 1 is the logical relationship diagram of the system modules of the present invention;
[0063] Figure 2 It is a schematic diagram of the physical object of the present invention;
[0064] Figure 3 It is an example scenario diagram for evaluating the child's learning conscientiousness;
[0065] Figure 4 It is an example scenario diagram for parents to feedback and monitor the child's learning status;
[0066] Figure 5 It is a block diagram of the composition of the behavior evaluation and feedback system;
[0067] Figure 6 It is an example scenario diagram of the voice interaction system;
[0068] Figure 7 It is a block diagram of the composition of the voice interaction system. Detailed implementation manners
[0069] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail an educational evaluation type desktop pet substituting for a tutor and a control method thereof according to the present invention with reference to the accompanying drawings.
[0070] As Figure 1 Figure 2 shown, an educational evaluation type desktop pet substituting for a tutor according to the present invention monitors the child's behavior in real time through computer vision, speech recognition and artificial intelligence technologies, and provides feedback and interaction functions. It includes a behavior detection module, a speech recognition module, an action control module, a data recording and analysis module, and a controller module.
[0071] The described behavior detection module is used to monitor the child's sitting posture, attention, distance from the screen, and other behaviors in real time. It uses OpenCV and MediaPipe to detect the child's sitting posture and attention, and uses the YOLO model to detect interfering objects (such as mobile phones). The behavior detection module includes a camera and a sensor module. The sensor module includes a distance sensor, an accelerometer, and a gyroscope. The camera uses computer vision technology to detect the child's posture and interfering objects (such as mobile phones). The camera is an in-built high-definition camera that supports 720P or 1080P resolution and is used to capture the child's facial expressions, body movements, and learning environment images. It also monitors the movement and posture changes of the pet itself in real time. When the child picks up the pet, the pet can make corresponding reactions. At the same time, it also assists in detecting the child's actions during the learning process, such as whether the child shakes the body frequently. The distance sensor is used to detect the distance between the child and the screen, and the accelerometer and gyroscope are used to detect the child's actions, such as whether the child's sitting posture is correct (whether the spine is straight), whether the child is focused on learning (such as whether the eyes are looking at the book or the screen), the degree of concentration (whether doing other things unrelated to learning), and healthy eye use (whether the distance from the screen is too close during learning). The functions described above are achieved through the behavior detection algorithm of the present invention: using the camera to capture the child's posture, obtaining the real-time video stream of the camera through OpenCV, then using the pre-trained pose detection, face detection, and hand detection models in MediaPipe, combining with YOLO for object detection, and then performing image processing and analysis.
[0072] The described voice interaction module is used to recognize the child's voice. It includes a microphone, a speaker, and a display screen. The microphone is used for voice interaction, the speaker is used for voice feedback and playing music, and the display screen is used to display the pet's expressions and animations and for interaction. The voice interaction module first uses Google Speech-to-Text to convert the child's voice input into text, then uses the GPT model for natural language processing to understand the user's intention and generate an appropriate response, and finally converts the text into voice output. It can detect the child's emotions (such as happy, frustrated) by recognizing the child's tone of voice, and then provide appropriate feedback (such as playing music or telling jokes) according to the emotional state to give emotional companionship and support. Preferably, the resolution of the display screen reaches 480×800 or above, which can well display the pet's expressions, actions, interaction content, and learning-related information, such as learning task prompts, knowledge point explanation pictures, etc. The microphone is an array composed of multiple high-sensitivity microphones, which can accurately identify the child's voice commands, effectively filter out environmental noise, and ensure the accuracy and clarity of voice interaction.
[0073] The action control module is used to control the electronic pet to generate actions according to the child's voice or emotion, supporting a predefined action library or dynamically generating actions. The child can interact with the desktop electronic pet. According to the child's voice commands and behavioral actions, the program controls the desktop pet to make corresponding responses, such as answering questions, playing games, giving emotional responses, etc. At the same time, the program can also intelligently adjust the interaction content and methods according to the child's interest preferences and usage habits, providing personalized companionship services. It uses speech recognition technology to understand the child's speech input and speech synthesis technology to generate the pet's speech feedback.
[0074] The data recording and analysis module stores the behavior data locally or in the cloud and uses data analysis tools to generate result reports. It records the child's behavior data (such as learning duration, sitting posture, attention, distance from the screen, etc.) and interaction data (such as speech interaction content, emotional state). Then it statistically analyzes the recorded data (such as average value, trend analysis), uses machine learning models to classify or predict behaviors. Finally, it visualizes the data: generates reports (such as bar charts, line charts, pie charts) for parents to view. Parents can better understand the child's growth situation through information such as behavior reports and learning time statistics. When the child completes certain learning tasks or performs well, virtual rewards are issued to the child through the application program in the pet to encourage the child to be positive.
[0075] The controller module includes a core processor, a power module, and a communication module. The core processor selects a high-performance and low-power ARM architecture processor, which has powerful computing capabilities and is used to process the data of the behavior detection module and implement communication with each module to ensure the smooth operation of the system. The power module is used to provide power for the desktop pet. The communication module includes a Wifi module and a Bluetooth module, which are connected to the home network or school network through Wifi or Bluetooth to achieve data synchronization, obtain online learning resources, and remote communication with the parent's mobile phone APP.
[0076] A control method for an educational evaluation desktop pet that substitutes for a tutor and combines education with entertainment according to the present invention includes the following steps:
[0077] Step 1: Behavior detection: The behavior detection module analyzes the child's behavior through pose detection and target detection, and sequentially performs the following steps:
[0078] Step 1.1: Pose detection uses a deep learning model to predict the positions of human key points in the image (such as shoulders, elbows, knees, etc.) to judge the user's pose;
[0079] Input: Image I
[0080] Output: Key point coordinates (x i , y i ), where i represents the i-th key point
[0081] Step 1.2: Determine the posture (such as sitting or standing) based on the relative positions of key points
[0082] For example, determine whether the spine is straight: If the spine angle exceeds the threshold, it is considered that the sitting posture is incorrect
[0083] Step 1.3: Use a large model to detect specific objects (such as mobile phones, books), and transform the object detection problem into a regression problem to directly predict the bounding box and class probability;
[0084] Input: Image I
[0085] Output: Bounding box (x, y, w, h) and class probability p
[0086] Loss function:
[0087]
[0088] where S 2 is the number of grids, B is the number of bounding boxes per grid, indicates whether the j-th bounding box of the i-th grid contains the target
[0089] Step 1.4: Estimate the distance between the user and the screen through the camera and the size of known objects;
[0090] Distance formula:
[0091]
[0092] where D is the distance, f is the focal length of the camera, W is the actual width of the object, and p is the pixel width of the object in the image.
[0093] Step 2: Voice interaction: Recognize the child's voice and give corresponding feedback and output voice reminders according to the voice; including the following specific steps:
[0094] Step 2.1: Voice recognition converts the voice signal into text. Use Mel-frequency cepstral coefficients (MFCC), a commonly used voice recognition feature extraction method, to extract the feature vector m, and then use the deep learning model Transformer for voice-to-text T conversion.
[0095] Input: Voice signal s(t)
[0096] Converted to: MFCC feature vector m.
[0097] Output: Text T
[0098] Step 2.2: Understand the child's intention and generate a response, using the Transformer model for natural language processing. The core of it is the self-attention mechanism, and the self-attention formula:
[0099]
[0100] where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the vector.
[0101] Step 2.3: Convert the text into speech signals for speech synthesis
[0102] Step 3: Action control: Generate the actions of the desktop electronic pet according to the child's input or emotional state, and perform the following steps in sequence:
[0103] Step 3.1: Predesigned the actions of the desktop electronic pet (such as dancing, waving, nodding), and store them in the action library;
[0104] Step 3.2: Generate actions according to the child's instructions or environmental status through rules;
[0105] Step 3.3: Control the desktop electronic pet to execute actions through hardware.
[0106] Step 4: Data analysis and visualization: After recording the child's behavior and the interaction data with the electronic pet, perform the following steps in sequence:
[0107] Step 4.1: Statistically calculate the mean, maximum, minimum, standard deviation and other statistics of the learning duration, and analyze the trend of the behavior data (such as the change of the learning duration)
[0108] Step 4.2: Use a classification model (decision tree) to classify the behavior (such as excellent, good, bad), and use a regression model (linear regression) to predict the future behavior trend of the child.
[0109] The decision tree classification selects the best splitting point through information gain:
[0110]
[0111] where H(D) is the entropy of the dataset D
[0112] The linear regression minimizes the loss function:
[0113]
[0114] Step 4.3: Conduct data visualization, and use the visualization tool Matplotlib to generate a bar chart of the learning duration, a line chart of the sitting posture angle, a pie chart of the behavior classification, etc.
[0115] Example 1:
[0116] Taking the experiment on evaluating the child's learning conscientiousness as an example, the application method of the present invention will be specifically described below.
[0117] As Figure 3 shown, 1 is the pet shell, 2 is the camera device, 3 is the child's learning interface, 4 is the collection of key human body points, 5 is the learning evaluation display module, 6 is the parental remote communication module, and 7 is the voice module. When the child is studying at the desk and completes the established experimental content, it is divided into the following steps: The behavior detection module first takes multi-angle shots through the camera to obtain images of key human body points such as the child's head, shoulders, and elbows, and uploads them to the data recording and analysis module. After image processing, it detects the child's learning posture and whether there is interference. Assuming the child has a correct sitting posture and studies conscientiously, the data recording and analysis module gives an excellent evaluation, and the desktop pet stretches out its hand to give the child a smiling face encouragement. Then, the collected images, videos, learning duration, learning evaluations and other data are uploaded to the cloud system for parents to view.
[0118] As Figure 4 Figure 5 shown, according to the child's learning progress, past learning habits and effects, parents select different learning modes and set the learning duration and content. After the child starts learning, when the parents are free, they can view the real-time status of the child's learning through the camera module. Assuming that the child is inattentive during learning, the data recording and analysis module immediately gives an evaluation of not being serious, sends the signal to the parental communication end through the core processor for message notification to the parents, and gives an evaluation in the data recording and analysis module. At the same time, the voice module issues a prompt statement to remind the child to study seriously.
[0119] Example 2:
[0120] Taking the experiment on the child's voice interaction with the pet as an example, the application method of the present invention will be specifically described below.
[0121] As Figure 6 Figure 7 shown, when the child encounters a knowledge point that he / she doesn't understand during learning and completes the established experimental content, it is divided into the following steps: First, when the child encounters a problem during learning, he / she asks the electronic pet through voice. At this time, the voice module of the electronic pet detects the child's voice signal, transmits the signal to the data recording and analysis module for problem analysis and processing. After processing, the corresponding text is generated and then converted into a voice signal to reply to the child. And relevant knowledge points are displayed on the display module, and the child can select relevant content to view according to his / her own needs. After the knowledge points corresponding to the questions asked by the child are analyzed by the core processor, the data is uploaded to the cloud system for parents to view, and parents can timely pay attention to the child's learning and mastery situation.
[0122] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. An educational evaluation-based desktop pet that combines education with entertainment to replace private tutoring, characterized in that, It includes a behavior detection module, a speech recognition module, an action control module, a data recording and analysis module, and a controller module; The behavior detection module is used to monitor the child's sitting posture, attention, and distance from the screen in real time. It uses OpenCV and MediaPipe to detect the child's sitting posture and attention, and uses the YOLO model to detect interfering objects; The speech interaction module is used to recognize the child's speech; The action control module is used to control the electronic pet to generate actions according to the child's speech or emotion; The data recording and analysis module stores the behavior data locally or in the cloud and uses data analysis tools to generate a result report; The controller module is used to process the data of the behavior detection module and implement communication with each module.
2. The entertainment-education evaluation desktop pet substituting for a tutor according to claim 1, characterized in that, The behavior detection module includes a camera and a sensor module. The sensor module includes a distance sensor, an accelerometer, and a gyroscope. The camera uses computer vision technology to detect the child's posture and interfering objects. The camera is an in-built high-definition camera that supports 720P or 1080P resolution and is used to capture the child's facial expressions, body movements, and learning environment images. It monitors the movement and posture changes of the pet itself in real time. The distance sensor is used to detect the distance between the child and the screen, and the accelerometer and gyroscope are used to detect the child's actions.
3. The entertaining education evaluation desktop pet replacing a tutor according to claim 1, characterized in that, The speech interaction module is used to recognize the child's speech and includes a microphone, a speaker, and a display screen. The microphone is used for speech interaction, the speaker is used for speech feedback and music playback, and the display screen is used to display the pet's expressions and animations and for interaction; The speech interaction module first uses Google Speech-to-Text to convert the child's speech input into text, then uses the GPT model for natural language processing to understand the user's intention and generate an appropriate response, and finally converts the text into speech output.
4. The entertainment-education evaluation desktop pet substituting for a tutor according to claim 1, characterized in that, The action control module, according to the child's speech instructions and behavioral actions, the program controls the desktop pet to make corresponding reactions. At the same time, the program intelligently adjusts the interaction content and method according to the child's interest preferences and usage habits to provide personalized companionship services. It uses speech recognition technology to understand the child's speech input and speech synthesis technology to generate the pet's speech feedback.
5. The entertaining education evaluation desktop pet substituting for a tutor according to claim 1, characterized in that, The data recording and analysis module stores the behavior data locally or in the cloud, uses data analysis tools to generate a result report, and records the child's behavior data and interaction data; then it statistically analyzes the recorded data, uses a machine learning model to classify or predict behaviors, and finally visualizes the data: generates a report for parents to view.
6. The entertaining education evaluation desktop pet substituting for a tutor according to claim 1, characterized in that, The controller module includes a core processor, a power module, and a communication module. The core processor selects a high-performance and low-power ARM architecture processor. The power module is used to provide power for the desktop pet; The communication module includes a Wifi module and a Bluetooth module, which are connected to the home network or school network through Wifi or Bluetooth to achieve data synchronization, obtain online learning resources, and remote communication with the parent's mobile phone APP.
7. A control method for an entertainment-education evaluation desktop pet that replaces a tutor as described in any one of claims 1-6, characterized in that, It includes the following steps: Step 1: Behavior Detection: The behavior detection module analyzes the child's behavior through pose detection and object detection; Step 2: Voice Interaction: Recognize the child's voice and give corresponding feedback according to the voice, and output voice reminders; Step 3: Action Control: Generate the actions of the desktop electronic pet according to the child's input or emotional state; Step 4: Data Analysis and Visualization: Record the child's behavior and the interaction data with the electronic pet and analyze them.
8. The control method of the entertaining and educational evaluation desktop pet substituting for a tutor according to claim 7, characterized in that, The said Step 1 includes the following specific steps: Step 1.1: The pose detection uses a deep learning model to predict the positions of human key points in the image to judge the user's pose; Input: Image I Output: Key point coordinates (x i , y i ), where i represents the i-th key point; Step 1.2: Judge the pose according to the relative positions of the key points; Step 1.3: Use a large model to detect specific objects, and transform the object detection problem into a regression problem, directly predicting the bounding box and class probability; Input: Image I Output: Bounding box (x, y, w, h) and class probability p Loss function: Among them, S 2 is the number of grids, B is the number of bounding boxes for each grid, indicates whether the j-th bounding box of the i-th grid contains the target Step 1.4: Estimate the distance between the user and the screen through the camera and the size of the known object; Distance formula: where D is the distance, f is the focal length of the camera, W is the actual width of the object, and p is the pixel width of the object in the image.
9. The control method of the entertainment-education evaluation desktop pet substituting for a tutor according to claim 7, characterized in that, The said Step 2 includes the following specific steps: Step 2.1: The speech recognition converts the speech signal into text, uses the Mel Frequency Cepstral Coefficient (MFCC) extraction method to extract the feature vector m, and then uses the deep learning model Transformer for the conversion from speech to text T; Input: Speech signal s(t); Converted to: MFCC feature vector m; Output: Text T; Step 2.2: Understand the child's intention and generate a response, use the Transformer model for natural language processing, the core of which is the self-attention mechanism, and the self-attention formula: where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the vector, Step 2.3: Convert the text into a voice signal for voice synthesis.
10. The control method of the entertainment education evaluation desktop pet replacing a tutor according to claim 7, characterized in that, The said Steps 3 and 4 include the following specific steps: Step 3.1: Predesign the actions of the desktop electronic pet in advance and store them in the action library; Step 3.2: Generate actions according to the child's instructions or environmental status through rules; Step 3.3: Control the desktop electronic pet to execute actions through hardware. Step 4: Data Analysis and Visualization: After recording the child's behavior and the interaction data with the electronic pet, the following steps are carried out in turn: Step 4.1: Statistically calculate the average value, maximum value, minimum value, and standard deviation statistics of the learning duration, and analyze the trend of the behavior data; Step 4.2: Use a classification model to classify the behavior, and use a regression model to predict the future behavior trend of the child; The decision tree classification selects the best splitting point through information gain: where H(D) is the entropy of the dataset D; The linear regression minimizes the loss function: Step 4.3: Conduct data visualization, and use the visualization tool Matplotlib to generate a bar chart of the learning duration, a line chart of the sitting posture angle, and a pie chart of the behavior classification.