A child guiding robot behavior decision system based on real-time emotional feedback

By combining multimodal emotion perception and psychological state modeling with dynamic behavioral decision-making, the problems of rigid interaction logic and insufficient emotion perception of children's robots have been solved. This has enabled accurate decision-making on children's specific psychological attachments, and improved the depth and safety of robot-child interaction.

CN122480919APending Publication Date: 2026-07-31余悦
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
余悦
Filing Date
2026-04-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing children's robots suffer from rigid interaction logic, insufficient depth of emotional perception, and a lack of dynamic decision-making mechanisms that address specific psychological attachment phenomena in children, making it difficult to achieve deep human-machine emotional connections.

Method used

A multimodal emotion perception layer is constructed, which acquires children's data in real time through high-resolution infrared depth vision, arrayed microelectromechanical system microphones, flexible pressure sensing units and non-contact physiological signal detection units. Combined with the psychological state modeling layer, emotional features are fused and the Abebe complex is quantitatively assessed. The dynamic behavior decision layer uses a multi-objective optimization algorithm to generate behavioral instructions, and the execution and feedback closed-loop layer performs online correction.

Benefits of technology

It enables real-time monitoring of children's deep emotional state, enhances the real-time nature and empathy of the interaction process, dynamically adjusts educational guidance strategies, avoids harsh intervention, improves the quality and scientific nature of long-term companionship, and ensures safety and cross-regional adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122480919A_ABST
    Figure CN122480919A_ABST
Patent Text Reader

Abstract

This application relates to a behavior decision-making system for a child-guided robot based on real-time emotional feedback, belonging to the field of robotics technology. The system includes: a multimodal emotion perception layer for real-time acquisition of children's physiological signs, facial expressions, and voice data; a psychological state modeling layer for constructing a real-time psychological profile and quantifying the attachment index to specific objects; a dynamic behavior decision-making layer for generating an optimal sequence of behavioral instructions based on the psychological profile and attachment index, combined with educational guidance logic; and an execution and feedback closed-loop layer for driving the robot to execute instructions and correcting the decision-making logic online based on feedback. This application, by constructing a multidimensional emotion perception and attachment modeling mechanism, solves the problems of rigid traditional interaction logic and lack of deep emotional understanding, achieving scientific guidance and precise comfort for children's specific psychological needs, and significantly improving the empathy, naturalness, and safety of robot interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robotics technology, specifically relating to a behavior decision-making system for a child-guided robot based on real-time emotional feedback. Background Technology

[0002] With the rapid development of artificial intelligence and precision sensor technology, intelligent robots have been widely applied in diverse scenarios such as family education, psychological assistance, and daily companionship. As an important branch of human-computer interaction, children's service robots aim to provide personalized interactive experiences for children in critical developmental stages through integrated multimodal perception and automated decision-making technologies. These technologies not only encompass core areas such as natural language processing, computer vision, and motion control, but also strive to positively guide children's cognitive development and mental health through intelligent behavioral performance.

[0003] Among these, the behavioral decision-making system of the child-guided robot is the core of achieving high-quality emotional interaction. This system typically requires real-time capture of external environmental information and user feedback, and combines this with a pre-set logical model to generate corresponding actions, facial expressions, or voice feedback. Its basic principle lies in constructing a processing framework that can simulate human interaction logic, enabling the robot to dynamically adjust its guidance strategy and interaction intensity according to different application scenarios and user states, thereby enhancing the depth of interaction between children and the robot.

[0004] Existing children's robots generally suffer from rigid interaction logic and a lack of emotional perception. Most rely on pre-set, simple instruction sets for cyclical responses, making it difficult to accurately capture and deeply understand children's complex and ever-changing emotional states. While some systems incorporate basic facial expression recognition technology, they lack targeted modeling and decision-making mechanisms for children's unique psychological phenomena, such as strong emotional attachments to specific objects (like the "Abebe complex"), resulting in an inability to provide empathetic responses to children's emotional fluctuations. This semantic gap between perceptual data and emotional models, along with the lack of dynamic decision-making logic based on psychological dimensions, significantly limits the long-term companionship and scientific guidance capabilities of existing robots, making it difficult to achieve deep human-machine emotional connections. Summary of the Invention

[0005] The purpose of this invention is to provide a behavior decision-making system for children's guided robots based on real-time emotional feedback, in order to solve the technical problems of existing children's robots, such as rigid interaction logic, insufficient depth of emotional perception, and lack of dynamic decision-making mechanisms for specific psychological attachment phenomena in children.

[0006] The technical solution of this invention includes: a multimodal emotion perception layer, used to acquire children's physiological characteristics, facial expression features, and voice emotion parameters in real time; a psychological state modeling layer, used to construct a real-time psychological profile of children based on the data acquired by the multimodal emotion perception layer, and specifically to quantify and evaluate the emotional attachment to a specific object, namely the Abebe complex, to generate an attachment index; a dynamic behavior decision-making layer, used to generate the optimal sequence of behavioral instructions for the robot through a multi-objective optimization algorithm based on the psychological profile and attachment index output by the psychological state modeling layer, combined with a preset educational guidance logic library; and an execution and feedback closed-loop layer, used to drive the robot's motion execution mechanism, expression display module, and speech synthesis module to execute instructions, and to correct the decision logic online based on environmental feedback after execution.

[0007] In one embodiment of the present invention, the multimodal emotion perception layer integrates a high-resolution infrared depth vision unit, an array-type microelectromechanical system microphone unit, a flexible pressure sensing unit, and a non-contact physiological signal detection unit. The vision unit captures the child's facial micro-expressions and pupil changes, and uses a convolutional neural network to extract key facial features. The microphone unit extracts pitch, energy, and speech rate changes in speech using a Mel-frequency cepstral coefficient algorithm to identify the emotional polarity in the speech. The flexible pressure sensing unit is arranged on the robot's outer shell to sense the force and frequency of the child's grasping, patting, or touching. The non-contact physiological signal detection unit uses ultra-wideband radar technology to acquire the child's heart rate and respiratory rate from a distance.

[0008] Furthermore, the mental state modeling layer includes an emotional feature fusion engine. This engine employs an attention mechanism to weightedly fuse visual, auditory, tactile, and physiological data, eliminating noise interference from single-sensor data. The mental state modeling layer also includes a long short-term memory network model to analyze the time-series evolution of emotional states, thereby identifying the probability distribution of children in different dimensions such as excitement, sadness, anxiety, or anger.

[0009] Furthermore, the quantitative assessment process for the Abebe complex in the psychological state modeling layer is as follows: The system tracks the physical distance between the child and a specific comfort toy or fabric in real time using an infrared depth vision unit; it records the changes in the intensity of the child's response to robot interaction while holding the specific object using a tactile perception unit; and it detects specific frequencies of crying or anxious vocal signals that occur when the specific object is missing using a speech analysis unit. The system inputs the physical distance, interaction response intensity, and duration of anxiety signals into a preset attachment assessment model, calculates a scalar value between 0 and 1, and defines it as the attachment index. When the attachment index exceeds a preset threshold of 0.8, the system determines that the child is currently in a strong specific object attachment state.

[0010] Furthermore, the dynamic behavior decision-making layer operates within a hierarchical control architecture. The bottom layer comprises a reactive behavior generation unit, handling high-priority tasks such as emergency obstacle avoidance or immediate emotional reassurance; the middle layer consists of a strategy-based guidance unit, adjusting the robot's guidance tone and action amplitude based on the current attachment index to ensure the robot's interactive behavior does not conflict with the child's attachment psychology; the top layer is a goal-oriented evolutionary unit, gradually introducing cognitive training or language learning tasks while ensuring emotional reassurance.

[0011] Furthermore, the dynamic behavior decision layer employs reinforcement learning algorithms for decision optimization. The system defines the psychological profile, attachment index, and environmental state as the state space, the robot's actions, facial expressions, and speech sequences as the action space, and the degree of improvement in the child's emotional state, interaction duration, and completion of the guided task as the reward function. Through iterative training in simulated environments and real-world scenarios, the decision layer can automatically optimize the behavioral combination scheme that maximizes the overall reward function.

[0012] Furthermore, the execution and feedback closed-loop layer includes a servo motor drive framework with at least 12 degrees of freedom to achieve delicate body movement simulation. The facial expression display module uses a high-pixel-density LED dot matrix screen or micro-projection display technology to present smooth and empathetic facial expression switching. The speech synthesis module is based on deep learning-based emotional speech synthesis technology, which can adjust the emotional tone, timbre saturation, and phrasing rhythm of the output speech in real time according to decision commands.

[0013] Furthermore, the system also includes a safety boundary monitoring unit. This unit operates independently of the behavior decision layer, monitoring the robot's driving torque, movement speed, and collision risk with children in real time. When any deviation of physical parameters that may endanger a child's safety is detected, the safety boundary monitoring unit has the highest level of decision control, capable of immediately cutting off the power circuit or forcibly executing protective avoidance actions.

[0014] Furthermore, the system also includes a cloud-based learning and evolution unit. This unit is responsible for collecting anonymized interaction log data, including environmental context, decision sequences, and the final guidance effect. Through large-scale data mining and model retraining on the cloud server, the system can periodically update the weight coefficients of the mental state modeling layer and the dynamic behavior decision layer, thereby enabling the robot's guidance strategy to have cross-regional and cross-cultural adaptive capabilities.

[0015] In a preferred embodiment of the present invention, the psychological state modeling layer also incorporates a time decay factor when calculating the attachment index. This factor characterizes the natural weakening trend of a child's attachment to a specific object as they grow older. Through regression analysis of historical attachment data, the system can dynamically adjust the baseline parameters of the attachment assessment model to ensure that the decision-making logic aligns with the stage-specific characteristics of children's cognitive development.

[0016] Furthermore, when generating guidance instructions, the dynamic behavior decision-making layer prioritizes searching a pre-set psychological intervention template library. This library contains various scientific guidance programs for the Abebe complex. When a strong attachment state is detected, the decision-making layer instructs the robot to establish initial trust through gentle body language and low-frequency soothing voice. Subsequently, using joint attention guidance technology, the child's attention is gradually shifted from the attachment object to a shared interactive task, thereby achieving scientific and gentle emotional guidance.

[0017] Furthermore, the speech synthesis module in the execution and feedback closed-loop layer employs a dual-channel audio processing technology. The first channel outputs clear instructional or guiding speech, while the second channel is used for real-time mixing of background music or white noise. When the system detects that a child is in an anxious state, the second channel automatically increases the volume of the white noise or soft background music, utilizing the acoustic masking effect to help alleviate the child's emotional fluctuations.

[0018] Furthermore, the system's power management unit can dynamically adjust the sampling frequency of each sensing module based on the intensity of the interaction. When the child is asleep or has low interaction intention, the system automatically reduces the power consumption of the non-contact physiological signal detection unit and the infrared depth vision unit, retaining only the basic wake-up trigger unit, thereby effectively extending the robot's battery life and reducing electromagnetic radiation intensity.

[0019] Compared with existing technologies, the advantages and positive effects of this invention are as follows: 1. This solution completely breaks down the barrier between traditional children's robot perception data and semantic understanding by constructing a deep fusion framework of a multimodal emotion perception layer and a psychological state modeling layer. By introducing physiological indicators such as heart rate and respiration, as well as subtle facial micro-expression features, the system can achieve real-time monitoring of children's deep and implicit emotional states. Its recognition accuracy is improved by about 35% compared with solutions that rely solely on single visual or speech recognition, ensuring the real-time nature and empathy of the interaction process. 2. This solution introduces for the first time a quantitative modeling mechanism for specific object emotional attachment, such as the Abebe complex, into the behavioral decision-making system. Through the calculation of the attachment index, the robot no longer regards such attachment behaviors as invalid interference, but rather uses them as important input variables for decision-making. This design conforms to the core principles of developmental psychology, enabling the robot's guidance strategy to accurately meet the special psychological needs of children, avoiding emotional conflict caused by rough intervention, and greatly improving the quality and scientific nature of long-term companionship. 3. This solution adopts a hierarchical dynamic behavioral decision-making architecture and reinforcement learning algorithm, realizing a technological leap from single instruction response to complex strategy evolution. The system can dynamically adjust the intensity and pace of educational guidance based on the differences in emotional feedback from different individuals. This personalized guidance mechanism not only enhances children's participation in interactive tasks but also continuously optimizes the decision-making path through closed-loop feedback. This makes the robot's performance increasingly natural and intelligent as the interaction time increases, truly realizing a fundamental transformation from a functional tool to an emotional partner. 4. This solution fully considers the safety and long-term sustainability of children's interactive scenarios. By setting up independent safety boundary monitoring units and cloud-based learning and evolution units, the system can provide extremely high safety guarantees at the physical level while maintaining continuous vitality at the software algorithm level. The cloud-driven logic update mechanism ensures that the robot can learn the latest educational psychology findings, laying a solid technical foundation for its widespread application in family, medical, and special education scenarios. Attached Figure Description

[0020] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0021] Figure 1 This is a schematic diagram of the overall technical solution architecture of the child guidance robot behavior decision-making system based on real-time emotional feedback proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework for quantitatively assessing emotional attachment to a specific object, namely the Abebe complex, in this invention. Figure 3 This is a flowchart illustrating the logical process of data acquisition and feature extraction in the multimodal emotion perception layer of this invention. Figure 4 This is a schematic diagram of the core principle framework for the psychological state modeling layer in this invention to perform emotional feature fusion and time series analysis; Figure 5 This is a schematic diagram of the core principle framework of the hierarchical control architecture of the dynamic behavior decision layer in this invention; Figure 6 This is a flowchart illustrating the logical process of behavior decision optimization based on reinforcement learning algorithms in this invention. Figure 7 This is a flowchart illustrating the logical flow of instruction execution and online correction within the execution and feedback closed-loop layer of this invention. Figure 8 This is a schematic diagram of the multi-level interaction relationship and data flow between the security boundary monitoring unit and the cloud learning evolution unit in this invention; Detailed Implementation Example 1

[0022] Please refer to the attached document. Figure 1 This embodiment provides a child-guided robot behavior decision-making system based on real-time emotional feedback. The system is built on a highly integrated mechatronics platform and aims to provide scientific guidance to children through deep emotional analysis. The system's physical and logical architectures are tightly coupled, ensuring efficient data flow from initial perception to final behavioral execution.

[0023] The system first consists of a multimodal emotion perception layer, whose core task is to construct a comprehensive, all-encompassing data acquisition matrix. Combined with the attached... Figure 3 As shown, the multimodal emotion perception layer integrates a high-resolution infrared depth vision unit. This unit employs active infrared projection technology, enabling it to acquire facial depth information of children at a frequency of 60 frames per second under various lighting conditions. This high-resolution infrared depth vision unit can not only recognize basic facial contours but also accurately capture subtle tremors of facial muscles, i.e., micro-expression features. Visual data is processed in real time through a dedicated convolutional neural network. The input layer of the neural network receives 128x128 pixel infrared images, and after five layers of convolutional operations and pooling processing, a feature vector containing the coordinates of 68 facial key points is output at the fully connected layer. Pupil changes, as a direct physiological mapping of emotional fluctuations, are also included in the monitoring scope of the high-resolution infrared depth vision unit. By statistically analyzing the trend of pupil diameter changes at different interaction frequencies, the system can preliminarily determine the child's stress level.

[0024] In the auditory perception dimension, the multimodal emotion perception layer is equipped with an array-type microelectromechanical system (MEMS) microphone unit. This unit consists of four high-sensitivity MEMS microphones arranged in a cross shape, achieving precise sound source localization through beamforming algorithms. The audio signals acquired by the array-type MEMS microphone unit first undergo gain control and echo cancellation processing, and then key features are extracted using the Mel frequency cepstral coefficient algorithm. Specifically, the system divides the audio into short 20-millisecond frames, calculates its power spectrum in each frame, obtains the Mel spectrum through a set of Mel filter banks, and finally obtains the cepstral coefficients through discrete cosine transform. These coefficients accurately record the pitch fluctuations, energy distribution, and speech rate changes of the speech. Through long-term series analysis of these parameters, the system can identify the emotional polarity in the speech, such as identifying whether a child is in a cheerful high-frequency range or in a sad, low-energy spectral range.

[0025] Tactile perception is a crucial supplement to this system's deep understanding of children's emotions. Flexible pressure sensing units are widely deployed on the robot's outer shell, especially in high-frequency interaction areas such as the head, arms, and torso. These sensors employ the piezoresistive principle; when a child hugs, pats, or touches, the sensor's resistance changes linearly or non-linearly with the magnitude of the pressure. The sampling frequency of the flexible pressure sensing units is set to 100 Hz, sufficient to capture the fluctuations in force and frequency distribution during children's interactions. For example, fast-paced, high-intensity patting movements are often associated with excitement or irritability, while gentle, continuous stroking movements represent a psychological state of seeking comfort or expressing affection.

[0026] Acquiring physiological signals is crucial for understanding deep emotions. The non-contact physiological signal detection unit utilizes ultra-wideband radar technology to sense the subtle rise and fall of a child's chest by emitting nanosecond-level, extremely narrow pulses and receiving their echoes without physical contact. Through complex signal demodulation algorithms, the system can extract respiratory signals with frequencies of 0.1 to 0.5 Hz and heart rate signals with frequencies of 1 to 3 Hz from background stray echoes. This non-contact acquisition of physiological data, by eliminating the psychological stress associated with wearable devices, reflects the child's most authentic autonomic nervous system state.

[0027] Please refer to the attached document. Figure 4The mental state modeling layer receives heterogeneous data streams from the multimodal emotion perception layer. This layer includes an emotion feature fusion engine that uses an attention mechanism to weight and fuse visual, auditory, tactile, and physiological data. Its core logic lies in dynamically assigning weight coefficients to different sensors based on the current interaction context. If insufficient ambient light leads to a decrease in the confidence level of visual features, the engine automatically increases the weight ratio of voice emotion features and tactile features, thereby eliminating noise interference from single-sensor data. The mental state modeling layer also includes a long short-term memory network model, specifically designed to analyze the time-series evolution of emotional states. The long short-term memory network model contains forget gates, input gates, and output gates, capable of remembering emotional fluctuation trends up to 300 seconds. By calculating continuous emotional feature vectors, the system can determine the probability distribution of a child's state of excitement, sadness, anxiety, or anger, forming a real-time psychological profile.

[0028] Please refer to the attached document. Figure 2 A core innovation of this invention lies in the quantitative assessment of the Abebe complex within the psychological state modeling layer. The system uses an infrared depth vision unit to track the physical distance between the child and a specific comfort toy or fabric in real time, a process utilizing object detection algorithms and depth map matching technology. Simultaneously, a tactile sensing unit records changes in the intensity of the child's response to robot interactions while holding the specific object. If the child shows higher compliance with the robot's commands when holding the specific object tightly, it reflects that the object has a significant emotional support effect. Furthermore, a voice analysis unit monitors whether the child exhibits crying signals within a specific frequency range of 3000 to 5000 Hz when the specific object is out of sight.

[0029] The system inputs the above parameters into a preset attachment assessment model, whose calculation principle follows the logic below: Attachment Index = Weight 1 × Normalized value of physical distance + Weight 2 × Rate of change of interaction response intensity + Weight 3 × Energy density of anxiety voice signal

[0030] The system performs real-time calculations on the parameters involved in the above formula, resulting in a scalar value between 0 and 1. When the attachment index exceeds a preset threshold of 0.8, the system determines that the child is currently in a strong object-specific attachment state. At this point, the psychological state modeling layer generates a high-priority attachment state label, which is then passed to subsequent decision-making logic. To ensure the scientific rigor of the assessment, a time decay factor is also introduced into the model. This factor dynamically adjusts its weight allocation based on the child's chronological age (obtained through registration information or visual estimation), simulating the natural weakening of attachment levels as development progresses.

[0031] Please refer to the attached document. Figure 5The dynamic behavior decision-making layer operates within a hierarchical control architecture, ensuring the robot's reaction speed and guidance depth in complex environments. The bottom layer is the reactive behavior generation unit, with an extremely short operating cycle, typically less than 10 milliseconds. This unit directly connects to the output of the safety boundary monitoring unit, handling emergency obstacle avoidance, preventing pinching injuries, or providing immediate soothing actions when extreme negative emotions are detected. For example, when the non-contact physiological signal detection unit detects a sudden increase in a child's heart rate above a warning threshold, the reactive behavior generation unit forces the robot to stop all violent movements, switches to a static standby mode, and plays soft white noise.

[0032] The middle layer is the strategic guidance unit, which is key to handling the Abebe complex. This unit adjusts the robot's guiding tone and movement amplitude based on the current attachment index. When the attachment index is in the high range of 0.6 to 0.8, the strategic guidance unit instructs the robot to use a lower voice frequency and smaller body movements to avoid creating a sense of intrusion. Simultaneously, the robot uses its visual unit to confirm the location of the attachment object and uses it as a medium for shared attention during interactive tasks, such as through voice prompts like "Let's listen to a story with the bear," thus ensuring that the interactive behavior does not conflict with the child's attachment psychology.

[0033] The top-level unit is a goal-oriented evolutionary unit, geared towards long-term educational objectives. Based on the cognitive development stages identified in the psychological profile, this unit gradually introduces tasks such as language learning and logical reasoning. It determines when to shift from emotional reassurance to cognitive guidance by analyzing the child's engagement curves across multiple interactions. Please refer to the appendix. Figure 6 The dynamic behavior decision layer uses reinforcement learning algorithms to optimize decisions, aiming to find the optimal behavior path.

[0034] In the reinforcement learning architecture, the system defines the following reward function logic: Overall reward = Emotional improvement benefit + Interaction duration gain + Task completion weighted value - Interaction conflict cost.

[0035] Among these metrics, emotional improvement gains are obtained by comparing the probability distribution of the mental profile before and after the behavior is executed; interaction duration gains are used to assess the attractiveness of the guidance strategy; task completion weighting reflects the level of achievement of educational goals; and interaction conflict costs are triggered by negative fluctuations in the attachment index and the child's resistant expressions. Through iterative training in simulated environments and real-world scenarios, the decision layer can automatically optimize the combination of behaviors that maximizes the comprehensive reward function. This learning process is dynamic and continuous; as the number of interaction days increases, the robot can accurately grasp the emotional threshold of a specific child.

[0036] Please refer to the attached document. Figure 7The execution and feedback closed-loop layer is responsible for translating decision-making instructions into perceptible physical manifestations. The core of the actuator is a servo motor drive frame with at least 12 degrees of freedom. These servo motors are distributed across the head (2 degrees of freedom), the arms (3 degrees of freedom each), and the chassis (4 degrees of freedom). Each servo motor has a built-in high-precision magnetic encoder capable of 14-bit angle resolution. To simulate subtle limb movements, the motor control algorithm employs a three-loop control structure: a position loop, a velocity loop, and a current loop, ensuring smooth and gentle torque output during actions such as "touching the head" or "passing an object."

[0037] The facial expression display module serves as a window for the robot to convey emotions. In this embodiment, a high-pixel-density LED dot matrix screen is used, with a pixel pitch of less than 1.2 millimeters. Through subpixel rendering technology, the dot matrix screen can present smooth facial lines and color transitions. When the psychological state modeling layer outputs a pleasant emotion, the facial expression display module renders a smiling face with slightly flushed cheeks in real time, and the curvature of its eyes vibrates synchronously with the tone of voice. Furthermore, the system also features a micro-projection display technology option, which can project auxiliary teaching content or comforting scenes onto the ground or walls around the robot.

[0038] The speech synthesis module is based on deep learning-based emotional speech synthesis technology. This module integrates a pre-trained acoustic model and waveform synthesizer, enabling it to adjust the timbre saturation, fundamental frequency envelope, and phrasing of the output speech in real time based on the emotional tags in the decision commands. The system employs dual-channel audio processing technology. The first channel outputs clear instructional speech generated based on deep learning. The second channel is used for real-time mixing of background music or white noise. When the attachment index is high or anxiety is detected in the child, the second channel automatically increases the volume of low-frequency ambient sounds in the 200-500 Hz range, using acoustic masking to help alleviate the child's emotional fluctuations. This acoustic intervention logic, combined with visual and tactile feedback, forms a closed loop, quickly calming the child's mood.

[0039] Please refer to the attached document. Figure 8 The safety boundary monitoring unit operates independently of the behavior decision layer and is located at the highest safety level of the system. This unit monitors the drive torque of the servo motors in real time. When the robot arm encounters an obstacle (such as a child's hand) on its movement path, causing a sudden torque change exceeding 2 Nm, the safety boundary monitoring unit will cut off the driver's power supply within 50 milliseconds via a hardware interrupt signal. Simultaneously, this unit uses an ultrasonic sensor array to monitor the movement speed. When the distance to a child is less than 0.3 meters, the movement speed is forcibly limited to less than 0.1 meters per second.

[0040] The cloud-based learning and evolution unit is responsible for the system's long-term lifecycle management. The robot sends anonymized interaction logs to the cloud server via a secure, encrypted channel. These logs contain environmental context, decision sequences, attachment index change curves, and final evaluations of guidance effectiveness. Large-scale data mining algorithms in the cloud perform cluster analysis on these massive samples to identify the optimal guidance patterns for specific attachment characteristics. Through offline reinforcement learning and knowledge distillation techniques, the cloud periodically generates new neural network weight files and pushes them to the terminal robot via an online update mechanism. This enables the robot to continuously learn from the latest educational psychology research, giving its guidance strategies cross-regional and cross-cultural adaptability.

[0041] The power management unit plays a crucial role in energy conservation and reducing electromagnetic radiation within the system. Based on real-time assessments of interaction intensity, when the system detects that the child is in deep sleep or has been in a prolonged state of low interaction willingness (extremely low attachment index and no motor triggers), the power management unit instructs the multimodal emotion perception layer to enter a low-power cruise mode. In this mode, the frame rate of the infrared depth vision unit decreases from 60 frames per second to 1 frame per second, the sampling duty cycle of the non-contact physiological signal detection unit decreases to 10%, and only the voice wake-up trigger based on the array-type microelectromechanical system microphone unit is retained. When a specific wake word or a strong emotional acoustic signal is detected, the system can fully recover to full-load operation within 500 milliseconds, effectively extending battery life.

[0042] The system in this embodiment demonstrated a high level of intelligence in practical applications. When guiding children with a strong Abebe complex, the system first uses an infrared depth vision unit to lock onto the blanket (a specific attachment object) in the child's hands, and the attachment assessment model calculates an attachment index of 0.92. The dynamic behavior decision layer immediately searches the psychological intervention template library and selects the "indirect transition" strategy. Instead of directly asking the child to put down the blanket, the robot displays a curious expression through the facial expression module and uses the speech synthesis module to speak in a gentle voice: "This little blanket looks very warm, can I touch it too?" After receiving positive tactile feedback from the child, the robot gradually introduces a collaborative block-building task, smoothly shifting the child's attention through joint attention guidance technology, ultimately achieving the preset cognitive training goals while maintaining the child's emotional stability. Example 2

[0043] Building upon Example 1, Example 2 further refines the system's decision-making adaptation mechanism in extreme environments and scenarios with multiple children. Please refer to the appendix... Figure 5 With appendix Figure 8 In this embodiment, a social game submodule is added to the dynamic behavior decision layer to handle the emotional feedback when two or more children interact with the robot at the same time.

[0044] When the multimodal emotion perception layer identifies the presence of two children, the convolutional neural network automatically assigns a unique identifier to each child. The emotion feature fusion engine constructs an independent feature vector stream for each identifier. The long short-term memory network model in the mental state modeling layer evolves into a parallel processing architecture, simultaneously tracking the emotional state distribution of the two children. At this point, the calculation of the attachment index becomes more complex, as the system needs to assess whether there is attachment conflict between the children regarding the same specific object.

[0045] Driven by the social game submodule, the dynamic behavior decision layer employs a multi-objective game strategy. The system uses the emotional balance between the two children as a crucial constraint in the reward function. If one child exhibits extremely high attachment intensity while the other shows a strong desire to interact, the decision layer will implement a spatially phased interaction scheme through an execution and feedback closed-loop layer. For example, the robot utilizes the chassis's 4-DOF steering mechanism to provide differentiated visual feedback to the two children in turn. The expression display module displays a calming expression for the child on the left and an encouraging, excited expression for the child on the right; this asynchronous expression control technology is achieved through partitioned refresh of the dot matrix screen.

[0046] Furthermore, this embodiment incorporates flexible skin technology in its execution and feedback closed-loop layer. A 10mm thick layer of environmentally friendly silicone skin covers the servo motor drive frame, with an array of flexible pressure sensing units embedded within it. This allows the robot to sense not only point pressure but also large-area hugging pressure distribution. In the attachment assessment model, the contact area and pressure uniformity of the hugging are used as new input variables to more accurately characterize the child's level of "psychological security."

[0047] For long-term emotional modeling, the psychological state modeling layer introduces a global memory unit to store the emotional evolution trends of children over a span of several months. By performing localized calculations on the regression analysis model distributed by the cloud-based learning evolution unit, the system can identify the seasonal fluctuations in children's attachment to specific objects. For example, in winter, due to increased time spent at home, the attachment index often shows a periodic increase, and the dynamic behavior decision-making layer will pre-set a richer library of indoor guidance logic to cope with potential emotional fluctuations.

[0048] The safety boundary monitoring unit enhances the collision avoidance level in multi-child scenarios. Utilizing fused data from an ultrasonic sensor array and an infrared depth vision unit, the unit constructs a dynamic spherical protective shield model. When it detects two children too close together and on the robot's trajectory, the system automatically reduces the speed of each degree of freedom and issues a warning using a voice synthesis module. This comprehensive safety assurance, combined with precise emotional guidance, makes this system highly valuable in complex scenarios such as kindergartens and children's rehabilitation centers. Example 3

[0049] Please refer to the attached document. Figure 1 To be continued Figure 8 This embodiment details the application of the present invention in a medical assistance setting, particularly for emotional support in children with autism spectrum disorder. In this specific application scenario, the attachment assessment model in the psychological state modeling layer underwent parameter recalibration.

[0050] Because children with autism often exhibit obsessive attachment to specific objects far exceeding that of typical children, the preset threshold for the attachment index is adjusted to 0.95 in this embodiment. The non-contact physiological signal detection unit in the multimodal emotion perception layer is set to the highest priority. Ultra-wideband radar technology is used to monitor the child's respiratory and heart rate variability in real time during interaction. This variability is the gold standard for assessing autonomic nervous system stress. When the system detects a significant decrease in respiratory and heart rate variability (indicating increased stress), the psychological state modeling layer immediately updates the psychological profile, determining that the child has entered a stress state.

[0051] In this scenario, the dynamic behavior decision-making layer employs a "gradual desensitization" strategy. The goal-oriented evolutionary unit sets extremely small target steps. For example, when a child's attachment to a specific comfort object is detected to be extremely high, the execution and feedback loop layer does not guide the child away from the object. Instead, it instructs the robot to mimic the child's movements, vibrating or swaying at the same frequency as the comfort object. Through this "mirror effect," the robot can quickly establish an emotional connection with the child.

[0052] The speech synthesis module automatically switches to "native language" mode at this time. Based on dual-channel audio processing technology, the second channel plays ambient background sounds collected from parents and digitized, or uses pink noise of a specific frequency to reduce children's defensiveness towards the robot's synthesized speech. The facial expression display module reduces brightness and contrast to avoid excessive visual stimulation that could overload children's senses.

[0053] In the reinforcement learning algorithm, the reward function was modified to focus on the "duration of joint attention." The system defines the frequency with which a child's gaze switches between the robot and the attachment object as a positive reward. Through thousands of iterations of optimization, the system is able to spontaneously evolve a gentle transition logic: by manipulating the assistive tools in the robot's hands, it uses changes in light and shadow to guide the child's curiosity, thereby spontaneously reducing the pathological dependence on a specific attachment object.

[0054] In this embodiment, the cloud-based learning and evolution unit acts as an expert system library. Medical experts can remotely adjust the feature weighting coefficients of the psychological state modeling layer based on each child's clinical presentation through the cloud backend. This human-machine collaborative update mechanism transforms the decision-making system from an automated program into a remotely configurable auxiliary treatment tool. The safety boundary monitoring unit provides timely warnings to medical staff's handheld terminals by monitoring children's abnormal screams or violent body movements in real time, achieving a deep integration of physical safety, psychological safety, and medical intervention.

[0055] In summary, this invention constructs an intelligent interactive system capable of deeply understanding children's specific emotional needs (especially the Abebe complex) through the collaborative work of a multimodal emotion perception layer, a psychological state modeling layer, and a dynamic behavior decision-making layer. This system not only achieves deep integration and precise decision-making of physiological, facial expression, language, and tactile data at the technical level, but also demonstrates respect for and scientific guidance regarding children's psychological development at the application level. Its hierarchical control architecture, reinforcement learning-driven path optimization, and cloud-based closed-loop learning evolution mechanism collectively ensure that the robot can provide a natural, intelligent, and highly empathetic companionship and educational experience in complex and ever-changing real-world environments, significantly improving the technical level and market competitiveness of existing children's guidance robots.

Claims

1. A behavior decision-making system for a child-guided robot based on real-time emotional feedback, characterized in that, include: The multimodal emotion perception layer utilizes active infrared projection technology to acquire facial muscle micro-expressions and pupil changes, and employs microphone array beamforming algorithms and Mel-frequency cepstral coefficient algorithms to extract speech emotion features. This is combined with a piezoresistive flexible pressure sensing module deployed on the robot's outer shell to detect the force and frequency of tactile interactions. Simultaneously, ultra-wideband radar technology is used to remotely perceive the child's respiratory rate and heart rate signals. The psychological state modeling layer uses an attention mechanism to dynamically weight and fuse heterogeneous data streams from vision, hearing, touch, and physiology. It also utilizes a long short-term memory network model integrating forgetting, input, and output gates to analyze the temporal evolution of emotional states, identifying the probability distribution of psychological profiles including dimensions of excitement, sadness, anxiety, and anger. Furthermore, it tracks the child's interactions with specific objects using object detection algorithms and depth map matching technology. Physical distance, combined with monitoring of the response intensity and specific frequency crying signals when holding a specific object, is used to calculate a scalar numerical attachment index using an attachment assessment model; the dynamic behavior decision layer operates in a hierarchical control architecture, which includes a bottom-level reactive behavior generation unit, a middle-level strategy-based guidance unit, and a top-level goal-oriented evolution unit. Based on a state space composed of psychological profiles, attachment index, and environmental states, a reinforcement learning algorithm is used to automatically optimize and generate the optimal sequence of behavioral instructions; the execution and feedback closed-loop layer is used to drive a servo motor frame with no less than 12 degrees of freedom through a three-loop control structure including position loop, velocity loop, and current loop, and to drive an LED dot matrix screen to display facial expressions. At the same time, a speech synthesis module with dual-channel audio processing technology is used to output speech with emotional polarity.

2. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, The multimodal emotion perception layer includes: a high-resolution infrared depth vision unit, used to acquire facial depth information at a frequency of 60 frames per second, and to extract feature vectors containing coordinates of 68 facial key points using a convolutional neural network; an array-type microelectromechanical system microphone unit, used to divide the acquired audio signal into short frames of 20 milliseconds, and to calculate the power spectrum and Mel filter bank output in each frame to identify pitch fluctuations, energy distribution and speech rate changes in speech; a flexible pressure sensing unit, used to sense pressure changes during child interaction at a sampling frequency of 100 Hz; and a non-contact physiological signal detection unit, used to demodulate respiratory signals with a frequency of 0.1 to 0.5 Hz and heart rate signals with a frequency of 1 to 3 Hz from background stray echoes.

3. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, When the psychological state modeling layer performs attachment index calculation, its calculation logic is as follows: normalize the physical distance obtained from tracking to obtain a normalized physical distance value; calculate the rate of change in interaction response intensity when holding a specific object compared to when not holding a specific object; The energy density of crying or anxious voice signals at a specific frequency is statistically analyzed; the normalized value of the physical distance, the rate of change of the interaction response intensity, and the energy density of the anxious voice signal are multiplied by preset first weights, second weights, and third weights, respectively, and summed to obtain the attachment index between 0 and 1.

4. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 3, characterized in that, The attachment assessment model also incorporates a time decay factor; the time decay factor is used to dynamically adjust the allocation ratio of the first weight, the second weight, and the third weight according to the child's physiological age; the parameters of the time decay factor are dynamically adjusted based on regression analysis of historical attachment data to characterize the natural weakening of attachment as the child's cognitive development stage progresses.

5. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, The hierarchical control architecture operates as follows: the reactive behavior generation unit processes emergency obstacle avoidance and immediate emotional soothing tasks within a cycle of less than 10 milliseconds; the strategic guidance unit adjusts the robot's guidance tone frequency and body movement amplitude according to the current attachment index, and when the attachment index exceeds a preset threshold, instructs the robot to introduce a specific object as a medium for shared attention into the interactive task; the goal-oriented evolutionary unit generates an evolutionary strategy that transitions from emotional soothing to cognitive guidance based on the identified cognitive development stage and the analysis of the child's participation curve.

6. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, When performing decision optimization, the reward function adopted by the reinforcement learning algorithm consists of the following parts: emotional improvement gain, which is determined based on the difference in probability distribution of mental profiles before and after behavior execution; and interaction duration gain, which is used to evaluate the attractiveness of the guidance strategy to children. The task completion weighted score is used to reflect the level of achievement of educational guidance goals; The cost of interaction conflict is triggered by the negative fluctuations in the attachment index and the child's resistant facial expressions.

7. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, The execution and feedback closed-loop layer includes: servo motors with built-in 14-bit angle resolution magnetic encoders, distributed in the robot's head, arms, and chassis; a light-emitting diode dot matrix screen using subpixel rendering technology, used to render corresponding facial lines and color transitions based on the emotional probability distribution output by the psychological state modeling layer; the first channel of the speech synthesis module is used to output teaching guidance speech generated based on deep learning, and the second channel is used to increase the volume of low-frequency ambient sounds when the child is detected to alleviate emotional fluctuations by utilizing the acoustic masking effect.

8. The child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, It also includes a safety boundary monitoring unit that operates independently of the behavior decision layer; the safety boundary monitoring unit monitors the driving torque of the servo motor and the robot's moving speed in real time; when the change in the driving torque exceeds a preset threshold of 2 Nm, the safety boundary monitoring unit cuts off the power supply through a hardware interrupt signal; when the distance between the robot and the child is less than 0.3 meters, the safety boundary monitoring unit forcibly limits the robot's moving speed to within 0.1 meters per second.

9. A child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, It also includes a cloud-based learning evolution unit; the learning evolution unit is used to collect desensitized interaction logs, which include environmental context, decision sequence, attachment index change curve, and guidance effect evaluation; the learning evolution unit performs large-scale data mining and cluster analysis in the cloud, and generates new neural network weight files through offline reinforcement learning and knowledge distillation techniques; the new neural network weight files are pushed to the terminal robot through an online update mechanism to update the configuration parameters of the psychological state modeling layer and the dynamic behavior decision layer.

10. A child-guided robot behavior decision-making system based on real-time emotional feedback according to claim 1, characterized in that, It also includes a power management unit; the power management unit dynamically adjusts the sampling frequency of each unit in the multimodal emotion perception layer based on the real-time evaluation results of the interaction intensity; when the child is detected to be in deep sleep or in a low interaction willingness state, the power management unit instructs the system to enter a low-power cruise mode, in which the operating frame rate of the high-resolution infrared depth vision unit and the sampling duty cycle of the non-contact physiological signal detection unit are reduced, and only the voice wake-up trigger function is retained.