Intelligent accompanying robot system based on multi-modal interaction and control method

By adopting a multimodal interaction intelligent companion robot system in smart toys, combining lightweight face recognition and speech emotion analysis, real-time and accurate emotion recognition and feedback are achieved, solving the shortcomings of existing smart toys in terms of interaction methods and emotion recognition, and reducing hardware costs and network dependence.

CN120023829APending Publication Date: 2025-05-23CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510444578.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing smart toys have shortcomings in interaction methods, emotional recognition and feedback mechanisms, and cannot provide natural and personalized emotional companionship. Due to computing power and power consumption limitations, it is difficult to achieve real-time AI processing and multimodal perception.

Method used

The intelligent companion robot system with multimodal interaction is adopted, combined with the ESP32-S3 chip for face recognition and speech emotion analysis, and the STM32F103RCT6 chip drives expression display, body movement and music feedback, and alternately runs face and speech recognition tasks through polling to achieve real-time emotional recognition and feedback.

Benefits of technology

It provides real-time and accurate emotional recognition and feedback, reduces network dependence and hardware costs, and realizes low-energy consumption and efficient multimodal intelligent interaction, which is suitable for emotional companionship and treatment of children, the elderly and special groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023829A_ABST
    Figure CN120023829A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent accompanying robot system based on multi-modal interaction and a control method, which adopt a dual-chip collaborative framework to realize multi-modal emotion interaction. According to the equipment, a quantized and optimized lightweight face recognition and voice detection model is locally operated in an off-line manner through an ESP32-S3 chip, and facial expressions and voice emotions of a user are analyzed in real time; meanwhile, an STM32F103RCT6 chip is adopted to be responsible for peripheral control, and the peripheral control comprises OLED expression display, body movement driven by a servo motor and a music feedback module, so that visual and dynamic emotion feedback is achieved. A task polling mechanism is utilized, face and voice recognition tasks are efficiently and alternately processed under limited computing resources, and low power consumption and high responsiveness of the system are ensured. The interactive naturalness and the user immersion of the intelligent toy are improved, and the intelligent toy has a good personalized emotion regulation effect and is suitable for the fields of child accompanying, psychological counseling, special crowd assistance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent robots, and specifically relates to an intelligent companion robot system integrating emotion recognition and adaptive feedback, which is suitable for family companionship, emotional auxiliary treatment and entertainment interaction scenarios. Background Art

[0002] In recent years, with the development of artificial intelligence (AI) and Internet of Things (IoT) human-computer interaction technology, the smart toy market has ushered in new changes, and smart companion robots have gradually become an important tool in the fields of family, medical care and education.

[0003] Traditional toys are mechanical in structure or mainly for children to interact with. They have a single interaction mode and lack the ability to recognize and respond to user emotions. They cannot meet the needs of modern and special groups (such as children with autism, patients with emotional disorders, etc.) for emotional companionship and intelligent interaction. Although existing smart toys have certain customization functions, such as performing simple tasks through voice commands, most products still have the following problems: there are still many deficiencies in interaction mode, emotion recognition and feedback mechanism. Traditional companion robots mainly rely on voice or button control and lack natural and intuitive gesture interaction capabilities, resulting in a user experience that is not smooth and natural enough. There is a lack of personalized emotion recognition. Existing smart toys need to rely on fixed commands for interaction and cannot recognize and respond based on the user's facial expressions or voice emotions, making it difficult to achieve personalized companionship functions. Voice and visual interactions are independent and cannot integrate multimodal perception. Currently, smart toys on the market usually use a single voice recognition or simple camera capture function, which cannot comprehensively analyze the user's facial expressions and voice intonation, resulting in a stiff and unnatural interactive experience. The feedback mechanism lacks dynamic adaptation: The feedback actions of existing robots (such as expression changes and music playback) are mostly preset modes, and lack the ability to dynamically adjust according to the user's real-time emotional state. Due to the computing power and power consumption limitations of embedded devices, it is difficult to achieve computing resources and power consumption for real-time AI processing, and it is challenging to run face recognition and voice recognition on smart toys at the same time. High-computing AI processing usually requires cloud support, and cloud computing will bring privacy issues and network dependence, affecting the stability and practicality of the device. High hardware cost and complexity: Some high-end companion robots use complex bionic skin and a large number of servos to achieve expression changes, resulting in high hardware costs and difficulty in popularization. In addition, the complex mechanical structure also increases the difficulty of maintenance and the failure rate.

[0004] In order to overcome the defects of the above-mentioned prior art, the present invention proposes an intelligent companion robot system and control method based on multimodal interaction. The invention can provide users with a more natural and real emotional companionship experience, and is suitable for children, the elderly and special groups who need emotional healing. Summary of the invention

[0005] Based on the above technical deficiencies, the present invention provides an intelligent, low-cost, highly interactive companion robot system and control method. The system can not only detect user expressions and voice emotions, but also provide real-time feedback through hand raising, expression switching and music playing. The multimodal interactive intelligent companion robot system and control method of the present invention include two main parts: the artificial intelligence processing unit is mainly responsible for collecting and processing the raw data from the camera and microphone to realize face recognition and voice emotion analysis. The peripheral control unit is mainly responsible for driving the expression display, body movements and music feedback of the companion robot.

[0006] The artificial intelligence processing unit uses the ESP32-S3 chip, which is mainly responsible for collecting and processing raw data from the camera and microphone to achieve face recognition and voice emotion analysis. In order to reduce power consumption and optimize real-time performance, this unit uses task polling to alternately run face recognition and voice recognition tasks.

[0007] The peripheral control unit uses the STM32F103RCT6 chip, which is mainly responsible for driving the expression display, body movements and music feedback of the companion robot. This unit exchanges data with the AI ​​processing unit through UART (or SPI), receives emotion recognition results and performs corresponding actions.

[0008] The peripheral module, the camera module uses a low-power OV2640 module, which is connected to ESP32-S3 through a DVP interface. The camera is responsible for real-time acquisition of user facial images, and the image data is processed to form an AI processing unit for face recognition and emotion analysis.

[0009] The peripheral module uses an I2S digital microphone or other low-power microphone module to collect user voice signals. After A / D conversion and auxiliary processing, the collected audio data is input into the voice recognition module for emotion recognition.

[0010] The peripheral module, the audio feedback module uses the MAX98357A digital audio power to drive the speaker to achieve emotional music playback, such as playing soothing music to relieve user stress, or playing cheerful music to respond to user emotional joy.

[0011] The peripheral module, the OLED display screen, is used to display the facial expressions of the companion robot in real time and switch the preset expression patterns according to the received emotional information.

[0012] The peripheral module and the motion drive module adopt a servo motor drive mechanism, so that the companion robot can realize actions such as raising hands, waving, and nodding, thereby providing physical feedback.

[0013] The peripheral module also includes a power supply management module, an infrared sensor module and other necessary auxiliary modules, forming a complete multi-modal emotional interaction system.

[0014] Beneficial Effects

[0015] Compared with the existing technology, the present invention has the following advantages and positive effects:

[0016] 1. Real-time and accurate emotion recognition and feedback, using ESP32-S3 for lightweight face recognition and voice emotion analysis, combined with the efficient control design of STM32F103RCT6, to achieve real-time collection and intelligent judgment of user facial expressions, voice emotions and voice input. By polling and running face and voice recognition tasks alternately, emotion recognition and response speed are kept at a low latency, thus providing a smooth interactive experience.

[0017] 2. Low energy and high efficiency, the system adopts offline local AI model to avoid cloud data transmission and calculation, reduce network dependence, optimize task scheduling, and effectively control power consumption. The dual chips have clear division of labor, ESP32-S3 bears the burden of AI analysis data, and STM32F103RCT6 focuses on peripheral control, reasonably sharing the computing burden, extending the device's extended time, and is suitable for battery-powered and portable applications.

[0018] 3. Multi-state intelligent interaction, integrating multiple input methods such as face recognition and voice emotion analysis, ensuring the collection of user emotions from multiple angles and levels, and extremely rich interaction modes. The peripheral feedback module realizes expression display, motion control and music therapy, and provides an immersive emotional companionship experience through visual, auditory and preset feedback.

[0019] 4. Wide application applicability. The multimodal interactive intelligent companion robot system and control method of the present invention are not only suitable for children's companionship and entertainment, psychotherapy, emotional activities of special groups and other fields, but also have corresponding market promotion value and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the overall flow chart of the present invention.

[0021] Figure 2 It is a flowchart of the lightweight MobileNetV3-Small (INT8 quantization) model of the present invention.

[0022] Figure 3 It is a flow chart of the voice wake-up (WakeWord) function of the present invention.

[0023] Figure 4.1It is a schematic diagram of the system circuit of the microcontroller STM32F103RCT6 in an embodiment of the present invention.

[0024] Figure 4.2 It is a main control schematic diagram of the microcontroller STM32F103RCT6 in an embodiment of the present invention.

[0025] Figure 5 4 is a circuit schematic diagram of other peripheral circuits in the embodiment of the present invention.

[0026] Figure 6 It is a PCB provided by the implementation of the present invention.

[0027] Figure 7 It is a front view of the physical object provided by the implementation of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to explain the present invention and is not used to limit the present invention.

[0029] The embodiments of the present invention relate to a multi-modal interactive intelligent companion robot system and a control method, as shown in the attached Figure 1 As shown, it includes two main parts: AI data processing module and main control module. The AI ​​chip processes the collected user facial information and voice information, performs emotion recognition extraction and processing; the main control chip is used to receive the emotion data results sent by the AI ​​chip, and control the external settings to perform corresponding emotion feedback interaction. The UART protocol is used to realize the data communication between ESP32-S3 and STM32F103RCT6, which can not only reduce the complexity of hardware design, but also ensure the reliability of data transmission. In the peripheral circuit, the user can also interact with the robot by pressing the pressure sensor on the robot's head, raise his hand and hug feedback and change his expression. The user can also control the robot to raise his hand and other actions through infrared gestures.

[0030] Furthermore, in the embodiment of the present invention, a lightweight MobileNetV3-Small (INT8 quantization) model is used to process face data. MobileNetV3-Small is the third generation lightweight convolutional neural network proposed by Google. It realizes efficient feature extraction through the inverted residual structure, SE attention mechanism and h-swish activation function, and further compresses the model and accelerates reasoning in combination with symmetric / asymmetric quantization strategies. It performs well in edge computing scenarios such as face recognition and mobile payment, with a typical delay of <15ms and a model size of <1MB, making it an ideal choice for resource-constrained devices. Among them, the inverted residual structure (InvertedResiduals) extracts features in low-dimensional space through depthwise separable convolution (Depthwise Separable Convolution), and combines the linear bottleneck layer to reduce parameter redundancy and improve computational efficiency. Depthwise separable convolution is to decompose the standard convolution into depth convolution (spatial feature extraction) and point-by-point convolution (channel feature fusion), and the amount of calculation is reduced to 1 / 8-1 / 9 of the traditional convolution. The SE attention mechanism (Squeeze-and-Excitation) compresses spatial information through global average pooling, and then dynamically adjusts channel weights through the fully connected layer to enhance the expressiveness of key features. The quantized version of this module uses h-sigmoid instead of traditional sigmoid to reduce the impact of nonlinear calculations on quantization errors. The hardware-friendly activation function h-swish shown in Formula 1 uses a piecewise linear function to approximate the Swish activation function, which not only retains nonlinear expressiveness, but also reduces computational complexity and is more friendly to INT8 quantization. NAS and NetAdapt optimization: The layer structure and number of channels are automatically determined through neural architecture search (NAS), and the NetAdapt algorithm is combined to optimize the allocation of computing resources layer by layer, further compressing the model while maintaining accuracy.

[0031]

[0032] Further, as attached Figure 2 As shown in the figure, the specific steps of using the lightweight MobileNetV3-Small (INT8 quantization) model to process face data on the ESP32-S3 chip are as follows:

[0033] Further explanation, data collection and preprocessing: collect the user's face image in real time through the camera module, crop the face area from the collected image according to the model requirements, and scale it to the size required by the model (such as 224×224 or 160×160). Normalize the image pixel values, adjust the data range to the interval required by the model, and convert the processed image data into the data type required by the model input, and convert it into INT8 format data (such as the input format obtained through the quantization step).

[0034] Further explanation, the quantization of the model: Use the MobileNetV3-Small architecture to pre-train in floating-point precision to ensure that the model has good face recognition performance. Use TensorFlow Lite Converter to convert the floating-point model to an INT8 quantized model, including the quantization of weights and activation values. During the conversion process, use key data sets to make adjustments to ensure that the accuracy loss of the determined model is minimal, and obtain the quantized .tflite model file. All calculation results in this file are in INT8 format, which is suitable for running on embedded devices built with computing power.

[0035] Further explanation, model deployment and initialization: Model storage stores the estimated .tflite model file to the device's external or external storage. Loading the model uses TensorFlow Lite for Microcontrollers (or ESP-DL library) on ESP32-S3 to load the simulation model, initialize the interpreter, and allocate the necessary memory graph. Configure the interpreter, configure the input and output tensors, and ensure that the input tensor data format is consistent with the image data after investment.

[0036] To further explain, inference execution: Prepare input data Fill the reserved image data into the model's input tensor. Make sure the input data has been converted to INT8 format and the data dimensions are consistent with the model requirements. Call the interpreter's inference function (such as interpreter.Invoke) to let the model process the input data. Get output results Read the inference results from the model output tensor, which is usually output as a probability vector representing the confidence of each category ("happy", "sad", "angry", etc.).

[0037] Further explanation, the output probability is post-processed, the result is normalized using methods such as Softmax, and the recognized emotion category is determined. The recognition result is passed to the main control unit (STM32F103RCT6).

[0038] Further, as attached Figure 3 As shown in the figure, the specific steps of the wakeword model to implement speech recognition function on ESP32-S3 are:

[0039] Further explanation, data acquisition and preprocessing: configure the I2S driver of ESP32-S3 to collect audio data in real time (usually sampled in 16-bit PCM format), set the sampling rate (such as 16kHz or 8kHz, depending on the model requirements). Data framing and windowing: interrupt the continuous audio data stream into frames of fixed length (such as 20 to 40 milliseconds per frame), and visualize each frame to apply a window function (Formula 2 Hamming window) to reduce edge effects. Calculate features such as MFCC Mel-frequency cepstral coefficients (Formula 3, Formula 4) for each frame of audio, which are usually the input of the wake-up word model. MFCC extraction can be implemented on ESP32-S3 using existing lightweight DSP libraries, or directly use the pre-compiled TFLM feature extraction module.

[0040]

[0041] C[x(n)]=F -1 [log(|F[x(n)]| 2 )] (3)

[0042]

[0043] Further explanation, model preparation and merging: Select a suitable wake-up word detection model (such as a small model based on a trained neural network) and use TensorFlow on a PC to ensure that the model can distinguish between wake-up words and background noise. Use TensorFlow Lite Converter to convert the trained floating-point model to an INT8 quantized model to adapt to the computing resources of ESP32-S3 and significantly reduce memory usage and inference latency. During the solution process, use key audio data for adjustment to ensure that the accuracy loss of the solved model is minimal. Get the converted .tflite model file, which contains the header weight and infringement map and is stored in INT8 format.

[0044] Further explanation, model deployment and initialization: Integrate TensorFlow Lite for Microcontrollers (TFLM) or ESP-DL library on ESP32-S3 for model loading and reasoning. Store the determined wake-up word model in the device payload or PSRAM. Use the TFLM API in the code to load the model file and allocate the input and output tensor graphs. Initialize the TFLM interpreter and set the input tensor shape of the model to ensure that it matches the MFCC feature size extracted in the repair phase. Fill the MFCC features obtained by grouping audio reconstruction into the model input tensor to ensure that the data format is INT8 and meets the input requirements of the model. Use interpreter.Invoke() (or the corresponding API) to run model reasoning and obtain the output tensor. Analyze the model output results (usually probabilistic warnings), determine whether the confidence of the trigger word exceeds the set threshold, and determine whether to trigger. To avoid false triggering, subsequent processing steps (such as sliding window statistics, smoothing, etc.) can be added.

[0045] Further explanation, trigger mechanism and system response: When the model determines that the wake-up word is detected, the system event is triggered through software, such as notifying the main control unit or starting the subsequent voice recognition process. According to the wake-up word recognition result, the signal can be sent to the STM32F103RCT6 control peripherals through UART or communication interface to realize the emotional interaction feedback of the overall system (such as changing expression, playing music, and performing actions).

[0046] Further explanation, optimization and debugging: Test the inference time on ESP32-S3 to ensure that the wake-up detection meets the real-time requirements. If the delay is too large, you can further optimize the code or try to adjust the model structure. When monitoring for a long time, set some modules of ESP32-S3 to low power mode, or use an event-driven sleep wake-up mechanism. False trigger rate control: By adjusting the wake-up threshold and subsequent processing auxiliary algorithm, reduce the probability of false triggering and improve recognition accuracy.

[0047] Furthermore, ESP32-S3 runs FreeRTOS, which can manage multiple tasks through a task scheduler. In round-robin task scheduling, each task has a time slice, and by letting each task take turns to execute, it avoids long-term CPU occupation. The usual steps of round-robin task scheduling are as follows: Tasks are scheduled in order of priority. The running time of each task is limited (time slice), and the time slice is changed to switch tasks. The execution of each task is usually non-blocking, which can avoid long waiting between tasks.

[0048] Further, as attached Figure 4.1 , Attachment Figure 4.2 and attached Figure 5As shown in the figure, the hardware schematic diagram of the main control chip system and each peripheral, the main control chip configures each peripheral: configures the required peripheral clocks (such as I2S, SPI, UART clocks) for the camera, audio module, OLED display, servo motor, infrared sensor, and pressure sensor, configures the appropriate working mode of the relevant GPIO pins (such as I2C / SPI data line, PWM output, etc.), initializes the working mode and communication protocol of the peripherals (such as I2C, SPI, PWM).

[0049] According to the ESP32-S3 module, voice recognition + facial emotion data analysis and emotion classification are performed to control the expression display, action and voice feedback of the intelligent companion robot. Different feedback strategies are triggered according to different emotions: when happy, a smiling face is displayed, fast music is played, and the arms are waved; when sad, a melancholy expression is displayed, soothing music is played, and the user is hugged; when angry, an angry expression is displayed and the chest is hugged. Furthermore, infrared sensor gestures recognize specific actions, the robot performs hugs and interactions with different expressions, and user touch actions trigger voice and expression change interactions.

[0050] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. An intelligent companion robot system and control method based on multimodal interaction, characterized in that: include: Emotion recognition module: AI chip processes collected user facial and voice information to perform emotion recognition; The main control chip is used to receive the emotion data sent by the AI ​​processing chip and control the external settings to perform corresponding feedback interactions; Data transmission module, used for data communication between the AI ​​processing chip and the main control chip; The interactive feedback module includes an expression display component, a voice / music playback component, and an action execution component, which are used to interact according to the user's emotional state.

2. The intelligent companion robot system and control method based on multimodal interaction according to claim 1, characterized in that: The main control module adopts a microcontroller STM32F103RCT6 for peripheral control.

3. The intelligent companion robot system and control method based on multimodal interaction according to claim 1, characterized in that: Taking into account the overall cost, power consumption, computing power requirements and more accurate processing of user facial data and voice data, the ESP32-S3 chip is used alone for face recognition. It has an AI hardware acceleration unit (MVP) and supports machine learning (ML) model inference.

4. According to the intelligent companion robot system and control method based on multimodal interaction according to claim 3, its emotion recognition module is characterized in that: the camera module adopts OV2640 module (SPI or DVP interface) and is matched with ESP32-S3 chip to process user facial data.

5. The intelligent companion robot system and control method based on multimodal interaction according to claim 3, wherein the emotion recognition module further comprises: The lightweight MobileNetV3-Small (INT8 quantization) model is used for processing. The user's facial data is obtained through the camera module to extract the key features of the eyes and mouth. The AI ​​processing chip classifies the emotions of the user's facial images through the neural network model and sends the emotion recognition results to the main control chip.

6. According to the intelligent companion robot system and control method based on multimodal interaction as described in claim 2, its voice recognition feature is: a wakeword model is used to implement a simple voice recognition function on the AI ​​processing chip ESP32-S3, and the digital microphone module selects INMP441 (I2S interface).

7. According to the multimodal interaction-based intelligent companion robot system and control method of claim 6, the AI ​​processing chip is characterized by: avoiding running multiple high-energy reasoning tasks at the same time, and running them alternately in a polling task manner. The AI ​​processing chip enters the voice recognition mode when the user's face is not detected to save computing power and avoid task conflicts.

8. According to the intelligent companion robot system and control method based on multimodal interaction as described in claim 2, its data transmission module is characterized by: using UART, SPI, and I2C communication protocols to enable low-latency and high-stability data transmission between the AI ​​processing chip and the main control chip.

9. According to claim 8, the intelligent companion robot system and control method based on multimodal interaction, its interactive feedback module is characterized by: including the APDS-9960 infrared gesture sensor, which is used for gesture recognition of specific actions, and the robot makes hugs and interactions with different expressions; the FSR402 resistive film pressure sensor is located on the head of the robot, and the user's touch action triggers voice and expression change interaction; the bus multi-level servo is used as the robot's joint control device, and the servo selects the SG90 module; two OLED display modules are combined to display the robot's expression.

10. According to claim 8, the intelligent companion robot system and control method based on multimodal interaction has the following interactive feedback characteristics: the main control chip adjusts the interaction strategy based on the AI ​​chip voice recognition + facial emotion analysis data, such as playing adjusted music when the user's anxiety state is detected, or triggering a hugging action when the user's emotions reverse.

Citation Information

Cited By

  • Exoskeleton control and interaction system and method for intelligent doll

    CN120985686A

  • AI petty intelligent interaction system and method

    CN121979388A