Interactive Robot Emotion Recognition via Multimodal Sensing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current interactive robots are limited in their ability to effectively interact with humans due to limitations in sensing and responding to environmental stimuli and user emotions.

Innovation Solution

An interactive robot system equipped with image and audio capturing devices, pressure and odor sensing, and output devices, utilizing a human-robot interaction method that includes modules for sensing, recognizing, analyzing, and executing responses based on neural network and deep learning algorithms to understand and respond to user emotions and environmental information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple sensing devices and complex algorithms are added to improve interaction capability, then the robot's ability to recognize and respond to user emotions and environmental stimuli is improved, but the device complexity increases

Engineering Contradiction:
Improveinteraction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The interaction system is divided into distinct functional modules: audio capturing module, image capturing module, pressure sensing module, odor sensing module, neural network analysis module, and response execution module. Each module handles a specific aspect of interaction, making the complex system manageable and maintainable while achieving comprehensive interaction capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The robot integrates multiple sensing devices (audio, image, pressure, odor) and output devices (audio output, expression output, movement output) into a single platform that can handle various types of interactions simultaneously. The neural network analysis module processes multiple input types (audio information, image information, sensing data) to generate comprehensive responses, making the system versatile across different interaction scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If neural network and deep learning algorithms are used to accurately understand user emotions, then the measurement precision of emotion recognition is improved, but the loss of time for processing increases

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of input data before neural network analysis. Audio information is captured and preprocessed, image information is captured and preprocessed, and sensing data is collected in advance. This preliminary action prepares the data in an optimized format, reducing the processing time required by the neural network while maintaining high recognition accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The interaction system operates continuously, with the neural network analysis module constantly processing incoming data streams from multiple sensors. The system maintains continuous monitoring of user emotions and environmental stimuli, eliminating idle processing time and ensuring real-time response capabilities through uninterrupted analysis

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10482886B2Interactive robot and human-robot interaction method
Publication Date: 2019.11.19 FU TAI HUA IND SHENZHEN
  • US10482886B2 patent drawing
  • US10482886B2 patent drawing
  • US10482886B2 patent drawing

AI summary

An interactive robot includes an image capturing device, an audio capturing device, an output device, and a processor. The processor is configured to obtain audio information captured by the audio capturing device and image information captured by the image capturing device, recognize a target from the audio information and the image information, confirm basic information and event information of the target and link the basic information with the event information, obtain key information from the event information of the target, implement a neural network analysis algorithm on the key information to confirm an emotion type of the target, search a preset public knowledge database according to the key information to obtain a relevant result, apply a deep learning algorithm on the relevant result and the emotion type of the target to determine a response, and execute the response through the output device.