Assisting Human Interaction via Action Decoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals with emotion recognition difficulties, such as those with Autism Spectrum Conditions, face challenges in social situations due to misinterpretation of idioms and nuanced human interactions, and there is a need for automated guidance in understanding human actions and intentions, particularly in professional and cultural contexts.

Innovation Solution

A system that receives and decodes action data from humans, including audio and visual cues, to generate user response data, using facial expression detection, speech analysis, and contextual information to provide socially contextual information and assist in responding appropriately, employing a device with video and audio capture capabilities and a remote server for processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated guidance systems are implemented to decode human actions and provide response recommendations, then social interaction understanding is improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improveaccuracy of interpreting human actions and intentionsVSAvoidcomplexity of decoding system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex task of understanding human interaction into distinct components: audio processing, visual processing, facial expression detection, speech analysis, and contextual interpretation. Each component handles a specific aspect of the input data and passes results to the next stage, making the overall system more manageable and interpretable while maintaining high accuracy in decoding human actions and intentions.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple data sources (audio, visual, contextual) are integrated to improve interpretation accuracy, then measurement precision increases, but information processing time and computational load increase

Engineering Contradiction:
Improveaccuracy of action interpretationVSAvoidprocessing time for decoding actions
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of audio and visual data streams by pre-processing audio signals for speech and tone detection, and pre-processing visual data for facial expression and body language detection before the main decoding stage. This preliminary action prepares the data in advance, reducing the computational burden during real-time interaction and minimizing processing delays while maintaining high interpretation accuracy.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the system provides detailed contextual information and response recommendations, then ease of operation for users with emotion recognition difficulties is improved, but information overload may occur for users without such needs

Engineering Contradiction:
Improveease of social interaction for users with ASCVSAvoidamount of information provided to user
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The system adapts the quantity and detail of information provided based on the specific needs of the user. For users with Autism Spectrum Conditions, the system provides detailed contextual information including decoded actions, intentions, and suggested responses. For other users, the system can provide summarized or optional detailed information, allowing each user to receive information at the appropriate level of detail for their specific requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10467916B2Assisting human interaction
Publication Date: 2019.11.05 BISHOP JONATHAN EDWARD
  • US10467916B2 patent drawing

AI summary

A method of, and system for, assisting interaction between a user and at least one other human, which includes receiving (202) action data describing at least one action performed by at least one human. The action data is decoded (204) to generate action-meaning data and the action-meaning data is used (206) to generate (208) user response data relating to how a user should respond to the at least one action.