Remote Collaboration Feedback Using Vocal Entrainment and Stress Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Remote collaborative teams face challenges in high workload situations due to stress and pressure, which impact communication and collaboration, as they lack the perceptual cues of face-to-face interactions, and existing systems fail to utilize the relationship between vocal entrainment and physiological stress for improving team coordination.

Innovation Solution

A system and method that processes audio and physiological signals to determine vocal entrainment and stress alignment between remote collaborators, generating feedback to enhance communication and collaboration through human-machine interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If remote collaborative communication is used, then team coordination and collaboration are enabled, but perceptual cues such as facial expressions and body language are lost

Engineering Contradiction:
Improveteam coordinationVSAvoidperceptual cues
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system uses audio signals as an intermediary to capture and transmit vocal entrainment patterns that serve as proxies for lost perceptual cues. By analyzing speech-related features and physiological data, the system mediates the transmission of emotional and cognitive state information between remote team members, compensating for the absence of visual and tactile feedback.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If vocal entrainment analysis is added to detect stress and improve coordination, then team coordination is improved, but system complexity increases

Engineering Contradiction:
Improveteam coordinationVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs multi-functionality by using a single audio signal processing framework that simultaneously extracts multiple speech-related features (pitch, intensity, spectral characteristics) and physiological features (heart rate, respiration rate, skin conductance) from the same audio input. This unified approach allows vocal entrainment analysis, stress detection, and coordination monitoring to be achieved through one integrated system rather than separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If real-time feedback is provided on vocal entrainment and stress alignment, then communication and collaboration are enhanced, but processing requirements and system complexity increase

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by continuously monitoring and analyzing vocal entrainment patterns and stress levels in real-time during communication interactions. By processing audio signals and physiological data as they occur, the system proactively detects changes in coordination and stress alignment, enabling immediate feedback and intervention before communication breakdowns occur, rather than analyzing patterns after the fact.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4386750B1System and method for real-time feedback of remote collaborative communication
Publication Date: 2025.11.12 HONEYWELL INTERNATIONAL INC
  • EP4386750B1 patent drawingFigure 1
  • EP4386750B1 patent drawingFigure 2
  • EP4386750B1 patent drawingFigure 3

AI summary

A system and method for providing real-time feedback of remote collaborative communication between users includes extracting speech-related features and physiological features from at least one of the users and using these features to determine a stress state of at least one user. In response to the determined stress state audio signals may be processed to manipulate one or more vocal features of the speech supplied from another user, and/or at least one device may supply feedback to another user that provides suggestions as to how to manipulate or more of their vocal features.