Separating Transcribed Text and Physiological Data in Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional videoconferencing technologies fail to provide reliable feedback on the intelligibility of speech and physiological state of remote participants, especially in telehealth settings, due to connectivity issues and sub-optimal call quality, making it difficult for professionals to assess comprehension and emotional responses.
Innovation Solution
A system and method that separately convey transcribed text and physiological data from video and audio data, using speech-to-text conversion, feature extraction, and synchronization to display and store this information in real-time, allowing for better comprehension of the remote participant's emotional and physiological state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional videoconferencing technologies are used, then basic communication is enabled, but reliable feedback on speech intelligibility and physiological state is not provided
Solution Approach 1:
The system segments the communication data into separate channels: video data, audio data, transcribed text, and physiological data are transmitted through independent pathways. This segmentation ensures that failure in one channel does not compromise the entire communication, thereby improving feedback reliability while reducing information loss.
Solution Approach 2:
The system introduces intermediate processing layers including speech-to-text conversion modules and physiological data extraction units that act as mediators between the raw audio/video signals and the final feedback output. These intermediaries enhance information extraction accuracy and ensure reliable transmission of comprehensive feedback.
2Loss of information
If video and audio data are transmitted together, then complete information is conveyed, but connectivity problems cause loss of transmitted information
Solution Approach 1:
The transmission system divides the data stream into separate segments for video, audio, text, and physiological data. Each segment can be transmitted independently through dedicated channels, ensuring that if one segment is lost due to connectivity issues, the others remain intact, thus maintaining both information completeness and transmission reliability.
Solution Approach 2:
The system prepares backup transmission paths and redundancy mechanisms in advance for each data type. Critical information such as transcribed text and physiological data are prioritized for retransmission upon connection restoration, cushioning against potential information loss during connectivity interruptions.
3Loss of information
If physiological data is extracted and synchronized with text, then comprehensive feedback is provided, but system complexity increases
Solution Approach 1:
The system merges multiple data streams (video, audio, text, physiological data) into a unified feedback structure that presents comprehensive information in an integrated manner. By combining these data types through synchronization and temporal alignment, the system provides complete feedback while managing complexity through structured integration rather than separate processing modules.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Transcribed text and physiological data of a remote video conference participant are transmitted to a local device separately from the video data, which depicts the remote party during a time interval. An image of the video data is captured at a time instant within the time interval. A value of a remote party feature is determined remotely using the video data. The remote party feature can be the remote party's heart rate at the time instant. The value of the feature is received onto the local device. Audio data captures sounds spoken by the remote party and is converted by the remote device into words of text. The audio data converted into a particular word was captured at the time instant. The particular word is received onto the local device. The particular word and the value of the feature are displayed in association with one another on the local device.