AR Video Conferencing with Text-Searchable Transcripts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems are cumbersome for reviewing meeting content, as recordings are not text-searchable and require platform compatibility among devices, making it difficult to extract relevant information from discussions.
Innovation Solution
A system comprising capturing and displaying devices connected to a server that performs voice, text, handwriting, and object recognition, generating digitized documents with transcribed text and timestamps, and using augmented reality glasses to superimpose additional information, ensuring compatibility and efficient information retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video recordings are used to review meeting content, then the meeting content can be recorded and stored, but the recordings are not text-searchable and require reviewing the entire video which is time-consuming
Solution Approach 1:
The patent creates a digital copy of the meeting content in the form of a transcript that mirrors the video recording. This transcript copy is text-searchable and can be reviewed without watching the entire video, allowing users to quickly locate specific discussions while maintaining synchronization with the original video content.
Solution Approach 2:
The patent introduces an intermediary processing system that converts video audio content into text format. This intermediary layer (transcription and recognition system) bridges the gap between video recording and text searchability, enabling efficient content retrieval without sacrificing the original video review capability.
2Loss of information
If multiple capturing devices are used to capture meeting content from different angles, then comprehensive meeting content can be recorded, but the system complexity and device compatibility requirements increase
Solution Approach 1:
The patent employs universal capture devices (smartphones, tablets, laptops, cameras) that can function as meeting recording devices without requiring specialized equipment. These multi-functional devices are already familiar to users and reduce system complexity while capturing comprehensive meeting content from various angles.
Solution Approach 2:
The patent divides the meeting capture task into multiple segments handled by different capturing devices positioned at various locations. Each device captures a specific portion or angle of the meeting, and the server integrates these segmented captures into a comprehensive recording, improving content completeness without requiring a single complex device.
3Device complexity
If all connecting devices must be compatible with the same technology platform, then system integration is simplified, but device compatibility requirements limit user choice and increase costs
Solution Approach 1:
The patent introduces a server as an intermediary that handles device compatibility issues. The server receives data from various capturing devices and displays data from various displaying devices without requiring the devices themselves to be compatible with each other. This intermediary layer abstracts the compatibility requirements away from the end devices.
Solution Approach 2:
The system uses universal communication protocols and data formats that allow different types of devices (smartphones, tablets, laptops, cameras, AR glasses) to interoperate through the server. This universal approach maintains system integration simplicity while greatly expanding device compatibility and user choice.
Data Source
AI summary
A system includes a plurality of capturing devices and a plurality of displaying devices. The capturing devices and the displaying devices can be communicatively connected to a server. The server can receive captured data from the capturing devices and transform the data into a digitized format. At least one of the capturing devices can record a video during a meeting and the capturing device can transmit the captured video to the server as a video feed. The video feed can show an area that includes handwritten text, e.g., a whiteboard. The server can receive the video feed from the capturing device and perform various processes on the video. For example, the server can perform a voice recognition, text recognition, handwriting recognition, face recognition and/or object recognition technique on the captured video.


