AI-Enhanced Video Conferencing Single Stream Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems face limitations in frame rate and data processing speed, which restrict the efficiency and quality of real-time audio, video, and data transmission in multi-party sessions, particularly when combining various information channels.
Innovation Solution
A computer-implemented method using single-stream technology with an AI interface that processes and enhances audio, video, and data streams by integrating AI service programs to improve the quality and efficiency of video conferencing, allowing for real-time transmission of enriched information to all participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If single-stream technology is used to combine all audio, video and data streams into a single stream, then the system simplicity and frame rate are improved, but the system cannot provide additional processed information (such as speech transcription, facial recognition) to enhance conference quality
Solution Approach 1:
The system segments the information processing by creating two separate streams: the original first individual stream containing raw audio/video/data, and a second individual stream containing AI-processed information. This allows both the original content and enhanced information to be transmitted independently without compromising frame rate of the primary stream while adding value through the secondary stream.
Solution Approach 2:
The AI interface acts as an intermediary component that receives the first individual stream, processes it through AI service programs, and generates the second individual stream with additional information. This intermediary architecture enables information enhancement without directly interfering with the performance characteristics of the original single-stream transmission.
2Loss of information
If AI service programs are introduced to process the individual stream and add information, then the quality and efficiency of video conference are improved, but the system complexity increases
Solution Approach 1:
The AI interface is designed as a universal component that can handle multiple types of AI service programs (speech recognition, facial recognition, transcription, etc.) through a single standardized interface. This multi-functionality approach allows the system to enhance information quality with various processing capabilities while maintaining a relatively simple and unified system architecture.
Solution Approach 2:
The AI interface serves as an intermediary layer between the conference server and participants, encapsulating the complexity of AI processing within this intermediate component. This allows the rest of the system to remain relatively simple while still benefiting from advanced AI processing capabilities for improving information quality.
3Loss of information
If the conference server transmits the first individual stream to the AI interface for processing, then additional information can be generated, but the transmission delay increases
Solution Approach 1:
The system performs AI processing in parallel with stream transmission by sending the first individual stream to the AI interface while simultaneously broadcasting it to participants. The AI-processed second stream is generated and transmitted concurrently, minimizing additional delay by not sequencing the operations serially.
Solution Approach 2:
The conference server maintains continuous transmission of the first individual stream to participants while simultaneously feeding it to the AI interface for processing. This continuous parallel operation ensures that information is transmitted without interruption while AI processing occurs in the background, minimizing impact on real-time transmission quality.
Data Source
Figure 1
Figure 2
AI summary
A computer-implemented videoconferencing method for transmitting information using streaming technology is proposed. The method according to the invention comprises the following steps: receiving at least audio and video streams, and preferably also data streams, from a software-implemented conference server (10), which combines these streams into a first single stream (15); transmitting the first single stream (15) to an AI interface (20) and to an AI service program (30a, 30b, 30c); receiving information (32) from the AI interface (20), wherein this information (32) is generated by analyzing the content of the first single stream (15) by the at least one AI service program (30a, 30b, 30c); and forwarding this information (32) from the AI interface (20) to the conference server (10).Insertion of this information (33) or part of this information (33) by the conference server (10) into the first single stream (15), creating a second single stream (35), and transmission of the second single stream (35) from the conference server (10) to the participant terminals (40) of the video conference.;