AI Platform for Real-Time VoIP Audiovisual Frame Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current VoIP and AVoIP systems lack integration with artificial intelligence, limiting their ability to provide real-time processing and analysis of audiovisual content during calls, such as automated speech recognition, trigger detection, and scenario identification, which hampers enhanced functionality and data-driven decision-making.

Innovation Solution

A fully integrated AI platform is incorporated with VoIP and AVoIP infrastructure, processing audiovisual content in real-time by applying AI to each frame of the media stream, enabling features like automated speech recognition, trigger detection, and scenario identification, and providing real-time recommendations and analytics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If AI processing is applied to each frame of audiovisual content in real-time, then enhanced functionality and data-driven decision-making are achieved, but system complexity and processing requirements increase

Engineering Contradiction:
Improveenhanced functionalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the AI processing into discrete frame-level operations, where each audiovisual frame is independently processed by specific AI algorithms (speech recognition, trigger detection, scenario identification). This segmentation allows the complex AI platform to handle enhanced functionality through modular, manageable units rather than monolithic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The AI platform is designed as a universal system that performs multiple functions simultaneously - speech-to-text conversion, trigger word detection, sentiment analysis, and scenario identification - all within a single integrated architecture. This multi-functionality resolves the contradiction by consolidating diverse AI capabilities into one system rather than requiring separate systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If real-time AI processing is applied to audiovisual streams, then automated speech recognition and trigger detection are enabled, but processing time and computational resources increase

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-configuring AI models and algorithms before real-time processing begins. Speech recognition models, trigger detection rules, and scenario identification parameters are prepared in advance, allowing the system to process incoming audiovisual frames immediately without delay for model initialization or configuration during the call.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The AI processing operates continuously throughout the audiovisual stream without interruption or batch processing delays. Each frame is processed as it arrives, maintaining continuous speech recognition, trigger detection, and scenario identification throughout the entire call duration, ensuring no processing time is lost to间断性 operations.

Inventive Principle:
Principle #20Continuity of useful action

3Loss of information

If AI processing is applied to each frame prior to transfer, then real-time transcription and analytics are provided, but network bandwidth and transmission load increase

Engineering Contradiction:
Improveinformation completenessVSAvoiddata transmission volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the essential processed information from each audiovisual frame - such as transcribed text, detected triggers, and identified scenarios - rather than transmitting the entire processed frame data. This extraction approach provides complete information about the call content while minimizing the volume of data that needs to be transmitted across the network.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The AI processing system acts as an intermediary between the audiovisual stream and the network transmission. It processes frames locally, extracts key information, and transmits only the extracted data rather than the full audiovisual content, thereby reducing network load while maintaining information completeness for transcription and analytics purposes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11050807B1Fully integrated voice over internet protocol (VoIP), audiovisual over internet protocol (AVoIP), and artificial intelligence (AI) platform
Publication Date: 2021.06.29 DIALPAD INC
  • US11050807B1 patent drawing
  • US11050807B1 patent drawing
  • US11050807B1 patent drawing

AI summary

An AI platform is fully integrated with existing VoIP/AVoIP telephony infrastructure. In the course of providing VoIP/AVoIP audiovisual calls, a VoIP/AVoIP media stream of audiovisual content is processed, and transferred between endpoints. AI processing is applied to each frame of the transferred audiovisual content, in real-time while the audiovisual call is occurring. For example, automated speech recognition can be performed on the content, in which the speech of the audiovisual content is converted to text. The audiovisual call can further be automatically transcribed to a text file in real-time. Another example is the automatic detection of the occurrence of specific triggers during calls. Additional enhanced functionality is automatically provided as a result of applying the AI processing to the transferred audiovisual content. For example, in response to detecting the occurrence of a specific trigger, a corresponding directive can be automatically output on a screen of a calling device.