Multi-Stage Language Analysis for Robocall Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication networks face challenges in accurately identifying and mitigating undesirable voice calls, such as robocalls, due to language ambiguity, which complicates the determination of whether a call is spam or not.
Innovation Solution
The implementation of a multi-stage language identification process using artificial intelligence and machine learning techniques to detect the language of voice calls, involving techniques like speech-to-text analysis and comparison of audio characteristics, to enhance the accuracy of identifying robocalls and spam calls.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multi-stage language identification process is implemented, then accuracy of identifying robocalls is improved, but device complexity increases
Solution Approach 1:
The language identification process is divided into multiple stages: first stage uses speech-to-text analysis to convert audio to text, second stage compares audio characteristics across different languages, and third stage makes the final language determination. This segmentation allows each stage to focus on specific analysis tasks, improving overall accuracy while making the complex process more manageable and implementable
Solution Approach 2:
The patent introduces intermediate processing steps between audio input and final language identification. Speech-to-text conversion serves as an intermediary that transforms audio data into text data for comparison, while audio characteristic extraction acts as another intermediary that prepares language-specific features for analysis. These intermediaries bridge the gap between raw audio and language identification algorithms
2Measurement precision
If speech-to-text analysis and audio characteristic comparison are used, then language detection accuracy is improved, but processing time increases
Solution Approach 1:
The system performs preliminary speech-to-text conversion and audio characteristic extraction during the call setup phase or while the call is being established, rather than waiting until the call is fully active. This preliminary processing allows language identification to occur before the caller and callee become fully engaged, reducing the perceived processing time while maintaining high accuracy through thorough multi-stage analysis
Data Source
AI summary
A system described herein may receive audio associated with a voice call; determine an amount of speech in the audio associated with the voice call; determine, based on the amount of speech, an expected length of a transcript, associated with a particular language, of the voice call; generate or receive the transcript, associated with the particular language, of the voice call; identify a length of the transcript, associated with the particular language, of the voice call; compare the length of the transcript to the expected length of the transcript; determine, based on comparing the length of the transcript to the expected length of the transcript, whether the voice call is associated with the particular language; and output an indication of whether the voice call is associated with the particular language.


