Audio Image Analysis for Diminished Capacity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying diminished capacity in voice call participants rely on manual intervention, which is time-consuming and prone to inaccuracy, necessitating an automated solution for analyzing audio data to detect signs of diminished capacity and protect customers' assets.

Innovation Solution

A system utilizing a deep learning classification model to analyze audio data from voice calls, generating a diminished capacity score and applying a security protocol to customer accounts when the score exceeds a threshold, ensuring accurate and efficient detection and protection against ill-advised transactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis by phone agents is used to identify diminished capacity, then human judgment can be applied, but the process is time-consuming and inaccurate

Engineering Contradiction:
Improveaccuracy of diminished capacity detectionVSAvoidtime required for manual analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical system of manual human analysis with an automated audio analysis system that uses signal processing and machine learning algorithms to detect speech characteristics indicating diminished capacity. The system extracts acoustic features from audio recordings and applies trained models to automatically identify customers exhibiting signs of diminished capacity, eliminating the need for manual listening and analysis by phone agents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the audio analysis process to automatically identify and flag customers with diminished capacity without requiring human intervention. The automated system processes audio recordings, generates diminished capacity scores, and triggers appropriate workflows independently, freeing phone agents from time-consuming manual analysis while maintaining detection accuracy.

Inventive Principle:
Principle #25Self-service

2Productivity

If manual intervention is used to review audio files, then contextual understanding can be applied, but productivity is reduced due to time-consuming analysis

Engineering Contradiction:
Improvenumber of calls analyzed per monthVSAvoidaccuracy of capacity determination
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces manual review processes with automated audio analysis systems that process speech patterns, acoustic features, and linguistic characteristics to determine diminished capacity. The system uses machine learning models trained on labeled data to consistently evaluate audio recordings, enabling high-volume processing of thousands of calls per month while maintaining reliable and accurate capacity determinations through objective, data-driven analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates feedback mechanisms where the automated analysis results are reviewed and refined based on outcomes. The machine learning models are continuously improved using feedback from flagged cases and ground truth data, ensuring that productivity increases do not compromise reliability. The system provides feedback loops that allow for validation and adjustment of diminished capacity scores.

Inventive Principle:
Principle #23Feedback

3Speed

If automated audio analysis is implemented, then processing speed increases, but system complexity increases

Engineering Contradiction:
Improveanalysis processing speedVSAvoidcomplexity of analysis system
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the audio analysis system into distinct modular components: audio recording collection, signal processing and feature extraction, machine learning model analysis, score generation, and workflow triggering. Each module performs a specific function and can be independently developed, tested, and maintained. This modular architecture enables fast processing speed while managing system complexity through clear separation of concerns and standardized interfaces between components.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If deep learning models are used for analysis, then accuracy improves, but computational resources required increase

Engineering Contradiction:
Improveaccuracy of classificationVSAvoidcomputational processing power
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by performing feature extraction and preprocessing of audio data before applying deep learning models. The system extracts relevant acoustic features such as pitch, tone, speech rate, and pauses from raw audio recordings and prepares structured input data in advance. This preliminary processing reduces the computational burden on deep learning models during classification, enabling high accuracy while optimizing computational resource usage by providing pre-processed, meaningful features rather than raw audio data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12039531B2Systems and methods for identifying diminished capacity of a voice call participant based on audio data analysis
Publication Date: 2024.07.16 FMR CORP
  • US12039531B2 patent drawing
  • US12039531B2 patent drawing
  • US12039531B2 patent drawing

AI summary

Systems and methods are described herein for identifying diminished capacity of a voice call participant based on audio data analysis. A server generates an audio image based on an audio file comprising speech data corresponding to a first user on a voice call. The server determines a diminished capacity score based on one or more characteristics of the audio image analyzed using a deep learning classification model. The server applies a security protocol to an account associated with the first user when the diminished capacity score is at or above a threshold value. The server receives a transaction request corresponding to the account from the first user. The server processes the transaction request based on the security protocol applied to the account.