Online Profile Verification via Speech-Video Text Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for authenticating users in online communication are inadequate in verifying the authenticity of images associated with online profiles, leading to potential identity theft and fraud, as they often rely on limited biometric verification or device-specific data without comprehensive validation of the user's identity.

Innovation Solution

A method involving the receipt of a picture with a string of text, where the individual records a video pronouncing the text, which is analyzed for audio and visual matching using automated and human verification, with geolocation and crowd-sourced voting to confirm the authenticity of the image, and the use of a device with processors to determine and certify the identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated mechanisms are used to analyze video and determine audio-text matching, then verification efficiency is improved, but verification accuracy may deteriorate due to limitations in automated speech recognition

Engineering Contradiction:
Improveverification efficiencyVSAvoidverification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent combines automated speech recognition analysis with human reviewer verification in a hybrid system. The automated mechanisms perform initial screening to improve efficiency, while human reviewers provide final accuracy validation, merging both approaches to resolve the contradiction between speed and precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary human review layer between automated analysis and final verification certification. This intermediary step allows automated systems to handle routine cases efficiently while enabling human judgment to correct automated errors, thus balancing productivity and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple verification steps including crowd-sourced voting are implemented, then verification reliability is improved, but system complexity increases

Engineering Contradiction:
Improveverification reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the verification process into distinct modular components: automated video analysis, speech recognition, human reviewer validation, and crowd-sourced voting. Each module operates independently with defined inputs and outputs, making the complex system manageable and maintainable while achieving high reliability through multiple validation layers.

Inventive Principle:
Principle #1Segmentation

3Object-affected harmful factors

If geolocation data is collected and verified to confirm user location, then fraud detection capability is improved, but user privacy concerns increase

Engineering Contradiction:
Improvefraud detection capabilityVSAvoiduser privacy
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies geolocation verification selectively rather than universally. Geolocation data collection is triggered only in specific contexts where fraud risk is elevated or where location verification is necessary for the service functionality, rather than collecting all user data by default, thus balancing fraud detection with privacy preservation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9721079B2Image authenticity verification using speech
Publication Date: 2017.08.01 CHEN STEVE Y
  • US9721079B2 patent drawing
  • US9721079B2 patent drawing
  • US9721079B2 patent drawing

AI summary

Verifying the identity of a person claiming to be represented by a picture by way of providing a string of text (randomly generated or generated by another person seeking verification of same) to be recited by the claimant. The string of text is recited in a video which is received by an intermediary server at a network node, or by a person seeking such verification. Automated processes may be utilized to compare the audio and video received to the picture and string of text sent. Further, comparisons to previously received audio, video, and strings of text, as well as the same available from third parties, may be used to determine fraud attempts. Viewers of the person's profile may also vote on the authenticity of a profile, thereby raising or lowering a certification confidence level, with their votes weighted more heavily towards those who have high confidence levels.