Voice-Based Social Network With Synchronized Speech-Text Overlay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing social networks lack the ability to seamlessly integrate voice-based interactions, where users cannot listen to posts in the user's voice, and they lack smart interactive sharing features using AI techniques.

Innovation Solution

A novel voice-based social network that uses Speech-To-Text technology to transcribe user voice posts, allowing them to be composed, explored, and shared with matched dictated text, along with optional picture or video elements, and provides advanced interfaces for recommendation, connection, and messaging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional text-based social networks are used, then users can share information easily, but users cannot listen to posts in the user's voice

Engineering Contradiction:
Improvevoice-based interactionVSAvoidvoice information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent replaces the mechanical typing system with an acoustic field system. Users speak into their devices and the speech is converted to text through speech recognition, allowing voice-based information sharing while maintaining text display functionality. This substitution enables auditory interaction without requiring physical text input.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces speech recognition technology as an intermediary between the user's voice and the social network system. This intermediary converts acoustic signals to text, enabling the system to process and display voice-based posts while maintaining compatibility with existing text-based interfaces.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If audio-based social networks are used, then users can share voice clips, but the title or description does not necessarily match with the user's voice

Engineering Contradiction:
Improveaudio sharingVSAvoidtext-audio synchronization
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent implements a feedback mechanism where the system transcribes the user's spoken words and displays the transcribed text alongside the audio clip. This feedback loop allows users to verify that the text matches their speech and enables real-time correction, ensuring synchronization between audio and text representations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary transcription of the audio content before finalizing the post. By pre-processing the audio to generate text representations, the system can present both the audio clip and its corresponding text to the user for review and correction, ensuring accuracy before publication.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If existing social networks are used, then users can share posts, but the system lacks smart interactive sharing using AI techniques

Engineering Contradiction:
Improvesharing efficiencyVSAvoidAI integration
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-service functionality where the system automatically generates captions, tags, and descriptions from the audio content using AI processing. Rather than requiring users to manually create all post elements, the system serves itself by extracting relevant information from the audio and presenting it for user approval, significantly reducing the effort required for post creation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameters of information extraction by using AI models to analyze audio characteristics, sentiment, and content automatically. This parameter transformation from raw audio to structured data (captions, tags, categories) enables smart interactive sharing while managing complexity through automated processing rather than manual configuration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12300247B2Voice-based social network
Publication Date: 2025.05.13 VERMA PRAMOD KUMAR
  • US12300247B2 patent drawing
  • US12300247B2 patent drawing
  • US12300247B2 patent drawing

AI summary

This invention presents a novel voice-based social network, where users can compose, explore, and share voice posts. Each voice post is composed of audio, text with dictation or transcription from speech, and other optional elements such as picture, video, contact, etc. During the composition step, the user speaks to the microphone, and the system generates text using the text-to-speech method. Users optionally attach a picture or video and category. Each voice post is visualized as a text on the top of the picture as an overlay. Text is highlighted with a synced part-of the speech. Users can explore posts using search interfaces using keywords and categories. Users can also comment using voice posts. This system also provides advanced interfaces such as recommendation interface where users can see related posts, connection interface where users can connect each other, message interface where users can communicate with each other via voice messages.