Webcast Server Voice Interaction Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In real-time interactive webcast systems, audience members who type slowly or are unable to input text face difficulties in expressing their opinions effectively, leading to a poor user experience and limited audience coverage.

Innovation Solution

A method and device for playing voice data in webcast systems that allows audience members to interact with the anchor through voice, where voice data from multiple electronic devices is received, sequenced, segmented, and pushed to all connected devices, enabling all users to hear and respond vocally, thereby enhancing interaction and user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text input is used for audience interaction, then communication precision is improved, but accessibility deteriorates for users who type slowly or cannot input text

Engineering Contradiction:
Improvecommunication precisionVSAvoidaccessibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent replaces the mechanical text input system with an acoustic field-based voice recognition system. Users speak their comments directly into the device, and the system converts the acoustic signals into text for display in the live room, eliminating the need for manual typing while maintaining communication precision through accurate voice-to-text conversion.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces voice recognition technology as an intermediary between the user and the communication system. The voice recognition module acts as a mediator that converts acoustic input from users into text format, enabling users who cannot type directly to still participate in the interaction without loss of communication precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If voice interaction is enabled for all users, then accessibility is improved, but system complexity increases due to voice data processing requirements

Engineering Contradiction:
ImproveaccessibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a voice recognition service provider as an intermediary that handles the complex voice data processing tasks. The recognition server acts as a mediator between the user's acoustic input and the live room display, managing the complexity of voice-to-text conversion externally rather than within the client device, thus improving accessibility without significantly increasing local system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses voice recognition technology to create a textual copy of the user's spoken words. Instead of requiring complex real-time audio processing and playback systems for every user, the system captures the acoustic signal, converts it to text through recognition services, and displays the text copy in the live room, simplifying the overall system architecture while maintaining accessibility.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11379180B2Method and device for playing voice, electronic device, and storage medium
Publication Date: 2022.07.05 BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
  • US11379180B2 patent drawing
  • US11379180B2 patent drawing
  • US11379180B2 patent drawing

AI summary

A method for playing voice, which is applied to a webcast server, the method comprising: receiving voice data sent by at least one first electronic device for obtaining a voice data set, the first electronic device having a first preset authority, and the voice data set comprising at least one piece of the voice data; receiving audio-video data sent by a second electronic device, the second electronic device having a second preset authority, the audio-video data comprising the voice data selected for playback, wherein the voice data selected for playback comprises any voice data of the voice data set clicked for playback; pushing the audio-video data to each first electronic device.