Voice Signal Processing for Live Streaming Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

During live streaming, users face inconvenience due to the need for frequent manual operation of live streaming application functions, such as microphone connection and gift-giving, which can be cumbersome and error-prone, especially in noisy environments.

Innovation Solution

A method and system that utilizes voice signals to automate the operation of live streaming application functions by recognizing a preset wake-up word, allowing users to control the terminal through voice commands without manual intervention, with a dual wake-up system for improved accuracy in noisy conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually operate live streaming application functions frequently, then the functions can be executed, but user convenience deteriorates and operation complexity increases

Engineering Contradiction:
Improveuser convenienceVSAvoidoperation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical operations (clicking buttons, selecting functions) with voice-based acoustic control. Users issue voice commands that are captured by the microphone, processed through wake-up word detection and intent recognition, and translated into application actions, eliminating the need for manual interaction with the interface.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically detecting wake-up words, recognizing user intent, and executing commands without requiring manual confirmation or intervention. The voice processing system autonomously handles function activation, parameter adjustment, and command execution based on spoken instructions.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If voice recognition is used to control live streaming functions, then user convenience improves, but recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvehands-free operationVSAvoidvoice recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the system provides visual or auditory confirmation when wake-up words are detected and when commands are successfully recognized. This allows users to verify that their voice input was correctly interpreted, and provides opportunities for correction if recognition errors occur.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary wake-up word detection before full command processing. The wake-up word recognition acts as a preliminary filter that activates the more complex intent recognition system only when necessary, reducing false triggers and improving overall recognition reliability in noisy environments.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If manual operations are performed frequently during live streaming, then various functions can be activated, but time consumption increases

Engineering Contradiction:
Improvefunction activation capabilityVSAvoidoperation time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent replaces time-consuming manual navigation and selection operations with direct voice command execution. Users can activate functions, adjust settings, and control the live streaming application through spoken instructions, significantly reducing the time required to perform repeated operations during live streaming sessions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If voice commands are used for live streaming control, then operational efficiency improves, but system complexity increases

Engineering Contradiction:
Improveoperational efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the voice processing system into distinct functional modules: wake-up word detection, intent recognition, command parsing, and execution control. This modular architecture manages system complexity by organizing functions into separate, manageable components that can be independently optimized and maintained.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11688389B2Method for processing voice signals and terminal thereof
Publication Date: 2023.06.27 BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
  • US11688389B2 patent drawing
  • US11688389B2 patent drawing
  • US11688389B2 patent drawing

AI summary

Disclosed is a method for processing a voice signal applicable to a terminal. The method can include: receiving a first voice signal; sending the first voice signal to a server in response to determining that the first voice signal includes a preset wake-up word; and receiving a second voice signal in response to receiving an acknowledgement result from the server, and responding to an interaction instruction corresponding to the second voice signal in a live streaming room.