Cloud Service for Real-Time Audio and Video Noise Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Online communication sessions, such as conference calls, often face disruptions due to unwanted background noise and undesirable images, which can disrupt the flow of conversation and user experience.
Innovation Solution
A cloud-based service provider uses machine-learning models to identify and remove undesirable sounds and images from communication data in real-time by analyzing audio and video streams, filtering out unwanted portions before transmission to improve user satisfaction and reduce network bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If users mute their microphones to reduce background noise, then unwanted sounds are removed, but the flow of conversation is disrupted as users must become unmuted to provide input
Solution Approach 1:
The system extracts and removes only the unwanted background noise portions from the audio stream while preserving the speaker's voice. The machine learning model identifies and separates harmful sounds (dog barking, doorbells, etc.) from useful speech, allowing continuous audio transmission without requiring users to mute their microphones.
Solution Approach 2:
A cloud-based service acts as an intermediary between the user's microphone and the communication session. This intermediary receives audio data, processes it through machine learning models to remove unwanted sounds, and returns cleaned audio to the session, eliminating the need for users to manually mute or unmute their microphones.
2Reliability
If the system removes unwanted sounds from communication data, then user satisfaction is improved, but network bandwidth requirements increase due to processing and retransmission of cleaned data
Solution Approach 1:
The audio cleaning process is performed preliminarily at the cloud service before data reaches the communication session. By pre-processing and removing unwanted sounds in advance, the system reduces the amount of data that needs to be transmitted and processed again at the receiving end, optimizing bandwidth usage.
Solution Approach 2:
The system changes the parameter of audio data quality by transforming raw audio into cleaned audio through machine learning processing. This parameter transformation allows the same network infrastructure to deliver higher quality audio without requiring proportional increases in bandwidth, as the cleaning process efficiently targets only harmful portions.
Data Source
AI summary
This disclosure describes techniques implemented partly by a communications service for identifying and altering undesirable portions of communication data, such as audio data and video data, from a communication session between computing devices. For example, the communications service may monitor the communications session to alter or remove undesirable audio data, such as a dog barking, a doorbell ringing, etc., and/or video data, such as rude gestures, inappropriate facial expressions, etc. The communications service may stream the communication data for the communication session partly through managed servers and analyze the communication data to detect undesirable portions. The communications service may alter or remove the portions of communication data received from a first user device, such as by filtering, refraining from transmitting, or modifying the undesirable portions. The communications service may send the modified communication data to a second user device engaged in the communication session after removing the undesirable portions.


