Voice Data Compensation Using Machine Learning Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Internet-based communications often experience voice dropouts and degradations due to network congestion, leading to unclear speech, especially when speakers have poor diction, heavy accents, or speak unfamiliar languages, making it difficult for participants to understand each other.
Innovation Solution
A method and system that detect voice data loss or degradation in real-time, using a server or client device to compensate by either predicting and restoring missing phonemes or words from historical data or inserting noise, and further improving intelligibility by replacing unintelligible portions with substitute voice data based on idealized models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice data is transmitted over network, then communication connectivity is achieved, but voice dropouts and degradations occur due to network congestion
Solution Approach 1:
The system performs preliminary actions by detecting voice data loss or degradation in real-time and immediately compensating for it. The compensation process includes detecting the loss, determining prediction probability, and restoring missing phonemes or inserting noise before the degraded voice data is fully processed, thereby maintaining communication reliability while preserving connectivity.
2Measurement precision
If voice data compensation is performed using historical voice data, then speech intelligibility is improved, but system complexity increases due to prediction probability calculation
Solution Approach 1:
The system changes parameters by calculating prediction probability as a quantitative metric to determine whether to use historical voice data for compensation. This parameter-based decision-making approach balances speech intelligibility improvement with controlled system complexity, as the prediction probability threshold provides a clear criterion for when to apply the more complex historical data compensation method.
3Reliability
If noise is inserted to compensate for voice data loss, then compensation is performed, but speech clarity may deteriorate
Solution Approach 1:
The system uses prediction probability as a parameter to decide between two compensation methods: inserting noise or using historical voice data. By setting a prediction probability threshold, the system ensures that noise insertion (which may deteriorate clarity) is only used when prediction confidence is low, thereby maintaining speech clarity while still providing reliable compensation for voice data loss.
Data Source
AI summary
A method comprises: obtaining, at an apparatus, first voice data from a first user device associated with a first speaker participant in a communication session; detecting voice data loss or degradation in the first voice data; determining whether prediction probability of correctly compensating for the voice data loss or degradation is greater than a predetermined probability threshold; if the prediction probability is greater than the predetermined probability threshold, first compensating for the voice data loss or degradation using historical voice data received by the apparatus prior to receiving of the first voice data, the first compensating producing first compensated voice data; if the prediction probability is not greater than the predetermined probability threshold, second compensating for the voice data loss or degradation by inserting noise to the first voice data to produce second compensated voice data; and outputting the first compensated voice data or the second compensated voice data.


