Voice Alignment via Loss and Discontinuity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice alignment methods for abnormal voice recognition in communications networks suffer from significant errors, requiring multiple algorithms and processes to achieve accuracy, which is inefficient and often inaccurate due to limitations in user privacy protection policies.

Innovation Solution

A voice alignment method that detects voice loss and discontinuity in test voices before alignment, selecting an appropriate algorithm based on detection results to align the test voice with the original voice, using techniques such as inserting or deleting silent statements to synchronize start and end time domain locations, thereby improving alignment efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing voice alignment methods are used to align original voice and test voice, then voice alignment can be performed, but the alignment accuracy is low and requires multiple algorithms and processing times

Engineering Contradiction:
Improvevoice alignment accuracyVSAvoidnumber of algorithms and processing steps
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing detection of voice loss and discontinuity before the alignment operation. The detection unit identifies whether the test voice has voice loss or discontinuity issues, and this detection result guides the selection of appropriate alignment algorithms. This preliminary detection step prevents盲目 application of multiple complex algorithms, thereby improving alignment accuracy while reducing unnecessary processing complexity.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple algorithms and processing times are used to overcome alignment errors, then alignment accuracy may improve, but processing efficiency decreases

Engineering Contradiction:
Improvevoice alignment accuracyVSAvoidvoice alignment efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements dynamics by dynamically selecting alignment algorithms based on the detection results. Instead of statically applying multiple algorithms in fixed sequences, the system adaptively chooses the most appropriate algorithm according to the specific voice conditions (whether voice loss or discontinuity is detected). This dynamic approach maintains high alignment accuracy while significantly improving processing efficiency by avoiding unnecessary algorithm applications.

Inventive Principle:
Principle #15Dynamics

3Reliability

If user privacy protection policies are strictly followed, then user privacy is protected, but the ability to recognize abnormal voices is restricted

Engineering Contradiction:
Improveuser privacy protectionVSAvoidabnormal voice recognition capability
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies segmentation by dividing the voice analysis into separate functional components: a detection unit that identifies voice loss and discontinuity, and an alignment unit that performs alignment based on detection results. This segmentation allows the system to operate within privacy constraints by processing voices locally without requiring access to sensitive user data, while still maintaining effective abnormal voice recognition capability through the specialized detection and alignment mechanisms.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3764361B1Method and apparatus for aligning voices
Publication Date: 2023.08.30 HUAWEI TECH CO LTD
  • EP3764361B1 patent drawingFigure 1~2
  • EP3764361B1 patent drawingFigure 3~5
  • EP3764361B1 patent drawingFigure 6~7

AI summary

This application relates to artificial intelligence, and provides a voice alignment method, including: obtaining an original voice and a test voice, where the test voice is a voice generated after the original voice is transmitted over a communications network; performing loss detection and/or discontinuity detection on the test voice, where the loss detection is used to determine whether the test voice has a voice loss compared with the original voice, and the discontinuity detection is used to determine whether the test voice has voice discontinuity compared with the original voice; and aligning the test voice with the original voice based on a result of the loss detection and/or the discontinuity detection, to obtain an aligned original voice and an aligned test voice, where the result of the loss detection and/or the discontinuity detection is used to indicate a manner of aligning the test voice with the original voice. According to the voice alignment method provided in this application, the voice alignment method is determined based on the result of the loss detection and/or the discontinuity detection, and voice alignment may be performed based on a specific status of the test voice by using a most suitable method, thereby improving efficiency of the voice alignment.