Spoken Text Fluency Analysis for Repetition and Pause Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in accurately processing spoken language due to its non-fluent factors such as repetition, pause, and redundancy, which affect the fluency and quality of spoken text information.
Innovation Solution
A method and apparatus that determine word features and correlation features in spoken text information using neural networks, employing Hadamard products and relative position encoding to improve fluency assessment and correction, utilizing 2D and 3D feature extraction to enhance accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spoken text information is processed directly, then processing speed is maintained, but accuracy deteriorates due to non-fluent factors
Solution Approach 1:
The patent segments the spoken text into individual words and analyzes each word's features separately. The word feature module determines word features, correlation feature module determines correlation features, and fluency effect module determines fluency effects. This segmentation allows accurate identification of non-fluent factors while maintaining manageable processing complexity through modular analysis.
Solution Approach 2:
The patent introduces multiple feature dimensions including word features, correlation features, and fluency effects to analyze spoken text. By transforming the text analysis from a single-dimension approach to a multi-dimensional feature space, the system achieves higher accuracy in identifying non-fluent factors while the structured dimensionality provides clear processing pathways.
2Reliability
If non-fluent factors are identified and corrected, then fluency quality improves, but processing time increases
Solution Approach 1:
The patent performs preliminary analysis by determining word features and correlation features before final fluency assessment. This preliminary action allows the system to pre-identify potential non-fluent factors and their correlations, enabling more efficient processing in subsequent steps and reducing overall processing time while maintaining high fluency quality.
Solution Approach 2:
The patent maintains continuous analysis through the fluency effect module that processes words sequentially while maintaining context. The continuous determination of fluency effects based on accumulated word features and correlation features enables efficient real-time processing without requiring complete text analysis before producing results, thus reducing processing time while ensuring reliability.
3Measurement precision
If word features and correlation features are determined separately, then processing accuracy improves, but computational complexity increases
Solution Approach 1:
The patent divides the computational task into separate modules: word feature module, correlation feature module, and fluency effect module. Each module processes specific aspects independently, which improves accuracy by allowing specialized processing while the modular architecture manages computational complexity through clear task division and potential parallel processing opportunities.
Data Source
AI summary
Spoken language processing method and apparatus, a device, and a storage medium, which relate to the field of artificial intelligence and, in particular, to the fields of deep learning, natural-language understanding, intelligent customer service, and the like. The specific implementation solution includes: determining a word feature of a word in spoken text information; determining a correlation feature of the word in the spoken text information; and determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word.


