Spoken Text Fluency Analysis for Repetition and Pause Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in accurately processing spoken language due to its non-fluent factors such as repetition, pause, and redundancy, which affect the fluency and quality of spoken text information.

Innovation Solution

A method and apparatus that determine word features and correlation features in spoken text information using neural networks, employing Hadamard products and relative position encoding to improve fluency assessment and correction, utilizing 2D and 3D feature extraction to enhance accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spoken text information is processed directly, then processing speed is maintained, but accuracy deteriorates due to non-fluent factors

Engineering Contradiction:
Improveprocessing accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the spoken text into individual words and analyzes each word's features separately. The word feature module determines word features, correlation feature module determines correlation features, and fluency effect module determines fluency effects. This segmentation allows accurate identification of non-fluent factors while maintaining manageable processing complexity through modular analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple feature dimensions including word features, correlation features, and fluency effects to analyze spoken text. By transforming the text analysis from a single-dimension approach to a multi-dimensional feature space, the system achieves higher accuracy in identifying non-fluent factors while the structured dimensionality provides clear processing pathways.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If non-fluent factors are identified and corrected, then fluency quality improves, but processing time increases

Engineering Contradiction:
Improvefluency qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis by determining word features and correlation features before final fluency assessment. This preliminary action allows the system to pre-identify potential non-fluent factors and their correlations, enabling more efficient processing in subsequent steps and reducing overall processing time while maintaining high fluency quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuous analysis through the fluency effect module that processes words sequentially while maintaining context. The continuous determination of fluency effects based on accumulated word features and correlation features enables efficient real-time processing without requiring complete text analysis before producing results, thus reducing processing time while ensuring reliability.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If word features and correlation features are determined separately, then processing accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvefeature identification accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent divides the computational task into separate modules: word feature module, correlation feature module, and fluency effect module. Each module processes specific aspects independently, which improves accuracy by allowing specialized processing while the modular architecture manages computational complexity through clear task division and potential parallel processing opportunities.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12437748B2Spoken language processing method and apparatus, and storage medium
Publication Date: 2025.10.07 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12437748B2 patent drawing
  • US12437748B2 patent drawing
  • US12437748B2 patent drawing

AI summary

Spoken language processing method and apparatus, a device, and a storage medium, which relate to the field of artificial intelligence and, in particular, to the fields of deep learning, natural-language understanding, intelligent customer service, and the like. The specific implementation solution includes: determining a word feature of a word in spoken text information; determining a correlation feature of the word in the spoken text information; and determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word.