Joint Stochastic Deterministic Formatting for Voice Dictation Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice dictation often contains ambiguous components, such as number sequences, which can be interpreted in various contexts, leading to inadequate formatting in written documents, as existing technologies lack effective disambiguation methods.
Innovation Solution
A joint deterministic-stochastic configuration is employed, combining manually engineered formatting rules with a stochastic model to interpret and disambiguate ambiguous dictations, using a deterministic grammar-based formatter integrated with online learning to customize disambiguation based on user-specific behaviors and formatted examples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a deterministic grammar-based formatter is used to format voice dictation, then formatting rules are clearly defined and consistent, but ambiguous components cannot be disambiguated leading to inadequate formatting
Solution Approach 1:
The patent combines deterministic grammar-based formatting with stochastic modeling to create a hybrid system. The deterministic component provides consistent formatting rules while the stochastic component handles disambiguation of ambiguous voice dictation components, resolving the contradiction between formatting consistency and disambiguation accuracy.
Solution Approach 2:
A stochastic model acts as an intermediary between the voice dictation input and the deterministic formatter. This intermediary analyzes ambiguous components and provides disambiguated interpretations that the deterministic formatter can then apply consistent formatting rules to, solving both the need for consistency and accurate disambiguation.
2Adaptability or versatility
If manually engineered formatting rules are used, then formatting style is controlled and predictable, but adaptability to user-specific behaviors is limited
Solution Approach 1:
The system transitions from static manually engineered rules to a dynamic stochastic model that can adapt to user-specific behaviors. The stochastic model learns from user interactions and adjusts its disambiguation decisions accordingly, providing adaptability while maintaining a relatively simple system configuration through online learning.
Solution Approach 2:
The stochastic model performs online learning from user-formatted examples to automatically adapt to individual user behaviors without requiring manual reconfiguration. This self-service capability allows the system to customize disambiguation for each user while keeping the overall system configuration simple.
3Measurement precision
If context-based disambiguation is implemented, then accurate formatting is achieved, but processing time increases
Solution Approach 1:
The stochastic model is pre-trained on user-formatted examples offline, so that when actual voice dictation is processed, the disambiguation can be performed quickly using the pre-learned patterns. This preliminary action reduces the processing time during actual use while maintaining high formatting accuracy.
Solution Approach 2:
The patent replaces complex mechanical context analysis with a stochastic model that uses probabilistic patterns learned from data. This substitution allows for faster processing while maintaining accurate context-based disambiguation, as the stochastic model can evaluate multiple contexts in parallel using learned probability distributions.
Data Source
AI summary
Methods and apparatus for speech recognition on user dictated words to generate a dictation and using a discriminative statistical model derived from a deterministic formatting grammar module and user formatted documents to extract features and estimate scores from the formatting graph. The processed dictation can be output as formatted text based on a formatting selection to provide an integrated stochastic and deterministic formatting of the dictation.


