WFST-Based Spoken Text to Written Text Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Spoken language, due to its informality and lack of standardization, poses challenges in machine translation and text conversion, leading to inaccurate translations and hindered text exchange.
Innovation Solution
A text conversion method and device utilizing a weighted finite-state transducer (WFST) model database to differentiate between non-spoken and spoken morphemes, specifically identifying and removing spoken morphemes with characteristics like inserted, repeated, and amending morphemes to convert spoken text into written text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition technology is used to convert voice to text, then voice input is converted into text, but the recognition result retains spoken language characteristics making it unsuitable for formal occasions
Solution Approach 1:
The patent introduces a spoken text processing module as an intermediary between speech recognition and machine translation. This module contains a spoken text database and a spoken text processing unit that identifies and removes spoken language characteristics (such as filler words, repetitions, and colloquial expressions) from the recognized text, transforming it into standardized written text suitable for formal translation tasks
2Productivity
If spoken text is directly used for machine translation, then translation can be performed, but translation accuracy deteriorates due to non-standardized spoken language features
Solution Approach 1:
The patent applies preliminary text processing before machine translation by pre-identifying and removing spoken language characteristics from the recognized text. The spoken text processing unit prepares the text in advance by eliminating filler words, correcting repetitions, and standardizing expressions, so that the subsequent machine translation operates on cleaned, standardized text, improving translation accuracy without sacrificing productivity
Data Source
AI summary
The method includes acquiring a target spoken text, where the target spoken text includes a non-spoken morpheme and a spoken morpheme; determining, from a target weighted finite-state transducer (WFST) model database, a target WFST model corresponding to the target spoken text, where output of a state that is corresponding to the spoken morpheme and that is in the target WFST model is empty, and output and input of a state that is corresponding to the non-spoken morpheme and that is in the target WFST model are the same; and determining, according to the target WFST model, a written text corresponding to the target spoken text, where the written text includes the non-spoken morpheme and does not include the spoken morpheme.


