Medical Speech Recognition Using Segmented Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for recording and transferring medical records by caregivers often result in misunderstandings due to variations in recording customs, leading to potential mismanagement of patient care.
Innovation Solution
An intelligent medical speech automatic recognition method and system that utilizes a processing unit to train models from generic, medical, and textbook data to transform spoken medical records into complete sentence writing characters, eliminating the need for handwriting or typing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If caregivers record medical records using traditional methods (telephone, paper, electronic file), then the recording process is simple and accessible, but misunderstandings occur due to variations in recording customs and lack of unified format
Solution Approach 1:
The patent segments the speech recognition process into multiple specialized models: a first model for converting speech to phonetic symbols, a second model for converting phonetic symbols to text, and a third model for post-processing and formatting. This segmentation allows each model to specialize in specific medical terminology and contexts, improving overall accuracy while maintaining manageable complexity through modular architecture
Solution Approach 2:
The system performs preliminary actions by pre-training multiple specialized models on medical-specific data before actual speech recognition. The models are trained in advance on medical textbooks, medical records, and domain-specific vocabulary, so that when speech is input, the pre-prepared models can immediately and accurately process medical terminology without requiring complex real-time decision-making
2Measurement precision
If speech recognition converts speech to text directly without specialized medical training, then the system is simpler to implement, but the recognition accuracy for medical terminology is insufficient
Solution Approach 1:
The patent applies preliminary action by training multiple specialized models in advance on medical-specific datasets including medical textbooks, medical records, and domain-specific vocabulary. This pre-training ensures high recognition accuracy for medical terminology is achieved before actual use, rather than attempting to achieve it during real-time speech processing
Solution Approach 2:
The training process is segmented into multiple independent model training stages: first model trained on general speech-to-phonetic conversion, second model on phonetic-to-text conversion with medical vocabulary, and third model on post-processing and formatting. This segmentation allows parallel training of specialized components, reducing overall training time while achieving high precision through cumulative specialization
3Reliability
If multiple specialized models are trained for medical speech recognition, then the recognition accuracy for medical terminology is improved, but the system complexity and data processing requirements increase
Solution Approach 1:
The patent segments the complex recognition task into three specialized models with distinct functions: first model for speech-to-phonetic conversion, second model for phonetic-to-text conversion, and third model for post-processing. Each model is trained on specific medical data, allowing high accuracy while managing complexity through clear functional separation and modular architecture
Solution Approach 2:
The patent introduces phonetic symbols as an intermediary representation between speech input and text output. The first model converts speech to phonetic symbols, the second model converts phonetic symbols to text, and the third model performs post-processing. This intermediary layer simplifies the overall transformation process and allows each model to focus on a specific conversion task, reducing system complexity while improving reliability
Data Source
AI summary
An intelligent medical speech automatic recognition method includes performing a first model training step, a second model training step, a voice receiving step, a signal pre-treatment step and a transforming step. The first model training step is performed to train a generic statement data and a medical statement data of a database to establish a first model. The second model training step is performed to train a medical textbook data of the database to establish a second model. The voice receiving step is performed to receive a speech signal. The signal pre-treatment step is performed to receive the speech signal from the voice receiver and transform the speech signal into a to-be-recognized speech signal. The transforming step is performed to transform and recognize the to-be-recognized speech signal into a complete sentence writing character according to the first model and the second model.


