Voice Macro Dictation System for Medical Report Structuring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in processing dictated information, particularly in medical reports, due to lack of structure, complexity in form filling, and difficulty in linking relevant data to external databases, leading to errors and inefficiencies in information retrieval and privacy concerns.

Innovation Solution

A method and system that processes dictated information into a dynamic form using voice macros and Extended Markup Language (XML) to create a report template with predefined work-type fields, allowing authors to dictate in any order and linking relevant data to an external database without parsing or coding, enhancing recognition accuracy and data management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If form filling dictation is used to improve recognition accuracy, then speech recognition accuracy improves, but device complexity increases due to template transformation requirements

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtemplate transformation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically performs template transformation and form filling without requiring manual intervention or complex configuration tools. The speech recognition engine self-adapts to the form structure and automatically fills fields based on recognized speech, eliminating the need for separate template transformation processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-loads and stores form templates in a standardized internal format, so that when speech recognition is performed, the templates are already prepared and ready for automatic filling. This preliminary preparation eliminates the need for complex real-time template transformation during the recognition process.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If plain text dictation is used to simplify the dictation process, then ease of operation improves, but information retrieval difficulty increases due to lack of structure

Engineering Contradiction:
Improvedictation simplicityVSAvoidinformation retrieval difficulty
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system segments the dictated text into structured fields corresponding to the form template. Each recognized speech element is automatically mapped to its appropriate form field, creating a structured document that maintains the simplicity of plain text dictation while enabling easy information retrieval through field-based access.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer between speech recognition and form filling. This layer automatically structures the recognized text according to the form template, serving as a mediator that converts unstructured speech output into structured form data without requiring manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Stability of the object's composition

If heavily formatted reports are used to organize information, then information structure improves, but processing complexity increases due to lack of standard structure

Engineering Contradiction:
Improveinformation structureVSAvoidprocessing complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The system uses a universal form template structure that can accommodate multiple types of reports and information formats. The standardized template design allows different kinds of information to be organized in a consistent manner, enabling easy processing and retrieval without requiring complex report-specific structures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameter of document structure from variable and report-specific to fixed and standardized through the use of templates. By defining a standard structure with specific fields and formats, the system enables consistent processing across different reports while maintaining the ability to organize diverse information types.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If manual parsing and coding tools are used to extract information, then information extraction capability improves, but productivity decreases due to time-consuming processes

Engineering Contradiction:
Improveinformation extraction capabilityVSAvoidreport creation speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system replaces manual parsing and coding operations with automated speech recognition and form filling processes. The speech recognition engine directly extracts information from spoken input and automatically populates form fields, eliminating the need for manual text parsing and coding while maintaining high information extraction capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs information extraction automatically through the speech recognition and form filling process, without requiring separate manual parsing or coding steps. The recognized speech is self-directed to the appropriate form fields based on the template structure, enabling rapid information extraction and report creation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8712772B2Method and system for processing dictated information
Publication Date: 2014.04.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8712772B2 patent drawing
  • US8712772B2 patent drawing

AI summary

A method and system for processing dictated information into a dynamic form are disclosed. The method comprises presenting an image (3) belonging to an image category to a user, dicatating a first section of speech associated with the image category, retrieving an electronic document having a previously defined document structure (4) associated with the first section of speech, this associating the document structure (4) with the image (3), wherein the document structure comprises at least one text field, presenting at least a part of the electronic document having the document structure (4) on a presenting unit (5), dictating a second section of speech and processing the second section of speech in a speech recognition engine (6) into dicatated text and associating the dictated text with the text field.