AI Speech and Video Processing for Clinical Documentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Medical professionals spend significant time manually documenting and processing information from medical procedures, reducing their availability for patient care.

Innovation Solution

An AI platform that integrates image and speech data processing using natural language processing and image classification to automate the extraction and structuring of clinical information, including quality-of-care indicators, for use in generating reports and populating electronic medical records.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual documentation methods are used, then information can be recorded, but medical professionals spend excessive time on documentation tasks

Engineering Contradiction:
Improvedocumentation efficiencyVSAvoidtime available for patient care
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical documentation processes with an automated AI system that uses natural language processing, image classification, and pattern recognition to extract and structure clinical information from speech and video data, eliminating the need for manual dictation and report generation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service documentation by automatically processing raw clinical data from microphones and cameras, extracting relevant information, and generating structured reports without requiring physician intervention for routine documentation tasks

Inventive Principle:
Principle #25Self-service

2Productivity

If automated AI processing is implemented, then documentation efficiency improves, but system complexity increases

Engineering Contradiction:
Improveinformation processing speedVSAvoidAI platform complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The AI platform is divided into distinct functional modules: speech-to-text conversion, natural language processing, image classification, pattern recognition, and report generation. Each module handles a specific aspect of data processing, making the overall complex system manageable through functional segmentation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an AI processing layer as an intermediary between raw clinical data collection and final report generation. This intermediary layer handles the complex processing tasks using trained models and algorithms, shielding clinicians from the underlying system complexity while delivering simplified structured outputs

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12412647B2AI platform for processing speech and video information collected during a medical procedure
Publication Date: 2025.09.09 UTECH PRODUCTS INC
  • US12412647B2 patent drawing
  • US12412647B2 patent drawing
  • US12412647B2 patent drawing

AI summary

An AI based platform for processing information collected during a medical procedure. A method includes capturing images and speech during a medical procedure; processing the images using a trained classifier to identify image-based quality-of-care indicators (QIs); converting the speech into text; parsing the text into sentences; performing a search and replace on predefined text patterns in the sentences; identifying text-based QIs in the sentences; classifying sentences into sentence types using a trained model; updating sentences by integrating the image-based QIs with text-based QIs; and outputting structured data that includes sentences organized by sentence type.