Speech Sequence Modeling for Physiological State Indicators

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for evaluating physiological and psychological states from speech utterances face challenges in accurately converting sequences of speech records into meaningful indicators due to the abstract nature of embedding vectors and the lack of direct association with measurable physical or physiological quantities, especially in medical contexts where ground truth is often uncertain and based on cursory physician examinations.

Innovation Solution

A system utilizing machine learning models, including neural networks and Transformer models, to convert sequences of speech records into indicators of physiological, psychological, and emotional states by embedding vectors, aggregating session data, and providing confidence levels, which are trained on large datasets to adapt to individual voices and provide continuous updates on patient conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If speech records are converted into embedding vectors using sequence-to-sequence conversion, then the speech data can be processed and stored in a compact format, but the embedding vectors become abstract and lose direct association with measurable physiological quantities

Engineering Contradiction:
Improvedata processing complexityVSAvoidphysiological state measurement accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer between the abstract embedding vectors and the target physiological states. This intermediary is implemented through a dual-encoder architecture where the first encoder creates initial embeddings and the second encoder refines them into clinically meaningful representations. The intermediary bridge connects these representations to measurable physiological quantities through trained mappings that preserve both the compactness of embeddings and the interpretability of physiological indicators.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the speech processing pipeline into distinct functional components: a first encoder for initial feature extraction, a second encoder for refined processing, and separate evaluation modules for different physiological states. This segmentation allows each component to specialize in specific transformations, maintaining the benefits of embedding compression while enabling precise mapping to measurable physiological parameters through independent optimization of each segment.

Inventive Principle:
Principle #1Segmentation

2Productivity

If conventional sequence-to-sequence conversion is used for speech evaluation, then the system can process speech data efficiently, but it fails to accurately capture individual voice characteristics and provides unreliable diagnostic indicators

Engineering Contradiction:
Improvespeech processing efficiencyVSAvoiddiagnostic indicator reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by training separate encoder models tailored to specific diagnostic tasks and voice characteristics. Each encoder is optimized for particular physiological states or voice types, allowing the system to maintain high processing efficiency while adapting to individual voice characteristics. The local quality of each encoder's output is optimized for its specific diagnostic purpose, improving overall reliability without sacrificing throughput.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent utilizes parameter changes by dynamically adjusting encoder configurations and transformation parameters based on the specific speech sample and diagnostic context. The system adapts embedding dimensions, transformation matrices, and evaluation parameters in real-time to match individual voice characteristics and clinical requirements, thereby improving reliability while maintaining efficient processing through automated parameter optimization.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If embedding vectors are used to represent speech data, then the system can reduce data dimensionality and improve processing speed, but the abstract nature of embeddings makes them difficult to interpret and validate against ground truth

Engineering Contradiction:
Improvedata dimensionalityVSAvoidphysiological state information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms that continuously validate embedding representations against ground truth physiological data. The system uses feedback loops to compare predicted physiological states with actual measurements, adjusting the embedding transformations in real-time. This feedback ensures that dimensionality reduction does not lose critical physiological information and that the abstract embeddings remain aligned with measurable quantities through iterative refinement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies another dimension by introducing additional transformation layers that map embedding vectors into clinically interpretable spaces. The dual-encoder architecture creates multiple dimensional representations: the first embedding layer for compact storage, the second embedding layer for clinical interpretation, and an additional validation dimension for physiological correlation. This multi-dimensional approach preserves information while reducing complexity through structured dimensional transformation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12555595B2Converting a sequence of speech records of a human subject into a sequence of indicators of a physiological state of the subject
Publication Date: 2026.02.17 CORDIO MEDICAL LTD
  • US12555595B2 patent drawing
  • US12555595B2 patent drawing
  • US12555595B2 patent drawing

AI summary

A system includes a memory and a processor. The memory is configured to store a machine learning (ML) model trained using a plurality of sequences of speech records of humans having each at least one of a respective known physiological state, psychological state and emotional state. The processor is configured to (i) receive a sequence of speech records of a human subject, (ii) apply the trained ML model to infer from the sequence of speech records of the human subject a sequence of one or more indicators indicative of at least one: a physiological state, a psychological state, and an emotional state of the human subject, and (iii) make the indicators available.