Medical Code Vector Embeddings for Automated Similarity Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated systems fail to effectively capture the complexity and similarity of medical information encoded using medical codes, as they lack the ability to interpret the relationships between medical codes due to the absence of grammar rules and require human knowledge for similarity determination.

Innovation Solution

Generating medical code vector embeddings using a medical embedding model trained with machine learning, which transforms medical codes into multi-dimensional vectors in a configurable space, allowing for the aggregation of vectors to determine similarities and differences between medical information sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If medical codes are used to encode medical information, then information can be standardized and processed automatically, but the system cannot capture similarity relationships between different codes without human knowledge

Engineering Contradiction:
Improveautomated processing of medical informationVSAvoidsimilarity relationships between medical codes
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent introduces vector embeddings as an intermediary representation between discrete medical codes and similarity analysis. Each medical code is transformed into a multi-dimensional vector that captures semantic relationships, allowing automated systems to compute similarities through vector operations without requiring human medical knowledge.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms one-dimensional discrete medical codes into multi-dimensional continuous vector spaces. This dimensional transformation enables the representation of similarity relationships by positioning codes with similar meanings closer together in the vector space, allowing automated detection of relationships that were invisible in the original code format.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If human medical knowledge is used to determine similarity between medical codes, then accurate similarity assessment is achieved, but the process requires manual intervention and cannot be fully automated

Engineering Contradiction:
Improvesimilarity assessment accuracyVSAvoidautomated analysis capability
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The patent enables the system to determine code similarities autonomously by training embedding models on medical data. The model learns similarity relationships from patterns in the data itself, allowing the system to self-assess similarities without requiring external human medical knowledge for each comparison.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the representation parameters of medical codes from discrete alphanumeric strings to continuous multi-dimensional vectors. This parameter transformation allows numerical computation of similarities using standard vector operations, enabling automated precision measurement while maintaining accuracy through learned representations.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If vector embeddings are generated for medical codes, then automated similarity analysis becomes possible, but computational complexity and resource requirements increase

Engineering Contradiction:
Improveautomated similarity analysisVSAvoidcomputational system complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent performs embedding generation as a preliminary step that transforms medical codes into vector representations before analysis. Once codes are converted to vectors, similarity analysis becomes a straightforward computational operation, separating the complex transformation step from the simpler analysis step.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates vector copies of medical codes that preserve semantic information in a computationally convenient format. These vector representations serve as simplified proxies for the original codes, enabling efficient automated analysis without requiring complex rule-based systems.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10891352B1Code vector embeddings for similarity metrics
Publication Date: 2021.01.12 OPTUM INC
  • US10891352B1 patent drawing
  • US10891352B1 patent drawing
  • US10891352B1 patent drawing

AI summary

Aggregate vectors corresponding to non-textual information/data are provided in a multi-dimensional space. A computing entity access a plurality of instances of medical information comprising medical codes. The computing entity generates one or more medical sentences from the plurality of instances of medical information. Each medical sentence comprises one or more medical codes. The computing entity generates an embedding vector dictionary comprising a plurality of multi-dimensional vectors based on a medical embedding model trained using machine learning and the one or more medical sentences. Each multi-dimensional vector corresponds to a medical code. The computing entity generates a plurality of aggregate vectors based on the embedding vector dictionary and analyzes at least a portion of the plurality of aggregate vectors to identify two or more aggregate vectors that are similar or different based on a distance between the two or more aggregate vectors in the multi-dimensional space.