LLM Evaluation via Embedding Space Quality Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The performance of generative large language models (LLMs) is not fully evaluated, leading to nonsensical or inaccurate responses due to issues with training data, seed data, and inability to accurately make linkages, with users lacking quantitative measures of response reliability.

Innovation Solution

A system and method for evaluating generative LLMs by obtaining prompts from a selected subject matter domain, supplying them to the LLM via an API, receiving responses, transforming them into an embedding space, determining quality measurements using domain-based metrics, and generating a performance evaluation of the LLM based on these measurements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If generative LLMs are used to provide human-like responses, then user engagement and conversational capability are improved, but response reliability and accuracy deteriorate due to nonsensical or fictitious content generation

Engineering Contradiction:
Improveconversational capabilityVSAvoidresponse accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism by transforming LLM responses into embedding vectors and comparing them against reference embeddings from trusted sources. This feedback loop quantifies response quality and enables continuous evaluation of LLM reliability, allowing users to assess the trustworthiness of generated content.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary evaluation system that acts as a mediator between the LLM and the user. This system uses embedding transformations and domain-based metrics to assess response quality before presenting it to the user, providing an additional layer of reliability verification without disrupting the conversational flow.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If LLMs are trained on vast amounts of unstructured web data, then model generativity and language coverage are improved, but data quality control and training reliability worsen due to scraping from unreliable sources

Engineering Contradiction:
Improvelanguage coverageVSAvoidtraining data quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent replaces traditional mechanical quality control methods with an embedding-based evaluation system. Instead of manually verifying training data quality, the system uses vector transformations and domain metrics to automatically assess the reliability of generated responses, enabling scalable quality control for vast amounts of unstructured data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the evaluation parameters from raw text comparison to embedding space analysis. By transforming both training data and generated responses into vector representations, the system can quantify quality differences and assess training data reliability through mathematical metrics rather than subjective textual analysis.

Inventive Principle:
Principle #35Parameter changes

3Object-generated harmful factors

If LLM architecture and parameters are kept opaque for closed solution protection, then proprietary advantages are maintained, but user understanding and trust in response reliability deteriorate

Engineering Contradiction:
Improveproprietary protectionVSAvoidtransparency information
Core Design Contradiction:
Object-generated harmful factorsVSLoss of information

Solution Approach 1:

The patent introduces an intermediary evaluation layer that does not require exposing the LLM's internal architecture or parameters. This mediator system assesses response quality through embedding transformations and domain metrics, providing transparency about response reliability without revealing proprietary model details.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the need for architectural transparency with a behavioral evaluation system. Instead of requiring users to understand the model's internal mechanics, the system evaluates and reports on response quality through embedding-based metrics, substituting structural opacity with functional transparency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Device complexity

If traditional evaluation methods are used for LLMs, then evaluation simplicity is maintained, but measurement precision of response quality deteriorates due to lack of quantitative metrics

Engineering Contradiction:
Improveevaluation complexityVSAvoidresponse quality measurement
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the measurement parameters from subjective textual assessment to objective embedding-based metrics. By transforming responses into vector spaces and applying domain-specific metrics, the system achieves precise quantitative measurement of response quality while maintaining relatively simple evaluation procedures.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent substitutes traditional manual evaluation methods with an automated embedding-based measurement system. This replacement enables precise quantitative assessment of response quality through mathematical operations on vector representations, eliminating the need for complex human judgment while improving measurement accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250111237A1System and method for evaluating generative large language models
Publication Date: 2025.04.03 AT&T INTELLECTUAL PROPERTY I L P
  • US20250111237A1 patent drawing
  • US20250111237A1 patent drawing
  • US20250111237A1 patent drawing

AI summary

Aspects of the subject disclosure may include, for example, a device that facilitates obtaining a plurality of prompts from a selected subject matter domain of a database configured to measure an effectiveness of a generative large language model (LLM) to distinguish variances between each prompt of the plurality of prompts; supplying the plurality of prompts to the LLM; receiving respective responses to each of the prompts from the LLM; transforming each of the prompts and respective responses to each of the prompts into an embedding space; determining, by applying domain-based metrics to the embedding space, a quality measurement of each respective response to produce a plurality of quality measurements; and generating, according to the plurality of quality measurements, a performance of the LLM. Other embodiments are disclosed.