Zero-Shot NLP Model for Visual Abnormality Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Variations in language, tone, and word choices in human descriptions of mechanical defects in text data streams lead to reduced accuracy and increased false positives in natural language processing models, making it difficult to accurately diagnose issues with complex items like laptops or automobiles.

Innovation Solution

A pre-trained zero-shot learning model is used to identify and quantify distinguishable characteristics in textual data streams by generating high similarity defined characteristic phrases through n-grams analysis, which are weighted and combined to produce a relevance score, reducing content variation and improving detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If natural language processing models (such as Bag of Words, Topic Model, LSTM) are used to identify visual abnormalities from language data streams, then the system can process unstructured text data, but the accuracy is reduced and false positives increase due to variations in language, tone, and word choices

Engineering Contradiction:
Improveability to process unstructured language dataVSAvoidaccuracy in detecting visual abnormalities
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms the input parameters by converting variable language expressions into fixed visual characteristic parameters through image generation. Different word choices and language variations all map to the same visual parameters (color, shape, texture), thereby resolving the accuracy problem while maintaining versatility in processing various language inputs

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary image generation step between language processing and visual abnormality detection. The text-to-image generation model acts as a mediator that translates linguistic variations into a common visual representation, allowing accurate detection without being affected by language diversity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If traditional NLP techniques are used to analyze customer call logs or surveys, then the system can handle diverse human descriptions, but the variations in descriptions result in reduced accuracy and increased false positives

Engineering Contradiction:
Improveability to handle diverse human descriptionsVSAvoidaccuracy in problem diagnosis
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system changes the parameter space by converting linguistic parameters (word choices, tone, phrasing) into visual parameters (color, shape, texture patterns). This transformation ensures that diverse human descriptions all map to the same underlying visual characteristics, improving reliability while maintaining adaptability to handle diverse inputs

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a visual copy or representation of the described abnormality through text-to-image generation. Instead of directly analyzing variable language text, the system generates a standardized visual copy of the defect that can be reliably compared against known abnormality patterns, reducing false positives while handling diverse descriptions

Inventive Principle:
Principle #26Copying

3Measurement precision

If visual inspection is performed to accurately diagnose mechanical defects, then the accuracy of problem identification is high, but the process time and costs increase

Engineering Contradiction:
Improveaccuracy in defect identificationVSAvoiddiagnosis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical visual inspection process with an automated computational system. The text-to-image generation model combined with abnormality detection algorithms substitutes the manual visual inspection mechanism, achieving comparable accuracy while eliminating the time loss associated with human inspection processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary text-to-image generation and analysis before any physical inspection occurs. By pre-processing the language input and generating the visual representation upfront, the system enables rapid automated diagnosis that eliminates the need for time-consuming sequential visual inspection steps

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If manual visual inspection is required to confirm abnormalities, then the accuracy of diagnosis is high, but the productivity and throughput of the system decrease

Engineering Contradiction:
Improveaccuracy in abnormality detectionVSAvoidthroughput of diagnosis process
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent substitutes manual visual inspection mechanics with automated computational mechanics. The system uses text-to-image generation followed by automated abnormality detection algorithms to replace the manual inspection process, thereby maintaining high accuracy while dramatically increasing productivity and system throughput by enabling parallel processing of multiple cases simultaneously

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12061875B2Determining indications of visual abnormalities in an unstructured data stream
Publication Date: 2024.08.13 DELL PROD LP
  • US12061875B2 patent drawing
  • US12061875B2 patent drawing
  • US12061875B2 patent drawing

AI summary

A corpus of textual data records, labeled by experts as corresponding to a defined characteristic, that comprise descriptions of problems with an item are collected. A language model generates a plurality of n-grams from the corpus. Frequently occurring n-grams are analyzed using a zero-shot learning model to determine similarity to the defined characteristic. N-grams highly similar to the defined characteristic may be selected as defined phrases. N-grams highly similar to another characteristic may also be selected to reduce false positives. The zero-shot model may also be used to determine a weighting factor for each defined phrase for each record. A relevance score is determined for a record by multiplying the weighting factors for each phrase that has a similarity score relative to the record above a threshold based on the expert labeling. The relevancy score may be used to automatically diagnose problems with the item.