LLM Defect Prediction via Text Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI and machine learning techniques in semiconductor manufacturing are limited in their ability to predict defects due to reliance on traditional methods that cannot handle complex systems and require vast amounts of balanced data, while also being inadequate for privacy-preserving training with numerical sensor data.

Innovation Solution

A mechanism that encodes numerical sensor data into a non-numerical format, such as a grammar-based sequence of characters, to facilitate the training and inferencing of a large language model for process control in semiconductor manufacturing, allowing for privacy-preserving and effective defect prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional neural networks or statistical algorithms are used for defect prediction, then the model can process numerical sensor data, but the system requires vast amounts of balanced data and manual feature identification, limiting scalability and adaptability

Engineering Contradiction:
Improvedefect prediction capabilityVSAvoidscalability across different equipment and processes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary encoding layer that converts numerical sensor data into text representations. This intermediary format serves as a bridge between the raw numerical data and the large language model, enabling the model to process manufacturing data without requiring direct numerical input. The encoding mechanism translates process parameters into descriptive text that preserves critical information while making it compatible with LLM architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent fundamentally changes the data representation parameter from numerical format to text format. By transforming sensor data through encoding into textual descriptions, the system leverages the LLM's strength in processing natural language. This parameter change enables the model to handle heterogeneous data from different equipment and processes uniformly, improving scalability without requiring retraining for each new equipment type.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If generative large language models are used for defect prediction, then the model can handle complex systems and reduce manual setup, but the model struggles with numerical data processing and out-of-vocabulary numerals

Engineering Contradiction:
Improveautomatic process controlVSAvoidnumerical data accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The encoding mechanism acts as an intermediary that preserves numerical precision while making data compatible with LLM processing. Instead of feeding raw numbers directly to the model, the system encodes them into text representations that maintain the semantic meaning and precision of the original numerical values. This intermediary layer resolves the conflict between LLM's text-processing strength and numerical data requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical numerical processing approach with a linguistic text-based approach. Rather than relying on the model's numerical computation capabilities, the system substitutes numerical data with its textual representation, leveraging the LLM's sophisticated language understanding to handle both the data values and their contextual meanings simultaneously.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If proprietary sensor data is used for training the model, then the model can be customized for specific manufacturing processes, but the data privacy and security are compromised

Engineering Contradiction:
Improvecustomization for specific processesVSAvoiddata exposure and privacy loss
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential information needed for training from the proprietary sensor data through the encoding process. By transforming data into text representations that capture only the necessary patterns and relationships, the system separates the useful training information from the sensitive proprietary details. This extraction allows model customization while minimizing data exposure.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The encoding mechanism creates a textual copy or representation of the numerical sensor data that preserves the essential information patterns without exposing the raw proprietary values. This copy can be used for training and analysis while the original sensitive data remains protected, enabling process customization without direct data exposure.

Inventive Principle:
Principle #26Copying

4Reliability

If standard neural networks are used for quality prediction, then the model can be trained on existing data, but every use case requires separate feature identification and model selection, impacting scalability

Engineering Contradiction:
Improvequality prediction accuracyVSAvoidmodel setup and configuration
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal LLM-based framework that can handle multiple quality prediction use cases across different manufacturing processes and equipment types. The single model architecture, once trained on encoded data from various sources, serves multiple functions and can be applied to different scenarios without requiring separate model development for each use case, significantly reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The LLM framework provides self-service capability by automatically adapting to different manufacturing processes through its pre-trained language understanding and the encoding mechanism. Rather than requiring manual feature identification and model reconfiguration for each new process, the system self-adjusts by processing encoded data representations, eliminating the need for extensive manual setup and configuration.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12271825B2Predictive system for semiconductor manufacturing using generative large language models
Publication Date: 2025.04.08 LYNCEUS SAS
  • US12271825B2 patent drawing
  • US12271825B2 patent drawing
  • US12271825B2 patent drawing

AI summary

A method for process control in association with a production system. The method leverages a large language model (LLM) that has been trained or fine-tuned on production data in a manner that avoids or minimizes use of numerical sensor data. In particular, during training, historical sensor data is received. In lieu of using the historical sensor data to train the model directly, the data is first encoded into a grammar-based sequence of characters before it is applied to train or fine-tune the model.