Product Visual Inspection with Foundation Model Auto-Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI-based visual inspection systems face challenges due to the time-consuming and costly process of creating labeled datasets, deviations from training data conditions, resource intensity, scalability issues, and difficulty in adapting to changing requirements, leading to inaccuracies and bottlenecks in industrial product inspection.

Innovation Solution

A computer-implemented method using a foundation model based on artificial intelligence, combined with a federated learning approach, where the inspection model is generated centrally and deployed to clients, utilizing auto-labeled data and process-specific information to enable efficient and accurate visual inspection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised learning with labeled data is used for AI-based visual inspection, then inspection accuracy can be achieved, but the time and cost for creating labeled datasets increases significantly

Engineering Contradiction:
Improveinspection accuracyVSAvoidtime for data labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service through automated label generation using foundation models. The AI system generates its own training labels without human intervention, allowing the inspection model to train on automatically generated labeled data, thus eliminating the time-consuming manual labeling process while maintaining inspection accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The foundation model performs preliminary action by pre-generating labeled training data before the actual inspection task. This preliminary labeling action creates a ready-to-use training dataset that eliminates the need for time-consuming manual annotation during deployment, enabling faster model training and deployment

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual labeling is performed for large datasets, then training data quality improves, but resource requirements and system complexity increase

Engineering Contradiction:
Improvetraining data qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system replaces the mechanical process of manual human labeling with an automated AI-based foundation model. This substitution eliminates the need for human annotators and complex coordination systems, reducing system complexity while maintaining or improving training data quality through consistent automated labeling

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The foundation model acts as an intermediary between raw image data and the inspection model training process. It generates intermediate labeled data that serves as high-quality training material, simplifying the overall system architecture by removing the need for manual labeling infrastructure and processes

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If Vision Transformer models are used for image processing, then scalability to larger datasets is achieved, but computational complexity increases

Engineering Contradiction:
ImprovescalabilityVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system changes key parameters of the Vision Transformer implementation by using foundation models with optimized architectures and pre-trained weights. This allows the model to scale to larger datasets while managing computational complexity through parameter optimization, transfer learning, and efficient training strategies

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If foundation models are used for visual inspection, then adaptability to changing requirements improves, but initial model development resources increase

Engineering Contradiction:
Improveadaptability to changing requirementsVSAvoidinitial resources
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The foundation model performs preliminary action by pre-learning general visual patterns and features from large datasets before being adapted to specific inspection tasks. This preliminary training invests resources upfront but enables rapid adaptation to new requirements with minimal additional resources, improving long-term adaptability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The foundation model provides universality by learning general-purpose visual understanding that can be applied across multiple different inspection tasks and product types. This multi-functionality allows a single model to adapt to changing requirements without requiring complete retraining, reducing ongoing resource requirements

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4607456A1Computer-implemented method and device for visually inspecting a product
Publication Date: 2025.08.27 SIEMENS AG
  • EP4607456A1 patent drawingFigure 1
  • EP4607456A1 patent drawingFigure 2~3
  • EP4607456A1 patent drawing

AI summary

A computer-implemented method for the visual inspection of a product using an inspection device, comprising a detection means and a processor with a memory, wherein the following steps are carried out: a) providing at least one first data structure with first image data, which comprises at least one first identifier for image data for processing by a model based on artificial intelligence, b) providing a foundation model based on artificial intelligence, comprising a foundation identifier mask, which describes the foundation model for processing by a model based on artificial intelligence and is formed by a neural network, and the foundation identifier mask comprises device data, which describe the inspection device and/or properties of the inspection device,c) Applying the image data of the at least one first data structure to the founding model using the at least one first identifier and the founding identifier mask, d) Providing a second data structure with a recognition target comprising a manufacturing model with manufacturing data describing the product and/or the manufacturing of the product during model training, e) Generating and training an artificial intelligence-based inspection model with the image data of the at least one first data structure using the founding identifier mask and the second data structure, f) Providing at least one third data structure with third image data that does not have image identifiers, g) Applying the inspection model to the image data of the at least one third data structure and determining whether the recognition target has been achieved by evaluating the manufacturing model for the product.