Facility Data Classification Using Automated Cluster Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Facilities in the life sciences sector face challenges in efficiently and accurately labeling vast amounts of data for machine learning and artificial intelligence applications due to manual annotation being time-consuming, costly, and prone to errors and biases, leading to inefficient utilization of resources and unreliable data quality.

Innovation Solution

An automated system retrieves unlabeled data, applies Bayesian optimization and machine learning algorithms to identify topics and clusters, generates prioritized keywords, and uses language learning models to label clusters, training a classification model for efficient and high-quality data labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation by domain experts is used to label data, then data quality and accuracy are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedata labeling accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical annotation process with an automated machine learning-based system. The system uses pre-trained models to automatically label data, substituting human expert labor with computational processes that can scale efficiently without time loss

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates synthetic labeled data by generating fake data samples that mirror real data patterns. These synthetic samples are automatically labeled and used to train models, copying the structure of manual annotation without requiring actual human experts for each data point

Inventive Principle:
Principle #26Copying

2Measurement precision

If manual annotation by domain experts is used to label data, then data quality is improved, but human resource costs increase

Engineering Contradiction:
Improvedata labeling qualityVSAvoidhuman resources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent replaces human expert annotators with automated machine learning systems that perform labeling functions. The system uses algorithms to automatically assign labels to data, eliminating the need for continuous human resource input while maintaining labeling quality through model-based approaches

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service labeling where the machine learning model automatically annotates data without requiring human intervention. The model trains on available data and performs self-labeling, reducing dependency on human resources for routine annotation tasks

Inventive Principle:
Principle #25Self-service

3Measurement precision

If manual annotation is used to handle huge volume of data, then labeling accuracy is maintained, but scalability fails

Engineering Contradiction:
Improvelabeling accuracyVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual annotation with automated machine learning systems that can process huge volumes of data at scale. The system uses distributed computing and optimized algorithms to maintain labeling accuracy while handling large data volumes that would be impossible for human experts to process manually

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary data processing and pre-training of models on subsets of data before full-scale labeling. This preliminary action prepares the system to efficiently handle large volumes of data, establishing a foundation that enables scalable processing without sacrificing accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12572567B2Systems and methods for constructing a classification model using data associated with a facility
Publication Date: 2026.03.10 HONEYWELL INTERNATIONAL INC
  • US12572567B2 patent drawing
  • US12572567B2 patent drawing
  • US12572567B2 patent drawing

AI summary

Various embodiments described herein relate to systems and methods for constructing a classification model using data associated with a facility. In this regard, unlabeled data associated with the facility is retrieved. Based on an application of a first algorithm on the unlabeled data, parameters in the unlabeled data are determined. Using the parameters, machine learning algorithms identify topics and clusters in the unlabeled data. Also, a prioritized list of keywords is generated for each of the topics. Then, the prioritized list of keywords is input to a language learning model. The language learning model outputs a keyword from the list of keywords. If the keyword meets a predefined threshold, a cluster of the clusters is labeled using the keyword. Using first set of labeled clusters, the classification model is trained. The trained classification model provides recommendations related to the facility which are also renderable on a user interface as well.