Process Data Indexing for Similarity Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current industrial process data analysis systems struggle to efficiently retrieve similar situations, downstream events, and upstream events, limiting operators' ability to make informed decisions due to the lack of effective similarity search capabilities and inefficient data preprocessing methods.

Innovation Solution

A system and method that index process data prior to querying, allowing for weightable search parameters and efficient comparison of search instructions with indexed data, enabling the retrieval of similar situations without re-indexing and emphasizing important variables, regardless of the underlying process status.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the entire historical data is preprocessed whenever a query is performed, then comprehensive analysis is achieved, but significant computation effort is required at query time

Engineering Contradiction:
Improveanalysis comprehensivenessVSAvoidcomputation effort
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent precomputes and stores aggregate statistics (mean, standard deviation, min, max) for all historical data during an offline indexing phase, rather than processing the entire dataset at query time. This preliminary action enables fast similarity searches by comparing query parameters against precomputed aggregates, resolving the contradiction between analysis comprehensiveness and computation effort.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If preprocessing is dependent on query properties, then accurate matching is achieved, but preprocessing cannot be reused for similar queries

Engineering Contradiction:
Improvematching accuracyVSAvoidpreprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a universal index structure that stores aggregate statistics for all possible query combinations. The precomputed aggregates can serve multiple different query types without reprocessing, making the preprocessing universally applicable to various query scenarios while maintaining matching accuracy through the stored statistical parameters.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If multivariate query is regarded as a single entity, then simplified processing is achieved, but individual variable importance cannot be emphasized

Engineering Contradiction:
Improvequery processing complexityVSAvoidvariable weighting capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the multivariate query into individual variable components, each with its own weight parameter. This segmentation allows the system to process each variable separately with customized importance weights while maintaining overall query coherence, resolving the contradiction between processing simplicity and adaptability to emphasize important variables.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10789257B2System and method for similarity search in process data
Publication Date: 2020.09.29 TRENDMINER NV
  • US10789257B2 patent drawing
  • US10789257B2 patent drawing
  • US10789257B2 patent drawing

AI summary

An industrial process analysis system is disclosed The system comprises a process data connection device for the acquisition of process data from process data source, an input system for receiving at least one search instruction, an indexing system for indexing the process data to create a set of indexed process data a data processing device for processing the at least one search instruction to create a search parameter set and comparing distances of members of the search parameter set with corresponding members of the indexed process data to obtain a similarity value; and an output device to display results based on the similarity value.