Automated Metadata Generation for Image and Text Data Mining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of data mining in massive datasets is the complexity of extracting contextually relevant information from large volumes of data, particularly in image and text datasets, where manual examination is impractical due to their scale, and existing methods struggle with noise, dimensionality, and the need for automated techniques that can efficiently reduce dataset size and enhance data visualization.
Innovation Solution
An automated metadata system is developed to generate feature vectors from text and image data, using techniques like bigram/trigram proximity matrices for text and grey level co-occurrence matrices for images, which are used to create digital objects that facilitate Boolean searches and reduce the dataset size, enabling efficient data mining and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual examination of data is used, then data quality and relevance can be assessed, but the process becomes impractical due to the large scale of massive datasets
Solution Approach 1:
The patent creates digital objects that are copies or representations of the actual data objects. These digital objects contain metadata and feature vectors that capture the essential characteristics of the data, allowing automated examination without manually processing the entire massive dataset. The digital objects serve as proxies that enable efficient filtering and retrieval.
Solution Approach 2:
The patent extracts relevant information from the data by generating feature vectors and metadata that capture the essential characteristics. This extraction process separates the important information from the bulk data, allowing the system to work with condensed representations rather than the full datasets, thereby reducing examination time while maintaining quality assessment capability.
2Productivity
If automated metadata generation is implemented, then data processing efficiency improves, but the complexity of the system increases
Solution Approach 1:
The patent segments the data processing task into distinct components: extracting features, generating metadata, creating digital objects, and performing searches. This segmentation allows each component to be optimized independently and makes the overall complex system more manageable. The feature extraction, metadata generation, and object creation are separate modules that work together systematically.
Solution Approach 2:
The patent introduces digital objects as intermediary structures between the raw data and the search/query interface. These digital objects contain both the extracted features and the generated metadata, serving as a mediator that simplifies the interaction between users and the massive dataset. The digital objects act as a bridge that reduces the complexity of directly querying the underlying data.
3Measurement precision
If feature extraction and metadata generation are performed on all data, then search accuracy improves, but the computational resources and time required increase significantly
Solution Approach 1:
The patent performs feature extraction and metadata generation as preliminary actions before the actual search or query process. By pre-processing the data and creating digital objects with extracted features in advance, the system prepares the data structure that enables fast, accurate searches without requiring computational resources during the search operation itself. This preliminary action separates the computationally intensive extraction phase from the query execution phase.
4Speed
If the dataset size is reduced through automated filtering, then data retrieval speed increases, but the risk of losing relevant information increases
Solution Approach 1:
The patent incorporates feedback mechanisms where the generated metadata and feature vectors are used to evaluate and refine the filtering process. The system can assess the quality of filtered results and adjust the filtering criteria accordingly, ensuring that relevant information is not lost while maintaining fast retrieval speeds. The feedback loop allows continuous optimization of the balance between speed and completeness.
Data Source
AI summary
A tangible computer readable medium encoded with instructions for automatically generating metadata, wherein said execution of said instructions by one or more processors causes said “one or more processors” to perform the steps comprising: a. creating at least one feature vector for each document in a dataset; b. extracting said one feature vector; c. recording said feature vector as a digital object; and d. augmenting metadata using said digital object to reduce the volume of said dataset, said augmenting capable of allowing a user to perform a search on said dataset.


