Metadata Format Conversion for Source Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Metadata annotated in crawled documents is not effectively reflected back to the original data sources, leading to format discrepancies and loss of additional information added by search applications.
Innovation Solution
A computer-implemented method and system that includes a metadata handler to convert the metadata format of internal documents to match the original data sources, allowing the metadata to be posted back to the original documents, thereby maintaining consistency and additional information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If metadata is annotated in internal documents during search application processing, then additional information and enrichment are added to the documents, but the metadata format differences prevent effective reflection back to original data sources
Solution Approach 1:
The patent transforms metadata from the internal document format back to the original data source format by changing format parameters. The system identifies the original format characteristics and converts the enriched metadata annotations to match those parameters, enabling successful reflection back to the source while preserving the added information.
Solution Approach 2:
The patent creates a copy of the metadata information that can exist in both formats. It maintains the original metadata structure from the data source while overlaying additional enrichment information, effectively copying and adapting the data to serve dual purposes without losing information from either source.
2Reliability
If search applications automatically add fields to internal documents, then additional information is captured, but format discrepancies prevent synchronization with original documents
Solution Approach 1:
Instead of converting original documents to internal format and losing track of the original structure, the patent inverts the approach by starting with the internal document metadata and transforming it back to match the original data source format. This inversion enables synchronization while preserving automatically added fields.
Solution Approach 2:
The patent introduces a metadata handler as an intermediary component that mediates between the internal document format and the original data source format. This handler translates and reconciles the two formats, allowing automatic fields to be preserved and synchronized without direct format conflict.
3Loss of information
If manual operations add metadata to internal documents, then user-added tags are preserved, but format conversion loses the enriched information
Solution Approach 1:
The patent performs preliminary identification of user-added metadata characteristics before format conversion. By pre-analyzing the metadata structure and enrichment information, the system prepares the transformation process to preserve these elements, preventing information loss during the conversion back to the original format.
Solution Approach 2:
The patent implements a feedback mechanism where the metadata handler continuously monitors the conversion process and adjusts transformations to preserve user-added tags. The system receives feedback about which metadata elements are user-added versus automatically generated and applies appropriate preservation strategies during format conversion.
Data Source
AI summary
A computer-implemented method, a computer program product, and a computer system for reflecting metadata annotated in crawled documents to original data sources. In response to one or more internal documents in an application being annotated with metadata, the computer system converts a metadata format handled in the application to a metadata format handled in one or more data sources. The computer system posts, to one or more original documents in the one or more data sources, the metadata in the metadata format handled in the one or more data sources.


