Inverted Index Transposition for Data Restoration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Restoring original data from compressed data using inverted indexes is challenging due to transposed indexes not following the sequence of word appearance, and high-frequency words are often excluded to manage index size, making it difficult to accurately restore the original data.
Innovation Solution
An information processing device uses bitmap-type inverted indexes, static, and dynamic dictionaries to transpose and convert compression codes back into their original sequence, allowing for the restoration of original text data by referring to the indexes generated from text data, which associate morphemes or words with their positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If inverted indexes are used for compression and search, then search speed is improved, but the ability to restore original data is worsened because indexes are transposed and do not follow word appearance sequence
Solution Approach 1:
A transpose buffer is introduced as an intermediary data structure that stores compression codes in the order they appear in the original text. During restoration, the system reads from the transpose buffer rather than directly from the transposed inverted index, thereby recovering the original sequence information that was lost during indexing. This mediator preserves the beneficial transposed structure for fast search while enabling accurate data restoration.
2Quantity of substance
If high-frequency words are excluded from inverted indexes to manage size, then index size is reduced, but data restoration accuracy is worsened
Solution Approach 1:
The system segments words into two categories: high-frequency words that are excluded from the inverted index to save space, and low-frequency words that are included. The transpose buffer stores compression codes for all words in their original sequence, ensuring that even excluded high-frequency words can be restored correctly by maintaining positional information for the entire text.
3Productivity
If compression codes are transposed for efficient storage, then storage efficiency is improved, but the complexity of restoring original sequence is worsened
Solution Approach 1:
During the compression phase, the system performs a preliminary action by storing compression codes in a transpose buffer in their original text order before creating the transposed inverted index. This preliminary organization of data in sequence order simplifies the restoration process, as the system can directly read from the pre-organized transpose buffer without needing to perform complex reordering operations during restoration.
Data Source
AI summary
A non-transitory computer-readable recording medium stores therein a data generation program that causes a computer to execute a process including: referring to each index in which a morpheme, which is generated from text data and which is included in the text data, is associated to position of the morpheme in the text data; and arranging, in sequence of positions in the text data, morphemes associated in the indexes.


