Inverted Index Transposition for Data Restoration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Restoring original data from compressed data using inverted indexes is challenging due to transposed indexes not following the sequence of word appearance, and high-frequency words are often excluded to manage index size, making it difficult to accurately restore the original data.

Innovation Solution

An information processing device uses bitmap-type inverted indexes, static, and dynamic dictionaries to transpose and convert compression codes back into their original sequence, allowing for the restoration of original text data by referring to the indexes generated from text data, which associate morphemes or words with their positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If inverted indexes are used for compression and search, then search speed is improved, but the ability to restore original data is worsened because indexes are transposed and do not follow word appearance sequence

Engineering Contradiction:
Improvesearch speedVSAvoidoriginal data restoration capability
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

A transpose buffer is introduced as an intermediary data structure that stores compression codes in the order they appear in the original text. During restoration, the system reads from the transpose buffer rather than directly from the transposed inverted index, thereby recovering the original sequence information that was lost during indexing. This mediator preserves the beneficial transposed structure for fast search while enabling accurate data restoration.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If high-frequency words are excluded from inverted indexes to manage size, then index size is reduced, but data restoration accuracy is worsened

Engineering Contradiction:
Improveindex sizeVSAvoiddata restoration accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system segments words into two categories: high-frequency words that are excluded from the inverted index to save space, and low-frequency words that are included. The transpose buffer stores compression codes for all words in their original sequence, ensuring that even excluded high-frequency words can be restored correctly by maintaining positional information for the entire text.

Inventive Principle:
Principle #1Segmentation

3Productivity

If compression codes are transposed for efficient storage, then storage efficiency is improved, but the complexity of restoring original sequence is worsened

Engineering Contradiction:
Improvestorage efficiencyVSAvoidrestoration process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

During the compression phase, the system performs a preliminary action by storing compression codes in a transpose buffer in their original text order before creating the transposed inverted index. This preliminary organization of data in sequence order simplifies the restoration process, as the system can directly read from the pre-organized transpose buffer without needing to perform complex reordering operations during restoration.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10915559B2Data generation method, information processing device, and recording medium
Publication Date: 2021.02.09 FUJITSU LTD
  • US10915559B2 patent drawing
  • US10915559B2 patent drawing
  • US10915559B2 patent drawing

AI summary

A non-transitory computer-readable recording medium stores therein a data generation program that causes a computer to execute a process including: referring to each index in which a morpheme, which is generated from text data and which is included in the text data, is associated to position of the morpheme in the text data; and arranging, in sequence of positions in the text data, morphemes associated in the indexes.