The invention discloses a semantic-based text content index automatic identification method, which relates to the technical field of
information retrieval, and comprises the following steps: initializing sparse projection and LSH signature, performing iterative optimization by using a Lagrange duality form and gradient update, adjusting hash digits, obtaining a fragment index through a k-d tree, and constructing an
inverted index. According to the method, compression is performed through
Delta coding, an index map is constructed based on Jaccard similarity, compression is performed through WebGraph, CSNMF is used in combination with Z-
Laplacian regularization, a low-rank basis matrix and a low-rank coding matrix are generated, a compressed
inverted index is reconstructed after iterative optimization, and reconstructed
inverted index entries are generated. According to the method, through multi-resolution
hash table initialization, joint feature optimization, local adaptive quantization and low-rank index reconstruction, the semantic expression ability and the compression effect of an index structure are improved, the index precision and efficiency are improved, and intelligent identification of index content is achieved.