The application discloses a cross-document type retrieval method and
system based on differential nested coding. The method includes six steps: document preprocessing, document type identification,
differential coding, index construction and optimization, online retrieval and result fusion. First, the document is standardized and preprocessed to extract structured information. Second, the BERT classifier is used to analyze the document features and identify the document type. Then, according to the document type, the corresponding
encoder is selected to generate the original coding vector. Next, the differential nested coding index is constructed and optimized. In the online retrieval stage, the user query is preprocessed, the relevant document type is predicted, and the approximate nearest neighbor retrieval is performed. Finally, the retrieval results are sorted by a stepwise weight fusion strategy, and the top-100 retrieval results are returned. Through
differential coding and index optimization, the application realizes efficient retrieval of different types of documents, and has significant practical value and application prospect.