This invention discloses a method for constructing a dataset through
multimodal data fusion and extraction in the field of
solid waste disposal. The method includes: first, converting the original document into a sequence of multimodal objects; second, constructing a document-level context index and entity
knowledge base, extracting entity relationships, and constructing a global
knowledge base and index; next, using the global
knowledge base and index, performing semantic classification and dynamic routing on the multimodal object sequence, performing multimodal divide-and-conquer extraction by assembling
contextual information packets, and outputting discrete
record fragments; finally, generating feature identifiers based on experimental conditions, reconstructing cross-
modal data through a confidence priority mechanism, and generating a standard dataset. This invention effectively solves the attention drift and task
confusion problems when large language models process complex long documents, overcomes the challenge of entity referencing resolution in cross-
paragraph contexts, and overcomes the difficulties of information incompleteness and alignment in multimodal table and image data, significantly improving the quality of dataset construction.