The invention discloses a water conservancy design archive retrieval
system and method based on a local lightweight
large model, and the method comprises the steps: S1, constructing a Python automatic preprocessing
assembly line, extracting texts for PDF and Word multi-format archives, correcting
metadata, and outputting standardized data; s2, constructing a full-
text retrieval and semantic retrieval dual-mode cross-
document retrieval service by relying on a Weavi ate local vector
database and a lightweight text embedding model; s3, analyzing a user query intention through a local
large model, synchronously triggering
metadata accurate retrieval and content semantic retrieval, and generating a structured result; and S4, integrating the core module into a
local area network Web platform, adopting Docker
containerization deployment, and combining an RBAC permission model and JWT
authentication to guarantee security. The
system comprises a preprocessing module, a cross-
document retrieval module, an
intelligent agent module and a background management module, and
collaboration is achieved through a standardized API. According to the method, the problem of archive fragmentation is solved, multi-mode retrieval breaks through keyword limitation, an
intelligent agent reduces manual intervention, a localized architecture prevents secret-related leakage, background management adapts to an existing I T environment, and full-process intelligent archive service is provided for water conservancy design.