The invention belongs to the technical field of computers, and discloses a news detail page
XPath automatic extraction method and
system based on a
large model, and the method comprises the steps: determining a
large model base through multi-aspect evaluation; the method comprises the following steps: preprocessing an
HTML page, designing a structured prompt template and constructing a thinking chain CoT
data set; utilizing a Lora technology
fine tuning model, and adopting a smooth
loss function and a dual reward mechanism of fusion format and content evaluation to optimize; and performing format calibration on the
XPath and the
JSON output by the model, and extracting and verifying the validity of the URL of the detail page by using the calibrated
XPath. Compared with a traditional template-based
information extraction method, the method does not depend on the stability of a webpage structure, has higher generalization ability and
fault tolerance, remarkably improves the
information extraction precision, enhances the complex webpage adaptability, improves the overall robustness and universality of a
system, and is suitable for large-scale popularization and application. And a stable extraction effect can still be kept in news websites with frequent structure change.