The invention discloses a webpage
data extraction method based on a large
language model, which comprises the following steps of: generating an
Xpath sequence webpage grabber by utilizing the large
language model, and
processing diversified and variable network environments through a two-stage framework: in the first stage, inversely checking and removing
HTML (
Hypertext Markup Language)
noise by utilizing LLM (Language
Language Model)
information extraction capability, and in the second stage, inversely checking and removing
HTML (
Hypertext Markup
Language Model)
noise; self-adaptive generation of an
Xpath action sequence is carried out according to the hierarchical structure of the
HTML; according to the combination of an external evaluation mechanism and a
local evaluation mechanism of LLMs, a plurality of
Xpath action sequences generated on different webpages in
one stage are integrated, and a general grabber specific to a website is generated. According to the method, a baseline method is always exceeded under zero sample setting, higher efficiency is shown in large-scale webpage
information extraction tasks, the method can quickly adapt to different website and task requirements, dependence on LLMs is reduced when similar tasks are processed, and therefore the efficiency of
processing a large number of webpage tasks is improved.