The invention belongs to the technical field of
computer data processing, and discloses a heterogeneous
resource pool oriented automatic data
crawling and semantic
parsing method and
system.The method comprises the steps that S1, a
database, an interface, a document and a webpage sample are parsed, and field neighborhoods, interface parameters and
metadata are extracted; generating a
feature vector and forming an access rule and a context
fingerprint; s2, constructing a collection task under the constraint of rules and fingerprints, calculating the priority according to two factors of update rate and
semantic consistency, and performing distributed execution; s3, unifying the data into an intermediate representation, and generating
semantic mapping based on
fingerprint limitation candidates; and S4, verifying mapping consistency, detecting conflicts, coding the conflicts into conflict cause vectors, and feeding back correction rules, weights and mapping. The
system is composed of an access engine, a collection driver and a
fusion center, and the access engine, the collection driver and the
fusion center complete rule and
fingerprint generation, task scheduling and collection, intermediate result
verification and conflict feedback
closed loop.