The invention provides an efficient and accurate network
threat intelligence automatic extraction method, and aims to solve the problems of incomplete entity recognition, inaccurate relation reasoning, easy model illusion and the like when an existing intelligence extraction technology faces CTI texts with multi-source isomerism, complex safety terms and implicit relation expression. According to the method, the advantages of a
deep learning model and a large
language model are fused, and the structured understanding ability of complex
threat intelligence is comprehensively improved. According to the specific technical scheme, firstly, multi-
source data from security reports, technical blogs and the like are processed in a unified mode through a text
standardization and entity preliminary screening module, and the basic quality of
information extraction is improved; secondly, an entity-driven
attention model is introduced,
threat entity
semantics are recovered through external knowledge enhancement and an entity-to-attention mechanism, a preliminary relation is recognized, and the extraction accuracy is improved; thirdly, capturing potential
attack chain logic and implicit association by adopting an example retrieval
mechanism based on relational logic driving and combining the analogy reasoning capability of a large
language model; and finally, through a
decision fusion and arbitration mechanism, consistency comparison, conflict
verification and deletion completion are carried out on results of the deep model and the
large model, so that the accuracy, integrity and robustness of network
threat intelligence extraction are remarkably improved.