The invention relates to the technical field of
network security, in particular to a
risk identification system based on text and image bimodal fusion, which is characterized in that for text
modal information, enhanced coding is carried out on webpage texts by constructing a Chinese sensitive
knowledge graph fusing voice, fonts and
semantics so as to accurately identify variation sensitive words; webpage visual features are extracted from the image
modal information and fused with an OCR text; a cross-
modal attention alignment mechanism is introduced to realize deep semantic interaction and complementary
verification of text and image features, and finally,
risk classification is performed on the fused features, and a reliable result is output by adopting confidence coefficient calibration. According to the method, the identification precision of risks such as variation content and disguised secret words is remarkably improved, and the robustness and accuracy of
risk identification are improved.