A risk identification system based on text and image bimodal fusion

By constructing a Chinese sensitive semantic knowledge graph and a cross-modal interaction mechanism, the problems of insufficient recognition of Chinese homophone substitution and code words in existing technologies have been solved, achieving efficient identification of Chinese risky websites and improving recognition accuracy and robustness.

CN121960712BActive Publication Date: 2026-06-12JILIN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-03-26
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing methods for identifying risky websites mainly rely on single-modal information analysis, which makes it difficult to deal with variations such as homophone substitution, similar-looking substitution, and coded language in Chinese. Furthermore, they lack deep semantic association and complementary verification, resulting in insufficient ability to identify complex and hidden risky content.

Method used

A Chinese sensitive semantic knowledge graph is constructed, and text modal information is enhanced and encoded through semantics, speech and glyphs. Semantic alignment and fusion of image and text features are achieved through cross-modal interaction mechanism, and risk identification results are output by combining confidence calibration.

Benefits of technology

It significantly improves the accuracy of identifying variant content and disguised coded language in Chinese scenarios, enhances the robustness and generalization ability of the model, reduces the dependence on labeled data, and improves the accuracy and practical value of risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960712B_ABST
    Figure CN121960712B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network security, and specifically provides a risk identification system based on text and image bimodal fusion. For text modal information, a Chinese sensitive knowledge graph fusing voice, character shape and semantics is constructed to enhance coding of webpage text, so that variant sensitive words can be accurately identified. For image modal information, webpage visual features are extracted and fused with OCR text. A cross-modal attention alignment mechanism is introduced to realize deep semantic interaction and complementary verification of text and image features. Finally, the fused features are classified in terms of risks, and reliable results are output by using confidence calibration. The application significantly improves the identification accuracy of risks such as variant content and disguised codes, and improves the robustness and accuracy of risk identification.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Method for detecting and positioning multi-mode media image-text synchronous forgery

    CN120125979A

  • Multi-modal financial risk identification method and device based on image and text

    CN121682708A