Efficient multi-modal content security perception method based on autoregressive feature compression

CN120197221AActive Publication Date: 2025-06-24INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510671968.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing multimodal content perception methods are difficult to take into account the efficiency of storage and computing and the flexibility of perception. Especially in the context of the large number of AIGC content, it is difficult to effectively identify bad content outside the distribution.

Method used

An efficient multimodal content security perception method based on autoregressive feature compression is adopted. By constructing a multimodal alignment model and autoregressive model, unified classification and retrieval, content security perception of text and images is supported, and the volume and retrieval time of the search library are reduced through feature compression.

Benefits of technology

It realizes both efficiency and flexibility in content security issues, reduces the storage requirements and search time of the search library, supports dynamic adjustment of the search length, is suitable for low-resource scenarios, and significantly reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197221A_ABST
    Figure CN120197221A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient multi-modal content security perception method based on autoregression feature compression, and belongs to the technical field of artificial intelligence. The method comprises the following steps: step 1, constructing a multi-modal alignment model, aligning text and image feature spaces, and introducing a classifier to carry out content security classification on each modal; 2, constructing an autoregression model, and carrying out efficient representation compression on the multi-modal features obtained by the multi-modal alignment model; and step 3, performing content security retrieval based on the multi-modal alignment model and the autoregression model, and dynamically adjusting the feature length after efficient representation compression according to retrieval time consumption to realize efficient and accurate content security judgment. According to the method, high efficiency and flexibility are taken into account in the aspect of content security, and common image and text modalities are taken into account at the same time.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Feature coding and decoding method, coding and decoding device training method, device and medium

    CN116012662A

  • Multi-mode video retrieval method and device based on multiple encoders, equipment and medium

    CN116737996A

  • Knowledge fusion method and system for multi-source heterogeneous multi-modal data

    CN118690838A

  • Multi-mode guided high-fidelity image compression method and system and medium

    CN119906827A

  • System and method for registering multi-modality images

    US20180308245A1