Image-text harmful information identification method and device, electronic equipment and storage medium
Through the strategy of decoupling the back-end alignment of the front-end, image and text feature vectors are extracted and cross-modal fusion is performed, which solves the problem of insufficient accuracy of multimodal information recognition in the prior art and realizes more efficient detection of harmful information.
Patent Information
- Application Number
- CN202510334786.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing methods of identifying harmful information mainly focus on single-modal data processing, making it difficult to effectively identify multimodal information containing images and text, especially the image content with implicit semantics is difficult to accurately classify.
The front-end decoupling and back-end alignment strategy is adopted to extract image and text feature vectors respectively, and information fusion is performed through a cross-modal attention mechanism to generate graphic and text fusion features and semantic fusion features, and finally input harmful information classification model for detection.
It realizes deep semantic alignment of graphic and text information, improves the accuracy and generalization ability of multimodal harmful information recognition, and enhances the ability to understand complex graphic and text content.
Smart Images

Figure CN120264053A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information processing technologies, and more particularly, to a method, apparatus, electronic device, and storage medium for identifying harmful graphic and text information. Background Art
[0002] With the development of the Internet and social media, graphic and text content has become the main form of information dissemination. However, the complexity and diversity of graphic and text information also bring the risk of harmful information dissemination, such as hate speech, violent content, pornographic content, and other bad information. These contents not only affect the network environment but may also have a negative impact on social stability and user mental health. Therefore, how to accurately and quickly identify harmful graphic and text information has become an important research direction in the field of network content review.
[0003] Current harmful information identification methods mainly focus on the processing of unimodal data such as text or images. Traditional text classification methods rely on keyword matching or rule bases and are difficult to handle complex semantic expressions and new implicit expressions. Although deep learning models have improved the text semantic understanding ability, they cannot process multimodal information containing image content, making text-based identification methods ineffective for harmful information that relies on image expression (such as text embedded in pictures, metaphorical images). Computer vision technology has been widely used in image content review, such as using deep learning models such as CNN and ViT to detect harmful images such as nudity and violence. However, the discriminative ability of images themselves is limited, and it is difficult to understand the implicit semantics therein. For example, for some pictures with suggestive but not directly harmful content, it is difficult to accurately classify them only through visual features. Summary of the Invention
[0004] Embodiments of the present disclosure at least provide a method, apparatus, electronic device, and storage medium for identifying harmful graphic and text information. By adopting a strategy of front-end decoupling and back-end alignment, image and text features are extracted separately at the front end, and cross-modal attention mechanisms are used at the back end for information fusion, thereby improving the ability to identify harmful information. While ensuring the feature expression ability, deep semantic alignment of graphic and text information is achieved, effectively improving the accuracy and generalization ability of multimodal harmful information identification.
[0005] Embodiments of the present disclosure provide a method for identifying harmful graphic and text information, including:
[0006] Obtain the graphic and text data to be identified, extract the image feature vector corresponding to the image modality data in the graphic and text data to be identified, and extract the text feature vector and semantic feature vector corresponding to the text modality data in the graphic and text data to be identified;
[0007] Fuse the image feature vector with the text feature vector to generate a text-image fusion feature, and fuse the image feature vector with the semantic feature vector to generate a semantic fusion feature;
[0008] Fuse the text-image fusion feature with the semantic fusion feature to generate the target feature corresponding to the text-image data to be recognized, and input the target feature into a pre-trained harmful information classification model to determine the harmful information detection result corresponding to the text-image data to be recognized.
[0009] In an optional implementation, extracting the image feature vector corresponding to the image modality data in the text-image data to be recognized specifically includes:
[0010] Segment the image modality data into multiple image patches, and flatten each image patch into a corresponding image vector;
[0011] Map the image vector to an embedding space through linear projection, and add corresponding position encoding to generate an image embedding sequence;
[0012] Input the image embedding sequence into a pre-trained encoder, and perform multi-level feature extraction and interaction through a multi-head self-attention mechanism and a feed-forward neural network to generate the image feature vector.
[0013] In an optional implementation, extracting the text feature vector and the semantic feature vector corresponding to the text modality data in the text-image data to be recognized specifically includes:
[0014] Perform optical character recognition processing on the text modality data to extract the text content corresponding to the text modality data;
[0015] Perform word segmentation and sub-word encoding on the text content to convert it into a token sequence, map the token sequence to an embedding space, and add corresponding position encoding to generate a text embedding sequence;
[0016] Input the text embedding sequence into a pre-trained text model for context feature extraction, and capture the semantic relationship between tokens in the token sequence through a multi-head self-attention mechanism and a feed-forward neural network;
[0017] Extract the feature representation output by the encoder corresponding to the text model as the text feature vector.
[0018] In an optional implementation, the semantic feature vector is extracted based on the following steps:
[0019] Input the text content into a preset large text model;
[0020] Perform natural language understanding on the text content through the text large model to generate the semantic feature vector at the semantic level.
[0021] In an alternative embodiment, fusing the image feature vector and the text feature vector to generate a text-image fusion feature specifically includes:
[0022] Concatenate the image feature vector and the text feature vector to generate a joint feature vector;
[0023] After performing dimensionality reduction or normalization processing on the joint feature vector, use the attention mechanism to perform feature fusion on the joint feature vector to generate the text-image fusion feature.
[0024] In an alternative embodiment, fusing the image feature vector and the semantic feature vector to generate a semantic fusion feature specifically includes:
[0025] After aligning the image feature vector and the semantic feature vector, input them into a preset cross-modal attention layer to determine the interaction weight between the image feature vector and the semantic feature vector;
[0026] Perform context modeling on the image feature vector and the semantic feature vector respectively through the multi-head self-attention mechanism, and capture the cross-modal semantic correlation features between the image feature vector and the semantic feature vector;
[0027] Perform weighted summation on the semantic correlation features according to the interaction weight to determine the joint cross-modal feature, and after fusing the joint cross-modal feature through a preset encoder, generate the semantic fusion feature.
[0028] In an alternative embodiment, input the target feature into a pre-trained harmful information classification model to determine the harmful information detection result corresponding to the text-image data to be recognized, specifically including:
[0029] Input the target feature into the harmful information classification model;
[0030] Perform binary classification of harmful and harmless on the target feature through the harmful information classification model to determine whether there is a harmful information detection result;
[0031] If the harmful information detection result is that there is harmful information, perform multi-classification of harmful information categories on the harmful information through the harmful information classification model to determine the harmful type corresponding to the harmful information.
[0032] The embodiments of the present disclosure also provide a text-image harmful information recognition device, including:
[0033] A feature extraction module, used to obtain the image and text data to be identified, extract the image feature vector corresponding to the image modality data in the image and text data to be identified, and extract the text feature vector and semantic feature vector corresponding to the text modality data in the image and text data to be identified;
[0034] A feature fusion module, used for fusing the image feature vector with the text feature vector to generate an image-text fusion feature, and fusing the image feature vector with the semantic feature vector to generate a semantic fusion feature;
[0035] The detection module is used to fuse the image-text fusion feature and the semantic fusion feature to generate a target feature corresponding to the image-text data to be identified, input the target feature into a pre-trained harmful information classification model, and determine the harmful information detection result corresponding to the image-text data to be identified.
[0036] An embodiment of the present disclosure also provides an electronic device, including: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the above-mentioned method for identifying harmful information in images and texts, or steps in any possible implementation of the above-mentioned method for identifying harmful information in images and texts are executed.
[0037] The embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for identifying harmful information in images and texts described above, or the steps in any possible implementation of the method for identifying harmful information in images and texts described above, are executed.
[0038] The embodiments of the present disclosure also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned method for identifying harmful information in images and texts, or the steps in any possible implementation of the above-mentioned method for identifying harmful information in images and texts.
[0039] A method, device, electronic device, and storage medium for identifying harmful graphic and text information provided by an embodiment of the present disclosure. Obtain the graphic and text data to be identified, extract the image feature vector corresponding to the image modality data in the graphic and text data to be identified, and extract the text feature vector and semantic feature vector corresponding to the text modality data in the graphic and text data to be identified; fuse the image feature vector with the text feature vector to generate a graphic and text fusion feature, and fuse the image feature vector with the semantic feature vector to generate a semantic fusion feature; fuse the graphic and text fusion feature with the semantic fusion feature to generate a target feature corresponding to the graphic and text data to be identified, and input the target feature into a pre-trained harmful information classification model to determine the harmful information detection result corresponding to the graphic and text data to be identified. Adopt a strategy of front-end decoupling and back-end alignment, extract features of images and texts separately at the front end, and use a cross-modal attention mechanism for information fusion at the back end, thereby improving the ability to identify harmful information. While ensuring the feature expression ability, deep semantic alignment of graphic and text information is achieved, effectively improving the accuracy and generalization ability of multi-modal harmful information identification.
[0040] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides detailed descriptions as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of the specification. These drawings show embodiments that conform to the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 Shows a flowchart of a method for identifying harmful graphic and text information provided by an embodiment of the present disclosure;
[0043] Figure 2 Shows a flowchart of another method for identifying harmful graphic and text information provided by an embodiment of the present disclosure;
[0044] Figure 3 Shows a schematic diagram of a device for identifying harmful graphic and text information provided by an embodiment of the present disclosure;
[0045] Figure 4 Shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some, rather than all, of the embodiments of the present disclosure. Components of the embodiments of the present disclosure generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided herein is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of the present disclosure.
[0047] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it does not require further definition and explanation in subsequent figures.
[0048] The term "and / or" in this document merely describes an associated relationship and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this document means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0049] Through research, it has been found that current harmful information recognition methods mainly focus on the processing of single-modal data such as text or images. Traditional text classification methods rely on keyword matching or rule bases and are difficult to handle complex semantic expressions and new implicit expressions. Although deep learning models have improved the ability to understand text semantics, they cannot process multi-modal information containing image content, making text-based recognition methods ineffective for harmful information that relies on image expression (such as text embedded in pictures, metaphorical images). Computer vision technology has been widely applied to image content review, such as using deep learning models such as CNN and ViT to detect harmful images such as nudity and violence. However, the discriminative ability of images themselves is limited and it is difficult to understand the implicit semantics therein. For example, for some pictures with suggestiveness but not directly showing harmful content, it is difficult to accurately classify them only through visual features.
[0050] Based on the above research, the present disclosure provides a method, apparatus, electronic device, and storage medium for identifying harmful graphic and text information. The method includes obtaining graphic and text data to be identified, extracting an image feature vector corresponding to the image modality data in the graphic and text data to be identified, and extracting a text feature vector and a semantic feature vector corresponding to the text modality data in the graphic and text data to be identified; fusing the image feature vector and the text feature vector to generate a graphic and text fusion feature, and fusing the image feature vector and the semantic feature vector to generate a semantic fusion feature; fusing the graphic and text fusion feature and the semantic fusion feature to generate a target feature corresponding to the graphic and text data to be identified, and inputting the target feature into a pre-trained harmful information classification model to determine a harmful information detection result corresponding to the graphic and text data to be identified. By adopting a strategy of front-end decoupling and back-end alignment, image and text are respectively feature-extracted at the front end, and cross-modal attention mechanism is used for information fusion at the back end, thereby improving the ability to identify harmful information. While ensuring the feature expression ability, deep semantic alignment of graphic and text information is achieved, effectively improving the accuracy and generalization ability of multi-modal harmful information identification.
[0051] To facilitate the understanding of this embodiment, first, a method for identifying harmful graphic and text information disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the method for identifying harmful graphic and text information provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such computer devices include, for example: terminal devices or servers or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the method for identifying harmful graphic and text information may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0052] See Figure 1 As shown in the flowchart of a method for identifying harmful graphic and text information provided in the embodiments of the present disclosure, the method includes steps S101 to S103, where:
[0053] S101. Obtain graphic and text data to be identified, extract an image feature vector corresponding to the image modality data in the graphic and text data to be identified, and extract a text feature vector and a semantic feature vector corresponding to the text modality data in the graphic and text data to be identified.
[0054] In the specific implementation, for the image feature vector, the image modality data is divided into multiple image blocks, and each image block is flattened into the corresponding image vector; the image vector is mapped to the embedding space through linear projection, and the corresponding position code is added to generate the image embedding sequence; the image embedding sequence is input into the pre-trained encoder, and multi-level feature extraction and interaction are performed through the multi-head self-attention mechanism and feedforward neural network to generate the image feature vector. For the text feature vector, optical character recognition is performed on the text modality data to extract the text content corresponding to the text modality data; the text content is converted into a word sequence by word segmentation and subword encoding, and the word sequence is mapped to the embedding space, and the corresponding position code is added to generate the text embedding sequence; the text embedding sequence is input into the pre-trained text model for context feature extraction, and the semantic relationship between words in the word sequence is captured through the multi-head self-attention mechanism and feedforward neural network; the feature representation output by the encoder corresponding to the text model is extracted as the text feature vector. For the semantic feature vector, the text content is input into the preset text large model; the text large model is used to perform natural language understanding on the text content to generate a semantic feature vector at the semantic level.
[0055] Here, the image and text data to be identified is multimodal data containing images and text content, and its data source can be social media, news platforms, user-uploaded content, etc. The image and text data to be identified can include image modal data: such as photos, screenshots, advertising pictures, etc.; text modal data: such as text information, accompanying text, comments, etc. in the picture. After obtaining the data, it is necessary to extract features from different modal data for subsequent fusion analysis.
[0056] Among them, the image feature vector is a high-dimensional numerical representation extracted from the image data, which can capture the color, texture, shape, structure and other information of the image. The text feature vector is used to represent the semantic information of the text modal data and is extracted through a deep learning model. The semantic feature vector is used to capture the deep semantic relationship of the text. Unlike the text feature vector, it emphasizes the understanding of the context.
[0057] In the specific implementation, the image data is converted to a standard size to meet the model input requirements (for example, scaled to 224×224 pixels), and the pixel values are normalized to make the data distribution more even and improve the model generalization ability. If there is text in the image, optical character recognition (OCR) processing is required to extract the text information separately.
[0058] Specifically, first, the input image is segmented into image patches of a fixed size, and each patch is flattened into a vector. Then, these vectors are mapped to the embedding space through linear projection, and positional encoding is added to preserve spatial information. Next, the embedded sequence is input into a pre-trained Transformer encoder, and multi-level feature extraction and interaction are performed through the multi-head self-attention mechanism and the feed-forward neural network. Finally, the feature representation output by the encoder (such as the class token or the aggregation of all patch tokens) is extracted as the global or local feature vector of the image for downstream tasks of harmfulness classification.
[0059] Here, the image is divided into multiple small blocks (such as a 16×16 grid) to make it more manageable. Each image patch is flattened into a vector to reduce the computational complexity, and the image patch is mapped to an embedding vector space of a fixed dimension through linear projection.
[0060] Among them, a convolutional neural network (CNN) or a vision Transformer (ViT) model is used for feature extraction. The CNN method extracts hierarchical features of the image through multiple convolutional layers and pooling layers. The ViT method learns the relationships between image patches through the multi-head self-attention mechanism. After being processed by the feed-forward neural network (FFN), the final image feature vector is obtained. Finally, the extracted image feature vector is a set of high-dimensional numerical vectors representing the important features of the image.
[0061] Furthermore, first, the input text is tokenized and sub-word encoded, and converted into a sequence of tokens. Then, the token sequence is mapped to embedding vectors, and positional encoding is added to capture the sequence order information. Next, the embedded sequence is input into a pre-trained RoBERTa model, and context features are extracted through multiple layers of Transformer encoders. The multi-head self-attention mechanism and the feed-forward neural network are used to capture the semantic relationships between tokens. Finally, the feature representation output by the encoder (such as the [CLS] token or the aggregation of all tokens) is extracted as the global or local feature vector of the text for downstream tasks of harmfulness classification.
[0062] Here, if the text modal data is embedded in an image (such as the text on a poster), OCR processing needs to be performed first to convert the text in the image into analyzable text. The text is split into word or sub-word units. For example, the text "harmful information" is split into ["you", "hai", "xin xi"]. Then, sub-word encoding is performed using BERT or SentencePiece to reduce the vocabulary size, and positional encoding is added to each token to preserve the order information.
[0063] Among them, the input text embedding sequence is fed into a pre-trained text model (such as BERT, RoBERTa), and the relationships between words in the text are analyzed through the multi-head self-attention mechanism to extract deep semantic information. Finally, the feature vector of the text is obtained as the numerical representation of the text.
[0064] Furthermore, a large-scale pre-trained language model such as GPT-4, T5, or BERT is used to analyze the text. These models can be trained based on a large corpus to understand the implicit meaning, sentiment, intention, etc. of the text.
[0065] Here, the hierarchical semantic relationship of the text is extracted through the Transformer structure. Combining the attention mechanism, key entities, sentiment tendencies, and context relationships in the text are focused on. A semantic feature vector is generated, which represents the deep semantic information of the text in a high-dimensional space.
[0066] Optionally, the image pre-trained feature extractor can be a ResNet-based network. The text pre-trained feature extractor can be a ResNet-based network. The text large model can be other current mainstream text large models.
[0067] In this way, the image feature vector represents image content such as shape, color, object, etc.; the text feature vector represents text modal information, emphasizing syntactic and semantic structures; the semantic feature vector represents the deep semantics of the text, focusing on context and sentiment information. These feature vectors will be used for multi-modal fusion later to improve the accuracy of harmful information recognition.
[0068] S102. Fuse the image feature vector and the text feature vector to generate a text-image fusion feature, and fuse the image feature vector and the semantic feature vector to generate a semantic fusion feature.
[0069] In a specific implementation, the image feature vector and the text feature vector respectively represent the independent feature information of the image and the text. To better understand the correlation between the image and the text, it is necessary to fuse these two types of features to obtain the image-text fusion feature. Concatenate the image feature vector and the text feature vector to generate a joint feature vector; after performing dimensionality reduction or normalization on the joint feature vector, use the attention mechanism to perform feature fusion on the joint feature vector to generate the image-text fusion feature.
[0070] Here, since the feature vectors of the image and the text are usually in different spatial dimensions (for example, the image feature may be 1024-dimensional, while the text feature may be 768-dimensional), it is necessary to perform dimension alignment. Use a fully connected layer or a 1×1 convolution to reduce the dimension of the high-dimensional feature so that the dimensions of the image and text features are the same. Through Batch Normalization or L2 normalization, the numerical ranges of different modality features are made similar for fusion.
[0071] After that, use a direct concatenation operation to connect the image feature vector and the text feature vector head to tail to form a longer joint feature vector, or use a fully connected layer for linear transformation to allow the model to learn the weight relationship between different modalities and enhance the fusion effect.
[0072] Preferably, in order to capture the correlation between the image and the text at a deeper level, a self-attention mechanism or a cross-modal attention mechanism can be used to calculate the attention weights, and the image feature vector and the text feature vector are fused through the attention weights.
[0073] Furthermore, the semantic feature vector represents the deep semantic information of the text, while the image feature vector mainly describes the visual content. The fusion of the two can enhance the model's understanding of the relationship between the implicit information in the text and the image. After aligning the image feature vector and the semantic feature vector, input them into a preset cross-modal attention layer to determine the interaction weights between the image feature vector and the semantic feature vector; respectively perform context modeling on the image feature vector and the semantic feature vector through the multi-head self-attention mechanism, and capture the cross-modal semantic correlation features between the image feature vector and the semantic feature vector; perform weighted summation on the semantic correlation features according to the interaction weights to determine the joint cross-modal feature, and after fusing the joint cross-modal feature through a preset encoder, generate the semantic fusion feature.
[0074] Here, since the semantic feature vector may be extracted by large models such as GPT-4 and T5, its dimension may be different from that of the image feature, so dimension alignment is required. Dimensionality reduction transformation (such as a fully connected layer) or dimensionality increase expansion (such as linear mapping) can be used to make the dimensions of the two feature vectors the same. Use a cross-modal attention mechanism to calculate the attention weight matrix. Through the attention mechanism, the model can focus on the influence of the semantic feature on the image feature.
[0075] Specifically, first, image features (such as patch tokens or global features extracted by ViT) and text features (such as token embeddings extracted by RoBERTa) are used as inputs respectively. Then, through the cross-modal attention layer, the interaction weights between the image features and the text features are calculated. The multi-head self-attention mechanism is used to perform context modeling on the image and text features respectively, and capture the cross-modal semantic associations. Then, the interacted features are weighted and fused to generate a joint cross-modal representation. Finally, the fused features are further refined through multiple layers of Transformer encoders, and the semantically aligned image-text joint features are output for downstream tasks of harmful classification.
[0076] In this way, the image-text fusion features establish direct associations between images and texts through methods such as direct concatenation, linear transformation, attention mechanism, and cross-modal transformers. The semantic fusion features combine the deep semantics of the text and the image content through cross-modal attention, bidirectional interaction, adaptive weighting, etc., to enhance understanding. Finally, the obtained image-text fusion features and semantic fusion features can provide a more comprehensive multi-modal information representation, providing high-quality input data for subsequent harmful information classification.
[0077] S103. Fuse the image-text fusion feature and the semantic fusion feature to generate a target feature corresponding to the text and image data to be recognized, and input the target feature into a pre-trained harmful information classification model to determine a harmful information detection result corresponding to the text and image data to be recognized.
[0078] In a specific implementation, the purpose of fusing the image-text fusion feature and the semantic fusion feature is to construct a more comprehensive target feature representation for more accurate harmful information classification. The image-text fusion feature and the semantic fusion feature respectively contain direct relevant information between images and texts and deep association information between text semantics and images. To improve the classification performance, these two features need to be fused to generate the final target feature.
[0079] Specifically, referring to Figure 2 shown in the figure, it is a flowchart of another method for identifying harmful image-text information provided by an embodiment of the present disclosure. The method includes steps S1031 to S1033, where:
[0080] S1031. Input the target feature into the harmful information classification model.
[0081] S1032. Through the harmful information classification model, perform binary classification for harmful and harmless on the target feature to determine whether there is a harmful information detection result.
[0082] S1033. If the detection result of whether there is harmful information is that there is harmful information, then perform multi-classification for the harmful information according to the harmful information category through the harmful information classification model to determine the harmful type corresponding to the harmful information.
[0083] Here, the graphic-text fusion feature and the semantic fusion feature are spliced or weighted and fused to generate the final target feature, and the target feature is input into a classifier (such as a Softmax classifier) for binary classification (harmful / harmless) or multi-classification (harmful category recognition) of harmful information. The output result is whether it is harmful and the specific harmful category. The harmful classification categories for graphic-text descriptions usually include targeted harm, general offense, and decadent culture, etc.
[0084] In this way, through the design of front-end decoupling and back-end alignment, efficient collaborative processing of image and text modalities is achieved: the front-end uses pre-trained models to extract image and text features respectively, and the back-end fully mines the complementary information of graphic and text modalities through feature splicing and cross-modal fusion (including information interaction at the feature layer and semantic layer), improving the accuracy and robustness of harmful information recognition. Combining optical character recognition technology and text large models further enhances the ability to understand complex graphic and text content, and finally accurately determines harmful information and its category through a classifier, effectively improving the detection accuracy of harmful content.
[0085] A method for identifying graphic and text harmful information provided by an embodiment of the present disclosure obtains graphic and text data to be identified, extracts the image feature vector corresponding to the image modality data in the graphic and text data to be identified, and extracts the text feature vector and semantic feature vector corresponding to the text modality data in the graphic and text data to be identified; fuses the image feature vector and the text feature vector to generate a graphic-text fusion feature, and fuses the image feature vector and the semantic feature vector to generate a semantic fusion feature; fuses the graphic-text fusion feature and the semantic fusion feature to generate the target feature corresponding to the graphic and text data to be identified, and inputs the target feature into a pre-trained harmful information classification model to determine the detection result of harmful information corresponding to the graphic and text data to be identified. Adopting a strategy of front-end decoupling and back-end alignment, feature extraction is performed on images and texts respectively at the front-end, and information fusion is performed using a cross-modal attention mechanism at the back-end, thereby improving the ability to identify harmful information. While ensuring the feature expression ability, deep semantic alignment of graphic and text information is achieved, effectively improving the accuracy and generalization ability of multi-modal harmful information recognition.
[0086] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0087] Based on the same inventive concept, an apparatus for identifying graphic and text harmful information corresponding to the method for identifying graphic and text harmful information is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the apparatus in the embodiments of the present disclosure is similar to that of the above-mentioned method for identifying graphic and text harmful information in the embodiments of the present disclosure, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.
[0088] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an apparatus for identifying graphic and text harmful information provided by an embodiment of the present disclosure. As Figure 3 shown in
[0089] a feature extraction module 310, configured to obtain graphic and text data to be identified, extract an image feature vector corresponding to image modality data in the graphic and text data to be identified, and extract a text feature vector and a semantic feature vector corresponding to text modality data in the graphic and text data to be identified.
[0090] a feature fusion module 320, configured to fuse the image feature vector and the text feature vector to generate a graphic and text fusion feature, and fuse the image feature vector and the semantic feature vector to generate a semantic fusion feature.
[0091] a detection module 330, configured to fuse the graphic and text fusion feature and the semantic fusion feature to generate a target feature corresponding to the graphic and text data to be identified, input the target feature into a pre-trained harmful information classification model, and determine a harmful information detection result corresponding to the graphic and text data to be identified.
[0092] Descriptions of the processing flows of the modules in the apparatus and the interaction flows between the modules can refer to the relevant descriptions in the above method embodiments, and will not be elaborated here.
[0093] An image-text harmful information recognition device provided by an embodiment of the present disclosure obtains image-text data to be recognized, extracts an image feature vector corresponding to the image modality data in the image-text data to be recognized, and extracts a text feature vector and a semantic feature vector corresponding to the text modality data in the image-text data to be recognized; fuses the image feature vector and the text feature vector to generate an image-text fusion feature, and fuses the image feature vector and the semantic feature vector to generate a semantic fusion feature; fuses the image-text fusion feature and the semantic fusion feature to generate a target feature corresponding to the image-text data to be recognized, and inputs the target feature into a pre-trained harmful information classification model to determine a harmful information detection result corresponding to the image-text data to be recognized. By adopting a strategy of front-end decoupling and back-end alignment, image and text are respectively feature-extracted at the front end, and cross-modal attention mechanism is used for information fusion at the back end, thereby improving the recognition ability of harmful information. While ensuring the feature expression ability, deep semantic alignment of image-text information is achieved, effectively improving the accuracy and generalization ability of multi-modal harmful information recognition.
[0094] Corresponding to Figure 1 the image-text harmful information recognition method in [[]], an embodiment of the present disclosure further provides an electronic device 400, as Figure 4 shown, which is a schematic structural diagram of the electronic device 400 provided by an embodiment of the present disclosure, including:
[0095] a processor 41, a memory 42, and a bus 43; the memory 42 is used to store execution instructions, including an internal memory 421 and an external memory 422; here, the internal memory 421 is also called the main memory, which is used to temporarily store operation data in the processor 41 and data exchanged with the external memory 422 such as a hard disk. The processor 41 exchanges data with the external memory 422 through the internal memory 421. When the electronic device 400 runs, communication between the processor 41 and the memory 42 is carried out through the bus 43, so that the processor 41 executes Figure 1 the steps of the image-text harmful information recognition method in [[]].
[0096] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the image-text harmful information recognition method described in the above method embodiment. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0097] An embodiment of the present disclosure further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, they can execute the steps of the image-text harmful information recognition method described in the above method embodiment. For details, please refer to the above method embodiment, and details will not be repeated here.
[0098] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0099] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0100] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0101] In addition, in each embodiment of the present disclosure, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0102] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0103] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for identifying harmful graphic and text information, characterized in that Including: Obtain the text and image data to be recognized, extract the image feature vectors corresponding to the image modality data in the text and image data to be recognized, and extract the text feature vectors and semantic feature vectors corresponding to the text modality data in the text and image data to be recognized; Fuse the image feature vectors and the text feature vectors to generate text-image fusion features, and fuse the image feature vectors and the semantic feature vectors to generate semantic fusion features; Fuse the text-image fusion features and the semantic fusion features to generate target features corresponding to the text and image data to be recognized, and input the target features into a pre-trained harmful information classification model to determine the harmful information detection result corresponding to the text and image data to be recognized.
2. The method according to claim 1, wherein Extracting the image feature vectors corresponding to the image modality data in the text and image data to be recognized specifically includes: Segment the image modality data into multiple image patches, and flatten each image patch into a corresponding image vector; Map the image vectors to an embedding space through linear projection, and add corresponding position encodings to generate an image embedding sequence; Input the image embedding sequence into a pre-trained encoder, and perform multi-level feature extraction and interaction through a multi-head self-attention mechanism and a feed-forward neural network to generate the image feature vectors.
3. The method according to claim 1, characterized in that Extracting the text feature vectors and semantic feature vectors corresponding to the text modality data in the text and image data to be recognized specifically includes: Perform optical character recognition processing on the text modality data to extract the text content corresponding to the text modality data; Perform word segmentation and sub-word encoding on the text content to convert it into a token sequence, map the token sequence to an embedding space, and add corresponding position encodings to generate a text embedding sequence; Input the text embedding sequence into a pre-trained text model for context feature extraction, and capture the semantic relationships between tokens in the token sequence through a multi-head self-attention mechanism and a feed-forward neural network; Extract the feature representation output by the encoder corresponding to the text model as the text feature vector.
4. The method according to claim 3, wherein Extract the semantic feature vectors based on the following steps: Input the text content into a preset large text model; Through the large text model, perform natural language understanding on the text content to generate the semantic feature vectors at the semantic level.
5. The method according to claim 1, characterized in that, Fusing the image feature vectors and the text feature vectors to generate text-image fusion features specifically includes: Concatenate the image feature vectors and the text feature vectors to generate a joint feature vector; After performing dimensionality reduction or normalization processing on the joint feature vector, use an attention mechanism to perform feature fusion on the joint feature vector to generate the text-image fusion features.
6. The method according to claim 1, wherein Fusing the image feature vectors and the semantic feature vectors to generate semantic fusion features specifically includes: After aligning the image feature vectors and the semantic feature vectors, input them into a preset cross-modal attention layer to determine the interaction weights between the image feature vectors and the semantic feature vectors; Perform context modeling on the image feature vector and the semantic feature vector respectively through the multi-head self-attention mechanism, and capture the cross-modal semantic association features between the image feature vector and the semantic feature vector; Perform weighted summation on the semantic association features according to the interaction weights to determine the joint cross-modal features, and generate the semantic fusion features after fusing the joint cross-modal features through a preset encoder.
7. The method according to claim 1, wherein Input the target features into a pre-trained harmful information classification model to determine the harmful information detection result corresponding to the to-be-identified text and image data, specifically including: Input the target features into the harmful information classification model; Perform binary classification of harmful and harmless on the target features through the harmful information classification model to determine whether there is a harmful information detection result; If the result of whether there is a harmful information detection result is that there is harmful information, perform multi-classification of harmful information categories on the harmful information through the harmful information classification model to determine the harmful type corresponding to the harmful information.
8. An apparatus for identifying harmful graphic and text information, characterized in that, Including: A feature extraction module, configured to obtain to-be-identified text and image data, extract an image feature vector corresponding to image modal data in the to-be-identified text and image data, and extract a text feature vector and a semantic feature vector corresponding to text modal data in the to-be-identified text and image data; A feature fusion module, configured to fuse the image feature vector and the text feature vector to generate a text and image fusion feature, and fuse the image feature vector and the semantic feature vector to generate a semantic fusion feature; A detection module, configured to fuse the text and image fusion feature and the semantic fusion feature to generate target features corresponding to the to-be-identified text and image data, input the target features into a pre-trained harmful information classification model, and determine the harmful information detection result corresponding to the to-be-identified text and image data.
9. An electronic device, characterized in that, Including: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the text and image harmful information recognition method according to any one of claims 1 to 7 are executed.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the text and image harmful information recognition method according to any one of claims 1 to 7 are executed.
Citation Information
Cited By
Bill identification method, apparatus and device, and storage medium
CN120564218A
Bill identification method, device and equipment and storage medium
CN120564218B
Multi-modal picture understanding method and device based on cross-modal mark fusion
CN120611154A
Multimodal picture understanding method and device based on cross-modal label fusion
CN120611154B