Method and device for identifying full-modal harmful information and electronic equipment
Through task routing and multi-dimensional feature representation combined with full-modal confidence fusion and voting fusion methods, information loss and misjudgment problems in cross-modal and multi-modal harmful information recognition are solved, and high-precision harmful information recognition is achieved.
Patent Information
- Application Number
- CN202510334784.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively deal with harmful information that is fusion across modal and multimodal, resulting in information loss or misjudgment, and lacks a unified fusion strategy.
Through task routing, multimodal data is allocated to the corresponding modal processing branch, modal features of the data are extracted, and harmful information identification models based on identity, content, attribute categories or semantics are used, combined with full-modal confidence fusion and voting fusion methods to achieve high-precision harmful information identification.
It realizes high-precision harmful information identification for multimodal data, and improves detection accuracy and coverage.
Smart Images

Figure CN120264052A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of content review, and in particular, to a method, an apparatus, and an electronic device for identifying all-modal harmful information. Background Art
[0002] With the development of Internet technology, a large amount of user-generated content (UGC) has been generated on various online platforms such as social media, short video platforms, and news portals. This content covers multiple modalities such as text, images, audio, and video, and spreads rapidly worldwide.
[0003] Although single-modal harmful information detection technology has made certain progress in its respective field, when faced with cross-modal and multi-modal fusion harmful information, single-modal harmful information detection methods are difficult to effectively process cross-modal content, such as pictures with inciting text, videos combining audio and pictures to express extreme ideas, etc. The existing technology lacks a unified fusion strategy and is difficult to establish effective cross-modal associations between text, images, audio, and video, resulting in information loss or misjudgment. Summary of the Invention
[0004] Embodiments of the present disclosure at least provide a method, an apparatus, and an electronic device for identifying all-modal harmful information. Through task routing, the input multi-modal data is allocated to the corresponding modal processing branches, and harmful information identification is respectively performed on text, images, audio, and video. For the identification results of each modality, multi-dimensional feature representation is adopted, and combined with the all-modal confidence fusion method and the voting fusion method, high-precision harmful information identification is achieved.
[0005] Embodiments of the present disclosure provide a method for identifying all-modal harmful information, including:
[0006] Obtain multi-modal data to be identified, and allocate the data to be identified to the corresponding modal processing branches according to the corresponding modal types;
[0007] For each of the modal processing branches, extract the data modal features corresponding to the data to be identified, and input them into at least one harmful information identification model that meets the preset task requirements to determine the harmful information identification results, where each of the harmful information identification models has the ability to identify harmful information based on identity, or based on content, or based on attribute categories, or based on semantics;
[0008] Determine the identification result triples of the modality, task, and confidence corresponding to the harmful information identification results of each modality, and perform confidence fusion and voting fusion on the identification result triples of all modalities to determine the harmful information fusion discrimination result.
[0009] In an alternative embodiment, the data to be recognized at least includes audio data, image data, video data, and text data. Before allocating the data to be recognized to the corresponding modality processing branch according to the corresponding modality type, the method further includes:
[0010] Separating the audio track data and video frame data corresponding to the video data;
[0011] Performing image data preprocessing on the video frame data and the image data, including at least image normalization, image denoising, and image enhancement;
[0012] Performing audio data preprocessing on the audio track data and the audio data, including at least audio normalization, audio noise reduction, and audio segmentation;
[0013] Performing preprocessing on the text data, including at least word segmentation, stop word removal, and stemming.
[0014] In an alternative embodiment, extracting the data modality features corresponding to the data to be recognized and inputting them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition result, specifically including:
[0015] For the audio track data and the audio data, extracting audio features through Mel-frequency cepstral coefficients or a deep learning model;
[0016] Inputting the audio feature data into at least one of a preset voiceprint recognition model, a speech transcription model, and an audio harmful attribute classification model;
[0017] Using the voiceprint recognition model to determine the voiceprint identity harmful recognition result according to the audio features;
[0018] Using the speech transcription model to transcribe the corresponding speech content according to the audio features to determine the speech keyword harmful recognition result;
[0019] Using the audio attribute classification model to determine whether the audio features match the harmful information categories corresponding to the preset task requirements to determine the audio harmful attribute category recognition result.
[0020] In an alternative embodiment, extracting the data modality features corresponding to the data to be recognized and inputting them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition result, specifically further including:
[0021] For the video frame data and the image data, visual features are extracted through a convolutional neural network, and the visual features are input into at least one of a preset face recognition model, an image harmful attribute classification model, and an image harmful target detection model;
[0022] The face recognition model determines the identity of a specific person based on the visual features to obtain a harmful face identity recognition result;
[0023] The image harmful attribute classification model determines whether the visual features match the harmful information categories corresponding to the preset task requirements to obtain an image harmful attribute category recognition result;
[0024] The image harmful target detection model determines whether the visual features match the harmful targets corresponding to the preset task requirements to obtain an image harmful target recognition result.
[0025] In an optional implementation, data modality features corresponding to the data to be recognized are extracted and input into at least one harmful information recognition model that meets the preset task requirements to obtain a harmful information recognition result. Specifically, it further includes:
[0026] For the text data, semantic features are extracted through a bag-of-words model or a word embedding model, and the semantic features are input into at least one of a preset intention recognition model, a text harmful attribute classification model, and a natural language understanding model;
[0027] The intention recognition model identifies the type of harmful intention based on the semantic features to obtain a text harmful intention recognition result;
[0028] The text harmful attribute classification model determines whether the visual features match the harmful information categories corresponding to the preset task requirements to obtain a text harmful attribute category recognition result;
[0029] The natural language understanding model performs semantic analysis based on the semantic features to obtain a semantic harmful recognition result.
[0030] In an optional implementation, confidence fusion is performed on the recognition result triples of the full modality, specifically including:
[0031] Corresponding weighting coefficients are respectively assigned to the audio data, the image data, the video data, and the text data;
[0032] According to the weighting coefficients, weighted summation is performed on the recognition result triples corresponding to each modality data to obtain a confidence fusion discrimination result.
[0033] In an alternative embodiment, vote fusion is performed on the recognition result triples for all modalities, specifically including:
[0034] Correspondingly assign harmful thresholds to the audio data, the image data, the video data, and the text data respectively;
[0035] Set an indicator function for screening that the recognition result triples are greater than the corresponding harmful thresholds, sum the recognition result triples corresponding to each modality data after processing through the indicator function, and determine the vote fusion discrimination result.
[0036] The embodiments of the present disclosure also provide an apparatus for recognizing all-modal harmful information, including:
[0037] A task routing module, configured to obtain multi-modal data to be recognized, and allocate the data to be recognized to the corresponding modality processing branches according to the corresponding modality types;
[0038] A task processing module, configured to, for each of the modality processing branches, extract the data modality features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition result, where each of the harmful information recognition models has the ability to recognize harmful information based on identity, or based on content, or based on attribute categories, or based on semantics;
[0039] An all-modal fusion module, configured to determine recognition result triples corresponding to the modalities, tasks, and confidence levels of the harmful information recognition results for each modality, perform confidence fusion and vote fusion on the recognition result triples for all modalities, and determine the harmful information fusion discrimination result.
[0040] The embodiments of the present disclosure also provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the above-mentioned method for recognizing all-modal harmful information, or the steps in any possible implementation manner of the above-mentioned method for recognizing all-modal harmful information are executed.
[0041] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the above-mentioned method for recognizing all-modal harmful information, or the steps in any possible implementation manner of the above-mentioned method for recognizing all-modal harmful information are executed.
[0042] An embodiment of the present disclosure also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the above-mentioned method for identifying all-modal harmful information, or the steps in any possible implementation manner of the above-mentioned method for identifying all-modal harmful information are implemented.
[0043] An all-modal harmful information identification method, device, and electronic device provided by an embodiment of the present disclosure obtain multi-modal data to be identified, and allocate the data to be identified to corresponding modal processing branches according to the corresponding modal types; for each of the modal processing branches, extract the data modal features corresponding to the data to be identified, and input them into at least one harmful information identification model that meets the preset task requirements to determine the harmful information identification result, where each of the harmful information identification models has the ability to identify harmful information based on identity, or based on content, or based on attribute category, or based on semantics; determine the identification result triplets of the modal, task, and confidence corresponding to the harmful information identification results of each modal, and perform confidence fusion and voting fusion on the identification result triplets of all modalities to determine the harmful information fusion discrimination result. Through task routing, the input multi-modal data is allocated to the corresponding modal processing branches, and harmful information identification of text, images, audio, and video is performed respectively. For the identification results of each modality, multi-dimensional feature representation is adopted, and combined with the all-modal confidence fusion method and the voting fusion method, high-precision harmful information identification is achieved.
[0044] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0045] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 Shows a flowchart of an all-modal harmful information identification method provided by an embodiment of the present disclosure;
[0047] Figure 2 Shows a flowchart of another all-modal harmful information identification method provided by an embodiment of the present disclosure;
[0048] Figure 3The figure shows a schematic diagram of an identification device for all-modal harmful information provided by an embodiment of the present disclosure;
[0049] Figure 4 The figure shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some, rather than all, of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure that is required to be protected, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0051] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0052] As used herein, the term "and / or" merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0053] Through research, it is found that although the single-modal harmful information detection technology has made certain progress in its respective fields, when facing cross-modal and multi-modal fusion harmful information, the single-modal harmful information detection method is difficult to effectively process cross-modal content, such as pictures with inciting text, videos combining audio and pictures to express extreme ideas, etc. The prior art lacks a unified fusion strategy and is difficult to establish an effective cross-modal association between text, images, audio, and video, resulting in information loss or misjudgment.
[0054] Based on the above research, the present disclosure provides a method, apparatus, and electronic device for identifying all-modal harmful information. The method includes obtaining multi-modal data to be identified, and allocating the data to be identified to corresponding modal processing branches according to the corresponding modal types. For each modal processing branch, data modal features corresponding to the data to be identified are extracted and input into at least one harmful information identification model that meets the preset task requirements to determine the harmful information identification result, where each harmful information identification model has the ability to identify harmful information based on identity, or based on content, or based on attribute category, or based on semantics. Identification result triples corresponding to the modal, task, and confidence of the harmful information identification result for each modality are determined, and confidence fusion and voting fusion are performed on the identification result triples of all modalities to determine the harmful information fusion discrimination result. Through task routing, the input multi-modal data is allocated to corresponding modal processing branches, and harmful information identification of text, images, audio, and video is performed separately. For the identification results of each modality, multi-dimensional feature representation is adopted, and combined with the all-modal confidence fusion method and the voting fusion method, high-precision harmful information identification is achieved.
[0055] To facilitate the understanding of this embodiment, first, a method for identifying all-modal harmful information disclosed in the embodiments of the present disclosure is introduced in detail. The execution subject of the method for identifying all-modal harmful information provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such computer devices include, for example: terminal devices or servers or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the method for identifying all-modal harmful information may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0056] See Figure 1 As shown in the figure, it is a flowchart of a method for identifying all-modal harmful information provided in the embodiments of the present disclosure. The method includes steps S101 to S103, where:
[0057] S101. Obtain multi-modal data to be identified, and allocate the data to be identified to corresponding modal processing branches according to the corresponding modal types.
[0058] In a specific implementation, during the process of identifying all-modal harmful information, first, multi-modal data to be identified is obtained, and the data is allocated to corresponding modal processing branches according to the corresponding modal types. Multi-modal data usually comes from different input sources, mainly including text data (such as social media posts, forum comments, instant messaging messages, etc.), audio data (such as voice calls, voice messages, podcast content, etc.), image data (such as social media pictures, screenshots, advertising posters, etc.), and video data (such as short videos, surveillance videos, live streams, etc.).
[0059] Here, the data acquisition methods include but are not limited to user uploads (such as content submitted on social media platforms, forums, etc.), online scraping (such as crawling web content, obtaining streaming media data through API interfaces, etc.), device collection (such as microphone recording, camera shooting, sensor collection, etc.), and database storage (such as platform logs, historical chat records, etc.).
[0060] As a possible implementation, since data of different modalities has different encoding formats and storage methods, the system needs to first parse and preprocess the data for subsequent feature extraction and analysis. Refer to Figure 2 shown in the flowchart of another method for identifying all-modal harmful information provided by an embodiment of the present disclosure. The method includes steps S1011 to S1014, where:
[0061] S1011: Separate the audio track data and video frame data corresponding to the video data.
[0062] S1012: Perform image data preprocessing on the video frame data and the image data, including at least image normalization, image denoising, and image enhancement.
[0063] S1013: Perform audio data preprocessing on the audio track data and the audio data, including at least audio normalization, audio noise reduction, and audio segmentation.
[0064] S1014: Perform preprocessing on the text data, including at least word segmentation, stop word removal, and stemming.
[0065] In specific implementation, for text data processing, first parse the text encoding format (such as UTF-8, GBK, etc.) and identify the text language (such as Chinese, English, etc.), and then perform character-level and word-level cleaning (such as removing HTML tags, special symbols, stop words, etc.). For audio data processing, first parse the audio file format (such as MP3, WAV, AAC, etc.), extract the audio waveform data, and perform noise reduction processing, and then perform voice activity detection (VAD) to remove silent segments. For image data processing, first parse the image format (such as JPEG, PNG, BMP, etc.), and then perform image normalization (such as resizing, converting color space), and finally perform image denoising and enhancement (such as filtering, sharpening, etc.). For video data processing, first parse the video encoding format (such as MP4, AVI, H.264, etc.), and then extract key frame images and audio track data and perform video content analysis (such as inter-frame deduplication, shot segmentation, etc.).
[0066] Furthermore, after data parsing and preprocessing are completed, the system needs to identify the modal type of the data and allocate the data to the corresponding modal processing branch. For example, text files (such as TXT, JSON, CSV) are classified as text modality; audio files (such as WAV, MP3) are classified as audio modality; image files (such as PNG, JPG) are classified as image modality; video files (such as MP4, AVI) are classified as video modality.
[0067] Optionally, determine the data type through the MIME type (such as text / plain, audio / mpeg, etc.), or further subdivide the modality through content analysis (such as audio signal detection, OCR recognition, frame rate detection, etc.). It should be noted that if the data contains multiple modalities (such as a video containing both an audio track and image frames), it is split into different modalities and processed separately.
[0068] Here, the text processing branch is used for text feature extraction and text harmful information recognition; the audio processing branch is used for speech transcription, voiceprint recognition, and audio harmful content detection; the image processing branch is used for face recognition, object detection, image attribute classification, etc.; the video processing branch is used for video frame analysis, shot detection, video object recognition, etc.
[0069] Among them, for data with composite modalities, such as video, it may be assigned to multiple branches at the same time. For example, the audio track data is assigned to the audio processing branch, and the video frame data is assigned to the image processing branch.
[0070] Finally, after the data is distributed to different modality processing branches, each branch performs specific feature extraction and harmful information recognition respectively, and then the recognition results are aggregated for final confidence fusion and voting fusion. For example, assume a video contains possible harmful information. The audio branch detects sensitive speech content, the image branch detects prohibited signs, and the text branch extracts potential illegal content from the subtitles. These recognition results will be combined, and the final determination result of harmful information will be obtained through subsequent fusion strategies (such as weighted confidence calculation, voting mechanism).
[0071] S102. For each of the modality processing branches, extract the data modality features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements, and determine the harmful information recognition result, where each of the harmful information recognition models has the ability to recognize harmful information based on identity, or based on content, or based on attribute category, or based on semantics.
[0072] In specific implementation, each modality processing branch is responsible for extracting data features of a specific modality and inputting them into a harmful information recognition model that meets the task requirements. This process involves multi-modal data feature extraction (feature acquisition for text, audio, image, video); harmful information recognition model matching (classification based on identity, content, attribute category, semantics); harmful information determination (fusing the results of multiple models to determine the final determination). Since data of different modalities have different structures and information characteristics, corresponding modality features need to be extracted respectively for input into a suitable harmful information recognition model.
[0073] Here, for the data of the text modality, semantic features are extracted through a bag-of-words model or a word embedding model, and the semantic features are input into at least one of a preset intention recognition model, a text harmful attribute classification model, and a natural language understanding model; the intention recognition model identifies the harmful intention type according to the semantic features to determine the text harmful intention recognition result; the text harmful attribute classification model determines whether the visual features match the harmful information category corresponding to the preset task requirements to determine the text harmful attribute category recognition result; the natural language understanding model performs semantic analysis according to the semantic features to determine the semantic harmful recognition result.
[0074] Among them, the text data includes social media posts, comments, messages, articles, etc. The feature extraction methods include: extraction of lexical features, such as TF-IDF (term frequency-inverse document frequency), word vectors (Word2Vec, FastText), BERT pre-trained word vectors; extraction of syntactic features, such as dependency syntactic analysis, part-of-speech tagging, sentence length statistics, etc.; extraction of sentiment features, such as positive and negative sentiment tendency analysis, subjectivity determination, etc.; entity recognition, such as extraction of user identity information (such as person names, place names, organization names), illegal words, sensitive topics, etc.; semantic feature extraction, such as using Transformer models (BERT, RoBERTa) for deep semantic encoding.
[0075] Here, for the data of the audio modality, audio features are extracted through Mel-frequency cepstral coefficients or deep learning models; the audio feature data is input into at least one of a preset voiceprint recognition model, a speech transcription model, and an audio harmful attribute classification model; through the voiceprint recognition model, specific task timbre determination is performed based on the audio features to determine the harmful recognition result of the voiceprint identity; through the speech transcription model, speech transcription recognition of the corresponding speech content is performed based on the audio features to determine the harmful recognition result of the speech keywords; through the audio attribute classification model, it is determined whether the audio features match the harmful information category corresponding to the preset task requirements to determine the harmful attribute category recognition result of the audio.
[0076] Among them, the audio data mainly involves speech content, intonation, background sound, etc. The feature extraction methods include speech transcription: converting audio to text (ASR, such as Whisper, Wav2Vec2.0); spectral analysis: MFCC (Mel-frequency cepstral coefficients), STFT (short-time Fourier transform); timbre features: voiceprint recognition (used to distinguish different speakers), emotion recognition (such as anger, fear, etc.); noise detection: distinguishing normal environmental sounds from abnormal audio (such as gunshots, explosions).
[0077] Here, for the data of the image modality, visual features are extracted through a convolutional neural network, and the visual features are input into at least one of a preset face recognition model, an image harmful attribute classification model, and an image harmful target detection model; through the face recognition model, specific person identity determination is performed based on the visual features to determine the harmful recognition result of the face identity; through the image harmful attribute classification model, it is determined whether the visual features match the harmful information category corresponding to the preset task requirements to determine the harmful attribute category recognition result of the image; through the image harmful target detection model, it is determined whether the visual features match the harmful target corresponding to the preset task requirements to determine the harmful target recognition result of the image.
[0078] Among them, the image data covers social media pictures, advertisements, emoticons, etc. The feature extraction methods include object detection: detecting faces, objects, Logos, QR codes, etc. (YOLO, Faster R-CNN); texture features: SIFT (Scale-Invariant Feature Transform), HOG (Histogram of Oriented Gradients); content recognition: OCR (Optical Character Recognition) is used to detect text content and analyze prohibited words or slogans; porn and violence detection: ResNet, EfficientNet are used for classification and determination.
[0079] Here, for the data in the video modality, here, the video data integrates images and audio, and multi-modal features need to be extracted separately, including key frame extraction: detecting scene changes based on histogram changes to reduce the amount of calculation; action recognition: extracting human actions (such as fighting, detecting attack behaviors); voice content analysis: extracting the audio track and combining text analysis for comprehensive determination.
[0080] It should be noted that the audio track data separated from the video modality data can be processed in the same way as the audio modality, and the video frame data separated from the video modality data can be processed in the same way as the image modality.
[0081] Furthermore, the features extracted from each modality are input into a preset harmful information recognition model, and the model can be classified into types based on identity recognition, content recognition, attribute type recognition, and semantic recognition. For the model based on identity recognition, it is determined whether it belongs to a sensitive group through user identity information, voiceprint, fingerprint, etc. for recognition, and it is applied to the data in the audio, text, and image modalities. The tasks it processes mainly include: identifying specific users, blacklisted groups, etc. For the model based on content recognition, it is detected by directly matching prohibited words, image features, or audio features, and it is applied to the data in the audio, text, image, and video modalities. The tasks it processes mainly include identifying illegal remarks, pornographic content, violent content, etc.; for the model based on attribute category recognition, it is detected through the classification labels of the content (such as "politically sensitive" / "fraud"), and it is applied to the data in the text, image, and audio modalities. The tasks it processes mainly include identifying fraud information, political rumors, etc.; for the model based on semantic recognition, it judges potential implicit or distorted expressions through semantic understanding, and it is applied to the data in the text and audio modalities. The tasks it processes mainly include identifying implicit illegal expressions (such as secret codes, homophonic puns), etc.
[0082] Specifically, in identity-based harmful information recognition, for data in text modality, based on the user's historical remarks and social relationship network, it is determined whether the data comes from a risky user; for data in audio modality, through voiceprint recognition, it is confirmed whether a specific speaker is on the blacklist; for data in image modality, through face recognition, it is matched with sensitive personnel in the database. In content-based harmful information recognition, for data in text modality, keyword matching (such as "illegal transaction", "terrorist attack") and deep language model recognition (BERT, GPT) are adopted; for data in audio modality, keyword matching and audio fingerprint library detection (such as gunshots, explosions) are adopted; for data in image modality, pornographic, violent, and prohibited items (such as drugs, weapons) are detected. In attribute-category-based harmful information recognition, for data in text modality, the text is classified (such as "hate speech", "fraudulent information"), and a classification model (such as LSTM, BERT) is trained; for data in image modality, the image content is classified, such as advertisement fraud and false news illustrations. In semantics-based harmful information recognition, for data in text modality, knowledge graph analysis is used to analyze implicit expressions or a Transformer semantic analysis model is adopted; for data in audio modality, deformed language or suggestive content is detected through audio semantic analysis.
[0083] Here, the results of the harmful information recognition model may be judgments with different confidence levels. To improve accuracy, a multi-model fusion strategy is usually adopted. The judgment results of harmful information can be: category (such as "hate speech", "pornographic content"); confidence level (such as "80% certain"); evidence (such as "matched prohibited word 'terrorist attack'").
[0084] In this way, in the process of multi-modal harmful information recognition, data modality features (deep features of text, audio, image, video) are extracted, and a suitable harmful information recognition model (based on identity, content, attribute category, semantics) is matched. The final result is determined through multi-model fusion (confidence level fusion, voting mechanism). This process ensures that the system can accurately identify different types of harmful information and improve the detection accuracy and coverage.
[0085] S103. Determine the recognition result triple corresponding to the modality, task, and confidence level of the harmful information recognition result of each modality, and perform confidence level fusion and voting fusion on the recognition result triples of all modalities to determine the harmful information fusion discrimination result.
[0086] In specific implementation, the processing branches of different modalities will generate their own recognition results. To improve the overall recognition accuracy and robustness, the recognition results of multiple modalities need to be fused. The recognition model of each modality processing branch will output a recognition result in the format of a recognition result triple: (modality, task, confidence level).
[0087] Here, the modality indicates which modality the recognition result belongs to (text, audio, image, video, etc.); the task indicates the category of the recognition task, such as hate speech detection, pornographic recognition, fraud detection, etc.; the confidence indicates the confidence of the model in the recognition result (0 - 1 or 0% - 100%).
[0088] Specifically, due to the different data expression methods of different modalities, the confidence in the same task may vary across different modalities. Therefore, confidence fusion is required. Corresponding weighting coefficients are assigned to audio data, image data, video data, and text data respectively; based on the weighting coefficients, weighted summation is performed on the recognition result triplets corresponding to each modality data to determine the confidence fusion discrimination result.
[0089] Here, the confidence fusion discrimination result can be calculated through the following formula:
[0090] F s_all =W t *TL(M t ,T t ,C t )+W i *TL(M i ,T i ,C i )+W a *TL(M a ,T a ,C a )+W v *TL(M v ,T v ,C v )
[0091] Among them, F s_all represents the confidence fusion discrimination result; TL(M t ,T t ,C t ), TL(M i ,T i ,C i ), TL(M a ,T a ,C a ), and TL(M v ,T v ,C v ) represent the recognition result triplets corresponding to text, image, audio, and video modality data respectively; TL represents the triplet Triplet, M represents the modality, T represents the task, C represents the confidence, t represents Text, i represents the image, a represents the audio, v represents the video; W t 、W i 、Wa and W v represent the weighting coefficients of text, image, audio, and video, respectively.
[0092] It should be noted that the default weighting coefficients can be set to equal weights or empirically set according to actual needs. The confidence fusion discrimination result needs to be compared with a preset confidence score threshold to give a discrimination result of whether it is harmful.
[0093] Furthermore, the purpose of voting fusion is to make a decision based on the recognition results of multiple modalities to ensure the stability of the overall judgment. Harmfulness thresholds are assigned to audio data, image data, video data, and text data respectively; an indicator function is set to filter out the recognition result triples greater than the corresponding harmfulness thresholds, and the recognition result triples corresponding to each modality data are processed through the indicator function and then summed to determine the voting fusion discrimination result.
[0094] Here, the voting fusion discrimination result can be calculated by the following formula:
[0095] F h_all = I(TL(M t , T t , C t ) > θ t ) + I(TL(M i , T i , C i ) > θ i ) + I(TL(M a , T a , C a ) > θ a )
[0096] + I(TL(M v , T v , C v ) > θ v )
[0097] where F h_all represents the voting fusion discrimination result; TL(M t , T t , C t ), TL(M i , T i , C i ), TL(M a , T a , C a ), and TL(M v , T v , C vrespectively represent the recognition result triplets corresponding to the text, image, audio, and video modality data; TL represents the triplet Triplet, M represents the modality, T represents the task, C represents the confidence, t represents text, i represents image, a represents audio, v represents video; I(·) represents the indicator function, which takes the value of 1 when the condition is true and 0 otherwise; θ t 、θ i 、θ a 、θ v respectively represent the harmfulness thresholds corresponding to the text, image, audio, and video modality data.
[0098] It should be noted that the harmfulness threshold can be set accordingly according to actual needs, and no specific limitation is made here. The voting fusion discrimination result needs to be compared with the preset voting score threshold to give the discrimination result of whether it is harmful. Either the confidence fusion discrimination result or the voting fusion discrimination result can be selected as the final harmful information fusion discrimination result.
[0099] As a possible implementation manner, after confidence fusion and voting fusion, the final harmful information determination can adopt a dual fusion strategy. First, use confidence fusion to calculate a weighted comprehensive confidence, and then make a final judgment based on the result of voting fusion. If the voting fusion results are consistent (all modalities are judged the same), then directly output the result; if the voting fusion results are inconsistent, when the weighted comprehensive confidence is greater than the confidence threshold, output the result with a higher confidence. If the weighted comprehensive confidence is not greater than the confidence threshold, it is determined that the recognition result of harmful information is uncertain or needs to be selected manually.
[0100] A method for identifying all-modal harmful information provided by an embodiment of the present disclosure includes obtaining multi-modal data to be recognized, and allocating the data to be recognized to corresponding modality processing branches according to the corresponding modality types; for each of the modality processing branches, extracting the data modality features corresponding to the data to be recognized and inputting them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition result, where each of the harmful information recognition models has the ability to recognize harmful information based on identity, or based on content, or based on attribute category, or based on semantics; determining the recognition result triplets corresponding to the modality, task, and confidence of the harmful information recognition results of each modality, and performing confidence fusion and voting fusion on the recognition result triplets of all modalities to determine the harmful information fusion discrimination result. Through task routing, the input multi-modal data is allocated to the corresponding modality processing branches, and the harmful information of text, image, audio, and video is recognized respectively. For the recognition results of each modality, multi-dimensional feature representation is adopted, and combined with the all-modal confidence fusion method and the voting fusion method, high-precision harmful information recognition is achieved.
[0101] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0102] Based on the same inventive concept, an apparatus for identifying all-modal harmful information corresponding to the method for identifying all-modal harmful information is further provided in the embodiments of the present disclosure. Since the principle of solving problems by the apparatus in the embodiments of the present disclosure is similar to the above method for identifying all-modal harmful information in the embodiments of the present disclosure, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.
[0103] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an apparatus for identifying all-modal harmful information provided by an embodiment of the present disclosure. As Figure 3 shown in
[0104] The task routing module 310 is configured to obtain multi-modal data to be identified and allocate the data to be identified to corresponding modal processing branches according to the corresponding modal types.
[0105] The task processing module 320 is configured to, for each of the modal processing branches, extract the data modal features corresponding to the data to be identified and input them into at least one harmful information identification model that meets the preset task requirements, and determine the harmful information identification result, where each of the harmful information identification models has the ability to identify harmful information based on identity, or based on content, or based on attribute category, or based on semantics.
[0106] The all-modal fusion module 330 is configured to determine an identification result triple corresponding to the modal, task, and confidence of the harmful information identification result of each modal, perform confidence fusion and voting fusion on the identification result triples of all modalities, and determine the harmful information fusion discrimination result.
[0107] The description of the processing flow of each module in the apparatus and the interaction flow between the modules can refer to the relevant description in the above method embodiments, and will not be elaborated here.
[0108] An identification device for all-modal harmful information provided by an embodiment of the present disclosure obtains multi-modal data to be identified, and distributes the data to be identified to corresponding modal processing branches according to the corresponding modal types; for each modal processing branch, extracts the data modal features corresponding to the data to be identified, and inputs them into at least one harmful information identification model that meets the preset task requirements to determine the harmful information identification result, where each harmful information identification model has the ability to identify harmful information based on identity, or based on content, or based on attribute category, or based on semantics; determines the identification result triple corresponding to the modal, task, and confidence of the harmful information identification result of each modal, and performs confidence fusion and voting fusion on the identification result triple of all modalities to determine the harmful information fusion discrimination result. Through task routing, the input multi-modal data is distributed to the corresponding modal processing branches, and the harmful information of text, image, audio, and video is identified respectively. For the identification results of each modality, multi-dimensional feature representation is adopted, and combined with the all-modal confidence fusion method and the voting fusion method, high-precision harmful information identification is achieved.
[0109] Corresponding to Figure 1 In the all-modal harmful information identification method, an embodiment of the present disclosure also provides an electronic device 400, as Figure 4 shown, which is a schematic structural diagram of the electronic device 400 provided by an embodiment of the present disclosure, including:
[0110] A processor 41, a memory 42, and a bus 43; the memory 42 is used to store execution instructions, including an internal memory 421 and an external memory 422; the internal memory 421 here is also called the main memory, which is used to temporarily store the operation data in the processor 41 and the data exchanged with the external memory 422 such as a hard disk. The processor 41 exchanges data with the external memory 422 through the internal memory 421. When the electronic device 400 runs, the processor 41 communicates with the memory 42 through the bus 43, so that the processor 41 executes Figure 1 And Figure 2 The steps of the all-modal harmful information identification method in
[0111] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the all-modal harmful information identification method described in the above method embodiment. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0112] The embodiments of the present disclosure also provide a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the steps of the method for identifying all-modal harmful information described in the above method embodiments can be executed. For details, reference can be made to the above method embodiments and will not be elaborated herein.
[0113] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0115] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0117] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0118] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for identifying all-modal harmful information, characterized in that, Including: Obtain multi-modal data to be recognized, and allocate the data to be recognized to corresponding modal processing branches according to the corresponding modal types; For each of the modal processing branches, extract the data modal features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition results. Among them, each harmful information recognition model has the ability to recognize harmful information based on identity, or based on content, or based on attribute category, or based on semantics; Determine the recognition result triples corresponding to the modal, task, and confidence of the harmful information recognition results of each modality, and perform confidence fusion and voting fusion on the recognition result triples of all modalities to determine the harmful information fusion discrimination result.
2. The method according to claim 1, characterized in that, The data to be recognized includes at least audio data, image data, video data, and text data. Before allocating the data to be recognized to the corresponding modal processing branches according to the corresponding modal types, the method further includes: Separate the audio track data and video frame data corresponding to the video data; Perform image data preprocessing on the video frame data and the image data, including at least image normalization, image denoising, and image enhancement; Perform audio data preprocessing on the audio track data and the audio data, including at least audio normalization, audio noise reduction, and audio segmentation; Perform preprocessing on the text data, including at least word segmentation, stop word removal, and stemming.
3. The method according to claim 2, wherein Extract the data modal features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition results, specifically including: For the audio track data and the audio data, extract audio features through Mel Frequency Cepstral Coefficients or deep learning models; Input the audio feature data into at least one of a preset voiceprint recognition model, a speech transcription model, and an audio harmful attribute classification model; Through the voiceprint recognition model, determine the harmful recognition result of the voiceprint identity according to the audio features for specific task tone determination; Through the speech transcription model, transcribe the corresponding speech content according to the audio features to determine the harmful recognition result of the speech keywords; Through the audio attribute classification model, determine whether the audio features match the harmful information category to be recognized corresponding to the preset task requirements to determine the audio harmful attribute category recognition result.
4. The method according to claim 2, wherein Extract the data modal features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements to determine the harmful information recognition results, specifically further including: For the video frame data and the image data, extract visual features through a convolutional neural network, and input the visual features into at least one of a preset face recognition model, an image harmful attribute classification model, and an image harmful target detection model; Through the face recognition model, determine the harmful recognition result of the face identity according to the visual features for specific person identity determination; Determine whether the visual features match the harmful information categories identified corresponding to the preset task requirements through the image harmful attribute classification model, and determine the image harmful attribute category recognition result; Determine whether the visual features match the harmful targets identified corresponding to the preset task requirements through the image harmful target detection model, and determine the image harmful target recognition result.
5. The method according to claim 2, wherein Extract the data modality features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements, and determine the harmful information recognition result. Specifically, it further includes: For the text data, extract semantic features through a bag-of-words model or a word embedding model, and input the semantic features into at least one of a preset intention recognition model, a text harmful attribute classification model, and a natural language understanding model; Through the intention recognition model, perform harmful intention type recognition according to the semantic features, and determine the text harmful intention recognition result; Determine whether the visual features match the harmful information categories identified corresponding to the preset task requirements through the text harmful attribute classification model, and determine the text harmful attribute category recognition result; Perform semantic analysis according to the semantic features through the natural language understanding model, and determine the semantic harmful recognition result.
6. The method according to claim 2, wherein Perform confidence fusion on the recognition result triples of the full modality, specifically including: Allocate corresponding weighting coefficients to the audio data, the image data, the video data, and the text data respectively; Perform weighted summation on the recognition result triples corresponding to each modality data according to the weighting coefficients, and determine the confidence fusion discrimination result.
7. The method according to claim 2, characterized in that Perform voting fusion on the recognition result triples of the full modality, specifically including: Allocate corresponding harmfulness thresholds to the audio data, the image data, the video data, and the text data respectively; Set an indicator function for screening that the recognition result triples are greater than the corresponding harmfulness thresholds, and sum the recognition result triples corresponding to each modality data after processing through the indicator function to determine the voting fusion discrimination result.
8. An identification device for all-modal harmful information, characterized in that, It includes: A task routing module, configured to obtain multi-modal data to be recognized, and allocate the data to be recognized to corresponding modality processing branches according to the corresponding modality types; A task processing module, configured to, for each of the modality processing branches, extract the data modality features corresponding to the data to be recognized, and input them into at least one harmful information recognition model that meets the preset task requirements, and determine the harmful information recognition result, where each of the harmful information recognition models has the ability to recognize harmful information based on identity, or based on content, or based on attribute categories, or based on semantics; A full modality fusion module, configured to determine the recognition result triples corresponding to the modality, task, and confidence of the harmful information recognition results of each modality, perform confidence fusion and voting fusion on the recognition result triples of the full modality, and determine the harmful information fusion discrimination result.
9. An electronic device, characterized in that, It includes: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the method for identifying all-modal harmful information according to any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the method for identifying all-modal harmful information according to any one of claims 1 to 7 are performed.
Citation Information
Cited By
Sensitive data identification method and device, computer equipment and readable storage medium
CN121121233A
Identification method, device and equipment for multi-modal fusion research and judgment of pornographic scene and medium
CN121412730A