Abnormal recognition method and device, computing equipment and computer readable storage medium
An anomaly identification method that extracts content units and adjusts the correlation influence of published content solves the problem of insufficient robustness caused by independent model execution in existing technologies, and achieves efficient and accurate identification of multi-anomaly domain data.
Patent Information
- Application Number
- CN202511532734.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-02-03
AI Technical Summary
In existing technologies, anomaly category identification models for multi-anomaly domain data execute independently, lack dynamic communication, and cannot identify anomalies for different components, resulting in insufficient robustness.
By extracting content units from published content and dynamically adjusting the anomaly detection strategy based on the correlation and influence between various anomaly detection tasks, multiple anomaly detection tasks are executed for the content units to generate overall anomaly detection results.
It enables refined decomposition and efficient identification of multimodal published content, improving the accuracy and robustness of identification in complex and abnormal scenarios.
Smart Images

Figure CN121456740A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of anomaly recognition technology, and in particular to anomaly recognition method, anomaly recognition device, a computing device, and a computer-readable storage medium. Background Technology
[0002] With the development of science and technology, the release of content containing multimodal data is becoming increasingly widespread. As a result, higher requirements are being placed on the accuracy and comprehensiveness of anomaly identification during the review process of the released content.
[0003] Currently, anomaly category identification for multi-anomaly domain data mainly relies on the structure of multiple independent models or shared parameter models. Based on the complete anomaly labeling of the input data under each anomaly category, the anomaly category identification of the data is achieved by using the anomaly category-level classification results generated by each model.
[0004] However, in the aforementioned technical solutions, each identification model independently performs the task of identifying anomalies in the overall data. There is a lack of dynamic communication between the models, and they cannot specifically identify anomaly categories for different components of the data. This limits their ability to identify heterogeneous anomaly types and affects the robustness of the overall identification. Therefore, a more accurate and comprehensive data anomaly identification method is urgently needed. Summary of the Invention
[0005] In view of this, embodiments of this specification provide an anomaly identification method. One or more embodiments of this specification also relate to an anomaly identification device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, an anomaly identification method is provided, comprising:
[0007] Obtain the target published content and extract each content unit from the target published content;
[0008] Based on the correlation between the target content unit and various anomaly identification tasks, various anomaly identification tasks are performed on the target content unit to obtain the anomaly information corresponding to each anomaly identification task. The target content unit is any one of the content units.
[0009] For target identification tasks in various anomaly identification tasks, the anomaly identification results of the target published content are determined based on the anomaly information of each content unit in the target identification task.
[0010] According to a second aspect of the embodiments of this specification, an anomaly detection device is provided, comprising:
[0011] The extraction module is configured to acquire the target published content and extract each content unit from the target published content.
[0012] The execution module is configured to perform various anomaly recognition tasks on the target content unit based on the correlation influence between the target content unit and various anomaly recognition tasks, and obtain the anomaly information corresponding to each anomaly recognition task. The target content unit is any one of the content units.
[0013] The determination module is configured to determine the anomaly identification results of the target published content based on the anomaly information of each content unit in various anomaly identification tasks.
[0014] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0015] Memory and processor;
[0016] The memory is used to store computer-executable instructions, and the processor is used to execute the computer program / instructions, which, when executed by the processor, implement the steps of the above-described anomaly identification method.
[0017] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described anomaly identification method.
[0018] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described anomaly identification method.
[0019] One embodiment of this specification implements a data anomaly identification method, comprising: acquiring target published content and extracting each content unit from the target published content; performing various anomaly identification tasks on the target content unit based on the correlation influence between the target content unit and various anomaly identification tasks, and obtaining anomaly information corresponding to each anomaly identification task, wherein the target content unit is any one of the content units; and determining the anomaly identification result of the target published content based on the anomaly information of each content unit for the target identification task in the various anomaly identification tasks.
[0020] By extracting content units from the target published content and performing multiple anomaly recognition tasks on the target units within each content unit, the anomaly recognition strategy can be dynamically adjusted based on the correlation between the target content unit and the anomaly recognition task. This allows for the capture of key features of different anomaly domains at a fine-grained level at the content unit level. Simultaneously, by independently analyzing the anomaly information of the target content unit for each anomaly recognition task and then further synthesizing it to generate the overall recognition result of the target published content, the system achieves adaptation to the differences in multiple anomaly domains. This enhances the fusion and coordination between heterogeneous anomaly tasks, thereby improving the accuracy and robustness of recognizing multimodal published content in complex anomaly scenarios. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of an anomaly identification method in the prior art provided by one embodiment of this specification;
[0022] Figure 2 This is a flowchart of an anomaly identification method provided in one embodiment of this specification;
[0023] Figure 3 This is a schematic diagram illustrating the execution process of an anomaly identification method provided in one embodiment of this specification;
[0024] Figure 4 This is a flowchart illustrating the processing procedure of an anomaly identification method provided in one embodiment of this specification;
[0025] Figure 5 This is a schematic diagram of a client-side display for identifying abnormal published content, provided in one embodiment of this specification.
[0026] Figure 6 This is a schematic diagram of the structure of an anomaly detection device provided in one embodiment of this specification;
[0027] Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0028] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0029] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0030] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0031] Furthermore, it should be noted that the data involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0032] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0033] Large Language Model (LLM) is a natural language processing model based on deep learning. It is trained on a large-scale corpus and has the ability to generate and understand text.
[0034] A feed-forward network (FFN) is a neural network consisting of an input layer, hidden layers, and an output layer. Data flows in one direction and there are no recurrent or feedback connections.
[0035] Mixture of Experts (MoE) is an architecture that dynamically selects and combines multiple independent expert models through a gating mechanism to efficiently handle complex tasks and improve scalability.
[0036] Average pooling is an operation that calculates the average value of local regions of input data to reduce spatial dimensionality. It is often used for feature dimensionality reduction in convolutional neural networks.
[0037] Optical Character Recognition (OCR) is a technology that detects and extracts text information from images, and is often used for document digitization or scene text parsing.
[0038] Automatic Speech Recognition (ASR) is a technology that converts speech signals into text and is widely used in scenarios such as voice assistants and speech transcription.
[0039] Binary Cross-Entropy Loss (BCE) is a loss function that measures the difference between two probability distributions. In classification tasks, it is used to optimize the matching degree between the model's prediction results and the true labels.
[0040] A linear classifier is a simple and efficient model that achieves classification by calculating a linear combination of input features and weight vectors.
[0041] With the development of science and technology, the release of content containing multimodal data is becoming increasingly widespread. As a result, higher requirements are being placed on the accuracy and comprehensiveness of anomaly identification during the review process of the released content.
[0042] Currently, anomaly category identification for multi-anomaly domain data mainly relies on the structure of multiple independent models or shared parameter models. Based on the complete anomaly labeling of the input data under each anomaly category, the anomaly category identification of the data is achieved by using the anomaly category-level classification results generated by each model.
[0043] Specifically, in the existing technology, the mainstream modeling methods for multi-anomaly identification tasks can be roughly divided into three categories: parameter sharing-based methods, multi-task learning methods based on task decoupling, and multi-label modeling methods based on Mixture of Experts (MoE) mechanisms.
[0044] Among them, the parameter-sharing method predicts all labels by directly sharing a feature extractor and a multi-output classifier head, and its modeling assumption is that all labels can share the same semantic space.
[0045] The multi-task learning method based on task decoupling transforms multi-anomaly identification into multiple independent tasks, and trains multiple networks for multiple tasks, sharing gradients only in some layers.
[0046] In recent years, hybrid expert mechanisms have been widely adopted in multi-task learning and multi-anomaly identification, such as using the HSQ (Hybrid Sharing Query) model for anomaly identification. This method treats each label as a task, combines a task-specific expert network model with a shared expert MoE structure, and uses a gating mechanism to control feature flow and mitigate inter-task interference.
[0047] However, despite the technological advancements achieved by the HSQ model in image anomaly detection tasks, it still faces the following challenges in content anomaly detection tasks:
[0048] 1) Inability to handle the problem of missing labels and inconsistency between anomaly identification rules: The HSQ model assumes that each sample is supervised on all labels. In actual anomaly identification data, missing labels are common, and direct supervision will introduce pseudo-label interference.
[0049] 2) Poor collaborative modeling of anomaly detection models, unable to adapt to the differences in multimodal and multiple anomaly domains: HSQ adopts a label-level anomaly detection model design, assuming that all labels share the same input information. It is difficult to collaboratively focus on different key tokens according to the anomaly detection queue, which limits the model's ability to refine modeling in heterogeneous anomaly domains. Moreover, the model only processes images and cannot integrate text and image content, making it unsuitable for anomaly detection tasks involving multimodal content such as text and image notes.
[0050] Specifically, see Figure 1 , Figure 1 A schematic diagram of a prior art anomaly identification method according to an embodiment of this specification is shown, such as... Figure 1 As shown.
[0051] When identifying anomalies in published content, manually labeling the identification queues for each anomaly type often results in incomplete coverage of all anomaly types and inconsistent rules between different anomaly identification queues. For example, only some anomaly identification queues, such as the "Text Content Anomaly Identification Queue," "Video Style Anomaly Identification Queue," and "Content Cover Anomaly Identification Queue," have labeled results for image and text information included in the published content. Other queues, such as the "Inappropriate Content Anomaly Identification Queue" and the "Non-Original Video Notes Anomaly Identification Queue," lack labeled information. Furthermore, the judgment criteria differ between different anomaly identification queues. For instance, the published content might be judged "Pass" in the "Text Content Anomaly Identification Queue," but labeled as "Industrialized" and thus a violation in the "Video Style Anomaly Identification Queue."
[0052] That is, each identification model independently performs the task of identifying anomalies in the overall published content through an anomaly identification queue. There is a lack of dynamic communication between the anomaly identification queues, and it is impossible to identify different anomaly categories in different components of the published content. This results in a limited ability to identify heterogeneous anomaly types and affects the robustness of the overall identification.
[0053] To address the aforementioned issues, this specification provides an anomaly identification method. Instead of directly identifying anomalies in the entire published content, this method extracts content units from the acquired content. It determines the anomaly identification task based on the correlation between each content unit and various anomaly identification tasks, achieving finer-grained anomaly identification at the content unit level. The anomaly identification result is determined by the anomaly information obtained from each content unit through the target identification task. This comprehensively reflects the anomalies of the target published content under various anomaly domains, enabling refined decomposition of published content in different modalities and efficient identification of complex anomalies.
[0054] Specifically, this specification provides an anomaly identification method, and also relates to an anomaly identification device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0055] See Figure 2 , Figure 2 A flowchart of an anomaly identification method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0056] Step 202: Obtain the target published content and extract each content unit from the target published content.
[0057] The embodiments described in this specification can be applied to platforms, applications, etc., that can publish content. Examples include social media platforms, news platforms, online education platforms, and internal enterprise systems and government e-government platforms.
[0058] The target published content is information in the form of multimodal data submitted by users to the platform, such as text and image notes. The target published content can include multimodal data such as images, text, titles, Optical Character Recognition (OCR) content, and Automatic Speech Recognition (ASR) content, which serve as the input objects for anomaly detection tasks.
[0059] Optionally, the target published content can be parsed using multimodal encoding and decoding to obtain content units in a unified format. For example, a user's image and text notes uploaded to a social media platform can have their video footage processed by image recognition to obtain image recognition content, and their audio processed by speech recognition to obtain speech recognition content. These are then encoded based on the image recognition content, speech recognition content, and text content respectively, and finally decoded using a multimodal decoder to obtain the individual content units.
[0060] A content unit is the smallest semantic representation unit extracted from the target published content. For example, a single token in a token sequence generated by a multimodal encoder can encompass different forms such as image tokens and text tokens.
[0061] Optionally, the extraction of content units can be achieved by using a visual encoder to split the image region in the target published content into an image token sequence, and a text encoder to split the text into a text token sequence. Finally, a decoder (such as a Large Language Model (LLM) decoder) is used to achieve semantic alignment and fusion of cross-modal information, generating a unified token sequence as input data for subsequent anomaly detection.
[0062] Specifically, one option for obtaining the target published content is to directly receive the target published content sent by the front end; another option is to obtain the target published content from a cloud database or local storage.
[0063] Specifically, once the target published content is obtained, each content unit is extracted from the target published content. One possible approach is to directly extract the data from the target published content to obtain each content unit. Another possible approach is to extract the target published content according to each modality, obtain the content units under each modality, and then splice and align them to obtain each content unit of the target published content.
[0064] In this step, by acquiring the target published content and extracting each content unit, a data foundation is provided for subsequent anomaly identification of the target published content based on each content unit.
[0065] Step 204: Based on the correlation influence between the target content unit and various anomaly identification tasks, perform various anomaly identification tasks on the target content unit to obtain the anomaly information corresponding to each anomaly identification task, where the target content unit is any one of the content units.
[0066] Anomaly detection tasks are those performed during the process of identifying anomalies in target published content. Specifically, anomaly detection tasks target independent categories of anomalies, and the output of these tasks must conform to the annotation rules and anomaly domain boundaries of that category. Anomaly detection can include categories such as semantic anomalies, content mismatch anomalies, and content source anomalies. An anomaly domain, defined by preset feature parameters, behavioral patterns, or data distribution during the execution of the anomaly detection task, is a specific range of content that differs from the content that should be included in the target content unit. Specifically, the anomaly domain defines the judgment boundaries and rules of the anomaly detection task. Each anomaly detection task can correspond to a specific anomaly domain, and the output of that task must conform to the annotation rules and boundaries of the corresponding anomaly domain.
[0067] Optionally, the anomaly detection task can be targeted at a single content unit or at an anomaly detection queue consisting of content units.
[0068] Optionally, the division of anomaly recognition tasks should be consistent with the labeling and classification structure in each anomaly domain for anomaly recognition of published content. For example, for the anomaly recognition task of "video style", the probability values corresponding to the subcategories of "industrial style" and "movie style" can be output.
[0069] Various anomaly recognition tasks are a collection of anomaly recognition tasks covering multiple preset categories, and their number is directly related to the coverage of anomaly types.
[0070] Optionally, in the embodiments of this specification, the scope of various anomaly recognition tasks is defined by the annotation results of the entire anomaly domain for the target published content, such as a tag tree definition. For example, it may include different categories of anomaly recognition tasks such as "video style anomaly recognition task," "text content anomaly recognition task," and "original video anomaly recognition task."
[0071] Association influence is the strength of feature association between a target content unit and various anomaly identification tasks, used to quantify the contribution of a content unit to a specific anomaly domain.
[0072] Optionally, the correlation influence can be calculated based on the correspondence between the feature information of the content unit and the anomaly detection task. The feature information may include semantic feature information, structural feature information, etc.
[0073] Optionally, the degree of correlation influence can be determined by the routing weights between the target content unit and various anomaly detection tasks. For example, for the content unit of the image token, a higher weight is assigned in the execution of the "video style anomaly detection task" and a lower weight is assigned in the execution of the "text content anomaly detection task", thereby optimizing the targeting of the anomaly detection task through a gating mechanism.
[0074] The target content unit is any one of the content units extracted from the target published content, serving as the basic processing unit for various anomaly detection tasks. Optionally, the target content unit can be selected by sorting the content units and selecting them sequentially; it can also be randomly selected from the content units; or it can be selected from the content units according to a preset selection rule. For example, if the selection rule is "prioritize the execution of 'text content' anomaly detection tasks," then text token content units that can represent semantic features will be selected first, while image token content units will not be selected.
[0075] Anomaly information is an intermediate result obtained by performing various anomaly recognition tasks on the target content unit, which includes the feature output of the target content unit under a specific anomaly domain.
[0076] Specifically, the abnormal information is determined based on the abnormal identification of the target content unit under various abnormal domains according to various abnormal identification tasks.
[0077] Optionally, various anomaly recognition tasks can be executed by corresponding anomaly recognition models. That is, for a target content unit, obtaining the anomaly information corresponding to various anomaly recognition tasks may include calling various anomaly recognition models for the target content unit to obtain the anomaly information corresponding to various anomaly recognition tasks.
[0078] Optionally, the anomaly information corresponding to various anomaly identification tasks can be post-processed based on the degree of correlation influence. Specifically, the weight coefficients corresponding to various anomaly identification tasks can be determined based on the degree of correlation influence, and the anomaly information corresponding to various anomaly identification tasks can be weighted to obtain weighted anomaly information.
[0079] In this step, various anomaly identification tasks are dynamically selected and executed based on the correlation influence between the target content unit and various anomaly identification tasks, generating corresponding anomaly information. This provides a targeted data foundation for determining the anomaly identification results and optimizes the anomaly identification accuracy of the target published content under multiple anomaly domains.
[0080] Step 206: For the target identification task in various anomaly identification tasks, determine the anomaly identification result of the target published content based on the anomaly information of each content unit in the target identification task.
[0081] Anomaly detection results are generated for target identification tasks within various anomaly detection tasks. Based on the anomaly information of each content unit under the target identification task, the results are comprehensively generated for the target published content. Specifically, the anomaly detection results correspond to various anomaly detection tasks, meaning the output anomaly detection results include the anomaly detection results of the target published content under various anomaly detection tasks.
[0082] Optionally, the anomaly identification results can be presented in probabilistic form to quantify the likelihood of the target content violating regulations within a specific anomaly domain corresponding to the target identification task. Specifically, the anomaly identification results can be combined with a structured representation of the full anomaly domain label tree. For example, for a "video style anomaly identification task" as the target identification task, the corresponding anomaly identification results may include style types such as "industrial style," "cinematic style," and "artistic style." For this target identification task, the obtained anomaly identification results may include "Video Style Anomaly Identification Queue - Industrial Style - 0.21," "Video Style Anomaly Identification Queue - Cinematic Style - 0.85," and "Video Style Anomaly Identification Queue - Artistic Style - 0.06," indicating that the probability of the target content's video style being industrial is 21%, cinematic is 85%, and artistic is 6%.
[0083] The target identification task is any anomaly identification task selected from various anomaly identification tasks. Specifically, for the target identification task, the anomaly information of each content unit under the target identification task can be aggregated and analyzed, and the probability of the target published content violating the rules under a specific anomaly domain can be quantified in probabilistic form.
[0084] Optionally, the target recognition task can be selected based on the structured representation of the full anomaly domain label tree and the actual audit requirements; or it can be selected randomly.
[0085] Optionally, the anomaly identification result can be determined based on the anomaly information. This can be achieved by generating an anomaly information set based on the anomaly information of each content unit under the target identification task, and performing feature extraction calculation to obtain local feature information under the target identification task. Then, the anomaly identification result can be determined by average pooling calculation based on the local feature information.
[0086] In the embodiments of this specification, by extracting content units from the target published content and performing multiple anomaly recognition tasks on the target units within each content unit, the anomaly recognition strategy can be dynamically adjusted based on the correlation influence between the target content unit and the anomaly recognition task. This allows for the capture of key features of different anomaly domains at a fine-grained level of content units. Simultaneously, by independently analyzing the anomaly information of the target content unit for each type of anomaly recognition task, and then further synthesizing and generating the overall recognition result of the target published content, the system achieves adaptation to the differences of multiple anomaly domains, enhances the fusion and coordination between heterogeneous anomaly tasks, and thus improves the accuracy and robustness of recognizing multimodal published content in complex anomaly scenarios.
[0087] In one optional embodiment of this specification, based on the correlation influence between the target content unit and various anomaly detection tasks, various anomaly detection tasks are performed on the target content unit to obtain anomaly information corresponding to each anomaly detection task, including:
[0088] Based on the feature information of the target content unit, determine the degree of correlation between the target content unit and various anomaly identification tasks;
[0089] Based on the degree of correlation and impact, tasks to be executed are selected from various anomaly identification tasks;
[0090] Execute the task to be executed for the target content unit and obtain the corresponding exception information for the task to be executed.
[0091] The feature information of a target content unit is quantified data extracted from the target content unit to characterize its semantic, structural, or modal attributes. Its form may include semantic embedding vectors, such as word vectors of text tokens; image feature vectors, such as convolutional neural network (CNN) features of visual tokens, etc.
[0092] The tasks to be executed are a subset of tasks selected from various anomaly identification tasks and executed on the target content unit. Specifically, the selection of tasks to be executed is based on the degree of correlation and influence.
[0093] Optionally, the allocator can be configured to filter tasks to be executed from various anomaly identification tasks based on the degree of correlation influence, and the target content unit can be assigned to the tasks to be executed.
[0094] Optionally, the allocator can be a gated router G(·), and correspondingly, the association influence degree can be based on the target content unit of the gated router. The routing weights configured with the feature information:
[0095]
[0096] in, This indicates that the target content unit is the nth content unit.
[0097] Optionally, in order to reduce the amount of data transmission and processing in the actual anomaly identification process and improve the efficiency of anomaly identification, a certain filtering strategy can be adopted to filter the tasks to be executed for the target content unit.
[0098] One option is to set a preset threshold for the correlation impact, filtering out only tasks whose correlation impact exceeds the preset threshold. For example, in a video anomaly detection scenario, if the target published content includes a video with industrial-style visuals, the correlation impact of the target content unit determined based on the image token (e.g., "frame tone range value") of that video frame on each task is as follows: "Video Style Anomaly Detection Task": 0.85 (high); "Text Content Anomaly Detection Task": 0.12 (low); "Sensitive Word Detection Anomaly Detection Task": 0.23 (low). If the preset threshold is set to 0.7, only the "Video Style Anomaly Detection Task" will be selected as the task to be executed, while the remaining anomaly detection tasks will be filtered out due to insufficient correlation impact.
[0099] Another option is to sort the tasks by their degree of relevance from highest to lowest, and then select the top-preset number of tasks to be executed. This is a Top-k selection strategy, targeting the specific content units. Only the top k tasks to be executed are selected to form a set of tasks to be executed.
[0100] Specifically, after selecting the tasks to be executed, the tasks can be executed on the target content unit to obtain the corresponding exception information.
[0101] Optionally, based on the degree of correlation influence, the abnormal information corresponding to the task to be executed can be further calculated to obtain the feature information of the target content unit.
[0102] Optionally, since each task to be executed is assigned multiple content units, and the degree of correlation between each content unit and each task to be executed is determined according to the actual impact, the identification result of each task to be executed for the target published content is determined jointly based on the abnormal information of each content unit corresponding to each task to be executed.
[0103] In the embodiments of this specification, the correlation influence between the target content unit and various anomaly identification tasks is calculated based on the feature information of the target content unit, and a certain filtering strategy is used to filter the tasks to be executed. This effectively reduces the processing overhead of redundant tasks, while ensuring that the tasks to be executed with high correlation influence are executed first. This reduces the consumption of computing resources and improves the efficiency of anomaly identification. By executing the filtered tasks to be executed on the target content unit, the obtained anomaly information can more accurately reflect the anomaly characteristics of the target content unit under various anomaly identification tasks, thereby improving the accuracy and robustness of anomaly prediction.
[0104] In an optional embodiment of this specification, an execution task is performed on the target content unit to obtain the exception information corresponding to the execution task, including:
[0105] Execute the task to be executed for the target content unit and obtain sub-anomaly information;
[0106] Based on the degree of correlation between each specified content unit and the task to be executed, the sub-anomaly information obtained for each specified content unit is weighted and calculated to obtain the anomaly information corresponding to the task to be executed. Here, the specified content unit is the content unit that executes the task to be executed.
[0107] Sub-anomaly information is the raw, unprocessed anomaly output obtained after executing the task to be executed for the target content unit. It takes the form of the independent anomaly identification prediction results of each anomaly identification task to be executed for the target content unit, and includes the sub-category probability distribution under a specific anomaly domain.
[0108] Optionally, various anomaly recognition tasks can be performed by the corresponding anomaly recognition model, and the sub-anomaly information is determined by the corresponding anomaly recognition model through feedforward calculation.
[0109] Weighted calculation is based on the correlation influence between a specified content unit and the task to be executed. It assigns dynamic weights to each sub-anomaly information and performs aggregation calculations to generate more discriminative anomaly information. Specifically, weighted calculation can be combined with routing weights and task-specific rules. For example, for a target content unit that is a text token, when performing a "video style anomaly recognition task", the obtained sub-anomaly information could be "artistic style". The correlation influence weight of the text token for this "video style anomaly recognition task" is 0.3. Therefore, the anomaly information obtained after weighted calculation is "artistic style - 0.3".
[0110] A designated content unit is a set of multiple target content units assigned to a task during the screening and allocation process. The selection of designated content units is based on the degree of correlation and influence between each content unit and the task to be executed.
[0111] In the actual recognition process, when selecting tasks to be executed for target content units, it can also be understood as assigning the target content units to the selected tasks to be executed.
[0112] Optionally, each task to be executed can construct a corresponding exception detection queue and perform exception detection tasks for the exception detection queue. Specifically, the exception detection queue may include content units assigned to the task to be executed.
[0113] Specifically, when each specified content unit is assigned to a task to be executed, the task can perform an anomaly identification task for the specified content unit, obtain sub-anomaly information, and further perform weighted calculations based on the configured routing weights to obtain the anomaly information:
[0114]
[0115] Among them, FFN i Let represent the anomaly recognition model corresponding to the i-th anomaly recognition task.
[0116] In the embodiments of this specification, by performing weighted calculations on the sub-anomaly information of a specified content unit, the local features of multimodal content units in the heterogeneous anomaly domain can be effectively integrated. This not only preserves the fine-grained discrimination capability of the anomaly identification task for content units, but also enhances the fusion effect of cross-modal information through dynamic weight allocation. At the same time, the weighted calculation strategy improves the robustness and accuracy of complex anomaly scene identification by suppressing the interference of low-correlation sub-anomaly information, thereby reducing computational redundancy and optimizing the overall review efficiency in the multi-anomaly domain.
[0117] In an optional embodiment of this specification, after extracting the content units from the target published content, the method further includes:
[0118] For each content unit, perform a shared anomaly identification task to obtain shared anomaly information;
[0119] For target identification tasks within various anomaly detection tasks, based on the anomaly information for each content unit, the anomaly detection results for the target published content are determined, including:
[0120] For target identification tasks in various anomaly identification tasks, the anomaly identification results of the target published content are determined based on the anomaly information of each content unit and the shared anomaly information of the target identification task.
[0121] The shared anomaly detection task is a general anomaly detection task performed on each content unit to obtain anomaly information of common anomaly features across different modalities of content units. Examples include sensitive word detection, basic compliance checks, or general violation pattern recognition. The execution of the shared anomaly detection task covers all content units within the target published content.
[0122] Shared anomaly information is the collection of anomaly information generated by each content unit of the target published content under the shared anomaly identification task.
[0123] Specifically, extracting content units from the target published content. In this case, it is possible to target each content unit Execute shared anomaly detection task q sObtain the anomaly information corresponding to each content unit:
[0124]
[0125] Among them, FFN s It is used to perform the shared anomaly detection task q s A shared anomaly detection model.
[0126] Each content unit Corresponding error information x (n,s) Integrate and obtain shared anomaly information X s :
[0127]
[0128] Once shared anomaly information is obtained, the identification result of the target published content can be determined by combining the shared anomaly information with the anomaly information of each content unit from the target identification task.
[0129] Specifically, since the shared anomaly information includes the shared anomaly information of each content unit, that is, the shared anomaly information of all content units of the target published content, and the target identification task may be performed on each content unit or on a specific content unit, it is necessary to determine the correspondence between the shared anomaly information and each content unit in the anomaly information, and based on the correspondence, determine the anomaly identification result of the target published content.
[0130] Optionally, the target recognition task is performed on each content unit. Abnormal information x (n,i) The set is counted as From X s In, extract with X i Shared anomaly information of the same content unit is counted as Summing the results, we obtain the fusion anomaly information X. out ,Right now:
[0131]
[0132] Having obtained the fused anomaly information, determining the anomaly identification result of the target published content may include determining the anomaly identification result of the target published content based on the fused anomaly information.
[0133] Optionally, after obtaining fusion anomaly information X out In this case, the fusion of abnormal information X can be performed. out Further processing can be performed, for example, by performing average pooling (AvgPool) and then passing it through a linear classifier (Linear), to output the final anomaly detection result p of the target published content, i.e.:
[0134] p = Linear(AvgPool(X) out ))∈[0,1] L
[0135] By using average pooling to process fused anomaly information, the feature dimension of the model is reduced, the computational load of the linear classifier is lowered, and overfitting of anomaly prediction results can be suppressed by smoothing local noise. Further processing using the linear classifier and optimizing task-related weights through end-to-end training improves the accuracy of anomaly detection.
[0136] In the embodiments of this specification, by performing a shared anomaly identification task for each content unit and obtaining shared anomaly information, and by jointly determining the anomaly identification result based on the shared anomaly information and the anomaly information, the common anomaly features of multimodal content units, such as the extraction of common semantic features, and the collaborative optimization combined with the features of the target identification task are realized. This provides a global anomaly discrimination basis for cross-modal content units, enhances the ability to identify basic anomaly types, and retains the fine-grained discrimination advantage of various anomaly identification tasks for local anomaly features by jointly determining the anomaly identification result based on the shared anomaly information and the anomaly information of the target identification task. At the same time, the noise interference of various anomaly identification tasks is suppressed by the shared anomaly information, thereby reducing computational redundancy while improving the model generalization and robustness in multi-anomaly domains, and optimizing the accuracy and efficiency of anomaly identification for multimodal published content in complex scenarios.
[0137] In one optional embodiment of this specification, based on the correlation influence between the target content unit and various anomaly detection tasks, various anomaly detection tasks are performed on the target content unit to obtain anomaly information corresponding to each anomaly detection task, including:
[0138] Based on the correlation between the target content unit and various anomaly recognition tasks, for the target content unit, the anomaly recognition model corresponding to each anomaly recognition task is invoked to obtain the anomaly information corresponding to each anomaly recognition task.
[0139] An anomaly detection model refers to an independent network module designed for anomaly detection tasks under a specific anomaly domain. It is used to generate anomaly information for the anomaly detection task based on the feature information of the target content unit through neural network calculation.
[0140] An anomaly detection model can be understood as an expert in a specific anomaly domain, specializing in anomaly detection within that domain. In other words, each anomaly detection model corresponds one-to-one with a specific anomaly detection task; each task is executed by a particular anomaly detection model, and each model performs a specific anomaly detection task.
[0141] Alternatively, the anomaly detection model can employ various types of network models, such as feed-forward network (FFN), deep generative network (DBN), and graph neural network (GNN).
[0142] For the anomaly recognition models constructed for various anomaly recognition tasks, a Mixture of Experts (MoE) architecture can be implemented. By combining dynamic routing weights determined based on the degree of correlation influence, and selectively calling some experts (i.e. anomaly recognition models) to execute the corresponding anomaly recognition tasks according to the characteristics of the target content unit, efficient task processing can be achieved.
[0143] The following explanation uses a feedforward network model as an example.
[0144] For various anomaly recognition tasks with numerous anomaly categories, complex semantic differences, and different data modal focuses, corresponding anomaly recognition models can be constructed, namely:
[0145] ε={FFN1,FFN1,…,FFN Q}
[0146] Among them, FFN i Let represent the anomaly recognition model corresponding to the i-th anomaly recognition task.
[0147] Optionally, it can also be applied to the shared anomaly detection task q s Construct a shared anomaly detection model FFN s It is used to extract common anomaly features from various content units.
[0148] Correspondingly, the correlation influence between the target content unit and various anomaly detection tasks can be considered as the routing weight between the target content unit and the anomaly detection model:
[0149]
[0150] Then for the target content unit Calling the FFN anomaly detection model corresponding to various anomaly detection tasks i Obtain anomaly information x corresponding to various anomaly identification tasks (n,i) for:
[0151]
[0152] Shared anomaly detection model FFN s Perform a shared anomaly detection task q for each content unit. sThe shared exception information obtained is as follows:
[0153]
[0154] In the embodiments of this specification, by dynamically calling the anomaly recognition model based on the correlation influence, fine-grained feature extraction and resource optimization allocation under multiple anomaly domains are realized. The anomaly recognition model constructed for a specific anomaly recognition task can accurately capture the local semantic features of the target content unit under a specific anomaly domain, improve the discrimination accuracy of anomaly recognition, and at the same time improve the recognition efficiency and robustness in complex anomaly scenarios.
[0155] In an optional embodiment of this specification, before invoking the anomaly recognition model corresponding to each anomaly recognition task for the target content unit based on the correlation influence between the target content unit and various anomaly recognition tasks, and obtaining the anomaly information corresponding to each anomaly recognition task, the method further includes:
[0156] Obtain the published sample content and extract each sample content unit from the published sample content. The published sample content carries sample labels for the sample content units to annotate abnormal results under various anomaly identification tasks.
[0157] Based on the correlation between the target sample content unit and various anomaly recognition tasks, for the target sample content unit, the anomaly recognition model corresponding to various anomaly recognition tasks is called to obtain the anomaly information corresponding to various anomaly recognition tasks. The target sample content unit is any one of the sample content units.
[0158] For target identification tasks in various anomaly identification tasks, based on the anomaly information of each sample content unit in the target identification task, the predicted anomaly identification result of the sample published content is determined;
[0159] Based on the predicted anomaly identification results and sample labels, anomaly identification models corresponding to various anomaly identification tasks are trained.
[0160] Sample published content is example data used for training anomaly detection models. Its format is consistent with target published content and can include multimodal data (such as images, text, OCR, ASR, etc.) and corresponding anomaly annotation information.
[0161] Each sample content unit is a set of the smallest semantic representation units extracted from the sample published content. Its form is the same as the content unit and can include image tokens, text tokens, etc. For example, in the sample published content, video frames can be split into image token sequences, and text can be split into text token sequences. These tokens together constitute the basic input data for model training.
[0162] The sample labels for anomaly identification results are predefined supervisory signals for various anomaly identification tasks, based on the sample publication content. They take the form of sub-category labels under the corresponding anomaly domain. Specifically, for correct sub-categories, the sample label is marked as true ("1"), and for incorrect sub-categories, it is marked as false ("0"). For a specific anomaly identification task, only one correct sub-category label is included, meaning there is only one "1" and all others are "0," i.e., one-hot encoding.
[0163] Optionally, for a specific anomaly identification task, if there is no corresponding subcategory label, the sample label is marked as missing, denoted as "-100".
[0164] Alternatively, a mask vector m can be constructed to check for anomalies in the labeled results. i This indicates whether a subcategory label exists. If it exists, the label is 1; otherwise, it is 0.
[0165] The target sample content unit is the content unit selected from the published sample content and used as the processing object in the current training step. For example, if "frame tone range value" is selected as the target sample content unit during training, the anomaly recognition model of the corresponding category needs to be called based on its feature information to generate the anomaly information of that unit.
[0166] The predicted anomaly detection result is the probability distribution output by the model after performing various anomaly detection tasks on the target sample content unit, which is used to compare with the sample labels. For example, in the "video style anomaly detection task", the model may output "industrial style -0.19", "movie style -0.83", and "artistic style -0.08", which are compared with the sample labels "industrial style -0", "movie style -1", and "artistic style -0" to calculate the error of the anomaly prediction model.
[0167] Specifically, once the predicted anomaly identification result is determined, the anomaly identification model corresponding to various anomaly identification tasks can be trained based on the predicted anomaly identification result and sample labels.
[0168] Alternatively, the anomaly detection model can be trained by calculating the loss between the predicted anomaly detection result and the sample label. Various types of losses can be used for calculation, such as cross-entropy loss, classification loss, and contrastive loss.
[0169] Taking cross-entropy loss as an example, the loss between predicting anomaly identification results and sample labels can be expressed as:
[0170]
[0171] in, Here, BCE (Binary Cross Entropy) is the indicator function, and BCE is the cross-entropy loss.
[0172] Once the loss is determined, gradient descent algorithms (such as stochastic gradient descent, Adam optimizer, etc.) can be used to adjust the parameters of various anomaly detection models.
[0173] In the embodiments of this specification, sample published content is obtained and the predicted anomaly identification result is determined. Based on the predicted anomaly identification result and sample labels, the anomaly identification model is trained. The difference between the predicted anomaly identification result and the sample labels is quantified, the parameters of the anomaly identification model are dynamically updated, the anomaly identification model's ability to capture local semantic features under specific anomaly domains is strengthened, the training efficiency of the anomaly identification model is improved, and the robustness of the model to complex anomaly scenarios and the anomaly identification accuracy of multimodal content units are enhanced.
[0174] In one optional embodiment of this specification, obtaining sample publication content includes:
[0175] Obtain the published content of the sample, as well as the pre-built anomaly label tree, which includes labeled anomaly information under various anomaly identification tasks;
[0176] Based on the anomaly label tree, determine the sample labels for each sample content unit in the published sample content, which are used to annotate the anomaly information under various anomaly identification tasks.
[0177] An anomaly label tree is a tree-like system that organizes various anomaly recognition tasks and their corresponding labeled anomaly information in a structured hierarchical manner. Its root node covers the classification boundary of the entire anomaly domain, and child nodes are progressively refined to specific anomaly categories. Labeled anomaly information consists of predefined anomaly category labels for sample content units under a specific anomaly recognition task.
[0178] Specifically, for Q anomaly detection tasks of various types, each anomaly detection task q i (i = 1, 2, ..., Q) has a K... i Each anomaly message is labeled and denoted as a set. Then, an anomaly label tree structure can be constructed using the format "anomaly identification task - anomaly information labeling".
[0179] Among them, the first layer of the anomaly label tree consists of various anomaly identification tasks q. i Examples include "video style anomaly recognition task" and "text content anomaly recognition task"; the second-level nodes are the labeled anomaly information r under each anomaly recognition task. ij The data is concatenated into structured labeled anomaly information such as "Video Style Anomaly Recognition Task - Industrial Style" and "Text Content Anomaly Recognition Task - Few Anomalies" [l1,l2,…,l].L The total number of anomaly messages marked is [number missing].
[0180] Each second-level node (i.e., leaf node) represents a specific "anomaly identification result", which can be used as the target to construct the truth value for anomaly identification. In the output stage of anomaly identification results, it can be restored to a complete anomaly identification in the form of "anomaly identification task + labeled anomaly information" through structure mapping.
[0181] Once the anomaly label tree is obtained, the sample content units in the published sample content can be retrieved from the anomaly label tree, and the sample labels for anomaly information can be marked under various anomaly identification tasks.
[0182] In the embodiments of this specification, the structured organization of the anomaly label tree and the mapping of anomaly information enable the unified generation of sample labels corresponding to the sample distribution content under multiple anomaly domains. The hierarchical structure of the anomaly label tree ensures the consistency of labeling for various anomaly identification tasks, avoiding model deviations caused by missing labels or rule conflicts in traditional methods. At the same time, by dynamically adapting to the labeling requirements of multimodal content units, the anomaly prediction model enhances the recognition accuracy of complex anomaly scenarios, thereby reducing labeling costs while optimizing the review efficiency and accuracy under multiple anomaly domains.
[0183] In one optional embodiment of this specification, based on an anomaly label tree, sample labels are determined to annotate the sample published content with anomaly information under various anomaly identification tasks, including:
[0184] Based on the anomaly label tree, we can identify whether sample content units have labeled anomaly information under various anomaly identification tasks and obtain the labeling and identification results.
[0185] Based on the annotation and recognition results, the sample labels for anomaly information of sample content units are determined under various anomaly recognition tasks.
[0186] The annotation recognition result is a judgment result based on the anomaly label tree, determining whether the sample content unit has corresponding anomaly information under various anomaly recognition tasks. This result can be in the form of a binary indicator signal (e.g., "1" indicates the existence of a valid annotation, and "0" indicates a missing annotation). Specifically, the annotation recognition result verifies whether the sample content unit matches the sub-category node generation under a specific anomaly recognition task by traversing the hierarchical structure of the anomaly label tree. For example, if the sample content unit "frame tone range value" has a corresponding style sub-category annotation under the "video style anomaly recognition task," the annotation recognition result is "1"; if it does not have a corresponding anomaly level sub-category annotation under the "text content anomaly recognition task," the annotation recognition result is "0." The annotation recognition result is used for the generation of subsequent sample labels, ensuring that the model only learns from valid annotation tasks.
[0187] Alternatively, the labeling and recognition results can be represented by constructing a mask vector.
[0188] Specifically, based on the number Q of various anomaly detection tasks and the structure of the anomaly label tree, a queue-level mask vector [m1, m2, ..., m] of length Q is established. Q ]:
[0189]
[0190] Correspondingly, based on the annotation and recognition results, the sample label is determined. Represented as:
[0191] If there is a corresponding label for the i-th anomaly detection task, that is, the label detection result is "m i If the value is 1, then the correct labeling position in the anomaly identification task is set to "1", and the rest are "0", which is "one-hot" encoding.
[0192] If there is no corresponding label for the i-th anomaly detection task, that is, the label detection result is "m i If the value is 0, then the truth value of all labeled sample label positions in this anomaly detection task is set to -100, indicating that the label is masked and does not participate in the loss calculation.
[0193] Based on the annotation and recognition results, the sample label representation can be obtained as follows:
[0194]
[0195] In the embodiments of this specification, anomaly information is identified by using a hierarchical structure based on anomaly label trees, and the identification results are characterized by labeling. Furthermore, the labeling results are mapped to the sample label generation process. Sample labels are correctly determined for anomaly identification tasks with valid labels, while anomaly identification tasks with missing labels are masked. This achieves automatic filtering of noise interference from invalid tasks during the training phase of the anomaly identification model, ensuring the labeling consistency of multi-anomaly domain tasks, reducing redundant computational overhead, enhancing the discrimination accuracy of the anomaly identification model in complex anomaly scenarios, and improving the robustness and computational efficiency of the anomaly identification model in cross-modal and multi-task scenarios.
[0196] In one optional embodiment of this specification, extracting content units from the target published content includes:
[0197] Extract content information for each modality from the target published content;
[0198] The content information of each modality is encoded separately to obtain the content unit under each modality;
[0199] Semantic alignment and fusion are performed on the content units in each modality to obtain each content unit.
[0200] The content information of each modality is the raw data form separated from the target published content and existing in different perceptual forms. These forms encompass heterogeneous data such as images, text, audio, OCR text, and ASR speech transcription. Specifically, the content information of each modality can be obtained through preprocessing (such as splitting video frames into image sequences, speech-to-text conversion, etc.) and serves as the basic input for subsequent feature extraction and anomaly detection. For example, in a video review scenario, the target published content may include frame sequences from the image modality, subtitle text from the text modality, and background music from the audio modality.
[0201] Content units in each modality are the smallest semantic representation units generated after encoding content information for different modalities, and their form depends on the feature extraction method of the specific modality. For example, content units in the image modality can be visual tokens (such as CNN feature vectors), content units in the text modality can be word vectors (such as BERT embeddings), and content units in the audio modality can be spectral feature vectors.
[0202] Semantic alignment fusion is the process of aligning and integrating the representations of content units from different modalities within a unified semantic space. It aims to eliminate semantic differences between modalities and generate joint representations across modalities. Specifically, semantic alignment fusion can be achieved through a shared semantic embedding space (e.g., using cross-modal attention mechanisms) or feature mapping functions. For example, image tokens and text tokens can be mapped to the same vector space, and then a fused global representation can be generated through weighted aggregation or gating mechanisms for unified processing of anomaly detection tasks.
[0203] Specifically, extracting content information for each modality from the target published content involves breaking down the target published content into raw data for each modality. For example, separating image frame sequences from a video, extracting text from speech, and extracting OCR text from an image. For the content information of each modality, encoding operations are performed separately to obtain content units for each modality. For example, for image modality data, image tokens representing local visual information (such as frame tone range values and texture gradient values) are extracted using a convolutional neural network (CNN) or Transformer encoder, denoted as: Text modalities generate text tokens through word embedding models (such as pre-trained language models (Bidirectional Encoder Representations from Transformers, or BERT for short)), denoted as: The resulting set of content units is: Where N = N v +N t .
[0204] Based on the semantic correlation of each modal content unit, heterogeneous representations are mapped to a unified dimension through a shared embedding space or cross-modal interaction mechanism to obtain each content unit, which serves as the input to the anomaly detection model to capture the collaborative features of multimodal content at the semantic level.
[0205] In the embodiments of this specification, by constructing a mapping space with the same semantics and determining each content unit while retaining the independent features of each modality, the refined representation of multimodal content units and the fusion of cross-modal information are realized. This enhances the anomaly recognition model's ability to perceive heterogeneous data, improves the collaborative discrimination efficiency of multimodal features in anomaly recognition tasks, avoids noise interference caused by semantic differences between modalities through semantic alignment, optimizes the robustness of the anomaly recognition model in complex scenarios, and thus improves the accuracy and generalization ability of anomaly recognition in multiple anomaly domains while reducing feature extraction redundancy.
[0206] In an optional embodiment of this specification, after determining the anomaly identification result of the target published content, the method further includes:
[0207] If the anomaly identification results meet the preset publishing conditions, the target content will be published.
[0208] If the anomaly identification results do not meet the preset release conditions, the anomaly identification results and the target identification task will be fed back to the user.
[0209] Obtain the modified target published content input from the user and return to execute the steps for extracting each content unit from the target published content.
[0210] The preset publishing conditions are pre-defined security standards used to determine whether target content can be published. Their core basis is the probability distribution of anomalies under each anomaly identification task in the anomaly identification results. Specifically, if the anomaly identification result of the target content under various anomaly identification tasks is a probability of no anomaly (e.g., "text content anomaly - pass" probability < 0.05, "video style - industrial style" probability < 0.1), then the preset publishing conditions are met.
[0211] The user interface is the terminal device or interface through which users interact with the platform or device where they publish content. It can include receiving feedback on anomaly identification results, submitting modified content, and triggering a re-review process. Specifically, the user interface can display anomaly identification results for the target published content to the user through the platform interface, and prompt the user to correct the target published content if anomalies are found. For example, when a user's uploaded content is marked as violating the "soft advertising" task, the user interface will push modification suggestions to the user and receive the edited published content.
[0212] The modified target content is a new version of the original target content adjusted by the user based on feedback from anomaly identification results. Upon receiving the modified target content, the system must re-enter the content extraction and anomaly identification process to verify whether the modification has eliminated the violation anomaly. For example, if a user modifies the video color tone to address the "video style - industrial style" anomaly, the system needs to re-extract its image token and text token and execute related anomaly identification tasks such as the "video style anomaly identification task" to ensure that the modified content meets the preset publishing conditions.
[0213] In the embodiments described in this specification, dynamic judgment based on preset publishing conditions is used, and feedback is sent to the user terminal in case of anomalies. Upon receiving the modified target publishing content based on user feedback, anomaly identification is performed again. This achieves bidirectional collaborative optimization between anomaly identification results and user behavior, improves the efficiency of users correcting non-compliant content, ensures that the modified content meets the security requirements of multiple anomaly domains at both the semantic and structural levels, avoids misjudgment of anomalies in a single review, and thus optimizes the overall throughput and user satisfaction of the review system while ensuring content compliance.
[0214] The following is in conjunction with the appendix Figure 3 To be continued Figure 4 Taking the application of the anomaly identification method provided in this specification to published content as an example, the anomaly identification method will be further explained. Figure 3 This diagram illustrates the execution process of an anomaly detection method provided in one embodiment of this specification, based on... Figure 3 The content shown can be used to obtain, for example Figure 4 The method execution steps shown are as follows: Figure 4 The flowchart of an anomaly identification method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0215] Step 402: Obtain image and text information from the sample published content, where the text information includes title, content, voice identifier, and character identifier.
[0216] Step 404: Encode the image information to obtain image content units, and encode the text information to obtain text content units.
[0217] Step 406: Use the large language model decoder to decode and obtain each content unit.
[0218] Step 408: Input each content unit into the gating router and the shared anomaly detection model respectively, and the gating router will assign each content unit to the corresponding anomaly detection model.
[0219] Step 410: Input the shared anomaly information output by the shared anomaly recognition model and the anomaly information obtained by each anomaly recognition model into the content selection module, and determine the fused anomaly information based on the mask vector.
[0220] Step 412: Perform average pooling on the fused anomaly information and obtain the predicted labels of the anomaly recognition results corresponding to each anomaly recognition task through a linear classifier.
[0221] Step 414: Calculate the cross-entropy loss based on the ground truth values of the sample labels in the anomaly label tree and the actual values of the predicted labels, and train the anomaly recognition model.
[0222] In the embodiments of this specification, efficient anomaly recognition and model training in complex anomaly scenarios are achieved through collaborative processing of multimodal data and a dynamic task allocation mechanism. First, heterogeneous information such as images and text is extracted from the sample published content, and unified content units with unified representations are generated through encoding and decoding operations, laying the foundation for cross-modal feature fusion. Based on the dynamic allocation strategy of a gated router, content units are accurately matched to corresponding anomaly recognition models. Combined with shared anomaly features extracted by the shared anomaly recognition model, task-level fine-grained discrimination capability is preserved, enhancing the synergistic effect of cross-modal information. By masking missing labels with mask vectors, fused anomaly information is obtained, suppressing noise interference and improving the robustness of the anomaly prediction model in multiple anomaly domains. Relying on the structured labeling system of the anomaly label tree, combined with end-to-end optimization of cross-entropy loss, label consistency is ensured, strengthening the anomaly prediction model's ability to capture local semantic features. This significantly improves the accuracy of anomaly recognition and the training efficiency of the anomaly prediction model while reducing computational redundancy.
[0223] See Figure 5 , Figure 5 This diagram illustrates a client-side display for identifying abnormal content published according to an embodiment of this specification. Figure 5 As shown.
[0224] The user edited the content to be published on the client, including images and text. The images were generated using AI; the text included the phrase "Original hand-drawn landscape sketch, the weather was great, I went out to sketch with my easel and created a satisfactory painting." After editing the content, the user clicked the "Publish Content" button. An anomaly detection was then performed on the published content, identifying that the image was not original content but rather AI-generated, thus constituting a non-original content anomaly and failing to meet the preset publishing conditions. In cases where the preset publishing conditions are not met, an anomaly detection result page was displayed on the client, showing the corresponding anomaly information and guiding the user to make corrections.
[0225] Corresponding to the above method embodiments, this specification also provides embodiments of anomaly detection devices. Figure 6 A schematic diagram of an anomaly detection device according to one embodiment of this specification is shown. Figure 6 As shown, the device includes:
[0226] Extraction module 602 is configured to acquire target published content and extract each content unit from the target published content;
[0227] The execution module 604 is configured to perform various anomaly recognition tasks on the target content unit based on the correlation influence between the target content unit and various anomaly recognition tasks, and obtain the anomaly information corresponding to each anomaly recognition task. The target content unit is any one of the content units.
[0228] The determination module 606 is configured to determine the anomaly identification result of the target published content based on the anomaly information of each content unit in various anomaly identification tasks.
[0229] Optionally, the execution module 604 is further configured to: determine the correlation influence between the target content unit and various anomaly identification tasks based on the feature information of the target content unit; select tasks to be executed from various anomaly identification tasks based on the correlation influence; execute the tasks to be executed for the target content unit to obtain the anomaly information corresponding to the tasks to be executed.
[0230] Optionally, the execution module 604 is further configured to: execute the task to be executed for the target content unit and obtain sub-abnormal information; and perform weighted calculation on the sub-abnormal information obtained for each specified content unit based on the degree of correlation influence of each specified content unit on the task to be executed, so as to obtain the abnormal information corresponding to the task to be executed, wherein the specified content unit is the content unit for executing the task to be executed.
[0231] Optionally, the device further includes a sharing module configured to: perform a shared anomaly identification task for each content unit to obtain shared anomaly information; the determining module 606 is further configured to: for the target identification task in various anomaly identification tasks, determine the anomaly identification result of the target published content based on the anomaly information of each content unit and the shared anomaly information of the target identification task.
[0232] Optionally, the execution module 604 is further configured to: based on the correlation influence between the target content unit and various anomaly recognition tasks, for the target content unit, call the anomaly recognition model corresponding to various anomaly recognition tasks to obtain the anomaly information corresponding to various anomaly recognition tasks.
[0233] Optionally, the device further includes a training module configured to: acquire sample published content and extract each sample content unit from the sample published content, wherein the sample published content carries sample labels for the sample content units indicating abnormal results under various anomaly recognition tasks; based on the correlation influence between the target sample content unit and various anomaly recognition tasks, for the target sample content unit, call the anomaly recognition model corresponding to each anomaly recognition task to obtain the anomaly information corresponding to each anomaly recognition task, wherein the target sample content unit is any one of the sample content units; for the target recognition task in each anomaly recognition task, based on the anomaly information of each sample content unit for the target recognition task, determine the predicted anomaly recognition result of the sample published content; and train the anomaly recognition model corresponding to each anomaly recognition task based on the predicted anomaly recognition result and the sample labels.
[0234] Optionally, the training module is further configured to: acquire sample published content and a pre-built anomaly label tree, wherein the anomaly label tree includes labeled anomaly information under various anomaly recognition tasks; and based on the anomaly label tree, determine the sample labels of each sample content unit in the sample published content that are labeled with anomaly information under various anomaly recognition tasks.
[0235] Optionally, the training module is further configured to: identify whether sample content units have labeled abnormal information under various anomaly recognition tasks based on the anomaly label tree, and obtain the label recognition results; and determine the sample labels of sample content units labeled with abnormal information under various anomaly recognition tasks based on the label recognition results.
[0236] Optionally, the extraction module 602 is further configured to: extract content information of each modality from the target published content; encode the content information of each modality to obtain content units under each modality; and perform semantic alignment and fusion on the content units under each modality to obtain each content unit.
[0237] Optionally, the device also includes a feedback module configured to: publish the target content if the anomaly identification result meets the preset publishing conditions; and feed back the anomaly identification result and the target identification task to the user terminal if the anomaly identification result does not meet the preset publishing conditions; obtain the modified target content input by the user terminal and return to execute the step of extracting each content unit from the target content.
[0238] The above is a schematic scheme of an anomaly identification device according to this embodiment. It should be noted that the technical solution of this anomaly identification device and the technical solution of the above-described anomaly identification method belong to the same concept. For details not described in detail in the technical solution of the anomaly identification device, please refer to the description of the technical solution of the above-described anomaly identification method.
[0239] Figure 7 A structural block diagram of a computing device 700 according to one embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0240] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or Near Field Communication (NFC).
[0241] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0242] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0243] The processor 720 is used to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-described anomaly identification method.
[0244] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described anomaly identification method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described anomaly identification method.
[0245] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described anomaly identification method.
[0246] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-described anomaly identification method belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-described anomaly identification method.
[0247] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described anomaly identification method.
[0248] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-described anomaly identification method belong to the same concept. Details not described in detail in the computer program's technical solution can be found in the description of the technical solution of the above-described anomaly identification method.
[0249] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0250] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0251] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0252] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0253] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An anomaly identification method, characterized in that, include: Obtain the target published content and extract each content unit from the target published content; Based on the correlation between the target content unit and various anomaly identification tasks, the various anomaly identification tasks are executed for the target content unit to obtain the anomaly information corresponding to the various anomaly identification tasks, wherein the target content unit is any one of the content units; For the target identification task in the various anomaly identification tasks, the anomaly identification result of the target published content is determined based on the anomaly information of each content unit in the target identification task.
2. The method according to claim 1, characterized in that, The step of performing various anomaly identification tasks on the target content unit based on the correlation influence between the target content unit and various anomaly identification tasks, and obtaining the anomaly information corresponding to the various anomaly identification tasks, includes: Based on the feature information of the target content unit, the correlation influence between the target content unit and various anomaly identification tasks is determined; Based on the correlation influence, tasks to be executed are selected from the various anomaly identification tasks; The task to be executed is performed on the target content unit to obtain the abnormal information corresponding to the task to be executed.
3. The method according to claim 2, characterized in that, The step of executing the task to be executed for the target content unit and obtaining the exception information corresponding to the task to be executed includes: Execute the task to be executed for the target content unit to obtain sub-abnormal information; Based on the degree of correlation between each designated content unit and the task to be executed, the sub-abnormal information obtained for each designated content unit is weighted and calculated to obtain the abnormal information corresponding to the task to be executed, wherein the designated content unit is the content unit that executes the task to be executed.
4. The method according to claim 1, characterized in that, After extracting each content unit from the target published content, the method further includes: For each of the aforementioned content units, a shared anomaly identification task is performed to obtain shared anomaly information; For the target identification task in the various anomaly identification tasks, based on the anomaly information of each content unit in the target identification task, the anomaly identification result of the target published content is determined, including: For the target identification task in the various anomaly identification tasks, the anomaly identification result of the target published content is determined based on the anomaly information of each content unit and the shared anomaly information of the target identification task.
5. The method according to any one of claims 1-4, characterized in that, The step of performing various anomaly identification tasks on the target content unit based on the correlation influence between the target content unit and various anomaly identification tasks, and obtaining the anomaly information corresponding to the various anomaly identification tasks, includes: Based on the correlation between the target content unit and various anomaly identification tasks, for the target content unit, the anomaly identification model corresponding to each anomaly identification task is invoked to obtain the anomaly information corresponding to each anomaly identification task.
6. The method according to claim 5, characterized in that, Before obtaining the anomaly information corresponding to each anomaly identification task by calling the anomaly identification model corresponding to each anomaly identification task for the target content unit based on the correlation influence between the target content unit and various anomaly identification tasks, the method further includes: Obtain sample published content and extract each sample content unit from the sample published content, wherein the sample published content carries sample labels for the sample content units that annotate abnormal results under the various anomaly identification tasks; Based on the correlation influence between the target sample content unit and various anomaly identification tasks, for the target sample content unit, the anomaly identification model corresponding to each anomaly identification task is invoked to obtain the anomaly information corresponding to each anomaly identification task, wherein the target sample content unit is any one of the sample content units; For the target identification task in the various anomaly identification tasks, based on the anomaly information of each sample content unit in the target identification task, the predicted anomaly identification result of the sample published content is determined; Based on the predicted anomaly identification results and the sample labels, the anomaly identification models corresponding to the various anomaly identification tasks are trained.
7. The method according to claim 6, characterized in that, The acquisition of sample published content includes: Obtain the sample publication content and a pre-constructed anomaly label tree, wherein the anomaly label tree includes labeled anomaly information under various anomaly identification tasks; Based on the anomaly label tree, determine the sample labels for each sample content unit in the sample published content, which are labeled with anomaly information under the various anomaly identification tasks.
8. The method according to claim 7, characterized in that, The step of determining sample labels for annotating the sample publication content with abnormal information under various anomaly identification tasks based on the anomaly label tree includes: Based on the anomaly label tree, identify whether the sample content unit has labeled anomaly information under the various anomaly identification tasks, and obtain the labeling and identification results; Based on the annotation and recognition results, the sample labels for the sample content units that annotate abnormal information under the various anomaly recognition tasks are determined.
9. The method according to claim 1, characterized in that, The extraction of content units from the target published content includes: Extract content information for each modality from the target published content; The content information of each modality is encoded to obtain the content unit under each modality; Semantic alignment and fusion are performed on the content units under each modality to obtain each content unit.
10. The method according to claim 1, characterized in that, After determining the anomaly identification result of the target published content, the method further includes: If the anomaly identification result meets the preset publishing conditions, the target publishing content will be published. If the anomaly identification result does not meet the preset release conditions, the anomaly identification result and the target identification task will be fed back to the user terminal. Obtain the modified target published content input by the user, and return to execute the step of extracting each content unit from the target published content.
11. An anomaly detection device, characterized in that, include: The extraction module is configured to acquire target published content and extract content units from the target published content. The execution module is configured to perform various anomaly identification tasks on the target content unit based on the correlation influence between the target content unit and various anomaly identification tasks, and obtain the anomaly information corresponding to the various anomaly identification tasks, wherein the target content unit is any one of the content units; The determination module is configured to determine the anomaly identification result of the target published content based on the anomaly information of each content unit in the target identification task among the various anomaly identification tasks.
12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the anomaly identification method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, It stores a computer program / instruction that, when executed by a processor, implements the steps of the anomaly identification method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, It includes a computer program / instructions that, when executed by a processor, implement the steps of the anomaly identification method according to any one of claims 1 to 10.