Content moderation method and apparatus, device, storage medium, and program product
Patent Information
- Application Number
- PCT/CN2025/145779
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-12-25
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025145779_01102026_PF_FP_ABST
Abstract
Description
Content moderation methods, devices, equipment, storage media, and program products
[0001] This application claims priority to Chinese Patent Application No. 202510372528.X, filed on March 27, 2025, entitled “Content Review Method, Apparatus, Device, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to a content moderation method, apparatus, device, storage medium, and program product. Background Technology
[0003] Multimodal large models can process and integrate information from different modalities (such as text, images, and audio) to provide a more comprehensive and in-depth understanding, and are widely used in image content moderation and risk control (or risk management) technologies. Taking image content moderation as an example, multimodal large models determine whether sensitive or illegal content exists in an image by recognizing image information. However, with the continuous evolution of image styles and the increasingly complex forms of illegal content, some illegal content can easily bypass the detection mechanism of multimodal large models, thus being incorrectly judged as normal images, seriously affecting the accuracy of multimodal large models in reviewing illegal images. Summary of the Invention
[0004] This application provides a content moderation method, apparatus, device, storage medium, and program product, capable of moderizing images of various styles and in complex scenes, thereby improving the accuracy of image moderation results. The technical solution is as follows:
[0005] In a first aspect, a content moderation method is provided, the method comprising: obtaining image segmentation results of a target image to be moderated, the image segmentation results including multiple first image blocks; fusing the multiple first image blocks based on semantic features of the multiple first image blocks to obtain multiple second image blocks; and determining the moderation result of the target image based on global semantic features corresponding to the target image and local semantic features corresponding to the multiple second image blocks respectively, by means of a content moderation model, wherein the global semantic features are used to describe the semantic features of the target image as a whole, and the local semantic features are used to describe the semantic features of the corresponding second image blocks.
[0006] Therefore, in the target image segmentation stage, this application fuses multiple first image blocks from the image segmentation results of the target image to be reviewed, combining the semantic features of these first image blocks to obtain multiple second image blocks. Combining the semantic features of the first image blocks allows for better analysis of the relationships between them, effectively avoiding situations where a single first image block is not in violation, but the content of multiple adjacent first image blocks is related and constitutes a violation. Thus, before image review, by analyzing the semantic features of the image blocks and fusing them, a more accurate set of multiple second image blocks that are beneficial for image content review is determined, making this application applicable to the review of images with diverse styles, complex images, and images with local violations. In the content review stage, this application determines the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks, using a content review model. Thus, by comparing and analyzing global semantic features and individual local semantic features, as well as comparing and analyzing multiple local semantic features, the content moderation model can effectively identify complex image content, improving the robustness of the content moderation model in images of various styles and complex scenarios, as well as the accuracy of the moderation results, and reducing the false positive rate.
[0007] The content moderation method provided in this application can be applied to any field or scenario related to image content information recognition, such as illegal image recognition, image-based abnormal event recognition, and lesion area detection based on medical images.
[0008] In one possible implementation, determining the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks through a content moderation model includes: determining the image block association information of the plurality of second image blocks based on the target image and the plurality of second image blocks; inputting the global semantic features corresponding to the target image, the local semantic features corresponding to the plurality of second image blocks and the image block association information into the content moderation model to obtain the review result output by the content moderation model.
[0009] In one possible implementation, the image patch association information includes the positional distribution relationship of multiple second image patches in the target image.
[0010] By introducing image patch association information, the content moderation model can more accurately understand the spatial relationships between different parts of an image. Based on this understanding of spatial relationships, when analyzing the local semantic features of each second image patch, it can combine / refer to the local semantic features of other second image patches adjacent to that image patch, thereby better capturing the contextual information in the image and improving the ability to recognize subtle content in the image.
[0011] It should be noted that when reviewing the content of a target image, whether or not to input image patch association information (e.g., the positional distribution of the second image patches) for multiple second image patches can be flexibly chosen based on the review requirements and the characteristics of the image to be reviewed. For images to be reviewed in complex scenarios, introducing image patch association information can improve the accuracy of the review results and the model's generalization ability; while for images to be reviewed in simple scenarios, image patch association information can be omitted to simplify the review process and reduce computational costs and inference latency.
[0012] In one possible implementation, the review result of the target image includes at least one of the following:
[0013] The target image has an image type, which includes either a violation image or a normal image;
[0014] The type of violation in the target image;
[0015] The review scores of the plurality of second image blocks, each second image block corresponds to at least one review score, and each review score corresponds to a violation type;
[0016] The location information of the illegal region in the target image, wherein the illegal region includes at least one of the plurality of second image blocks.
[0017] In one possible implementation, the target image is a user-uploaded image to be edited. If the image type of the target image is a normal image, the review result also includes at least one sensitive area present in the target image.
[0018] After determining the review result of the target image through the content review model, the method further includes:
[0019] The system receives an image editing request from the user, the image editing request carrying operation information and region information, the operation information indicating the processing operation that the user expects to perform on the target image, and the region information indicating the target region in the target image that the user expects to perform the processing operation on;
[0020] Based on the region information, the operation information, and the at least one sensitive region, determine whether the target region meets the editing conditions;
[0021] If the target area does not meet the editing conditions, a prompt message will be output to indicate that the processing operation is prohibited in the target area. The target area not meeting the editing conditions means that the image after the processing operation is performed on the target area is a violation image.
[0022] In one possible implementation, determining whether the target region meets the editing conditions based on the region information, the operation information, and the at least one sensitive region includes:
[0023] If the target area overlaps with the at least one sensitive area, and the processing operation indicates that illegal content should be introduced into the target area, then the target area is determined to be inconsistent with the editing conditions.
[0024] Therefore, when a user uploads a target image, this application first performs content review to determine if the image violates any regulations. If the review determines the target image is normal, this application can also accurately identify potentially sensitive areas within the image that could lead to violations if subsequently modified, thus achieving effective supervision of editing operations. If the editing operation on the target image does not touch sensitive areas, such as removing tourists from a landscape photo, no intervention will be taken. Conversely, if the editing operation involves sensitive areas, such as modifying a border line on a map, the modification will be prohibited, thereby preventing the creation of violating images. Thus, through the above mechanism, this application's embodiments effectively ensure the compliance of output content while ensuring the availability of content review services, preventing the creation of violating images from the source.
[0025] In one possible implementation, determining the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks respectively includes: obtaining the global image features corresponding to the target image; obtaining the local image features corresponding to the plurality of second image blocks respectively to obtain a plurality of local image features; and determining the global semantic features and the local semantic features corresponding to the plurality of second image blocks respectively through a feature extraction model based on the global image features and the plurality of local image features.
[0026] Among them, the global image features are class tokens, and the local image features are patch tokens. The global semantic features and global image features are in one-to-one correspondence, both corresponding to the target image; multiple local semantic features and multiple local image features are in one-to-one correspondence, and each second image patch corresponds to one local image feature and one local semantic feature.
[0027] In one possible implementation, the global image features can be randomly generated, or they can be a fixed-dimensional vector, i.e., a class token, obtained by transforming the target image input through a specific embedding layer (e.g., word embedding). Similarly, for each of the multiple second image patches, each second image patch can also be transformed into a fixed-dimensional vector, i.e., a patch token, through a specific embedding layer.
[0028] In one possible implementation, the feature extraction model includes multiple network layers connected in sequence; the step of determining the global semantic features and the local semantic features corresponding to the multiple image patches based on the global image features and the multiple local image features through the feature extraction model includes: inputting the global image features and the multiple local image features into the multiple network layers to obtain semantic features output by the multiple network layers respectively; wherein, the semantic features output by each network layer include the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features respectively; and determining the global semantic features and the local semantic features corresponding to the multiple second image patches based on the semantic features output by the multiple network layers respectively.
[0029] In one possible implementation, the network layer includes an attention module and a feature extraction module. The attention module is used to determine the feature similarity between the global image features and the plurality of local image features. The feature extraction module is used to extract semantic information from the global image features and the plurality of local image features based on the feature similarity, thereby obtaining the semantic features corresponding to the global image features and the semantic features corresponding to the plurality of local image features respectively.
[0030] In one possible implementation, determining the global semantic features and the local semantic features corresponding to the multiple image patches based on the semantic features output by the multiple network layers respectively includes: determining the semantic features corresponding to the global image features output by the target network layer as the global semantic features, wherein the target network layer includes at least one of the multiple network layers; and determining the semantic features corresponding to the multiple local image features output by the target network layer as the local semantic features corresponding to the multiple second image patches respectively.
[0031] In one possible implementation, the target network layer is the last network layer among the plurality of network layers, or the target network layer includes a first network layer, a second network layer, and a third network layer, wherein the first network layer extracts semantic features based on pixel information of the image, the second network layer extracts semantic features based on entity information in the image, and the third network layer extracts semantic features based on entity relationships in the image.
[0032] The last network layer typically captures high-level, global features of the image (target image and multiple image patches), reflecting the overall meaning and contextual relationships of the image. Therefore, selecting the global semantic features and multiple local semantic features output by the last network layer ensures that the content moderation model fully utilizes the global dependencies and semantic information of each element in the target image to output correct image moderation results.
[0033] In one possible implementation, the target network layer comprises three network layers: the first network layer, the intermediate network layer, and the last network layer among multiple network layers. The intermediate network layer can be the K / 2th network layer among multiple network layers, where K is the total number of multiple network layers.
[0034] The first network layer typically captures basic, local features of the image (target image and multiple image patches). These features provide basic structural and textural information, essential for subsequent feature extraction and recognition tasks. Choosing the output of the first layer helps the content moderation model fuse and complement these basic features with semantic features from other layers when recognizing the content of the target image. The intermediate network layers typically capture mid-level features of the image, containing both semantic information and sufficient detail to support subsequent recognition tasks. Choosing the output of the intermediate network layers balances the semantic understanding and detail capture capabilities of the feature extraction model, helping the content moderation model perform accurate image recognition in complex scenarios.
[0035] In one possible implementation, obtaining the image block results of the target image to be identified includes:
[0036] Determine the segmentation strategy corresponding to the target image; perform region segmentation on the target image according to the segmentation strategy to obtain the plurality of first image blocks.
[0037] When determining the segmentation strategy for an image, the following factors can be considered: model input requirements, image content, and computational resources. Based on these factors, the segmentation strategy determined in this application can include any of the following: determining the block size, uniform segmentation, non-uniform segmentation, and overlapping segmentation.
[0038] Secondly, a content moderation device is provided, which has the function of implementing the content moderation method described in the first aspect. The content moderation device includes at least one module for implementing the content moderation method provided in the first aspect.
[0039] Thirdly, a computer device is provided, comprising a processor and a memory, the memory being used to store a computer program for executing the content moderation method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the content moderation method described in the first aspect.
[0040] In one possible implementation, the computer device may further include a communication bus for establishing a connection between the processor and the memory.
[0041] Fourthly, a computer-readable storage medium is provided, wherein the storage medium stores instructions that, when executed on a computer, cause the computer to perform the content moderation method described in the first aspect.
[0042] Fifthly, a computer program product comprising instructions is provided, which, when executed on a computer, causes the computer to perform the steps of the content moderation method described in the first aspect. Alternatively, a computer program is provided that, when executed on a computer, causes the computer to perform the steps of the content moderation method described in the first aspect.
[0043] The technical effects achieved by the second to fifth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description
[0044] Figure 1 is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0045] Figure 2 is a flowchart illustrating a content moderation method provided in an embodiment of this application;
[0046] Figure 3 is a schematic diagram of an image block segmentation result provided in an embodiment of this application;
[0047] Figure 4 is a flowchart illustrating another content moderation method provided in an embodiment of this application;
[0048] Figure 5 is a schematic diagram of a content review process based on the semantic features of a single network layer, provided by an embodiment of this application.
[0049] Figure 6 is a schematic diagram of a content review process based on semantic features of multiple network layers provided in an embodiment of this application;
[0050] Figure 7 is a schematic diagram of a process for locating illegal areas by recognizing image content according to an embodiment of this application;
[0051] Figure 8 is a schematic diagram of the structure of a content moderation device provided in an embodiment of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0053] To facilitate understanding, before explaining the content review method provided in this application, the terminology, application scenarios, and implementation environment involved in the embodiments of this application will be introduced first.
[0054] First, the terminology used in the embodiments of this application will be introduced.
[0055] 1. Transformer Architecture
[0056] As a deep learning architecture, the transformer is widely used in fields such as natural language processing (NLP), speech recognition, computer vision, and reinforcement learning, achieving significant results in these areas. For example, in NLP, deep learning models based on the transformer architecture (referred to as transformer models) can be used for tasks such as text classification, machine translation, named entity recognition, and sentiment analysis; in speech recognition, transformer models can be used for tasks such as speech signal recognition and speaker recognition; and in computer vision, transformer models can be used for tasks such as image classification, object detection, and image generation.
[0057] The Transformer architecture introduces a self-attention mechanism, which allows the model to dynamically focus on information at different positions in the input sequence when processing sequential data. This enables it to capture and learn the complex dependencies between elements within the input sequence, thus eliminating reliance on traditional recurrent neural network (RNN) or convolutional neural network (CNN) architectures. This self-attention mechanism makes the Transformer model more efficient and expressive when processing long sequences.
[0058] 2. Self-attention mechanism
[0059] Self-attention mechanisms work by calculating attention weights for each element (often called a "token") in the input sequence with respect to all other elements. These weights reflect the importance of each element in generating the current output. These weights are then used to perform a weighted summation of the input sequence, generating a context representation that incorporates information about the current element as well as information about other elements related to it.
[0060] For transformer models, the self-attention mechanism allows the model to refer to / consider the dependency of each element in the input sequence on other elements in the input sequence when computing the representation of each element.
[0061] Self-attention mechanisms can be implemented in various ways, including scaled dot product attention and multi-head attention. Scaled dot product attention calculates the similarity score of each query relative to all keys by multiplying the query and key for each element in the input sequence. This similarity score is then normalized using a normalized exponential function (softmax) to obtain attention weights. These attention weights are then used to weight and sum the value matrix to obtain the output representation of each element in the current sequence. Multi-head attention, on the other hand, allows the model to simultaneously focus on information from different locations. It divides the original input sequence into multiple groups (also called "heads"), each head undergoing a linear transformation using an independent weight matrix and calculating self-attention. In this way, each head can independently learn different attention weights, thereby enhancing the model's ability to focus on different parts of the input sequence.
[0062] 3. Class token
[0063] Class tokens are a special type of positional encoding in the transformer model. They are added to the image's embedding representation (such as an embedding vector) to extract global semantic features of the entire image during image processing.
[0064] 4. Patch token
[0065] A patch token is a token generated by processing each image patch after the input image is segmented into multiple small image patches (i.e., patches). For a given image, the multiple patch tokens corresponding to the image serve as input to the model, enabling the model to capture local detail information (e.g., local semantic information) of the image patches during the inference process.
[0066] Secondly, the application scenarios of the embodiments of this application will be introduced.
[0067] Multimodal big data models can process and integrate information from different modalities (such as text, images, and audio) to provide a more comprehensive and in-depth understanding, and are widely used in the field of multimodal content moderation and risk control (referred to as risk control). Multimodal content moderation technology refers to using various techniques such as computer vision, natural language processing, and audio analysis to review information from multiple modalities, including images, videos, audio, and text, in order to identify inappropriate or illegal content.
[0068] Taking image moderation as an example, the content moderation solutions provided in related technologies are based on the contrastive language-image pretraining (CLIP) model. The CLIP model first extracts features from the input image using an image encoder to generate image feature vectors; simultaneously, it uses a text encoder to extract features from predefined illegal text descriptions, generating illegal text feature vectors. Then, the CLIP model calculates the similarity between the image feature vectors and each illegal text feature vector. If a similarity exceeds a threshold, the image is considered to match the corresponding illegal text description, meaning the image may contain the illegal content defined by that text description. Finally, based on the similarity judgment, the CLIP model outputs the moderation result: whether the image contains illegal content or is a normal image (also known as a compliant image).
[0069] Because CLIP model training relies on large-scale image-text pair sample datasets, while these datasets cover a wide range of image and text content, they still struggle to encompass all possible complex and diverse scenes. Therefore, CLIP's performance may be affected when it encounters images not adequately represented in the training data. Furthermore, the CLIP model architecture aims to capture the global correspondence between images and text, rather than focusing on local details or complex contextual relationships, resulting in lower accuracy when processing complex images. In addition, CLIP's performance heavily depends on the correspondence between images and text; without sufficient accompanying textual descriptions, CLIP may fail to accurately understand the specific semantics of an image. Especially in complex image scenes, concise textual descriptions may not adequately convey all the key information within the image.
[0070] It is evident that limitations in training data, constraints in model architecture, and a lack of sufficient textual descriptions can all lead to unstable performance of the CLIP model when processing complex images.
[0071] In summary, current multimodal content moderation technologies mainly rely on models trained on real data. However, this approach is difficult to use when reviewing image information in the following three situations, resulting in lower accuracy of image review results.
[0072] (1) Content in various styles.
[0073] The images exhibit a diverse range of styles, from realistic natural landscapes to abstract artworks. In particular, some non-traditional image styles, such as anime / manga, ink painting, pixel art, and flat design, differ significantly in visual expression from traditional photographs or realistic paintings. This stylistic diversity requires content moderation models not only to recognize common content but also to adapt to and understand the elements and compositions within these non-traditional styles.
[0074] As an example, a two-dimensional comic book image might contain exaggerated expressions or actions uncommon in the real world. If a content moderation model judges content solely based on real-world logic, it might misclassify it as abnormal or inappropriate. Similarly, the blank spaces and evocative imagery in traditional Chinese ink paintings could confuse content moderation models that rely on pixel-level feature matching, making them unable to effectively identify such content.
[0075] (2) Multiple categories of illegal content.
[0076] There are many types of illegal content. These categories not only have unique visual characteristics, but are also often accompanied by complex contexts and cultural backgrounds, making automatic identification and classification extremely complicated.
[0077] (3) Only a small number of areas violated the rules.
[0078] In some cases, infringing content may only occupy a small portion of an image, while the rest is completely compliant. In such situations, traditional global analysis methods may overlook small-scale violations due to the overall compliance of the content, or lead to misjudgments because local features are not prominent enough.
[0079] As an example, a landscape photo might contain inappropriate content tucked away in a corner. If the content moderation model only analyzes the overall image, it might overlook that inconspicuous corner because most of the content is beautiful scenery. Alternatively, a long video might contain only a few seconds of inappropriate footage, with the rest being normal narration. In these cases, high-precision content moderation technology is needed to accurately capture and flag these inappropriate segments.
[0080] Based on this, this application provides a content moderation method. In the target image segmentation stage, this application fuses multiple first image blocks from the image segmentation results of the target image to be reviewed, combining the semantic features of these first image blocks to obtain multiple second image blocks. Combining the semantic features of the first image blocks allows for better analysis of the relationships between them, effectively avoiding situations where a single first image block is not in violation, but the content of multiple adjacent first image blocks is related and constitutes a violation. Thus, before image review, by analyzing the semantic features of the image blocks and performing image block fusion, more accurate and beneficial second image blocks for image content review are determined, making the technical solution provided by this application applicable to the review of images with diverse styles, complex images, and local violations. In the content review stage, this application determines the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks, using a content moderation model. Thus, by comparing and analyzing global semantic features and individual local semantic features, as well as comparing and analyzing multiple local semantic features, the content moderation model can effectively identify complex image content, improving the robustness of the content moderation model in images of various styles and complex scenarios, as well as the accuracy of the moderation results, and reducing the false positive rate.
[0081] The content review method provided in this application can be applied to any field or scenario related to image content information recognition, such as illegal image recognition, image-based abnormal event recognition, and lesion area detection based on medical images.
[0082] As an example, as described above, the content review method provided in the embodiments of this application can be used to determine whether there is any illegal content in the target image by identifying the content in the target image.
[0083] As another example, in the field of smart homes, such as home security monitoring, home robots, and smart assistants, the content review method provided in the embodiments of this application can be used to identify the content of the monitoring screen in home surveillance videos in order to detect whether there are potential dangers or inappropriate behaviors. For example, detecting whether a tap is left running or a child / elderly person falls down.
[0084] As another example, for medical images such as X-rays, computed tomography (CT), and magnetic resonance imaging (MRI), the content review scheme provided in the embodiments of this application can also be used to identify the content of the above medical images in order to detect abnormal or lesion areas.
[0085] It should be noted that the above application scenarios are merely examples. In actual applications, the content review method provided in this application embodiment can also be applied to other scenarios. This application embodiment does not limit these scenarios, and will not provide further examples here.
[0086] It should be understood that the "review" involved in the embodiments of this application is merely an explanation of the image content review process, using whether the image violates regulations as the review result. In different application scenarios, the content to be reviewed (or the points of focus during the review process) may differ for different images; moreover, "content review" may also be replaced with descriptions such as "information recognition" or "object detection." However, regarding the processing flow of the input image, if the image processing flow adopts the inventive concept of the embodiments of this application, any modifications and substitutions made thereto should be included within the protection scope of the embodiments of this application.
[0087] Finally, the implementation environment of the embodiments of this application will be described.
[0088] The content moderation method provided in this application can be applied to a single computer device, such as a terminal or server; it can also be applied to a cluster of multiple computer devices to achieve the same result through distributed components (belonging to different computer devices); and it can also be applied to a cloud platform to integrate with client terminal systems through application programming interfaces (APIs) or software development kits (SDKs) to provide content moderation services to client terminals. This application does not impose any limitations on this.
[0089] Please refer to Figure 1, which is a schematic diagram of a computer device according to an embodiment of this application. The computer device includes at least one processor 101, a communication bus 102, a memory 103, and at least one communication interface 104.
[0090] The processor 101 can be a general-purpose central processing unit (CPU), graphics processing unit (GPU), network processor (NP), neural-network processing unit (NPU), microprocessor, or one or more integrated circuits for implementing the scheme of this application, such as application-specific integrated circuit (ASIC), programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0091] The communication bus 102 is used to transmit information between the aforementioned components. The communication bus 102 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in Figure 1, but this does not mean that there is only one bus or one type of bus.
[0092] The memory 103 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile optical disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but not limited thereto. The memory 103 may exist independently and be connected to the processor 101 via the communication bus 102. The memory 103 may also be integrated with the processor 101.
[0093] Communication interface 104 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 104 includes a wired communication interface and may also include a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0094] As an example, processor 101 may include one or more CPUs, such as CPU0 and CPU1 as shown in FIG1.
[0095] As an example, a computer device may include multiple processors, such as processor 101 and processor 105 as shown in Figure 1. Each of these processors may be a single-core processor or a multi-core processor. Here, "processor" may refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0096] In some embodiments, the computer device may further include output devices and input devices. The output device communicates with the processor 101 and can display information in various ways. For example, the output device may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device communicates with the processor 101 and can receive user input in various ways. For example, the input device may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0097] In some embodiments, memory 103 is used to store program code 110 for executing the scheme of this application, and processor 101 can execute program code 110 stored in memory 103. The program code 110 may include one or more software modules, and the computer device can implement the content moderation method provided in the embodiment of FIG2 below through processor 101 and program code 110 in memory 103.
[0098] It should be noted that the application scenarios and implementation environments described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the emergence of new application scenarios and the evolution of implementation environments, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0099] Next, the content review method provided in the embodiments of this application will be explained in detail.
[0100] Figure 2 is a flowchart of a content moderation method provided in an embodiment of this application, exemplified by applying the method to the computer device shown in Figure 1. Referring to Figure 2, the method includes the following steps.
[0101] Step 201: Obtain the image block results of the target image to be reviewed. The image block results include multiple first image blocks.
[0102] The multiple first image blocks are obtained by dividing the target image into blocks, with each first image block representing a region within the target image. For example, when certain regions in the target image have significant features, these regions can be divided into smaller blocks to better capture these features; while in other regions where features are not obvious, they can be divided into larger blocks to reduce computational load.
[0103] It should be noted that there may be overlapping image regions between any two first image blocks among the multiple first image blocks, or there may be no overlapping image regions. This application embodiment does not impose any restrictions on this.
[0104] In some embodiments, step 201 can be implemented by: determining the segmentation strategy corresponding to the target image; performing region segmentation on the target image according to the segmentation strategy to obtain multiple first image blocks.
[0105] In one possible implementation, when determining the segmentation strategy for an image, the following factors can be considered: model input requirements, image content, and computational resources.
[0106] As an example, considering that multiple image patches may subsequently be input into feature extraction models and / or content moderation models for processing, the segmentation strategy can be determined based on the specific requirements of the feature extraction model regarding the input image size when segmenting the target image. For instance, if the feature extraction model requires a fixed-size input image, the target image needs to be segmented according to that fixed size to meet that requirement. Of course, if the feature extraction model can handle images of arbitrary sizes, then an appropriate segmentation strategy can be selected based on actual needs.
[0107] As another example, the content and attributes of the target image are also important factors in determining the segmentation strategy. For instance, when processing target images with significant local features, it may be necessary to divide the target image into smaller blocks to better capture these features. However, when processing target images with more obvious global features, overly detailed segmentation may not be necessary.
[0108] As another example, the partitioning strategy also needs to consider available computational resources. Dividing the target image into smaller blocks (increasing the number of image blocks) may increase computational cost because feature extraction is required for each image block. Therefore, when computational resources are limited, a trade-off between the number of blocks and computational efficiency may need to be struck.
[0109] Based on the above factors, the segmentation strategy determined in the embodiments of this application may include any of the following:
[0110] (1) Determine the block size: Based on the requirements of the feature extraction model and the characteristics of the target image, determine the appropriate block size. The block size should be small enough to capture the local features of the target image, but it should not be too small to avoid increasing unnecessary computation.
[0111] (2) Uniform segmentation: Divide the target image into several blocks of the same size. This strategy is suitable for situations where the image content is relatively uniformly distributed.
[0112] (3) Non-uniform segmentation: The target image is segmented non-uniformly according to the complexity of the target image content and the importance of local features. For example, when some regions in the target image have significant features, these regions can be divided into smaller blocks to better capture these features; while in other regions where features are not obvious, they can be divided into larger blocks to reduce the amount of computation.
[0113] (4) Overlapping blocks: In order to capture edge and detail features in the target image, an overlapping block strategy can be considered. That is, there is a certain overlap between adjacent image blocks, which can increase the robustness and accuracy of feature extraction.
[0114] Step 202: Based on the semantic features of multiple first image blocks, fuse the multiple first image blocks to obtain multiple second image blocks.
[0115] The semantic features of the first image block can be obtained through a feature extraction model, that is, by sequentially extracting features from the first image block through multiple network layers in the feature extraction model, the semantic features of the first image block can be obtained.
[0116] In this embodiment of the application, the semantic features of multiple first image blocks are obtained in a manner that may be the same as or different from the implementation of obtaining the global semantic features of the target image and the local semantic features corresponding to the multiple second image blocks as described later.
[0117] It should be noted that when fusing multiple first image blocks, adjacent image regions from two first image blocks can be stitched together to obtain a second image block. Alternatively, two or more first image blocks can be directly stitched together to obtain a second image block. In other words, the second image block can include half of the image region of first image block A and half of the image region of first image block B; of course, the second image block can also include first image blocks C, D, and E. This application does not impose any limitations on this.
[0118] It should be understood that the number of multiple second image blocks is less than or equal to the number of multiple first image blocks, and each second image block is not exactly the same as each of the first image blocks, and may contain all or part of the image information of at least one first image block.
[0119] Therefore, in the target image segmentation stage, the embodiments of this application, by combining the semantic features of the first image blocks, can better analyze the correlation between the various first image blocks, thereby effectively avoiding situations where a single first image block is not in violation, but the content of multiple adjacent first image blocks is related and constitutes a violation. Thus, before image review, by analyzing the semantic features of image blocks and performing image block fusion, more accurate and beneficial second image blocks for image content review are determined, making the technical solution provided by the embodiments of this application applicable to the review of images with diverse styles, complex images, and local violations.
[0120] Step 203: Based on the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks, the review result of the target image is determined by the content review model. The global semantic features are used to describe the semantic information of the target image as a whole, and the local semantic features are used to describe the semantic information of the corresponding second image blocks.
[0121] In some embodiments, the process of determining the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks can be as follows: obtaining the global image features corresponding to the target image; obtaining the local image features corresponding to the multiple second image blocks to obtain multiple local image features; and determining the global semantic features and the local semantic features corresponding to the multiple second image blocks based on the global image features and the multiple local image features through a feature extraction model.
[0122] In this embodiment, the global image feature is a class token, and the local image feature is a patch token. The global semantic feature and the global image feature are in one-to-one correspondence, both corresponding to the target image; multiple local semantic features and multiple local image features are in one-to-one correspondence, and each second image patch corresponds to one local image feature and one local semantic feature.
[0123] It's important to note that the class token is a learnable embedding vector that is continuously optimized during the training of the feature extraction model. Before training begins, the class token is typically initialized randomly. This means that the initial value of the class token is random, but it will be adjusted during subsequent training using the backpropagation algorithm.
[0124] In one possible implementation, the global image features can be randomly generated or a fixed-dimensional vector obtained by transforming the target image through a specific embedding layer (e.g., word embedding), i.e., a class token. Similarly, for each of the multiple second image patches, each second image patch can also be transformed into a fixed-dimensional vector through a specific embedding layer to obtain local image features, i.e., patch tokens.
[0125] In one possible implementation, the feature extraction model includes multiple network layers connected in sequence. The process of determining the global semantic features and the local semantic features corresponding to the multiple second image patches based on global image features and multiple local image features can be as follows: inputting the global image features and multiple local image features into multiple network layers to obtain the semantic features output by the multiple network layers respectively; wherein, the semantic features output by each network layer include the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features respectively; and determining the global semantic features and the local semantic features corresponding to the multiple second image patches based on the semantic features output by the multiple network layers respectively.
[0126] As an example, the feature extraction model is a transformer model. The decoder in the transformer model is used as the feature extractor to extract semantic features from global image features and multiple local image features, so as to obtain global semantic features and multiple local semantic features. The multiple network layers are stacked layers of the decoder in the transformer model.
[0127] In one possible implementation, each network layer includes an attention module and a feature extraction module. The attention module is used to determine the feature similarity between global image features and multiple local image features, and the feature extraction module is used to extract semantic information from global image features and multiple local image features based on feature similarity, thereby obtaining semantic features corresponding to global image features and semantic features corresponding to multiple local image features respectively.
[0128] As an example, the feature extraction model is a transformer model. Global image features (i.e., the class token mentioned above) and multiple local image features (i.e., patch tokens of multiple image patches) are used as input to the transformer model. They are processed through a multi-layered self-attention mechanism and a feedforward network to generate richer semantic representations, namely global semantic features and multiple local semantic features.
[0129] It should be noted that since the class token and patch token will be fed together into the encoder of the transformer model, and the transformer encoder requires all input vectors to have the same dimension, the dimension of the class token needs to match the dimension of the second image patch embedding vector when determining the global image features (i.e., class token) of the target image.
[0130] In one possible implementation, since each network layer extracts different semantic information, the process of determining the global semantic features and the local semantic features corresponding to the multiple second image patches based on the semantic features output by multiple network layers can be as follows: the semantic features corresponding to the global image features output by the target network layer are determined as the global semantic features; the semantic features corresponding to the multiple local image features output by the target network layer are determined as the local semantic features corresponding to the multiple second image patches. The target network layer includes at least one of the multiple network layers.
[0131] As an example, the target network layer is the last network layer among multiple network layers.
[0132] The last network layer typically captures high-level, global features of the image (target image and multiple second image patches). These features reflect the overall meaning and contextual relationships of the image, which are crucial for image recognition tasks. Therefore, selecting the global semantic features and multiple local semantic features output by the last network layer ensures that the content moderation model fully utilizes the global dependencies and semantic information of various elements in the target image to output correct moderation results.
[0133] As another example, the target network layer includes a first network layer, a second network layer, and a third network layer from multiple network layers. The first network layer extracts semantic features based on pixel information of the image, the second network layer extracts semantic features based on entity information in the image, and the third network layer extracts semantic features based on entity relationships in the image. For example, the target network layer can be the first network layer, the K / 2th network layer, and the last network layer from multiple network layers, where K is the total number of network layers.
[0134] The first network layer typically captures basic, local features of the image (target image and multiple image patches). These features provide basic structural and textural information, essential for subsequent feature extraction and recognition tasks. Choosing the output of the first layer helps the content moderation model fuse and complement these basic features with semantic features from other layers when moderring the target image. Intermediate network layers (e.g., the K / 2th layer) typically capture mid-level features of the image. These features contain some semantic information while retaining sufficient detail to support subsequent recognition tasks. Choosing the output of intermediate network layers balances the semantic understanding and detail capture capabilities of the feature extraction model, helping the content moderation model to accurately moderate images in complex scenarios.
[0135] It should be noted that the process of obtaining the semantic features of multiple first image blocks in step 202 above can also be implemented through a feature extraction model. The process of obtaining the semantic features of the first image blocks is similar to the process of obtaining the local semantic features of the second image blocks and the process of obtaining the global semantic features of the target image, so it will not be described again.
[0136] To facilitate understanding, the following section will use the vision transformer (ViT) model as an example to illustrate the process of determining global semantic features and multiple local semantic features.
[0137] It's worth noting that the ViT model applies the traditional transformer structure to image processing. Its core idea is to segment the image into a series of small blocks and treat these blocks as sequential data. It uses the transformer's powerful self-attention mechanism to directly capture global dependencies, thereby improving processing efficiency and results.
[0138] In one possible implementation, the process of using the ViT model to determine global semantic features and multiple local semantic features may include the following steps (1)-(4).
[0139] (1) Construct the input sequence.
[0140] After converting multiple second image patches into embedding vectors, a sequence of embedding vectors (i.e., patch tokens) containing multiple second image patches is obtained, denoted as (P1, P2, ..., P...). nAt the beginning of the sequence, the global image features of the target image (i.e., the class token, denoted as C) are added, thus constructing a new sequence containing the class token and the patch tokens corresponding to multiple second image patches, denoted as (C, P1, P2...P...). n ).
[0141] (2) Add location encoding.
[0142] Since the transformer model lacks spatial awareness, a positional encoding needs to be added to each image feature (i.e., the class token and multiple patch tokens mentioned above). This positional encoding can be fixed or trainable; it is typically added to the embedding vector to help the transformer model understand the location of each image feature within the target image.
[0143] (3) Input to the encoder of the transformer.
[0144] The constructed input sequence (C, P1, P2... P...) n The input sequence and its positional encoding are fed into the transformer encoder. The transformer encoder uses a self-attention mechanism to process the input sequence, allowing the class token to interact with the patch tokens of all image patches through the self-attention mechanism.
[0145] (4) Output global semantic features and multiple local semantic features.
[0146] After processing by the transformer encoder, the class token and patch tokens of multiple second image patches are extracted from its output. At this point, the class token already contains the global semantic information of the target image, and the patch token also contains the local semantic information of the corresponding second image patch.
[0147] As an example, suppose there is a 224*224 pixel RGB (a color standard, where R represents red, G represents green, and B represents blue) image that needs to be reviewed. The implementation process of extracting global image semantic features and multiple local semantic features according to the content review method provided in this application is as follows: First, based on the block division result of the target image and the semantic features of a single first image block, multiple second image blocks (e.g., small blocks of 16*16 pixels) are determined, resulting in 196 second image blocks. Each second image block is flattened into a one-dimensional vector and mapped to a 768-dimensional embedding space through a linear layer. At this point, an image block embedding vector sequence with a shape of [196, 768] is obtained. Next, a 768-dimensional class token is initialized and added to the beginning of the image block embedding vector sequence. Then, positional encoding is added to each embedding vector (including the class token). Finally, an input sequence with a shape of [197, 768] is obtained, where the first element is the class token, and the remaining elements are the embedding vectors of the second image blocks.
[0148] Furthermore, the input sequence is fed into a transformer encoder. After processing through multiple layers of self-attention mechanisms and feedforward networks, a class token and multiple patch tokens for second image patches can be extracted from the output. At this point, the class token already contains the global semantic information of the image, and the patch token also contains the local semantic information of the corresponding second image patch. Based on the class token and the patch tokens of multiple second image patches, tasks such as image content review can be performed.
[0149] In some embodiments, the content moderation model in step 203 can also be a transformer model.
[0150] In some embodiments, the feature extraction model and content moderation model described above can be integrated into the same neural network model, or they can be two independent neural network models. This application does not impose any limitations on this; for ease of explanation, this application only illustrates the feature extraction model and content moderation model as independent neural network models, but this does not constitute a limitation on the relationship between the feature extraction model and the content moderation model.
[0151] In one possible implementation, the process of determining the review result of the target image is as follows: input the global semantic features corresponding to the target image and the local semantic features corresponding to multiple second image blocks into the content review model to obtain the review result of the target image output by the content review model.
[0152] Therefore, global semantic features can capture the overall information and context of an image, helping to identify the subject or scene of the target image. Local semantic features, on the other hand, focus on the details of the second image patch, capturing potentially interesting information, illegal or sensitive content, etc., that may exist in the target image. By combining global and local semantic features, the content moderation model can more comprehensively understand image content, thereby improving the accuracy of image moderation results, enhancing the robustness of the model, increasing moderation efficiency, and adapting to diverse image moderation methods.
[0153] It should be understood that when the target image is a simple image, such as an image containing only one object, or an image with a clear background and a fixed object position, the content moderation model only needs to combine global semantic features and the local semantic features of multiple second image patches to analyze and determine whether the object violates regulations. In the above review process, the content moderation model only needs to focus on the semantic information related to the object in the image when reviewing the target image.
[0154] However, when the target image is an image containing a complex scene, such as a group photo containing multiple objects and people, the content moderation model needs to understand not only the local semantic features of each second image block (such as people, objects, background, etc.) but also analyze the relationship between these second image blocks in order to determine whether the content of these second image blocks is combined to form an illegal scene.
[0155] In another possible implementation, the process of determining the review result of the target image is as follows: based on the target image and multiple second image blocks, determine the image block association information of the multiple second image blocks; input the global semantic features corresponding to the target image, the local semantic features corresponding to the multiple second image blocks and the image block association information into the content review model to obtain the review result of the target image output by the content review model.
[0156] The image patch association information includes the positional distribution relationship of multiple second image patches in the target image. For example, for each second image patch, the image patch association information is used to indicate the neighboring image patches of that second image patch, and the positional distribution relationship between that second image patch and each of its neighboring second image patches in the target image. Of course, the image patch association information also indicates information such as the distance value between that second image patch and other second image patches; this embodiment of the application does not limit this.
[0157] As an example, referring to the schematic diagram of the second image block shown in Figure 3, when the target image is divided into four second image blocks, the image association information of these four second image blocks can be as follows: Image block 1 is located in the upper left of the target image, specifically to the left of image block 2 and above image block 3; Image block 2 is located in the upper right of the target image, specifically to the right of image block 1 and above image block 4; Image block 3 is located in the lower left of the target image, specifically below image block 1 and to the left of image block 4; Image block 4 is located in the lower right of the target image, specifically below image block 2 and to the right of image block 3. When the target image is divided into 9 second image blocks, the image association information of these 9 second image blocks can be as follows: Image block 1: adjacent to image block 2 and image block 4, and image block 1 is located to the left of image block 2 and above image block 4; Image block 2: adjacent to image block 1, image block 3 and image block 5, and image block 2 is located to the right of image block 1, above image block 5 and to the left of image block 3; Image block 3: adjacent to image block 2 and image block 6, and image block 3 is located to the right of image block 2 and above image block 6; and so on, image block 9: adjacent to image block 8 and image block 6, and image block 9 is located to the right of image block 8 and below image block 6.
[0158] It should be noted that, as shown in Figure 3, after image block fusion of multiple first image blocks, the multiple second image blocks may or may not have overlapping areas. In other words, the embodiments of this application aim to illustrate that the multiple second image blocks obtained by image block fusion are not completely identical; that is, any two second image blocks can be completely different or may have some identical content. The embodiments of this application do not impose any restrictions on this.
[0159] In some embodiments, if multiple second image blocks have overlapping regions, the image block association information may further include the overlapping direction and degree of overlap of the overlapping image blocks among the multiple second image blocks.
[0160] As an example, referring to the five second image blocks shown in Figure 3, the corresponding image block association information may include: image block 1 and image block 2 overlap in the horizontal direction, and the overlap information accounts for about 25% of the total information of image block 2 (that is, about 25% of the information in image block 2 is the same as that in image block 1).
[0161] By introducing image patch association information, the content moderation model can more accurately understand the spatial relationships between different parts of an image. Based on this understanding of spatial relationships, when analyzing the local semantic features of each second image patch, it can combine / reference the local semantic features of other second image patches adjacent to it, thereby better capturing the contextual information in the target image and improving the ability to recognize subtle content in the image.
[0162] As an example, suppose the target image to be reviewed includes a person, a car, and the surrounding environment; multiple second image blocks are included: second image block 1 containing the person's legs, and second image block 2 containing the car tires. Semantically, both second image blocks may appear normal. However, considering their positional distribution, when second image block 1 is below second image block 2, the person's legs may be trapped under the car tires. In this case, the content review model can identify a car accident scene based on the positional relationship of these two second image blocks. Conversely, when second image block 1 is in front of second image block 2, the person may be walking normally in front of the car. The content review model can then identify the target image as a normal image containing both vehicles and pedestrians based on the positional relationship of these two second image blocks.
[0163] Therefore, by inputting image patch association information from multiple second image patches, the content moderation model can more accurately understand the overall layout and relationships between objects in the target image, thus more accurately determining whether the target image contains inappropriate content. Furthermore, for different types of images, introducing image patch association information during content moderation allows the model to learn richer features, thereby improving its ability to recognize different images. Based on this, the content moderation model, by reviewing the target image's global semantic features, the local semantic features of multiple second image patches, and the image patch association information, can significantly improve review accuracy, enhance the model's generalization ability, increase review efficiency, and enhance the interpretability of image review results.
[0164] In summary, whether or not to input image patch association information (e.g., the positional distribution of image patches) for multiple second image patches during content review of a target image can be flexibly chosen based on the review requirements and the characteristics of the image to be reviewed. For images to be reviewed in complex scenarios, introducing image patch association information can improve the accuracy of the review results and the model's generalization ability; while for images to be reviewed in simple scenarios, image patch association information can be omitted to simplify the review process and reduce computational costs and inference latency.
[0165] In some embodiments, since the high-dimensional structures of the global image features (class token) and the local image features (patch tokens) of multiple second image patches are similar, the class tokens of multiple sample images can be used as training data to train the content moderation model. The information in the class token can be considered as the label of the corresponding sample image.
[0166] As explained earlier, a class token is a special token that represents the category information of the entire image. A patch token is a vector representation obtained by linear transformation after determining the second image patch based on the input image. Although class tokens and patch tokens differ in function and meaning, they share certain similarities in their high-dimensional structure. These similarities are mainly reflected in the following aspects:
[0167] (1) Vector representation: Both class tokens and patch tokens convert image information into high-dimensional vector representations. This vector representation enables the model to process and understand image data.
[0168] (2) Linear Transformation: Linear transformations are involved in the generation of both class tokens and patch tokens. For a class token, it may be a pre-defined, fixed vector, but in some cases it can be updated through linear transformations. For a patch token, it is a vector representation obtained by mapping image patches through linear transformations.
[0169] (3) Self-attention mechanism: In the ViT model, both class tokens and patch tokens participate in the computation of the self-attention mechanism. The self-attention mechanism allows the model to consider other tokens in the sequence when processing each token, thereby capturing long-distance dependencies between elements. This mechanism enables the model to better understand global information in the image.
[0170] (4) Position Encoding: Since the transformer model itself does not contain recurrent or convolutional structures, it cannot directly capture the sequence information. Therefore, position encoding is usually added to provide positional information before the class token and patch token are input into the model. Position encoding can be generated by a linear combination of sine and cosine functions and added to the token's embedding vector to ensure that the model can distinguish tokens at different positions.
[0171] Therefore, by leveraging the similarity of the high-dimensional structure of global image features (class token) and local image features (patch token) of multiple second image patches, and using class tokens from multiple sample images to train the content moderation model, the trained content moderation model can not only perform image moderation based on the global semantic features of the target image, but also refer to the local semantic features of multiple second image patches corresponding to the target image. This reduces the training cost while improving the robustness of the content moderation model.
[0172] In some embodiments, the review results of the target image determined by the content moderation model include the following four types:
[0173] (1) The image type of the target image, which includes illegal images or normal images.
[0174] The distinction between "violation image" and "normal image" is relative. The same image may be identified as a violation image by a content moderation model in some scenarios, but as a normal image in others, depending on the model's inherent violation judgment logic.
[0175] (2) Types of violations in the target image.
[0176] In other words, even if the target image is a violation image, the content moderation model can still output the violation type.
[0177] The violation type is a predefined type within the content review model, and it can be a first-level category type, or a second-level or third-level category type. This application's embodiments do not impose any restrictions on this.
[0178] As an example, the violation type could be advertising. Of course, other violation types can also be included, and can be flexibly set and adjusted according to the specific identification scenario and the user's personalized identification needs. This application embodiment does not impose any limitations on this.
[0179] (3) Review scores of multiple second image blocks, each second image block corresponds to at least one review score, and each review score corresponds to a violation type.
[0180] In other words, the content moderation model compares the local semantic features of each second image block with the violation types of multiple violation types, and outputs the moderation score of each second image block relative to each violation type.
[0181] It should be understood that for a given second image patch, the review score can correspond one-to-one with the violation type. When violation types have multiple levels of classification, a first-level violation type may correspond to multiple review scores under multiple second-level classifications.
[0182] (4) Location information of the violation area in the target image, wherein the violation area includes at least one of the multiple second image blocks.
[0183] In other words, the content moderation model can identify whether each second image block contains inappropriate content; and when it is determined that one or more second image blocks contain inappropriate content, it can output the location information of the inappropriate area.
[0184] As an example, assuming that the first-level categories include violation category 1 and violation category 2, and the second-level categories include violation category 1.1, violation category 1.2, violation category 2.1 and violation category 2.2, the content moderation model can determine the image moderation results as shown in Table 1 below for a second image block.
[0185] Table 1
[0186] Based on Table 1 above, from the perspective of the second image block, the content moderation model can directly output the moderation results for each second image block, that is, the last column in Table 1 is used as the output of the moderation results for the second image blocks. The content moderation model can also output the moderation score and detection threshold for each second image block, so that users can make secondary judgments based on the output moderation results, thereby achieving personalized content moderation and improving the accuracy of image moderation results.
[0187] Based on Table 1 above, from the perspective of the target image as a whole, the content moderation model can output the following review result for the target image: if all the review results of the multiple second image blocks corresponding to the target image are normal, the image type of the target image is: normal image. However, if at least one of the multiple second image blocks corresponding to the target image has a violation review result, the image type of the target image is: violation image.
[0188] It should be noted that Table 1 above is only an example. In actual applications, the violation categories may include only the first-level categories or multiple-level categories. Moreover, the review scores and detection thresholds can be flexibly adjusted according to the actual application scenarios and needs. This application embodiment does not impose any restrictions on this.
[0189] In some embodiments, the target image is an image uploaded by the user to be edited. If the review result of the target image determined by steps 201-203 is a normal image, the review result may also include at least one sensitive area present in the target image.
[0190] For certain positively sensitive information, such as landmarks, maps, and traffic signs, which have specific meanings and whose modification would significantly change their meanings, if the output of the edited content involves modifications to the aforementioned positively sensitive information during intelligent image enlargement or image editing of the target image, it may result in the target image being deemed to be in violation of regulations.
[0191] Based on this, after determining at least one sensitive region in the target image, the embodiments of this application can also detect editing operations on the target image based on the at least one sensitive region, thereby avoiding the situation where the target image is a normal image but is subsequently modified into an illegal image.
[0192] In one possible implementation, the content moderation method provided in this application further includes: receiving a user's image editing request, the image editing request carrying operation information and region information, the operation information indicating the processing operation the user expects to perform on the target image, and the region information indicating the target region in the target image where the user expects to perform the processing operation; determining whether the target region meets the editing conditions based on the region information, operation information, and at least one sensitive region; if the target region does not meet the editing conditions, outputting a prompt message to indicate that the target region is prohibited from performing the processing operation, the target region not meeting the editing conditions meaning that the image after performing the processing operation on the target region is a violation image.
[0193] If the target area overlaps with at least one sensitive area, and the processing operation indicates that illegal content is introduced into the target area, then the target area is determined to be non-compliant with the editing conditions.
[0194] Therefore, when a user uploads a target image, this embodiment of the application can perform content review on the target image through steps 201-203 to determine whether the target image violates any regulations. If the review determines the target image is normal, this embodiment of the application can also accurately identify potentially sensitive areas in the target image that could lead to violations if subsequently modified, thereby achieving effective supervision of editing operations. If the editing operation on the target image does not touch sensitive areas, such as removing tourists from a landscape photo, no intervention will be taken. Conversely, if the editing operation on the target image involves sensitive areas, such as modifying a border line on a map, the modification operation will be prohibited, thereby preventing the generation of violating images. Thus, through the above mechanism, this embodiment of the application ensures the availability of content review services while effectively guaranteeing the compliance of output content, preventing the generation of violating images from the source.
[0195] In summary, during the target image segmentation stage, this embodiment of the application fuses multiple first image blocks in the image segmentation result of the target image to be reviewed, combining the semantic features of the multiple first image blocks, to obtain multiple second image blocks. Combining the semantic features of the first image blocks allows for better analysis of the correlation between them, effectively avoiding situations where a single first image block is not in violation, but the content of multiple adjacent first image blocks is related and constitutes a violation. Thus, before image review, by analyzing the semantic features of the image blocks and performing image block fusion, more accurate and beneficial second image blocks for image content review are determined, making the technical solution provided by this embodiment applicable to the review of images with diverse styles, complex images, and local violations. During the content review stage, this embodiment of the application determines the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks, using a content review model based on the global semantic features and multiple local semantic features. Thus, by comparing and analyzing global semantic features and individual local semantic features, as well as comparing and analyzing multiple local semantic features, the content moderation model can effectively identify complex image content, improving the robustness of the content moderation model in images of various styles and complex violation scenarios, as well as the accuracy of the moderation results, and reducing the false positive rate.
[0196] Based on the content moderation method shown in Figure 2 above, the following example illustrates the interaction between the feature extraction model and the content moderation model to jointly complete the moderation of the target image. Assuming that multiple network layers in the feature extraction model are encoders and the number of network layers is k, the following example further explains the above content moderation scheme. This is intended to illustrate the overall implementation process of the content moderation scheme in this application embodiment and does not constitute a limitation on the content moderation scheme provided in this application embodiment.
[0197] Referring to Figure 4, the processing flow of the content moderation method provided in this application embodiment can be summarized as follows: For the target image to be reviewed, in the target image segmentation stage, the multiple first image blocks are first fused based on the semantic features of the multiple first image blocks corresponding to the target image to obtain multiple second image blocks. After determining the multiple second image blocks corresponding to the target image, the global image features corresponding to the target image and the local image features corresponding to the multiple second image blocks are obtained through embedding representation based on the target image and the multiple second image blocks. Then, the semantic information in the global image features and the multiple local image features is extracted through a feature extraction model to obtain the global semantic features of the target image and the local semantic features corresponding to the multiple second image blocks. Further, the global semantic features of the target image and the local semantic features corresponding to the multiple second image blocks are input into the content moderation model, and the review result of the target image is determined through the content moderation model.
[0198] As explained earlier, operations such as extracting semantic features from the target image and multiple second image patches, and determining the review result based on global semantic features and multiple local semantic features, can all be integrated into a single neural network model. As shown in the dashed box in Figure 4, the feature extraction model and the content review model can be integrated into a single neural network model. The feature extraction model can serve as the feature extraction module in this neural network model, used to extract semantic features; while the content review model can serve as the classification module in this neural network model, used to output the review result of the target image.
[0199] Figure 5 shows a semantic feature extracted based on a single network layer according to an embodiment of this application. The embodiment of this application can use the global semantic features and multiple local semantic features output by a single network layer as the main semantic information for content review of the target image.
[0200] Referring to Figure 5, the process of reviewing the target image can be as follows: Based on the target image to be reviewed, multiple second image patches that can be used for effective review are determined, each second image patch representing a local region in the target image. Then, a class token is generated as the global image feature of the target image (denoted as C); and local semantic features corresponding to multiple second image patches are generated (denoted as P1, P2, P3…P9); the input sequence (C, P1, P2, P3…P9) is input into the feature extraction model, and semantic features are extracted through multiple network layers in the feature extraction model. Encoding is performed layer by layer starting from the first network layer (denoted as encoder#1), and finally, comprehensive semantic features are obtained in the last network layer (denoted as encoder#K). The output of each network layer includes the global semantic features of the class token and the local semantic features of all patch tokens. The information from each layer interacts and merges through a self-attention mechanism, gradually improving the semantic information. Furthermore, the semantic features output by the last network layer (i.e., the global semantic features of the class token and the local semantic features of all patch tokens) are used as input to the content moderation model. The content moderation model scores the illegal content of each category and compares it with the preset detection threshold, thereby outputting the moderation result of the target image.
[0201] As an example, the review result of the target image may include: the target image is a violation image, or the target image is a normal image.
[0202] Figure 6 shows a semantic feature extracted based on multiple network layers provided in an embodiment of this application. The embodiment of this application can use the global semantic features and multiple local semantic features output by multiple network layers as the main semantic information for content review of the target image.
[0203] Referring to Figure 6, assuming the target image is an image of a girl cooking food in the kitchen, considering that the first network layer can extract the encoded pixel information of the bottom layer, the middle network layer can extract some individual information, and the last network layer can extract more specific and rich semantic information, the semantic features output by the first network layer (denoted as encoder#1), the middle network layer (e.g., encoder#K / 2), and the last network layer (denoted as encoder#K) can be used as input to the content moderation model, so as to output the moderation result of the target image through the content moderation model.
[0204] Each network layer outputs the semantic features of the class token and the semantic features of all patch tokens. The depth of the semantic features extracted by different network layers is different, but the number of semantic features output is exactly the same, which is equal to the sum of the number of target images and multiple second image patches.
[0205] It should be noted that, for the semantic features output by the feature extraction model shown in Figure 6, semantic feature number 0 is the global semantic feature of the target image, while semantic features numbered 1-9 are the local semantic features corresponding to the multiple second image patches. Furthermore, Figure 6 is merely an example using encoder#K / 2 as the intermediate network layer. In practical applications, the intermediate network layer can be any network layer other than the first and last network layers. Of course, in practical applications, the second network, the intermediate network layer, and the penultimate network layer can also be used as the target network to perform content moderation on the target image based on the output of this target network layer; this embodiment does not impose any limitations on this.
[0206] It should be understood that the only difference between Figure 6 and Figure 5 is the choice of network layer. The explanations of other technical features and the beneficial effects are the same as or similar to those in Figure 4, and will not be repeated here.
[0207] Figure 7 illustrates a semantic feature extracted from a single network layer according to an embodiment of this application. This embodiment can also use the global semantic features and multiple local semantic features output from a single network layer as the main semantic information for identifying illegal regions in a target image, thereby returning the precise location of the illegal region through a content moderation model. Referring to Figure 7, the semantic features output from the last network layer in the feature extraction model (i.e., the global semantic features of the class token and the local semantic features of all patch tokens) are input into the content moderation model. The content moderation model scores each category of illegal content and compares it with a preset detection threshold, thereby outputting the moderation result of the target image.
[0208] As an example, the review result of the target image shown in Figure 7 may include the precise location of the offending object (or the offending area). The offending area may contain one or more second image patches within the target image.
[0209] It should be understood that the content of the examples in Figures 4 to 7 above can also be incorporated into the relevant steps of the content review method shown in Figure 2 above. The relevant steps of the content review method shown in Figure 2 above can also be replaced, combined, etc. If the review process of the target image falls within the inventive concept of this application, the modifications and integrations made should be included within the protection scope of the technical solution of this application.
[0210] Next, the content review device involved in the embodiments of this application will be introduced.
[0211] Figure 8 is a schematic diagram of a content moderation device provided in an embodiment of this application. This content moderation device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be the device shown in Figure 1. Referring to Figure 8, the content moderation device includes: a first acquisition module 801, an image block fusion module 802, and an approval module 803.
[0212] The first acquisition module 801 is used to acquire the image block results of the target image to be reviewed. The image block results include multiple first image blocks. For detailed implementation process, please refer to the relevant description of step 201 in the embodiment shown in Figure 2, which will not be repeated here.
[0213] The image patch fusion module 802 is used to fuse multiple first image patches based on their semantic features to obtain multiple second image patches. For detailed implementation, please refer to the description of step 202 in the embodiment shown in Figure 2; it will not be repeated here.
[0214] The review module 803 is used to determine the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to multiple second image blocks, through a content review model. The global semantic features are used to describe the overall semantic features of the target image, and the local semantic features are used to describe the semantic features of the corresponding second image blocks. For a detailed implementation process, please refer to the relevant description of step 203 in the embodiment shown in Figure 2, which will not be repeated here.
[0215] In one possible implementation, the audit module 803 includes:
[0216] The information determination unit is used to determine the image block association information of the multiple second image blocks based on the target image and multiple second image blocks;
[0217] The review unit is used to input the global semantic features corresponding to the target image, the local semantic features corresponding to multiple second image blocks, and the image block association information into the content review model to obtain the review result output by the content review model.
[0218] In one possible implementation, the image patch association information is the positional distribution relationship of multiple second image patches in the target image.
[0219] In one possible implementation, the review result of the target image includes at least one of the following:
[0220] The image type of the target image, which can include either a violation image or a normal image;
[0221] The type of violation in the target image;
[0222] The review scores of multiple second image blocks, each second image block corresponds to at least one review score, and each review score corresponds to a violation type;
[0223] Location information of the violation area in the target image, wherein the violation area includes at least one of a plurality of second image blocks.
[0224] In one possible implementation, the target image is a user-uploaded image to be edited. If the target image is a normal image, the review result also includes at least one sensitive area present in the target image. The content review device further includes:
[0225] The receiving module is used to receive the user's image editing request. The image editing request carries operation information and region information. The operation information indicates the processing operation that the user expects to perform on the target image, and the region information indicates the target region in the target image that the user expects to perform the processing operation on.
[0226] The determination module is used to determine whether a target area meets the editing conditions based on area information, operation information, and at least one sensitive area.
[0227] The prompt module is used to output a prompt message if the target area does not meet the editing conditions, so as to indicate that the target area is prohibited from performing processing operations. The target area does not meet the editing conditions, which means that the image after the processing operation is performed on the target area is a violation image.
[0228] In one possible implementation, the module is specifically used for:
[0229] If the target area overlaps with at least one sensitive area, and the processing operation indicates that illegal content is introduced into the target area, then the target area is determined to be unsuitable for editing.
[0230] In one possible implementation, the content moderation device further includes:
[0231] The second acquisition module is used to acquire global image features corresponding to the target image; and acquire local image features corresponding to multiple second image blocks respectively to obtain multiple local image features;
[0232] The feature determination module is used to determine global semantic features and local semantic features corresponding to multiple second image blocks based on global image features and multiple local image features through a feature extraction model.
[0233] In one possible implementation, the feature extraction model includes multiple network layers connected in sequence; the feature determination module is specifically used for:
[0234] Global image features and multiple local image features are input into multiple network layers to obtain semantic features output by each network layer. The semantic features output by each network layer include the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features.
[0235] Based on the semantic features output by multiple network layers, global semantic features and local semantic features corresponding to multiple second image patches are determined.
[0236] In one possible implementation, the network layer includes an attention module and a feature extraction module. The attention module is used to determine the feature similarity between global image features and multiple local image features. The feature extraction module is used to extract semantic information from global image features and multiple local image features based on feature similarity, thereby obtaining the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features respectively.
[0237] In one possible implementation, the feature determination module is specifically used for:
[0238] The semantic features corresponding to the global image features output by the target network layer are determined as global semantic features. The target network layer includes at least one of multiple network layers.
[0239] The semantic features corresponding to the multiple local image features output by the target network layer are determined as the local semantic features corresponding to the multiple second image blocks.
[0240] In one possible implementation, the target network layer is the last network layer among multiple network layers, or the target network layer includes a first network layer, a second network layer, and a third network layer. The first network layer extracts semantic features based on the pixel information of the image, the second network layer extracts semantic features based on the entity information in the image, and the third network layer extracts semantic features based on the entity relationships in the image.
[0241] In one possible implementation, the first acquisition module 801 is specifically used for:
[0242] Determine the segmentation strategy corresponding to the target image;
[0243] The target image is segmented into regions according to the segmentation strategy to obtain multiple first image blocks.
[0244] In this embodiment, during the target image segmentation stage, the multiple first image blocks in the image segmentation result of the target image to be reviewed are fused together with their semantic features to obtain multiple second image blocks. Combining the semantic features of the first image blocks allows for better analysis of the correlation between them, effectively avoiding situations where a single first image block is not in violation, but the content of multiple adjacent first image blocks is related and constitutes a violation. Thus, before image review, by analyzing the semantic features of the image blocks and performing image block fusion, more accurate and beneficial second image blocks for image content review are determined, making the technical solution provided by this embodiment applicable to the review of images with diverse styles, complex images, and local violations. During the content review stage, this embodiment uses a content review model based on the global semantic features corresponding to the target image and the local semantic features corresponding to the multiple second image blocks to determine the review result of the target image. Thus, by comparing and analyzing global semantic features and individual local semantic features, as well as comparing and analyzing multiple local semantic features, the content moderation model can effectively identify complex image content, improving the robustness of the content moderation model in images of various styles and complex violation scenarios, as well as the accuracy of the moderation results, and reducing the false positive rate.
[0245] It should be noted that the content moderation device provided in the above embodiments is only illustrated by the division of the above functional modules when reviewing the content of the target image. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the content moderation device and the content moderation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0246] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the aforementioned content review method.
[0247] This application also provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the aforementioned content review method.
[0248] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0249] It should be understood that "multiple" as mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., do not necessarily imply that they are different.
[0250] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0251] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A content moderation method, characterized in that, The method includes: Obtain the image block results of the target image to be reviewed, wherein the image block results include multiple first image blocks; Based on the semantic features of the plurality of first image blocks, the plurality of first image blocks are fused to obtain a plurality of second image blocks; Based on the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks respectively, the review result of the target image is determined by the content review model. The global semantic features are used to describe the semantic features of the target image as a whole, and the local semantic features are used to describe the semantic features of the corresponding second image blocks.
2. The method as described in claim 1, characterized in that, The step of determining the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks, respectively, through a content moderation model, includes: Based on the target image and the plurality of second image blocks, determine the image block association information of the plurality of second image blocks; The global semantic features corresponding to the target image, the local semantic features corresponding to the plurality of second image blocks, and the image block association information are input into the content review model to obtain the review result output by the content review model.
3. The method as described in claim 2, characterized in that, The image block association information refers to the positional distribution relationship of the multiple second image blocks in the target image.
4. The method according to any one of claims 1-3, characterized in that, The review result of the target image includes at least one of the following: The target image has an image type, which includes either a violation image or a normal image; The type of violation in the target image; The review scores of the plurality of second image blocks, each second image block corresponds to at least one review score, and each review score corresponds to a violation type; The location information of the illegal region in the target image, wherein the illegal region includes at least one of the plurality of second image blocks.
5. The method as described in claim 4, characterized in that, The target image is a user-uploaded image to be edited. If the image type of the target image is a normal image, the review result also includes at least one sensitive area present in the target image. After determining the review result of the target image through the content review model, the method further includes: The system receives an image editing request from the user, the image editing request carrying operation information and region information, the operation information indicating the processing operation that the user expects to perform on the target image, and the region information indicating the target region in the target image that the user expects to perform the processing operation on; Based on the region information, the operation information, and the at least one sensitive region, determine whether the target region meets the editing conditions; If the target area does not meet the editing conditions, a prompt message will be output to indicate that the processing operation is prohibited in the target area. The target area not meeting the editing conditions means that the image after the processing operation is performed on the target area is a violation image.
6. The method as described in claim 5, characterized in that, The step of determining whether the target region meets the editing conditions based on the region information, the operation information, and the at least one sensitive region includes: If the target area overlaps with the at least one sensitive area, and the processing operation indicates that illegal content should be introduced into the target area, then the target area is determined to be inconsistent with the editing conditions.
7. The method according to any one of claims 1-6, characterized in that, Before determining the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks respectively through the content review model, the method further includes: Obtain the global image features corresponding to the target image; Obtain the local image features corresponding to the plurality of second image blocks respectively to obtain a plurality of local image features; Based on the global image features and the multiple local image features, the global semantic features and the local semantic features corresponding to the multiple second image blocks are determined by a feature extraction model.
8. The method as described in claim 7, characterized in that, The feature extraction model comprises multiple network layers connected in sequence; The step of determining the global semantic features and the local semantic features corresponding to the multiple second image patches respectively through a feature extraction model based on the global image features and the multiple local image features includes: The global image features and the multiple local image features are input into the multiple network layers to obtain semantic features output by the multiple network layers respectively; wherein, the semantic features output by each network layer include the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features respectively; Based on the semantic features output by the multiple network layers, the global semantic features and the local semantic features corresponding to the multiple second image patches are determined.
9. The method as described in claim 8, characterized in that, The network layer includes an attention module and a feature extraction module. The attention module is used to determine the feature similarity between the global image features and the plurality of local image features. The feature extraction module is used to extract semantic information from the global image features and the plurality of local image features based on the feature similarity, so as to obtain the semantic features corresponding to the global image features and the semantic features corresponding to the plurality of local image features respectively.
10. The method as described in claim 8 or 9, characterized in that, The step of determining the global semantic features and the local semantic features corresponding to the multiple second image patches based on the semantic features output by the multiple network layers includes: The semantic features corresponding to the global image features output by the target network layer are determined as the global semantic features, wherein the target network layer includes at least one of the plurality of network layers; The semantic features corresponding to the multiple local image features output by the target network layer are determined as the local semantic features corresponding to the multiple second image blocks.
11. The method as described in claim 10, characterized in that, The target network layer is the last network layer among the plurality of network layers, or the target network layer includes a first network layer, a second network layer and a third network layer, wherein the first network layer extracts semantic features based on pixel information of the image, the second network layer extracts semantic features based on entity information in the image, and the third network layer extracts semantic features based on entity relationships in the image.
12. The method according to any one of claims 1-11, characterized in that, The process of obtaining the image block results of the target image to be reviewed includes: Determine the segmentation strategy corresponding to the target image; The target image is segmented into regions according to the segmentation strategy to obtain the plurality of first image blocks.
13. A content moderation device, characterized in that, The device includes: The first acquisition module is used to acquire the image block results of the target image to be reviewed, and the image block results include multiple first image blocks; The image patch fusion module is used to fuse the multiple first image patches based on their semantic features to obtain multiple second image patches; The review module is used to determine the review result of the target image based on the global semantic features corresponding to the target image and the local semantic features corresponding to the plurality of second image blocks, respectively, through a content review model. The global semantic features are used to describe the semantic features of the target image as a whole, and the local semantic features are used to describe the semantic features of the corresponding second image blocks.
14. The apparatus as claimed in claim 13, characterized in that, The review module includes: An information determination unit is configured to determine image block association information of the plurality of second image blocks based on the target image and the plurality of second image blocks; The review unit is used to input the global semantic features corresponding to the target image, the local semantic features corresponding to the plurality of second image blocks respectively, and the image block association information into the content review model to obtain the review result output by the content review model.
15. The apparatus as claimed in claim 14, characterized in that, The image block association information refers to the positional distribution relationship of the multiple second image blocks in the target image.
16. The apparatus according to any one of claims 13-15, characterized in that, The review result of the target image includes at least one of the following: The target image has an image type, which includes either a violation image or a normal image; The type of violation in the target image; The review scores of the plurality of second image blocks, each second image block corresponds to at least one review score, and each review score corresponds to a violation type; The location information of the illegal region in the target image, wherein the illegal region includes at least one of the plurality of second image blocks.
17. The apparatus as claimed in claim 16, characterized in that, The target image is a user-uploaded image to be edited. If the image type of the target image is a normal image, the review result also includes at least one sensitive area present in the target image. The device further includes: A receiving module is configured to receive an image editing request from the user, the image editing request carrying operation information and region information, the operation information indicating the processing operation that the user expects to perform on the target image, and the region information indicating the target region in the target image that the user expects to perform the processing operation on; The determination module is used to determine whether the target area meets the editing conditions based on the area information, the operation information, and the at least one sensitive area; The prompt module is used to output a prompt message if the target area does not meet the editing conditions, so as to prompt that the processing operation is prohibited in the target area. The target area does not meet the editing conditions, which means that the image after the processing operation is performed on the target area is a violation image.
18. The apparatus as claimed in claim 17, characterized in that, The determining module is specifically used for: If the target area overlaps with the at least one sensitive area, and the processing operation indicates that illegal content should be introduced into the target area, then the target area is determined to be inconsistent with the editing conditions.
19. The apparatus according to any one of claims 13-18, characterized in that, The device further includes: The second acquisition module is used to acquire global image features corresponding to the target image; and acquire local image features corresponding to the plurality of second image blocks respectively, to obtain a plurality of local image features; The feature determination module is used to determine the global semantic features and the local semantic features corresponding to the multiple second image blocks respectively, based on the global image features and the multiple local image features, through a feature extraction model.
20. The apparatus as claimed in claim 19, characterized in that, The feature extraction model comprises multiple network layers connected in sequence; the feature determination module is specifically used for: The global image features and the multiple local image features are input into the multiple network layers to obtain semantic features output by the multiple network layers respectively; wherein, the semantic features output by each network layer include the semantic features corresponding to the global image features and the semantic features corresponding to the multiple local image features respectively; Based on the semantic features output by the multiple network layers, the global semantic features and the local semantic features corresponding to the multiple second image patches are determined.
21. The apparatus as claimed in claim 20, characterized in that, The network layer includes an attention module and a feature extraction module. The attention module is used to determine the feature similarity between the global image features and the plurality of local image features. The feature extraction module is used to extract semantic information from the global image features and the plurality of local image features based on the feature similarity, so as to obtain the semantic features corresponding to the global image features and the semantic features corresponding to the plurality of local image features respectively.
22. The apparatus as claimed in claim 20 or 21, characterized in that, The feature determination module is specifically used for: The semantic features corresponding to the global image features output by the target network layer are determined as the global semantic features, wherein the target network layer includes at least one of the plurality of network layers; The semantic features corresponding to the multiple local image features output by the target network layer are determined as the local semantic features corresponding to the multiple second image blocks.
23. The apparatus as claimed in claim 22, characterized in that, The target network layer is the last network layer among the plurality of network layers, or the target network layer includes a first network layer, a second network layer and a third network layer, wherein the first network layer extracts semantic features based on pixel information of the image, the second network layer extracts semantic features based on entity information in the image, and the third network layer extracts semantic features based on entity relationships in the image.
24. The apparatus according to any one of claims 13-23, characterized in that, The first acquisition module is specifically used for: Determine the segmentation strategy corresponding to the target image; The target image is segmented into regions according to the segmentation strategy to obtain the plurality of first image blocks.
25. A computer device, characterized in that, The computer device includes a processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1-12.
26. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a computer, implements the method described in any one of claims 1-12.
27. A computer program product, characterized in that, The computer program product stores computer instructions, which, when executed by a computer, implement the method described in any one of claims 1-12.