Image review method and related device

By using cross-modal feature alignment technology, image and text features are mapped to the same space. By utilizing the semantic expression and similarity threshold of harmful text features, the efficiency and cost issues of reviewing various types of illegal images are solved, and efficient and accurate image review is achieved.

CN117351334BActive Publication Date: 2026-07-24TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2023-10-12
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies require training separate review models for each type of inappropriate image, resulting in slow model updates and iterations, high costs, and difficulty in efficiently handling the ever-increasing stream of inappropriate image content.

Method used

By aligning cross-modal features, image features and text features are mapped to the same space. By utilizing the semantic expression and similarity threshold of harmful text features, the similarity between the image and the harmful text features is directly calculated, thus achieving cross-modal image review.

Benefits of technology

It reduces development and maintenance costs, improves the efficiency and accuracy of reviewing inappropriate images, enables rapid response to new inappropriate content, and reduces the difficulty of model training and iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351334B_ABST
    Figure CN117351334B_ABST
Patent Text Reader

Abstract

The application discloses an image auditing method and related equipment. In the image auditing method, the inventory image features corresponding to the semantic expressions associated with the harmful text features, so the features of the two modalities of the image modality and the text modality can overcome the heterogeneous gap, so that the similarity between the image features and the text features can be directly calculated and accurately calculated. In addition, because each harmful text feature belongs to multiple violation categories in terms of semantic expression, and each harmful text corresponds to a harmful threshold that can be used as a basis for judgment, the method does not need to train a corresponding auditing model for different harmful types of content as in the traditional way, and through a set of cross-modal auditing method, the auditing problem of various types of violation images can be solved, the development cost and operation difficulty are reduced, and the positive semantic content of image dissemination is efficiently ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to image review methods and related equipment. Background Technology

[0002] As more and more UGC (user-generated content) is published on internet platforms, internet regulators are requiring the review of UGC content to control the display of illegal or unqualified content on the platform and maintain a clean internet environment.

[0003] Because UGC content, especially image content, is extremely rich, and the backgrounds and purposes of its creators vary greatly, the reasons for images being considered inappropriate are also diverse, such as inappropriate image behavior, false advertising, sensitive references, or the presence of prohibited words. Therefore, currently, for each type of inappropriate image, it is necessary to extensively collect and label large amounts of image data, and then train individual models using different types of image data, such as cartoon behavior models, real-person behavior models, advertising label models, and sensitive text models, in order to obtain a review model capable of screening out various types of inappropriate images.

[0004] Faced with a constant stream of various illegal image contents, the traditional approach of training individual review models for each type of illegal image data is problematic. This leads to an increase in the number of models to train or lengthy iterations of older models, resulting in slow updates, high development and maintenance costs. Currently, relevant technologies do not offer effective solutions to these challenges. Summary of the Invention

[0005] This application provides an image review method and related equipment for efficiently and cost-effectively reviewing different types of harmful images.

[0006] The first aspect of this application provides an image review method, including:

[0007] The image to be examined is input into the image feature extraction model to output the image features of the image to be examined;

[0008] Calculate the similarity between the image features and each harmful text feature in the feature library; wherein, the harmful text features are associated with stock image features corresponding to semantic expressions, and each harmful text feature belongs to multiple violation categories in terms of semantic expression;

[0009] For the similarity of each of the aforementioned harmful text features, determine the relationship between the harmful threshold of the harmful text feature and the similarity.

[0010] If the similarity is greater than the harmful threshold, then the image to be reviewed is determined to match the semantic expression of the harmful text and is therefore considered a harmful image.

[0011] A second aspect of this application provides an electronic device, including:

[0012] Central processing unit, memory, and input / output interfaces;

[0013] The memory is either a short-term storage memory or a persistent storage memory;

[0014] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first aspect of the embodiments of this application or any specific implementation thereof.

[0015] A third aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect of this application or any specific implementation thereof.

[0016] A fourth aspect of this application provides a computer program product comprising instructions or a computer program, which, when run on a computer, causes the computer to perform the method described in the first aspect of this application or any specific implementation thereof.

[0017] As can be seen from the above technical solutions, the embodiments of this application have at least the following advantages:

[0018] Because harmful text features are associated with corresponding inventory image features that express semantic meaning, the features from both the image and text modalities can overcome the heterogeneity gap, allowing for direct and accurate calculation of the similarity between image and text features. Furthermore, since each harmful text feature belongs to multiple violation categories in terms of semantic expression, and each harmful text has its own corresponding harmful threshold for judgment, this method eliminates the need to train separate review models for different types of harmful content, as is the case with traditional methods. A single cross-modal review approach can solve the review problems of various types of violating images, reducing development costs and operational difficulties, thereby efficiently ensuring the dissemination of positive semantic content in images. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0020] It should be noted that although the steps in the flowcharts (if any) involved in the embodiments are drawn sequentially according to the arrows, unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0021] Figure 1 This is a schematic diagram of a system architecture for a method according to an embodiment of this application;

[0022] Figures 2 to 7 This is a flowchart illustrating the method of an embodiment of this application;

[0023] Figure 8 This is a schematic flowchart of the model training method in an embodiment of this application;

[0024] Figure 9 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] In the following description, expressions such as "one specific implementation" or "one specific example" describe a subset of all possible embodiments. However, it is understood that "one specific implementation" or "one specific example" can be the same or different subset of all possible embodiments and can be combined with each other without conflict. In the following description, the term "multiple" means at least two. When a certain value mentioned in this application reaches a threshold (if it exists), in some specific examples, it may include the former being greater than the latter. When "any" or "at least one" or similar expressions are mentioned, it specifically refers to any one of the listed examples or any combination of these examples.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0029] For ease of understanding and explanation, before providing a further detailed description of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0030] Image review system: A software system that uses algorithms and human intervention to classify and filter image content based on risk. Generally, after receiving an input image, the system uses artificial intelligence algorithms to label and categorize the content; then, risky content undergoes manual review, and if problems are confirmed, it is filtered and processed through methods such as deletion, removal, or making it visible only to the user.

[0031] Image embedding feature: Also known as image feature, it is a feature vector that reduces image information to one dimension through model algorithms. This vector compresses the image content information, making it easier to support subsequent tasks such as image classification and text description generation.

[0032] Text embedding features, also known as text features, are similar to image features. They are created by using model algorithms to reduce text information to a one-dimensional feature vector. This vector compresses the content information of the text, making it easier to support subsequent tasks such as text classification and similarity comparison.

[0033] UGC: User Generate Content.

[0034] AIGC: AI Generate Content, refers to content generated or created through artificial intelligence algorithms.

[0035] Harmful images: Images that are detrimental to the content ecosystem, such as images depicting inappropriate behavior, false advertising, sensitive references, or containing prohibited words. With the increasing abundance of UGC and AIGC content, the reasons for and forms of harmful image violations have become increasingly diverse, posing a significant challenge to content moderation.

[0036] To better implement the image review method of this application, the following is provided: Figure 1 The diagram shows the system architecture of this method. The system may include at least one terminal device 101 and a server 102. Different types of applications may be installed on the terminal device 101, such as instant messaging applications, live streaming applications, conferencing applications, etc. The terminal device 101 may be a smartphone, tablet, laptop, desktop computer, smart vehicle, etc. The server 102 may be used to store application data and image data generated by the different types of applications on the terminal device 101. The server 102 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, etc.

[0037] This image review method can be executed by either terminal device 101 or server 102, or by both, depending on the specific application scenario. No restrictions are placed here. Specifically, when the method is executed jointly by terminal device 101 and server 102, the application scenario could be as follows: terminal device 101 receives and uploads the image to be reviewed input by the user; server 102 uses an image feature extraction model to extract the image features; server 102 calculates the similarity between the image features and the features of each harmful text; then, by using this similarity and its corresponding harmful threshold, it can determine which harmful text's semantic expression the image to be reviewed contains. This can also indicate that the image is a harmful image with inappropriate or malicious content, and will face deletion or other filtering processes, thereby creating a clean and comfortable online environment.

[0038] The above application environment is merely an example for ease of understanding. It is understood that the embodiments of this application are not limited to the above application environment.

[0039] The method described in this application will be further explained in detail below.

[0040] Please see Figure 2The first aspect of this application provides a specific embodiment of an image review method, which includes the following operation steps 20 to 23:

[0041] 20. Input the image to be reviewed into the image feature extraction model to output the image features of the image to be reviewed.

[0042] This image encoder includes, but is not limited to, encoders based on pre-trained images such as ViT (VisionTransformer) and CNN, or image encoders trained in other specific domains or scenarios, to achieve feature representation of the input image. Generally, image features can be represented in vector form.

[0043] 21. Calculate the similarity between image features and the features of each harmful text in the feature library.

[0044] Among them, harmful text features are associated with corresponding stock image features in terms of semantic expression, and each harmful text feature belongs to multiple violation categories in terms of semantic expression. Harmful text features can be features extracted from harmful text by a text feature extraction model; in this embodiment, the reason for associating harmful text features with corresponding stock image features is that the image features used for training the image feature extraction model (Image Encoder) and the text features used for training the text feature extraction model (Text Encoder) are aligned and mapped in the same space according to semantic correspondence, which can be simply referred to as cross-modal alignment, enabling the finding of the corresponding other through one, breaking down the gap caused by heterogeneous structures or different modalities. The harmful type (or violation category) can specifically be any one of the following: poor behavioral expression, containing false advertising, sensitive targeting, containing prohibited words, etc.

[0045] Cross-modal alignment in this application refers to aligning sub-elements of multimodal data, such as aligning multiple objects in an image with phrases (or words) in a sentence. This establishes a correlation in feature representation between two different modalities: image modality and text modality, to facilitate image-text matching. In short, cross-modal alignment aims to uncover the correlations between sub-elements of multimodal data. On the other hand, image-text matching can be understood as cross-modal semantic association alignment retrieval, where an instance in one modality is used to retrieve semantically related instances in another modality. This retrieval process measures the semantic similarity or relevance between images and text. For example, given an image (a dog bleeding on grass), image-text matching aims to retrieve text corresponding to the semantics shown in the image (e.g., "dog on grass," "bleeding dog"), and vice versa.

[0046] The features in this application can be in vector form, or features other than vector features, such as scalar features. Network models such as clip-vit-large-patch14 and BLIP2 can extract the above-mentioned text features and image features.

[0047] 22. For the similarity of each harmful text feature, determine the relationship between the harmful threshold of the harmful text feature and the similarity.

[0048] Since each harmful text in the feature library differs in terms of its harmfulness and semantic coverage, each harmful text should have a specific harmful threshold to accurately determine whether the image under review matches the semantic expression of that harmful text. For example, the similarity in this application embodiment can be any vector distance metric such as cosine distance and Euclidean distance, or it can be a weighted result of multiple distance metrics. The specific method can be determined as needed and is not limited.

[0049] 23. If the similarity is greater than the harmful threshold, the image to be reviewed is determined to be a harmful image.

[0050] If the similarity is greater than the harmful threshold, it is determined that the image under review matches the semantic expression of the harmful text. In other words, if it matches, it means that the image under review has the same type of semantic expression as the harmful text, such as expressing the same sensitive type of content. Therefore, it is a harmful image and will face filtering processes such as deletion.

[0051] In summary, this application's embodiments, through cross-modal feature alignment, enable the mapping of image and text features—two different modalities—to the same space, overcoming the heterogeneity gap between different modalities. This allows for direct and accurate calculation of the similarity between image and text features, and establishes a connection between the image feature extraction model and the text feature extraction model. Furthermore, since each harmful text in the feature library belongs to multiple semantic types, and each harmful text corresponds to a harmful threshold that can be used as a judgment criterion, this method eliminates the need to train separate review models for different harmful types of content, as is the case with traditional methods. A single cross-modal review method can solve the review problems of various types of illegal images, reducing development costs and operational difficulties, thereby efficiently ensuring the dissemination of positive semantic content in images.

[0052] Based on the examples above, some specific possible implementation examples will be provided below. In practical applications, the implementation content of these examples can be combined or implemented separately as needed, depending on the corresponding functional principles and application logic.

[0053] Please see Figures 2 to 7This application provides another specific embodiment of an image review method, which includes the following steps:

[0054] 20. Input the image to be reviewed into the image feature extraction model to output the image features of the image to be reviewed.

[0055] 21. Calculate the similarity between image features and the features of each harmful text in the feature library.

[0056] In some specific examples, step 21 may include: when the feature content of the feature library is within a preset level, using instruction set acceleration or matrix operation acceleration library, calculating the similarity between the image feature and each harmful text feature in the library; when the feature content of the feature library exceeds the preset level, using the feature index in the library established by hierarchical retrieval technology, locating at least one harmful text feature that is semantically similar to the image feature, and calculating the similarity between the image feature and at least one harmful text feature respectively.

[0057] like Figure 5 As shown, in practical applications, the text feature extraction model (or text encoder) can extract features from each text in a harmful text database (also known as a black database) to build a harmful text feature database. This step can be a one-time operation, and the feature database can be repeatedly applied after it is built. In the review process, for the input image to be reviewed, its image features Ii are extracted using the image feature extraction model Image Encoder, and then the similarity S{i} is calculated between it and each text feature in the harmful text feature database. The feature similarity between the image and the text can be measured using vector distances such as cosine distance and / or Chebyshev distance. For example, the similarity calculation between image features and text features can be implemented according to the size of the harmful text feature database.

[0058] 1) When the library size is within the tens of thousands (preset level): Since the library sample is small, a traversal calculation method can be used, or the library can be accelerated by using instruction set or matrix operations, and the similarity between all data in the library and the requested data can be calculated.

[0059] 2) When the database size is between tens of thousands and tens of millions: At this time, it is necessary to use some similar nearest neighbor retrieval methods, such as HNSW index or Faiss index. The core idea is a retrieval method with partitioning or hierarchical design. Through a small number of similarity calculations, the most similar data is found from coarse to fine. For example, by constructing a feature index within the database, it is possible to locate the high similarity between harmful text and image features in a certain layer and find highly similar harmful text.

[0060] 3) When the database size is over ten million: When the scale is very large, while using the above-mentioned hnsw index or faiss index and other retrieval ideas, it is also necessary to comprehensively consider issues such as memory, efficiency or storage in order to cope with the challenges of slow retrieval speed or low accuracy.

[0061] In some specific examples, there are many ways to generate harmful text, including manually collecting keywords, manually describing key content, and generating specific text content using specific image or language models. The first two are relatively accurate and targeted, but require human intervention, while the latter model generation methods are more automated, but still require manual review and optimization. The specific method can be determined by the user.

[0062] 22. For the similarity of each harmful text feature, determine the relationship between the harmful threshold of the harmful text feature and the similarity.

[0063] Each harmful text feature requires a judgment threshold (such as a harmful threshold). Therefore, to make the threshold design more reasonable, the process of determining each harmful threshold may include: constructing a set of harmful images and a set of normal images, each with more than a preset number of images; for each harmful text feature, calculating the similarity between the harmful text feature and each harmful image feature and each normal image feature to obtain the similarity measurement interval of the harmful text feature, which includes the interval formed by the maximum similarity and the minimum similarity; searching for the target judgment threshold within the similarity measurement interval as the harmful threshold of the harmful text feature, whereby the target judgment threshold is the value that makes all image features achieve the highest evaluation index in terms of both the actual violation category and the predicted violation category measurement.

[0064] like Figure 7As shown, the first step is to construct a comprehensive set of harmful and normal images to provide a search space for subsequent threshold determination. For each harmful text feature in the database, it generates a normal similarity set and a harmful similarity set with all normal and harmful image features. The range [Smin, Smax] defined by the maximum similarity Smax and minimum similarity Smin in the normal and harmful similarity sets can be called the similarity measurement interval (i.e., threshold search space) of that harmful text feature. Ideally, an optimal threshold should be selected for each harmful text feature to reduce the possibility of harmless images being misclassified as harmful text or incorrectly classified as harmful images. Therefore, a comprehensive selection scheme needs to be designed to rationally and logically search for the optimal target judgment threshold within the [Smin, Smax] space while pursuing the optimal search evaluation index Object (highest evaluation index). Generally, at least one of the following evaluation metrics—precision, recall, F1-Score, and false alarm rate—can be used to design the optimal search evaluation metric, Object. This design philosophy can include pursuing high levels of precision and recall, and low levels of false alarm rate. Specifically, precision P = the number of correctly predicted samples TP / (total predictions TP + FP), recall R = TP / (TP + FN), and F1-Score = (2 * P * R) / (P + R). True Positive (TP) represents the number of samples that are actually positive but predicted as positive; True Negative (TN) represents the number of samples that are actually negative but predicted as negative; False Positive (FP) represents the number of samples that are actually negative but predicted as positive; and False Negative (FN) represents the number of samples that are actually positive but predicted as negative.

[0065] 23. If the similarity is greater than the harmful threshold, the image to be reviewed is determined to be a harmful image.

[0066] like Figure 6 As shown, the similarity matching result between image features and each harmful text feature in the blacklist needs to be compared with the harmful threshold of each harmful text feature to obtain the final review result. Specifically, as long as the similarity S{i} is greater than the harmful threshold {i} corresponding to the blacklist text feature, the image to be reviewed can be considered to have matched the semantic expression of the blacklist text Text{i}. In addition, an image may match multiple blacklist texts (i.e., harmful texts). In this case, the higher the similarity, the higher the probability of matching.

[0067] As a possible implementation, when an image hits multiple harmful texts, the method of this application embodiment may further include: if the similarity of the features of multiple harmful texts is greater than their respective harmful thresholds, then the semantics of the harmful text corresponding to the highest similarity among the features hit by the image to be examined can be determined.

[0068] As a possible implementation, considering the possibility of a certain misjudgment rate during the review process, a harmless threshold can be set for each harmful text feature to improve fault tolerance and handling. Therefore, the method in this embodiment can further include: if the similarity between the image feature and the harmful text feature is between the harmless threshold and the harmful threshold corresponding to the harmful text feature, an early warning is issued; wherein the harmless threshold is less than the harmful threshold, and the harmless threshold is determined based on an evaluation index generated from the sample image features (similar to the operation described above for determining the harmful threshold). This early warning can be a request for manual review; the issuance of this warning can be considered as a determination that the image to be reviewed is suspected of being harmful.

[0069] As one possible implementation, the harmful text feature library can be updated, deleted, or modified according to business needs, thereby achieving rapid response to the need for reviewing new illegal content. Therefore, before calculating the similarity between image features and each harmful text feature in the feature library (i.e., before step 21), the image review method of this application embodiment may further include: adding, deleting, or modifying harmful text in the feature library to obtain a feature library for comparison with image features. It should be noted that the order of execution between the process of "adding, deleting, or modifying harmful text in the feature library" and step 20 is not limited, and they can also be executed simultaneously, as required.

[0070] In summary, as Figure 3As shown, the embodiments of this application may include four operation modules: image feature extraction, similarity feature retrieval, review result judgment, and harmful text database operation. Image feature extraction: Based on a deep neural network model, image content is extracted and converted into vector features for subsequent cross-modal similarity comparison. Similarity feature retrieval: The image vector feature X is compared with pre-prepared harmful text features to obtain the similarity between the two. If a similarity match is found, the next step of judgment is performed. Because the model maps data from two different modalities to the same vector space, the similarity between image features and text features can be directly calculated. Review result judgment: Based on the relationship between the harmful threshold and the similarity obtained in the previous step, it is determined whether the image to be reviewed is a harmful image. Generally, if the similarity is greater than a certain harmful threshold, the image can be directly judged as harmful; or, if the similarity is less than the harmless threshold, it is considered harmless. Other cases are judged as suspected harmful, requiring human review. Harmful text database operation: The harmful text feature database can add, modify, or delete content according to business needs, thereby achieving rapid response to the need for reviewing new illegal content. Of course, in addition to considering the relationship between similarity and the judgment threshold (harmful threshold, harmless threshold) in the above judgment process, information such as the degree of harm of harmful text can also be comprehensively considered.

[0071] Based on the above explanation, in some specific examples, each harmful threshold points to one (or more) violation categories. Therefore, correspondingly, after step 22, this image review method further includes: when faced with multiple similarities greater than the corresponding harmful threshold, the violation category pointed to by the largest similarity or the violation category jointly pointed to by all similarities is taken as the harmful type to which the image to be reviewed belongs. Furthermore, for example, when faced with 6 similarities greater than their respective harmful thresholds, where 3 similarities jointly point to violation category 1 (containing prohibited words) and 2 similarities jointly point to violation category 2 (poor behavioral expression), then the relatively majority of the jointly pointed-to violation category 1 can be taken as the harmful type to which the image to be reviewed belongs, that is, the image to be reviewed can be considered a harmful image containing prohibited words.

[0072] Please see Figure 8 This application provides a specific embodiment of a model training method, which includes the following operation steps 81 to 82:

[0073] 81. Adjust the image features and / or text features in the sample set to obtain semantically corresponding image features and text features.

[0074] For cross-modal techniques, step 81 can be viewed as performing cross-modal feature alignment on each image feature of the image modality and each text feature of the text modality in the sample set, so that semantically corresponding image features and text features can be aligned and mapped in the same space. Here, text features include harmful text features.

[0075] 82. Using semantically corresponding image features and text features, train initial feature extraction models for images and text respectively to obtain image feature extraction models and text feature extraction models applicable to the first aspect or any specific implementation of the first aspect (i.e., the image review method described above).

[0076] This helps establish cross-modal semantic alignment between models that apply different modal features, making it easier to find features of another modality from the features of one modality, and enabling direct similarity comparison between features of two modalities, thereby breaking down the heterogeneity gap between multimodal models.

[0077] In some specific examples, step 81 above may include: constructing a loss function based on the similarity between each image feature and each text feature; adjusting the image features and / or text features with the aim of minimizing the loss function to obtain semantically corresponding image features and text features. This allows the features of the image modality and the features of the text modality to be aligned and mapped in the same space due to their semantic correspondence, enabling them to be conveniently and specifically found, and training related image feature extraction models (i.e., Image Encoder) and text feature extraction models (i.e., Text Encoder). In practical applications, adjusting the similarity between two features can be understood as adjusting at least one of the features, causing a change in the spatial distance between the features due to the change. This includes making two features expressing the same semantic meaning closer together and two features expressing different semantic meanings more distant, to prevent mismatch between the input query and the output result, such as contradictory semantic expressions.

[0078] like Figure 4As shown, during model training, the initial input of the text model is a set of text [text1, text2, ..., textN], and the output is a set of extracted text features [T1, T2, ..., TN]. The input of the image model is a set of images [image1, image2, ..., imageN], and the output is a set of extracted image features [I1, I2, ..., IN]. Here, text{i} and image{i} are a pair of data corresponding in semantic content, i.e., a text-image pair. Finally, to ensure that the semantic features of image-text pairs within the same pair are similar, while the semantic features of image-text pairs in different pairs are far apart, thus achieving the goal of aligning image-text features corresponding to semantic content, the following specific steps can be taken: A loss function can be constructed based on the similarity between each image feature and each text feature, and this loss function can be used to adjust the similarity between image feature and text feature pairs (which can be regarded as the process of modifying features). For example, the feature similarity (I{i} multiplied by T{i}) between image-text pairs corresponding to semantic content can be made to approach a preset value of 1, so that I{i} and T{i} in this case are close in space. Conversely, the feature similarity between image-text pairs that do not correspond to semantic content is in the opposite direction to the preset value of 1 (such as approaching 0), so that I{i} and T{i} in this case are far apart in space, until the loss function reaches its minimum value. In this way, the positional relationship between the two modal features of image and text in the common feature space (the same space) can be finally constructed, so as to promote the alignment of image features and text features corresponding to semantic content. Based on the pre-training principles described above (which may utilize attention mechanisms), the finally trained image and text models can establish cross-modal semantic alignment relationships between the image and text modalities, enabling accurate and direct calculation of similarity between extracted image and text features. Optionally, the aforementioned loss function can specifically be the SDM loss function composed of the similarity of each image-text pair.

[0079] It should be noted that the operational details of the model training method provided in this application can be found in the specific implementation described in the first aspect above, and vice versa. Therefore, similar parts will not be repeated.

[0080] In summary, the method of this application embodiment is an image review technology with strong generalization ability and low development and maintenance costs. Its main advantages are:

[0081] 1. Reduced R&D Costs for Image Review Algorithms: Traditional methods require developing multiple sub-models for image review, each involving data collection and annotation, model training and optimization, service deployment and maintenance, etc., which is time-consuming and labor-intensive. This method, however, only requires developing one cross-modal model system to handle various types of inappropriate content, significantly reducing algorithm development costs and complexity. Specifically, using a cross-modal approach, based on the similarity comparison between text feature vectors and image feature vectors, it can recall inappropriate images, improving the recall rate for such images.

[0082] 2. Rapid Response and Handling of New Violating Images: When new violating content is discovered, traditional methods either require training a new model or iterating on an old one, a process that is time-consuming, labor-intensive, costly, and slow. This method, however, maintains a harmful text feature library and adds new content to it, enabling rapid response and handling of new violating images at a low cost. In other words, this application, by establishing a threshold search and operation mechanism for the harmful text feature library, supports the rapid addition and deletion of relevant text content, thereby quickly responding to and recalling new violating images.

[0083] 3. Cross-modal review technology has better generalization ability: It can greatly improve the iteration efficiency of image review capabilities with fewer or even zero samples, and quickly process and resolve various AIGC violations, improve the security of AIGC content, promote the application of AIGC technology, and recall more diverse violations.

[0084] Please see Figure 9 The electronic device 900 of this application embodiment may include one or more central processing units (CPUs) 901 and a memory 905, wherein the memory 905 stores one or more applications or data.

[0085] The memory 905 can be volatile or persistent storage. The program stored in the memory 905 can include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the central processing unit 901 can be configured to communicate with the memory 905 and execute the series of instruction operations stored in the memory 905 on the electronic device 900.

[0086] Electronic device 900 may also include one or more power supplies 902, one or more wired or wireless network interfaces 903, one or more input / output interfaces 904, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0087] The central processing unit 901 can perform the operations performed by the first aspect or any specific method embodiment of the first aspect, which will not be described in detail here.

[0088] This application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect or any specific implementation thereof.

[0089] This application provides a computer program product containing instructions or computer programs, which, when run on a computer, causes the computer to perform the method described in the first aspect or any specific implementation thereof.

[0090] It is understood that, in the various embodiments of this application, the sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system (if it exists) and device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0092] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system or apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0093] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0095] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An image verification method, characterized in that, include: The image to be examined is input into the image feature extraction model to output the image features of the image to be examined; The similarity between the image features and each harmful text feature in the feature library is calculated using cross-modal alignment; wherein, the harmful text features are associated with stock image features corresponding to semantic expressions, and each harmful text feature belongs to multiple harmful types in terms of semantic expression; For the similarity of each of the aforementioned harmful text features, determine the relationship between the harmful threshold of the harmful text feature and the similarity. If the similarity is greater than the harmful threshold, then the image to be reviewed is determined to match the semantic expression of the harmful text and is therefore a harmful image.

2. The image verification method according to claim 1, characterized in that, The method of calculating the similarity between the image features and each harmful text feature in the feature library using cross-modal alignment includes: When the feature content of the feature library is within a preset range, the similarity between the image feature and each of the harmful text features in the library is calculated by using instruction set acceleration or matrix operation acceleration library and cross-modal alignment. When the feature content of the feature library exceeds the preset level, the feature index in the library established by the hierarchical retrieval technology locates at least one harmful text feature that is semantically similar to the image feature, and cross-modal alignment is used to calculate the similarity between the image feature and the at least one harmful text feature.

3. The image verification method according to claim 1, characterized in that, The process of determining each harmful threshold includes: Construct a set of harmful images and a set of normal images, both exceeding the preset number of images; For each harmful text feature, the similarity between the harmful text feature and each harmful image feature and each normal image feature is calculated to obtain the similarity measurement interval of the harmful text feature. The similarity measurement interval includes the interval formed by the maximum similarity and the minimum similarity. The target determination threshold within the similarity measurement interval is searched and used as the harmful threshold of the harmful text feature. The target determination threshold is the value that enables all image features to achieve the highest evaluation index in terms of the measurement of the true violation category and the predicted violation category.

4. The image verification method according to claim 1, characterized in that, The method further includes: If the similarity between the image feature and the harmful text feature is between the harmless threshold corresponding to the harmful text feature and the harmful threshold, then an early warning reminder will be issued; Wherein, the harmless threshold is less than the harmful threshold, and the harmless threshold is determined based on the evaluation index generated from the features of the sample image.

5. The image verification method according to claim 1, characterized in that, The method further includes: If the similarity of multiple harmful text features is greater than their respective harmful thresholds, then the image under review is determined to contain the harmful text semantics corresponding to the highest similarity among them.

6. The image verification method according to claim 1, characterized in that, Before calculating the similarity between the image features and each harmful text feature in the feature library, the method further includes: Add, delete, or modify harmful text in the feature library to obtain a feature library for comparison with the image features.

7. The image verification method according to claim 1, characterized in that, The similarity is a weighted result of multiple distance metrics, including the cosine distance and / or Euclidean distance between feature vectors.

8. The image verification method according to claim 1, characterized in that, Each of the aforementioned harmful thresholds points to a violation category; after determining the relationship between the harmful threshold of the harmful text feature and the similarity, the method further includes: When faced with multiple similarities exceeding the corresponding harmful threshold, the violation category pointed to by the largest similarity or the violation category pointed to by all the similarities is taken as the harmful type to which the image to be reviewed belongs.

9. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 8.