Image review methods and related equipment

CN117351335BActive Publication Date: 2026-08-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-12
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]面对层出不穷的各种违规图像内容,针对各类违规图像数据,单独训练一个个审核模型的上述传统模式,需要训练的模型增多或迭代旧模型用时长,容易出现更新迭代慢,开发和维护模型成本高的挑战

Benefits of technology

[0018] Converting images into corresponding text descriptions helps to uncover the semantic meaning of visual content in text, reducing the false recognition rate caused by purely visual review. The application of a cross-modal alignment model overcomes the heterogeneity gap between different modalities, enabling the accurate identification of semantically corresponding text features from image features, thus ensuring that the text description accurately and reliably represents the content of the image. Furthermore, unlike traditional methods that require training separate review models for different image content, this application solves the review problem for various types of inappropriate images using a single cross-modal alignment model, reducing development costs and operational complexity, and efficiently ensuring the dissemination of positive semantic content through images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351335B_ABST
    Figure CN117351335B_ABST
Patent Text Reader

Abstract

This application discloses an image review method and related equipment. The image review method includes: inputting the image features of the image to be reviewed into a cross-modal alignment model to output the corresponding text features; generating a corresponding text description based on the text features; and identifying the text description as harmful text, then identifying the image to be reviewed as harmful. Converting images into corresponding text descriptions helps reduce the false recognition rate caused by purely visual review. The application of the cross-modal alignment model overcomes the heterogeneity gap between different modalities, enabling the accurate identification of semantically corresponding text features from image features, thereby ensuring that the text description accurately and reliably represents the content of the image. This application can solve the review problem of various types of illegal images through a single cross-modal alignment model, reducing development costs and maintenance difficulty, and efficiently ensuring the dissemination of positive semantic content from images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to image review methods and related equipment. Background Technology

[0002] As more and more UGC (user-generated content) is published on internet platforms, internet regulators are requiring the review of UGC content to control the display of illegal or unqualified content on the platform and maintain a clean internet environment.

[0003] Because UGC content, especially image content, is extremely rich, and the backgrounds and purposes of its creators vary greatly, the reasons for images being considered inappropriate are also diverse, such as inappropriate image behavior, false advertising, sensitive references, or the presence of prohibited words. Therefore, currently, for each type of inappropriate image, it is necessary to extensively collect and label large amounts of image data, and then train individual models using different types of image data, such as cartoon behavior models, real-person behavior models, advertising label models, and sensitive text models, in order to obtain a review model capable of screening out various types of inappropriate (harmful) images.

[0004] Faced with a constant stream of various illegal image contents, the traditional approach of training individual review models for each type of illegal image data is problematic. This leads to an increase in the number of models to train or lengthy iterations of older models, resulting in slow updates, high development and maintenance costs. Currently, relevant technologies do not offer effective solutions to these challenges. Summary of the Invention

[0005] This application provides an image review method and related equipment for efficiently and cost-effectively reviewing different types of harmful images.

[0006] The first aspect of this application provides an image review method, including:

[0007] The image to be examined is input into the image feature extraction model to output the image features of the image to be examined.

[0008] The image features are input into a cross-modal alignment model to output the text features to be reviewed corresponding to the image features; the semantic expressions of the text features to be reviewed and the image features correspond to each other;

[0009] A corresponding text description is generated based on the features of the text to be reviewed, and the text description is used to express the content of the image to be reviewed.

[0010] If the text description is identified as harmful text, then the image to be reviewed is identified as a harmful image.

[0011] A second aspect of this application provides an electronic device, including:

[0012] Central processing unit, memory, and input / output interfaces;

[0013] The memory is either a short-term storage memory or a persistent storage memory;

[0014] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first aspect of the embodiments of this application or any specific implementation thereof.

[0015] A third aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect of this application or any specific implementation thereof.

[0016] A fourth aspect of this application provides a computer program product comprising instructions or a computer program, which, when run on a computer, causes the computer to perform the method described in the first aspect of this application or any specific implementation thereof.

[0017] As can be seen from the above technical solutions, the embodiments of this application have at least the following advantages:

[0018] Converting images into corresponding text descriptions helps to uncover the semantic meaning of visual content in text, reducing the false recognition rate caused by purely visual review. The application of a cross-modal alignment model overcomes the heterogeneity gap between different modalities, enabling the accurate identification of semantically corresponding text features from image features, thus ensuring that the text description accurately and reliably represents the content of the image. Furthermore, unlike traditional methods that require training separate review models for different image content, this application solves the review problem for various types of inappropriate images using a single cross-modal alignment model, reducing development costs and operational complexity, and efficiently ensuring the dissemination of positive semantic content through images. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0020] It should be noted that although the steps in the flowcharts (if any) involved in the embodiments are drawn sequentially according to the arrows, unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0021] Figure 1 This is a schematic diagram of a system architecture for the image review method according to an embodiment of this application;

[0022] Figures 2 to 5 This is a flowchart illustrating the image review method according to an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the Query-Transformer module in an embodiment of this application;

[0024] Figure 7 This is a schematic flowchart of the model training method in an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] In the following description, expressions such as "one specific implementation" or "one specific example" describe a subset of all possible embodiments. However, it is understood that "one specific implementation" or "one specific example" can be the same or different subset of all possible embodiments and can be combined with each other without conflict. In the following description, the term "multiple" means at least two. When a certain value mentioned in this application reaches a threshold (if it exists), in some specific examples, it may include the former being greater than the latter. When "any" or "at least one" or similar expressions are mentioned, it specifically refers to any one of the listed examples or any combination of these examples.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0030] For ease of understanding and explanation, before providing a further detailed description of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0031] Image review system: A software system that uses algorithms and human intervention to classify and filter image content based on risk. Generally, after receiving an input image, the system uses artificial intelligence algorithms to label and categorize the content; then, risky content undergoes manual review, and if problems are confirmed, it is filtered and processed through methods such as deletion, removal, or making it visible only to the user.

[0032] Image embedding feature: Also known as image feature, it is a feature vector that reduces image information to one dimension through model algorithms. This vector compresses the image content information, making it easier to support subsequent tasks such as image classification and text description generation.

[0033] Text embedding features, also known as text features, are similar to image features. They are created by using model algorithms to reduce text information to a one-dimensional feature vector. This vector compresses the content information of the text, making it easier to support subsequent tasks such as text classification and similarity comparison.

[0034] UGC: User Generate Content.

[0035] AIGC: AI Generate Content, refers to content generated or created through artificial intelligence algorithms.

[0036] Prompt: A prompt, a string of text input to the generative model, used to guide the model to generate a text description that better matches the image illustration.

[0037] Harmful images: Images that are detrimental to the content ecosystem, such as images depicting inappropriate behavior, false advertising, sensitive references, or containing prohibited words. With the increasing abundance of UGC and AIGC content, the reasons for and forms of harmful image violations have become increasingly diverse, posing a significant challenge to content moderation.

[0038] To better implement the image review method of this application, the following is provided: Figure 1 The diagram shows the system architecture of this method. The system may include at least one terminal device 101 and a server 102. Different types of applications may be installed on the terminal device 101, such as instant messaging applications, live streaming applications, conferencing applications, etc. The terminal device 101 may be a smartphone, tablet, laptop, desktop computer, smart vehicle, etc. The server 102 may be used to store application data and image data generated by the different types of applications on the terminal device 101. The server 102 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, etc.

[0039] This image review method can be executed by either terminal device 101 or server 102, or by both, depending on the specific application scenario. No restrictions are placed here. Specifically, when the method is executed jointly by terminal device 101 and server 102, the application scenario could be as follows: terminal device 101 receives and uploads the image to be reviewed input by the user; server 102 uses an image extraction model to extract the image features to be reviewed, and uses a cross-modal alignment model to output the corresponding text features to be reviewed. Then, it generates a text description corresponding to the text features to be reviewed, and uses this text description to determine whether the image to be reviewed is a harmful image.

[0040] The above application environment is merely an example for ease of understanding. It is understood that the embodiments of this application are not limited to the above application environment.

[0041] The method described in this application will be further explained in detail below.

[0042] Please see Figure 2The first aspect of this application provides a specific embodiment of an image review method, which includes the following operational steps:

[0043] 20. Input the image to be reviewed into the image feature extraction model to output the image features of the image to be reviewed.

[0044] Image features of the image to be reviewed can be extracted from the image through an image feature extraction model (also known as an image encoder). That is, the image feature extraction model can be used to extract the suggestive content from the image to be reviewed. Generally, image features can be represented in vector form. This image encoder includes, but is not limited to, encoders based on pre-trained images such as ViT (VisionTransformer) and CNN, or image encoders trained in other specific domains or scenarios, to achieve feature representation of the input image.

[0045] 21. Input the image features into the cross-modal alignment model to output the text features to be reviewed corresponding to the image features.

[0046] The semantic expressions between the text features and image features to be reviewed correspond because the image features and text features used in the training of the cross-modal alignment model (which can be called the image-text cross-modal feature alignment module) are aligned and mapped in the same space according to semantic correspondence, and have mutual mapping relationships. For details, please refer to the explanation of the alignment model training method later, which will not be repeated here. Among them, the harmful text features in the text features belong to multiple types of harmfulness in terms of semantic expression. For example, the harmful type can be any one of the following: poor behavior expression, false advertising, sensitive targeting, containing prohibited words, etc.

[0047] It is understandable that the image-text cross-modal feature alignment module is used to construct cross-modal connections between the visual modality and the language modality (i.e., image-text modality), aligning image modality information with text modality information. Specifically, cross-modal alignment in this application can refer to aligning sub-elements of multimodal data in a sample set, such as aligning multiple objects in an image with phrases (or words) in a sentence, that is, establishing the correlation between the feature representations of the two different modalities, image modality and text modality, in order to perform image-text matching; in short, cross-modal alignment is to explore the correlation between sub-elements of multimodal data. On the other hand, image-text matching can be understood as cross-modal semantic association alignment retrieval, that is, using an instance of one modality, retrieving semantically related instances in another modality. This retrieval process measures the semantic similarity or relevance between images and text; for example, given an image (a dog bleeding on the grass), image-text matching aims to query the text corresponding to the semantics shown in the image (such as "dog on the grass" or "bleeding dog"), and vice versa.

[0048] The features in this application can be features in vector form, or features other than vector features, such as scalar features.

[0049] 22. Generate corresponding text descriptions based on the features of the text to be reviewed.

[0050] The text description is used to illustrate the content of the image to be reviewed. The advantage of generating a text description here is that it replaces the purely visual representation of the image with textual expression, effectively preventing harmful images from being mistakenly identified as harmless images and widely disseminated due to their concealed composition or differing visual perceptions.

[0051] 23. If the text description is identified as harmful, the image to be reviewed will be identified as a harmful image.

[0052] After obtaining the text description corresponding to the image to be reviewed, it can be determined whether the text description is harmful text. If it is, it means that the image to be reviewed contains semantic expressions of the same type as the harmful text, such as expressing the same sensitive content, and is therefore considered a harmful image, facing a filtering process such as deletion. The aforementioned harmful text may contain illegal words collected or mined in advance. This violation may refer to content that is dangerous in meaning, targets sensitive topics, or deliberately creates contradictions, which are detrimental to a healthy online environment.

[0053] In summary, this application's embodiments convert images into corresponding text descriptions, which helps to uncover the semantic meaning of visual content in text and reduce the false recognition rate caused by purely visual review. Specifically, through cross-modal feature alignment, the features of both image and text modalities—image features and text features—can be mapped to the same space, overcoming the heterogeneity gap between different modalities. This allows image features to accurately find their semantically corresponding text features, ensuring that the text description accurately and reliably replaces the content expression of the image. Because the harmful text features used in the training of this cross-modal alignment model belong to multiple harmful types in terms of semantic expression, i.e., the training samples are sufficient, this application does not need to train separate review models for different harmful types of image content as in the traditional approach. A single cross-modal alignment model can solve the review problem of various types of illegal images, reducing development costs and operational difficulties, and efficiently ensuring the dissemination of positive semantic content from images.

[0054] Based on the examples above, some specific possible implementation examples will be provided below. In practical applications, the implementation content of these examples can be combined or implemented separately as needed, depending on the corresponding functional principles and application logic.

[0055] Please see Figures 2 to 6 This application provides another specific embodiment of an image review method, which includes the following steps:

[0056] 20. Input the image to be reviewed into the image feature extraction model to output the image features of the image to be reviewed.

[0057] 21. Input the image features into the cross-modal alignment model to output the text features to be reviewed corresponding to the image features.

[0058] 22. Generate corresponding text descriptions based on the features of the text to be reviewed.

[0059] In some specific examples, step 22 may include: determining the suggestive features assigned to the text features to be reviewed, the suggestive features being used to record the semantic tendency type of the harmful content, the semantic tendency type including violent tendency and / or false propaganda tendency; and generating a text description corresponding to the text features to be reviewed based on the semantic tendency guidance of the suggestive features.

[0060] Based on the above methods, as a preferred solution, such as Figure 4 As shown, prompts can be added to the large language model decoder to guide the generation of tokens, thereby influencing the generation of image description text. Crucially, by designing non-safe, sensitive, and prohibited prompts, such as those related to blood or sudden wealth, the identification and recall rates of illegal or malicious image content can be rapidly improved. This allows for quick adaptation to business scenario needs and regulatory requirements without the need for dedicated data collection, annotation, and training, enabling timely detection and identification of risky content. For example, images often contain many points of interest, so the desired text containing sensitive words can be generated in a guided manner. Specifically, if the prompt feature assigned to the text to be reviewed is "blood," then based on the semantic guidance of "blood" (leaning towards the violent direction of "blood"), a text description containing the words "blood" or "violence" can be generated, effectively capturing important points of interest and improving the review progress and recall rate of harmful images.

[0061] In some specific examples, the process of "determining the corresponding suggestive features of the text to be reviewed" may include: calculating the feature similarity between the text features to be reviewed and each suggestive feature; determining a preset number of feature similarities with relatively high similarities among the feature similarities, and using the suggestive features corresponding to the determined feature similarities as the suggestive features of the text features to be reviewed. For example, the feature similarity here may include cosine similarity and / or Euclidean distance, etc. Using multiple suggestive features can guide the image to generate a more accurate text description, improving the representation effect of image-to-text semantics.

[0062] In some specific examples, the process of constructing the aforementioned suggestive features includes: collecting existing texts pointing to various semantic tendencies, and replacing some or all of the existing texts with suggestive texts of equivalent semantic meaning to obtain replacement texts; generating text features corresponding to the existing texts and replacement texts respectively, as suggestive features. Specifically, in practical applications, a semantic meaning can be expressed by different text sentences. For example, "a rude person" and "an easily angered person" can both describe people with unstable emotions. Therefore, in order to enrich the corpus and address the review of different content from multiple perspectives, a substitution method can be used to obtain replacement texts with semantic similarity to the existing texts, thereby expanding more possible suggestive features and preventing some illegal content from passing the review due to missing corpus data.

[0063] 23. If the text description is identified as harmful, the image to be reviewed will be identified as a harmful image.

[0064] Based on the above description, in some specific examples, the embodiments of this application may further include: determining whether the text description hits a preset number of preset keywords (i.e., keyword matching or filtering), and / or predicting whether the text description belongs to the harmful category of description through a trained classification model (i.e., text classification); preset keywords refer to words that are predefined as semantically harmful; if the text description hits a preset number of preset keywords and / or belongs to the harmful category of description, then the text description is determined to be harmful text. Here, "hitting" can refer to an inclusion relationship or a similar relationship, such as the similarity between the words "violence" and "irritability"; when both keyword matching and text classification are used, the weighted result of the two can be calculated to comprehensively determine whether the text description belongs to harmful text, and this weight can be determined by practical experience.

[0065] In some specific examples, the preset keywords come from a keyword library. The process of building the keyword library includes: collecting descriptive text corresponding to each historical harmful image, and mining keywords from each descriptive text that have a recall rate for the harmful image exceeding a threshold to form the keyword library. Here, using a large number of historical harmful images as a reference helps to improve the speed and quality of keyword accumulation, such as accumulating a batch of high-quality sensitive keywords, solving the cold start problem of the keyword library, and ensuring the accuracy of text generation.

[0066] In some specific examples, the process of the above text description hitting a preset keyword includes: the similarity between words in the text description and the preset keyword exceeds the judgment threshold corresponding to the preset keyword. For example, in practical applications, the judgment threshold corresponding to each preset keyword can be determined based on the recall rate and / or precision rate of the keyword for harmful sample images in the sample set; generally, the optimal threshold that can accurately recall the most harmful sample images is selected as the judgment threshold, thereby reducing false judgments caused by thresholds that are too high or too low. Correspondingly, if the similarity between a word in the text description and a preset keyword is greater than or equal to the judgment threshold corresponding to that keyword, it can be considered that the word hits the keyword; of course, multiple words can be selected to hit the preset keyword before considering the image to be reviewed as harmful, or the keyword with the largest difference from the judgment threshold can be selected as the hit word in the text description, without specific restrictions.

[0067] In some specific examples, given that the determination of whether the text description of the image under review belongs to harmful text can be made through keyword matching and / or text classification, the harmful type of the image under review can be determined accordingly by using preset keywords and / or the direction of the classification model. Specifically, each preset keyword points to one (or more) violation types. Accordingly, after step 23, this embodiment of the application may also include any of the following operations: Operation 1. The violation type pointed to by the text description with the highest similarity to each preset keyword is taken as the harmful type of the image under review; Operation 2. The violation type predicted by the classification model and matched by the text description is taken as the harmful type of the image under review; Operation 3. The violation type jointly pointed to by the preset keywords and the classification model is taken as the harmful type of the image under review. Here, the proposal of these three operation methods provides multiple evaluation channels for determining the image violation type in practical application scenarios, which helps to prevent the limitations of approval errors caused by a single orientation determination.

[0068] In practical applications, the operation process of the above embodiments can be as follows: Figure 3 As shown, the process includes the following:

[0069] Image feature extraction: Based on deep neural network technology, the image to be examined is converted into image features, such as using a CNN model to extract the vector features of the image.

[0070] Image text description generation: The vector features of the image are converted into descriptive text using a large language model. The higher the quality of the generated descriptive text, the better the recall rate for inappropriate images. Preferably, additional prompt information is input into the large language model to encourage the model to generate higher-quality text that better meets the needs of content moderation.

[0071] Keyword filtering and / or text classification review: Filter the description text generated in the previous step using a pre-prepared keyword library. If the description text matches a keyword, proceed to the next step. Alternatively, a general text multi-classification model can be used for further filtering; if the description text is identified as harmful, proceed to the next step as well.

[0072] Overall Review Results: Keyword filtering and text classification are two review methods that can be used individually or together for better results. If used together, the results of both methods will be combined. Generally, for higher recall, if either method determines the image content is harmful (i.e., violates regulations), it will be classified as a harmful image. For higher precision, it can be classified as a harmful image only if both methods determine it is harmful; otherwise, it will be considered a harmless image with normal content.

[0073] Keyword mining and operation: The keyword library can be updated by adding or deleting content based on business needs, enabling rapid response to the review of new inappropriate content. Ideally, to improve the speed and quality of keyword accumulation, a large number of inappropriate images can be collected and corresponding descriptive texts generated. Then, keywords that can efficiently recall inappropriate images can be statistically mined from these texts.

[0074] In summary, as Figure 4 As shown, after the image encoder extracts the image embedding features from the input image to be reviewed, it can be input into the image-text cross-modal feature alignment module to obtain the corresponding text features to be reviewed. The text features to be reviewed are decoded by the large language model decoder to generate a token sequence. The token ID is output by the text description generation module as the text description of the image to be reviewed, thereby realizing the cross-modal conversion from image to text. Finally, the text description is reviewed through keyword matching and / or text classification to achieve content review, thereby outputting the review result of whether the image to be reviewed is a harmful image.

[0075] The aforementioned image-text cross-modal feature alignment module is used to construct the connection between visual and language modalities, aligning image-text modal information, including two processes: guiding image-text representation learning and image-text generation learning. The image-text cross-modal embedding feature vector output by the image-text cross-modal feature alignment module not only contains visual information but also visual information related to the text. The large language model decoder includes, but is not limited to, language models based on BERT, GPT, etc., such as OPT, T5, etc., or language models trained in other specific domains or scenarios, thereby achieving text content prediction and generation. Keyword matching includes, but is not limited to, existing keyword matching algorithms based on AC automata, with the keyword lexicon supporting online updates. The text classification or detection model, including, but not limited to, existing text classification or detection algorithms based on BERT, CNN, etc., is used to detect and identify the violation risks of text content. Finally, the comprehensive review module outputs a judgment result, i.e., whether the image under review is a violation or harmful image.

[0076] Please see Figures 4 to 7 This application provides a specific embodiment of a model training method, which includes the following operation steps 71 to 72:

[0077] 71. Obtain a sample set that includes multiple feature pairs.

[0078] Each feature pair includes image features and text features, and the image features and text features correspond to each other in semantic expression.

[0079] 72. Adjust the model parameters of the initial alignment model based on the image and text features in the sample set until the initial alignment model reaches the convergence condition to obtain the cross-modal alignment model.

[0080] This cross-modal alignment model can be used in the image review method described in the first aspect above. The convergence condition includes that the (positive) similarity between image features and text features within the same feature pair tends towards a preset value due to model parameter tuning, and that the (negative) similarity between image features and text features within different feature pairs deviates from the preset value due to model parameter tuning. In practical applications, model parameter tuning can change the image features and / or text features within a feature pair, causing a change in their similarity, which can be understood as a change in spatial distance. Therefore, as one implementation possibility, the goal can be to increase the aforementioned positive similarity and decrease the negative similarity by first adjusting the image features and / or text features (i.e., changing the features), and then using the adjusted features as sample pairs for model parameter tuning, thereby achieving the model convergence condition. That is, model optimization can be carried out from the sample source.

[0081] For example, a text feature text{i} and an image feature image{i} can form a pair of data, i.e., an image-text pair. Each image{i} can have multiple paired text{i}, but may only correspond to a small portion of the semantic expressions of the text{i}. To ensure that semantically corresponding image{i} and text{i} are paired up (and thus preserved), while semantically non-corresponding image{i} and text{i} are kept separate and unpaired, the following approach is taken: A loss function is constructed based on the similarity between each image feature and each text feature. This loss function is then used to adjust the similarity between each pair of image and text features. For example, the feature similarity (dot product of I{i} and T{i}) between semantically corresponding image and text pairs tends to a preset value of 1, making I{i} and T{i} spatially close. Conversely, the feature similarity between semantically non-corresponding image and text pairs is reversed from the preset value of 1 (e.g., tending to 0), making I{i} and T{i} spatially far apart. This process continues until the loss function reaches its minimum value, thereby ultimately constructing the positional relationship between the two modalities of image and text features in a common feature space (the same space), thus aligning semantically corresponding image and text features.

[0082] In some specific examples, during the image-to-text (image-text) representation learning process, the output text embedding features may have a heterogeneous gap from the image modality to the text modality. For example, if the output embedding dimension is not the required preset dimension, its applicability to the text modality may not be strong enough. Therefore, before step 72, the training process of the cross-modal alignment model may also include the following operation: projecting the text features corresponding to the semantic expression of the image features into text features of a preset dimension (which can be regarded as a dimension transformation). This projection can be implemented by a linear projection layer; the execution order of this dimension transformation operation can be before or after the alignment of image features and text features, which can be determined by the user and is not restricted here.

[0083] like Figure 4 , Figure 5As shown, the cross-modal alignment model (or image-text cross-modal feature alignment module) consists of two parts: a Transformer module and a linear projection layer. The Transformer module aligns information from the image and text modalities, representing an image-to-text representation learning process. Specifically, the Transformer module primarily functions as a query-transformer. The linear projection layer, composed of multiple fully connected layers, projects the text embedding features output from the Transformer module into fixed-dimensional embedding features. This represents an image-to-text generation learning process, and the projected features incorporate both the visual features of the input image and descriptive text features. Finally, the fixed-dimensional embedding features serve as the final text embedding features, which can be input into a large language model decoder to generate the corresponding text description (e.g., a string of text). In practical applications, linear projection layers may not be implemented or configured in cross-modal alignment models. This is because there is a possibility that the text embedding features output by the Transformer module are directly fixed in dimension (e.g., 128-dimensional), without the need for dimensional projection transformation. However, practical experience shows that the occurrence of such a one-step transformation to fixed-dimensional embedding is very small. This is because there may still be a heterogeneous gap between image modality and text modality. Therefore, linear projection layers should generally be configured and applied.

[0084] The Query-Transformer described above mainly consists of two Transformer sub-modules, which share the same self-attention layer. For example... Figure 6 As shown, the left side is the image Transformer submodule, which receives image embedding features from the image encoder and uses them to extract visual features. The right side is the text-related Transformer submodule, which constructs the relationship between the query (such as the content of the input image to be reviewed) and the text through attention masking. The query embedding output by the QueryTransformer module is a (learned) image-to-text representation.

[0085] It should be noted that during model training, the image encoder and the large language model can directly use the existing trained network model without needing to train and fine-tune the model parameters. That is, the model parameters remain frozen, and only the image-text pair dataset needs to be used to train and fine-tune the image-text cross-modal feature alignment module. Alternatively, the image encoder and the large language model can be trained and fine-tuned separately.

[0086] In some specific examples, the convergence condition can be characterized by a target loss function. Specifically, the target loss function can be constructed based on the similarity between each image feature and each text feature in the sample set. Therefore, the optimization objective of the cross-modal alignment model can be summarized as minimizing the target loss function by adjusting the similarity between image features and text features within the same feature pair to approach a preset value (adjusting the trend means increasing similarity), and adjusting the similarity between image features and text features within different feature pairs to move in the opposite direction to the preset value (adjusting the trend means decreasing similarity).

[0087] For example, the process of "constructing a target loss function based on the similarity between each image feature and each text feature in the sample set" can be regarded as the process of training and fine-tuning the cross-modal alignment model, including: for each pair of image features and text features, calculating the image-text comparison loss and / or image-text matching loss between each pair of features; wherein, the image-text comparison loss refers to the vector similarity between the image features and the text features on the output vector, and the image-text matching loss represents the semantic matching degree between the text generated based on the image features and the input image; and / or, calculating the text generation loss corresponding to each image feature, the text generation loss including the attention result between the image features and the text features; and weightedly fusing the image-text comparison loss, the image-text matching loss and the text generation loss to form the target loss function.

[0088] It should be noted that, Figure 6 The Query-Transformer module shown also involves processes such as Image-Text Matching, Feed Forward, Cross Attention, Image-Text Contrastive Learning, Bidirectional Multimodal Causal, and Image-Grounded Text Generation. Therefore, QueryTransformer can be trained using any of the following loss functions:

[0089] 1) Image-Text Contrastive Loss: For each image query, the CLS terminology (i.e., feature vector) output is compared with the CLS terminology of the text output. Each pair of similarities is used to calculate a weighted final image-text contrastive loss. Alternatively, the highest similarity among the similarities calculated for each query can be selected as the comparison result between that query and the text, and used to calculate the final image-text contrastive loss. Under this image-text contrastive loss function, the query embedding and the text do not "see" each other.

[0090] 2) Image-Text Matching Loss: Under this loss, the query and text can see each other, and a probability (logit) is obtained to represent whether the input image features and the output text features match. This matching judgment can be regarded as a binary classification (match or not match) task. This matching task can be reviewed manually or predicted by the model, whichever is determined according to the actual situation, without any restrictions.

[0091] 3) Image-Grounded Text Generation Loss: This loss is used to examine the difference between the output text features corresponding to image features and the actual text features corresponding to those image features. In this loss, attention can be calculated between queries, but the attention of text terms to the queries is not calculated. Furthermore, self-attention within the text uses a causal mask and requires calculation of the attention of all queries to the text. The attention results calculated here are used to construct the text generation loss, where the attention results can be expressed as... Q is the query statement, K is the keyword, V is the value, and d is the key. k It is the length of the feature vector corresponding to the query statement.

[0092] It should be noted that the operation of the second aspect of this application can be implemented by referring to the specific content described in the first aspect, and vice versa; therefore, similar parts will not be repeated.

[0093] As can be seen from the above description, this method is an image review technology with strong generalization ability and low development and maintenance costs. Its main benefits are:

[0094] 1. Reduced cost of image review algorithm development: Traditional methods require the development of multiple sub-models for image review. The development of each model involves data collection and annotation, model training and optimization, service deployment and maintenance, etc., which is time-consuming and labor-intensive. In contrast, this method only requires the development of a single cross-modal model system, which can recall various types of illegal images, greatly saving algorithm development costs and difficulty.

[0095] 2. Freezing the parameters of the pre-trained image encoder and large language model facilitates training and solves the problem of requiring large amounts of labeled training data. Only image and text data are needed to train or fine-tune the image-text cross-modal feature alignment module, achieving general image-text cross-modal feature alignment. This solves the problem of existing models requiring specific training or fine-tuning for particular scenarios or domains. Of course, the image encoder and large language model can also be used for model training and fine-tuning to improve their practical application performance.

[0096] 3. It can leverage the capabilities of large visual and text models to achieve zero-shot image-text generation with strong generalization ability; in addition, it supports flexible combinations of different image encoders and large language models, allowing users to directly and fully enjoy the benefits of large model advancements.

[0097] 4. It supports the ability to generate text for images guided by prompts. At the same time, combined with operational mechanisms such as online updates of keyword dictionaries, it can quickly and timely recall illegal image content, enhance response and handling capabilities, and avoid the high cost and inefficient iterative investment required to retrain dedicated models.

[0098] Please see Figure 8 The electronic device 800 of this application embodiment may include one or more central processing units (CPUs) 801 and a memory 805, wherein the memory 805 stores one or more applications or data.

[0099] The memory 805 can be volatile or persistent storage. The program stored in the memory 805 can include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the central processing unit 801 can be configured to communicate with the memory 805 and execute the series of instruction operations stored in the memory 805 on the electronic device 800.

[0100] Electronic device 800 may also include one or more power supplies 802, one or more wired or wireless network interfaces 803, one or more input / output interfaces 804, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0101] The central processing unit 801 can perform the operations performed by any specific method embodiment of the first or second aspect described above, and the specifics will not be repeated here.

[0102] This application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect or any specific implementation thereof.

[0103] This application provides a computer program product containing instructions or computer programs, which, when run on a computer, causes the computer to perform the method described in the first aspect or any specific implementation thereof.

[0104] It is understood that, in the various embodiments of this application, the sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system (if it exists) and device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system or apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An image verification method, characterized in that, include: The image to be examined is input into the image feature extraction model to output the image features of the image to be examined. The image features are input into a cross-modal alignment model, which establishes a correlation between image modality and text modality in feature representation for image-text matching, so that the cross-modal alignment model outputs the text features to be reviewed corresponding to the image features based on the correlation; the semantic expressions of the text features to be reviewed and the image features correspond. Generate a corresponding text description based on the features of the text to be reviewed. The text description is used to express the content of the image to be reviewed. The process includes: determining the suggestive features corresponding to the features of the text to be reviewed, wherein the suggestive features are used to record the semantic tendency type of the harmful content, and the semantic tendency type includes violent tendency and / or false propaganda tendency; generating a text description corresponding to the features of the text to be reviewed based on the semantic tendency type of the suggestive features; wherein determining the suggestive features corresponding to the features of the text to be reviewed includes: calculating the feature similarity between the features of the text to be reviewed and each suggestive feature; determining a preset number of feature similarities with relatively high similarities among the feature similarities, and using the suggestive features corresponding to the determined feature similarities as the suggestive features of the features of the text to be reviewed. If the text description is identified as harmful text, then the image to be reviewed is identified as a harmful image.

2. The image verification method according to claim 1, characterized in that, The process of constructing the suggestive features includes: Collect existing texts pointing to each of the aforementioned semantic tendency types, and replace some or all of the existing texts with illustrative texts of equivalent semantic meaning to obtain replacement texts; Generate text features corresponding to the existing text and the replacement text respectively, as the prompting features.

3. The image verification method according to claim 1, characterized in that, The method further includes: Determine whether the text description hits a preset number of preset keywords, and / or predict whether the text description belongs to the harmful category using a trained classification model; the preset keywords refer to words that are predefined as semantically harmful. If the text description matches a preset number of preset keywords and / or belongs to the harmful category of descriptions, then the text description is identified as harmful text.

4. The image review method according to claim 3, characterized in that, The preset keywords come from a keyword library, and the construction process of the keyword library includes: Collect descriptive text corresponding to each historical harmful image, and extract keywords with a recall rate exceeding a threshold from each descriptive text to form the keyword library.

5. The image verification method according to claim 3, characterized in that, The step of determining whether the text description matches a preset number of preset keywords includes: If the similarity between the words in the text description and the preset keywords exceeds the judgment threshold corresponding to the preset keywords, then it is determined that the text description matches the preset keywords.

6. The image verification method according to claim 3, characterized in that, Each of the preset keywords points to a violation type; After identifying the image to be reviewed as a harmful image, the method further includes any of the following operations: The violation type pointed to by the text description with the highest similarity to each of the preset keywords, or the violation type predicted by the classification model to be matched by the text description, is taken as the harmful type to which the image to be reviewed belongs; The violation type that is jointly pointed to by the preset keywords and the classification model is taken as the harmful type to which the image to be reviewed belongs.

7. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image auditing method and device, equipment and storage medium

    CN115443490A

  • Multi-modal pedestrian re-identification method based on multi-level cross-modal difference harmonic

    CN116682144A