Image auditing method and device, storage medium and program product

By obtaining input images, historical image libraries, and review rule libraries, and using multimodal large models for few-sample learning, the problem in existing technologies that models need to be fine-tuned to be applicable to content review scenarios is solved, achieving more efficient and high-quality image review.

CN120670608APending Publication Date: 2025-09-19GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510769928.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies do not fully utilize the few-sample learning value of historical cases in image review, resulting in the need for fine-tuning of the model to be applicable to content review scenarios, increasing training costs and time.

Method used

By obtaining input images, historical image libraries and audit rule libraries, a multimodal large model is used to generate image audit results based on target historical images and target audit rules, thus achieving few-sample learning.

Benefits of technology

By making full use of the value of few-sample learning, the applicability of large multimodal models in content review scenarios is improved, the training cost and time of image review are saved, and the review efficiency and quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670608A_ABST
    Figure CN120670608A_ABST
Patent Text Reader

Abstract

The invention discloses an image auditing method and device, a storage medium and a program product, and relates to the technical field of computers. The method comprises the following steps: acquiring an input image, a historical image library and an auditing rule library; obtaining at least two target historical images in a historical image library according to the input image and the historical image library; obtaining at least one target auditing rule in an auditing rule base according to the input image and the auditing rule base; and through the multi-modal large model, according to the input image, the at least two target historical images and the at least one target auditing rule, generating an auditing result of the input image, the auditing result of the input image being used for indicating that the input image is one of a violation image, a compliance image and an uncertain image. According to the method, the learning value of few samples is fully utilized, the open-source multi-modal large model can directly perform a security auditing task without fine tuning, the applicability of the multi-modal large model in a content auditing scene is improved, and the training cost of image auditing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image review method, device, storage medium, and program product. Background Art

[0002] Image review is to determine whether the image meets the established review rules based on the image content and obtain the image review result.

[0003] In related technologies, the target image's image semantics and text semantics are extracted and fused to obtain the target image's semantic features. Based on the target image's semantic features, the audit rules associated with the semantic features are retrieved from multiple pre-set audit rules. These retrieved audit rules and the target image are then input into a multimodal image-text model to output whether the target image violates a rule, the type of violation, and the reason for the violation. This type of method requires the use of retrieval enhancement technology to introduce sufficient similar reference cases and their penalty results into the model. This allows the pre-trained model to be directly used for content security audits through online few-sample learning, eliminating the need for additional training and development work.

[0004] However, the above method does not fully utilize the few-sample learning value of historical cases, resulting in the need for certain fine-tuning of the model to ensure its applicability in content review scenarios. Summary of the Invention

[0005] The present invention provides an image review method, device, storage medium, and program product. The technical solutions provided by the present invention are as follows:

[0006] According to one aspect of an embodiment of the present application, there is provided an image review method, the method comprising:

[0007] Obtain an input image, a historical image library, and an audit rule library; the historical image library includes image information corresponding to at least two historical images, the image information of the historical images includes a human review label for the historical images, the human review label is an audit label obtained through manual review, and the human review label is a compliance label or a violation label; the audit rule library includes at least one audit rule, the audit rule being used to describe a judgment condition for an image violation type;

[0008] Obtaining, based on the input image and the historical image library, at least two target historical images in the historical image library, wherein similarities between the target historical images and the input image satisfy a first condition;

[0009] According to the input image and the audit rule library, at least one target audit rule in the audit rule library is obtained, wherein the similarity between the target audit rule and the input image satisfies a second condition;

[0010] A multimodal large model is used to generate an audit result of the input image based on the input image, the at least two target historical images and the at least one target audit rule. The audit result of the input image is used to indicate that the input image is one of an illegal image, a compliant image, and an uncertain image.

[0011] According to one aspect of an embodiment of the present application, an image review device is provided, the device comprising:

[0012] A data acquisition module is configured to acquire an input image, a historical image library, and an audit rule library; the historical image library includes image information corresponding to at least two historical images, the image information of each historical image includes a human review label for each historical image, the human review label being a compliance label or a violation label obtained through manual review; the audit rule library includes at least one audit rule, the audit rule being used to describe a judgment condition for an image violation type;

[0013] an image determination module, configured to obtain, based on the input image and the historical image library, at least two target historical images in the historical image library, wherein similarities between the target historical images and the input image satisfy a first condition;

[0014] a rule determination module, configured to obtain, based on the input image and the audit rule library, at least one target audit rule in the audit rule library, wherein a similarity between the target audit rule and the input image satisfies a second condition;

[0015] A result generation module is used to generate an audit result of the input image based on the input image, the at least two target historical images and the at least one target audit rule through a multimodal large model, wherein the audit result of the input image is used to indicate that the input image is one of an illegal image, a compliant image, and an uncertain image.

[0016] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned image review method.

[0017] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned image review method.

[0018] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned image review method.

[0019] The technical solutions provided in the embodiments of the present application can bring the following beneficial effects:

[0020] By obtaining at least two target historical images in the historical image library and at least one target audit rule in the audit rule library, the audit result of the input image can be generated using a few-sample learning method. Compared with the method of only using the retrieved audit rules to conduct image audits in related technologies, the technical solution of this application fully utilizes the learning value of few samples, takes into account the timeliness of historical images and the alignment of audit rules, and allows the open source multimodal large model to directly perform security audit tasks without fine-tuning, thereby improving the applicability of the multimodal large model in content audit scenarios, saving the steps required for image auditing, reducing the training cost of image auditing, and improving the efficiency of image auditing. In addition, by comprehensively considering historical images and audit rules to analyze the illegal content of the input image, a more complete and sufficient retrieval and sorting method is provided, which enhances the confidence of the audit results of the input image and improves the quality of image auditing. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of a computer system provided by one embodiment of the present application;

[0022] Figure 2 This is a flowchart of an image review method provided by one embodiment of the present application;

[0023] Figure 3 This is a schematic diagram of a prompt text configuration template provided by an embodiment of the present application;

[0024] Figure 4 This is a schematic diagram of the process of generating audit information using a multimodal large model provided by an embodiment of the present application;

[0025] Figure 5 This is a schematic diagram of the overall process of image review provided by one embodiment of the present application;

[0026] Figure 6 This is a schematic diagram of a process for generating a target history black image and a target history white image provided by an embodiment of the present application;

[0027] Figure 7 This is a schematic diagram of a process for constructing an audit rule index library provided by an embodiment of the present application;

[0028] Figure 8This is a schematic diagram of a process for constructing a historical image library provided by one embodiment of the present application;

[0029] Figure 9 is a block diagram of an image review device provided by one embodiment of the present application;

[0030] Figure 10 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0032] Please refer to Figure 1 , which shows a schematic diagram of a computer system provided by an embodiment of the present application. The computer system may include: a terminal device 10 and a server 20.

[0033] There can be one or more terminal devices 10. The terminal device 10 can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, game console, e-book reader, multimedia player, wearable device, intelligent voice interaction device, smart home appliance, vehicle terminal, aircraft, etc.

[0034] The terminal device 10 may be installed with a client for a target application that has an image review function. A user may input an image into the target application, which then outputs a review result for the image, indicating whether the input image violates a violation. This application does not limit the type of target application. Optionally, the target application may be an application that requires downloading and installation, or a click-to-use application, which is not limited in this application.

[0035] The server 20 is used to provide background services for the client of the target application installed and running on the terminal device 10. For example, the server 20 can be the background server of the target application. The server 20 can be a standalone physical server, a server cluster consisting of multiple servers, or a cloud computing service center. Optionally, the server 20 provides background services for the target applications on multiple terminal devices 10 simultaneously. The terminal device 10 and the server 20 can communicate with each other via a network.

[0036] In an embodiment of the present application, an input image, a historical image library, and an audit rule library are obtained. The historical image library includes image information corresponding to at least two historical images. The image information of the historical images includes human-reviewed labels for the historical images. The human-reviewed labels are audit labels obtained through manual review and are either compliance labels or violation labels. The audit rule library includes at least one audit rule, which is used to describe the judgment criteria for the type of image violation. The present application is used to audit whether the input image violates a rule based on the historical image library and the audit rule library, and obtain an audit result for the input image. Based on the input image and the historical image library, at least two target historical images are obtained from the historical image library, where the similarity between the target historical images and the input image satisfies a first condition. Furthermore, based on the input image and the audit rule library, at least one target audit rule is obtained from the audit rule library, where the similarity between the target audit rule and the input image satisfies a second condition. Then, a multimodal large model is used to generate an audit result for the input image based on the input image, the at least two target historical images, and the at least one target audit rule. The audit result for the input image indicates whether the input image is a violation image, a compliance image, or an uncertain image.

[0037] Please refer to Figure 2 , which shows a flow chart of an image review method provided by an embodiment of the present application. The execution subject of each step of the method may be a computer device. The method may include at least one of the following steps 210 to 240:

[0038] Step 210: Obtain input images, a historical image library, and an audit rule library.

[0039] In some embodiments, the historical image library includes image information corresponding to at least two historical images. The image information of the historical images includes image features of the historical images, human review labels of the historical images, machine review labels of the historical images, review times of the historical images, and label analysis text of the historical images. The human review labels are review labels obtained through manual review, and the human review labels are either compliance labels or violation labels. The machine review labels are review labels obtained through image security algorithms, and the machine review labels are either violation labels or compliance labels. The label analysis text is used to indicate the analysis process of the human review labels of the historical images obtained through review.

[0040] Historical images refer to images of a historical time period that have been manually reviewed and algorithmically reviewed. The end time of the historical time period is any time before the current time. This application does not limit the length of the historical time period, and this application does not limit the starting time of the historical time period.

[0041] Historical images are manually reviewed to obtain human-reviewed labels. Human-reviewed labels can be either compliance labels or violation labels. Historical images are processed through image security algorithms to obtain machine-reviewed labels. Machine-reviewed labels can be either compliance labels or violation labels. If the human-reviewed or machine-reviewed label of a historical image is a violation label, the violation label can also be used to indicate the violation type of the historical image. The compliance label is used to indicate that the historical image is not in violation, i.e., the historical image does not contain any violation image content; the violation label is used to indicate that the historical image is in violation, i.e., the historical image contains any violation image content. The violation image content includes at least one violation content, and the violation type of the historical image is determined by the violation image content contained in the historical image.

[0042] Image features of historical images are extracted using a feature extraction model. These models can include mainstream cross-modal feature extraction models such as CLIP (Contrastive Language-Image Pre-Training), DINO (Self-Distillation with No Labels), and JINA-CLIP. Vector search tools (such as FAISS) or professional vector libraries (such as Milvus) are then used to index and store the image features of historical images in a historical image library.

[0043] The review time of a historical image refers to the time when the historical image was manually reviewed and received the historical image's human review label. Therefore, the review time of a historical image can also be referred to as the historical image's human review time. Alternatively, if the review time of a historical image is measured in days, the review time of the historical image refers to the date the historical image underwent manual review. If the review time of a historical image is measured in hours, minutes, or other time units, the review time of the historical image refers to the time when the manual review of the historical image was completed.

[0044] The label analysis text for historical images is output by a multimodal large model based on the historical image, its human-reviewed labels, its machine-reviewed labels, and at least one review rule. The label analysis text indicates the analysis process for obtaining the human-reviewed labels for the historical image. If the human-reviewed label for the historical image is compliant, the label analysis text indicates the reasoning behind the image's compliance. If the human-reviewed label for the historical image is non-compliant, the label analysis text indicates the reasoning behind the image's violation and determines the type of violation.

[0045] In some embodiments, the at least two historical images include at least one historical black image and at least one historical white image, the human-reviewed label of the historical black image is a violation label, and the human-reviewed label of the historical white image is a compliance label.

[0046] Historical black images are historical images that have been labeled as illegal by human reviewers, while historical white images are historical images that have been labeled as compliant by human reviewers. Based on the human review labels of historical images, the historical image library can be divided into a black library and a white library. The black library stores the image information corresponding to at least one historical black image, while the white library stores the image information corresponding to at least one historical white image.

[0047] In some embodiments, the audit rule library includes at least one audit rule, and the audit rule is used to describe the judgment conditions of the image violation type.

[0048] Audit rules describe the criteria for determining violations, and each rule corresponds to a specific violation. Manual operations can modify, delete, or add at least one existing audit rule to update the latest version of the audit rule library.

[0049] The image audit involved in this application is to audit the input image based on at least one audit rule in the audit rule library to determine whether there is any illegal content described in the audit rule in the input image.

[0050] In some embodiments, the input image is configured with a machine-verified label. This indicates that the input image has been processed by an image security algorithm. The machine-verified label for the input image can be either a compliance label or a violation label.

[0051] Optionally, the input image can be configured with a human review label. The review results of the input image obtained in this application will be used to compare with the human review label of the input image. Input images with discrepancies between the review results and the human review label will undergo a second manual quality inspection to identify possible misjudgments in the initial manual review, thereby improving the efficiency of image quality inspection and the quality of image review.

[0052] Optionally, the input image may not be configured with a human review label for the input image. The human review label for the input image will be determined based on the audit result of the input image obtained in this application. If the audit result is used to indicate that the input image is a compliant image, or the audit result is used to indicate that the input image is an illegal image, the audit result of the input image will be directly used as the human review label for the input image. If the audit result is used to indicate that the input image is an uncertain image, the input image needs to be manually reviewed, and the human review label for the input image is determined by manual review. This can reduce the cost of manual review and improve the efficiency of image review.

[0053] Step 220 : Obtain at least two target historical images in the historical image library according to the input image and the historical image library, wherein the similarity between the target historical images and the input image satisfies a first condition.

[0054] Based on the first condition, at least two target historical images are selected from at least two historical images in the historical image library. The at least two target historical images are historical images with a high degree of correlation with the input image. The at least two target historical images include at least one target historical black image and at least one target historical white image. The at least one target historical black image is a historical black image with a high degree of correlation with the input image, and the at least one target historical white image is a historical white image with a high degree of correlation with the input image. In other words, based on the first condition, at least one target historical black image is selected from the black image library, and based on the first condition, at least one target historical white image is selected from the white image library.

[0055] The specific process of determining at least two target historical images may refer to the following embodiment and will not be described in detail here.

[0056] Step 230: Obtain at least one target audit rule in the audit rule library based on the input image and the audit rule library, wherein the similarity between the target audit rule and the input image satisfies the second condition.

[0057] Based on the second condition, at least one target audit rule is selected from at least one audit rule in the audit rule library. The at least one target audit rule is an audit rule with a high degree of relevance to the input image. The specific process for determining the at least one target audit rule can be referenced in the following embodiment and is not described in detail here.

[0058] In step 240, a multimodal large model is used to generate an audit result of the input image based on the input image, at least two target historical images, and at least one target audit rule. The audit result of the input image is used to indicate whether the input image is one of an illegal image, a compliant image, and an uncertain image.

[0059] In some embodiments, step 240 includes at least one of sub-steps 241 - 242 .

[0060] Sub-step 241, generates prompt text by a prompt text generator based on the input image, at least two target historical images and at least one target review rule. The prompt text is used to instruct the multimodal large model to generate the review result of the input image according to the preset template.

[0061] The prompt text generator is used to generate prompt text according to a prompt text configuration template (preset template), based on an input image, at least two target historical images, and at least one target review rule. The image information corresponding to the input image, at least two target historical images, and at least one target review rule are input to the prompt text generator, which then outputs the prompt text. The input image here is an image configured with an organic review tag.

[0062] Figure 3 A schematic diagram of a prompt text configuration template is shown. The prompt text configuration template includes generation instruction text, image information of at least one target historical black image, image information of at least one target historical white image, an input image, at least one target review rule, and output instruction text. The generation instruction text instructs the multimodal macro model to generate the review result of the input image, and the output instruction text indicates the format of the output data of the multimodal macro model.

[0063] Optionally, an optimal prompt text may be generated through a prompt text optimization iterative framework.

[0064] Sub-step 242: Generate review information of the input image based on the prompt text using the multimodal large model.

[0065] The audit information of the input image includes at least one of the audit result of the input image, the violation type of the input image, and the audit analysis text of the input image.

[0066] The multimodal large language model here can be any publicly available large language model, such as a natural language model based on a transformer structure obtained by training with a large amount of data. The large amount of data can reach a sample level of more than 100 million, and this application does not limit this.

[0067] Optionally, if the generated prompt text is used to instruct the multimodal large model to generate an audit result of the input image, the output prompt text is used to indicate the format of the audit result of the input image. The prompt text is input to the multimodal large model, and the multimodal large model outputs the audit result of the input image.

[0068] Optionally, if the above-mentioned generated prompt text is used to indicate that the multimodal large model generates the audit information of the input image, the output prompt text is used to indicate the format of the audit information of the input image, such as Figure 3 As shown in the prompt text configuration template, the output prompt text format is {"Reasoning Analysis": "xxx", "Target Image Content Security Violation": "Yes / No / Unsure"}. The prompt text is input to the multimodal large model, which then outputs the review information for the input image.

[0069] The review result of the input image indicates whether the input image is a violation image, a compliance image, or an uncertain image. If the input image is a violation image, it means that the input image contains violation image content; if the input image is a compliance image, it means that the input image does not contain violation image content; if the input image is an uncertain image, it means that the input image may or may not contain violation image content.

[0070] The audit analysis text of the input image is an inference analysis text for generating audit information of the input image, and is used to perform inference analysis on whether the image content of the input image contains at least one illegal content described by an audit rule.

[0071] When the audit result of an input image indicates that the input image is a violation image, the audit information of the input image includes the audit result of the input image, the violation type of the input image, and audit analysis text of the input image. The audit analysis text of the input image is used to describe the reasoning and analysis process of the violation content described by at least one audit rule contained in the input image. The violation type of the input image can be determined based on the audit rule violated by the input image, that is, based on the violation image content contained in the input image.

[0072] In the case where the audit result of the input image is used to indicate that the input image is a compliant image, the audit information of the input image includes the audit result of the input image and the audit analysis text of the input image, where the audit analysis text of the input image is used to describe the reasoning and analysis process that the input image does not contain at least one illegal content described by the audit rule.

[0073] In the case where the audit result of the input image is used to indicate that the input image is an uncertain image, the audit information of the input image includes the audit result of the input image, the possible violation type of the input image and the audit analysis text of the input image, where the audit analysis text of the input image is used to describe the reasoning and analysis process that the input image may contain violation content described by at least one audit rule.

[0074] By generating prompt text first, the audit information of the input image generated by the multimodal large model can be output in the format of a preset template, so that the audit result of the input image can be quickly located from the audit information, or the violation type of the input image can be located, or the audit analysis text of the input image can be located, so that users can clearly understand the audit rules violated by the input image, and make the audit result of the input image explainable, thereby enhancing the confidence of the audit result of the input image, accelerating the image audit speed, and helping to improve the efficiency of image audit.

[0075] Figure 4The schematic diagram of the process of generating audit information by the multimodal large model is shown. First, the image information corresponding to the input image, TOPK2 historical black image and TOPK2 historical white image, and TOPN2 audit rules are input into the prompt text generator. The prompt text generator generates prompt text according to the prompt text configuration template. Here, TOPK2 historical black image is a historical black image that meets the first condition in at least one historical black image, TOPK2 historical white image is a historical white image that meets the first condition in at least one historical white image, and TOPN2 audit rule is an audit rule that meets the second condition in at least one audit rule. The prompt text is input into the multimodal large model, and the multimodal large model outputs the audit information of the input image.

[0076] The technical solution provided by the embodiment of the present application, by obtaining at least two target historical images in the historical image library and at least one target audit rule in the audit rule library, enables the audit result of the input image to be generated using a few-sample learning method. Compared with the method of only using the retrieved audit rules to perform image audit in the related art, the technical solution of the present application fully utilizes the learning value of few samples, considers the timeliness of historical images and the alignment of audit rules, and allows the open source multimodal large model to directly perform security audit tasks without fine-tuning, thereby improving the applicability of the multimodal large model in content audit scenarios, saving the steps required for image audit, reducing the training cost of image audit, and improving the efficiency of image audit. Moreover, by comprehensively considering historical images and audit rules to analyze the illegal content of the image content of the input image, a more complete and sufficient retrieval and sorting method is provided, which enhances the confidence of the audit results of the input image and improves the quality of image audit.

[0077] Figure 5 A schematic diagram of the overall process of image review is shown. The historical image library includes a black library and a white library. The black library includes image information of at least one historical black image, and the white library includes image information of at least one historical white image. The image information of the historical images includes image features of the historical images, human review labels of the historical images, machine review labels of the historical images, review times of the historical images, and label analysis texts of the historical images. Based on an input image configured with an machine review label, at least one target historical image is screened out from the historical image library. And based on an input image configured with an machine review label, at least one target review rule is screened out from the review rule library. Based on the input image configured with an machine review label, at least one target historical image, and at least one target review rule, a prompt text is generated, and the prompt text is input into a multimodal large model. The multimodal large model outputs the review information of the input image, including at least one of the review result of the input image, the violation type of the input image, and the review analysis text of the input image.

[0078] In some embodiments, step 220 includes at least one of sub-steps 221 - 224 .

[0079] Sub-step 221 , performing feature extraction on the input image to obtain image features of the input image.

[0080] The input image is input to the feature extraction model, and the feature extraction model outputs the image features of the input image.

[0081] Sub-step 222 , obtaining at least one target historical black image in the at least one historical black image according to the image features of the input image and the image features corresponding to the at least one historical black image.

[0082] At least one target historical black image is selected from the at least one historical black image according to a first condition, where the first condition refers to a similarity condition between image features of the input image and image features of the historical black image. The target historical black image is a historical black image that satisfies the first condition among the at least one historical black image.

[0083] The specific process of determining at least one target historical black image can be referred to the following embodiment and will not be described in detail here.

[0084] Sub-step 223 , obtaining at least one target historical white image in the at least one historical white image according to the image features of the input image and the image features corresponding to the at least one historical white image.

[0085] At least one target historical white image is selected from the at least one historical white image based on a first condition, where the first condition refers to a similarity condition between image features of the input image and image features of the historical white image. The target historical white image is a historical white image from the at least one historical white image that satisfies the first condition.

[0086] The specific process of determining at least one target historical white image may refer to the following embodiment and will not be described in detail here.

[0087] Sub-step 224 , obtaining at least two target history images according to at least one target history black image and at least one target history white image.

[0088] At least one target history black image and at least one target history white image are collectively referred to as at least two target history images.

[0089] By extracting at least one target historical black image from at least one historical black image, and extracting at least one target historical white image from at least one historical white image, the audit result of the input image will be generated based on the target historical black image and the target historical white image, providing a diverse historical sample for the multimodal large model, which helps to generate more accurate image audit results.

[0090] In some embodiments, for the process of determining at least one target historical black image, sub-step 222 includes at least one step from sub-steps A1 to A3.

[0091] Sub-step A1: calculating the similarity between the image features of the input image and the image features corresponding to at least one historical black image, and obtaining the image similarity corresponding to at least one historical black image.

[0092] Optionally, a calculation method such as Euclidean distance or cosine similarity can be used to calculate the similarity between the image features of the input image and the image features corresponding to at least one historical black image, thereby obtaining the image similarity corresponding to at least one historical black image. The image similarity corresponding to each historical black image refers to the similarity between the image features of the input image and the image features of the historical black image, and can also be understood as the visual similarity between the input image and the historical black image.

[0093] Exemplarily, the image similarity corresponding to the historical black image can be expressed as S_visual.

[0094] Sub-step A2: obtaining the image similarities corresponding to at least one historical black image, and sorting the image similarities in descending order to obtain the historical black images corresponding to the first number of image similarities.

[0095] The image similarities corresponding to the at least one historical black image are sorted in descending order of image similarity to obtain at least one sorted image similarity. The first number of historical black images can be obtained by obtaining the historical black images corresponding to the first number of image similarities ranked at the top of the at least one sorted image similarity.

[0096] The specific value of the first quantity is set by the technical personnel according to the image review requirements and is not limited in this application.

[0097] Sub-step A3: obtaining at least one target historical black image according to the image information corresponding to the first number of historical black images.

[0098] Optionally, the first number of historical black images may be directly determined as at least one target historical black image.

[0099] Optionally, at least one target historical black image may be further determined from the first number of historical black images according to the image information respectively corresponding to the first number of historical black images.

[0100] By extracting a first number of historical black images based on the similarity between image features, and then determining at least one target historical black image based on the first number of historical black images, the obtained target historical black image is a historical black image whose image similarity with the input image meets the first condition. In this way, the multimodal large model can specifically analyze the image content of the input image based on the target historical black image, thereby enhancing the targeted nature of image review and improving the accuracy of image review results.

[0101] In some embodiments, sub-step A3 includes at least one of sub-steps A31 to A33.

[0102] Sub-step A31 : performing optical character recognition on the input image and the first number of historical black images respectively to obtain text information of the input image and text information corresponding to the first number of historical black images respectively.

[0103] Optical character recognition (OCR) is performed on the input image using a character recognition tool to obtain text information of the input image. The text information of the input image refers to the text information originally existing in the input image. Optical character recognition is performed on the first number of historical black images using the character recognition tool to obtain text information corresponding to the first number of historical black images. The text information of the historical black images refers to the text information originally existing in the historical black images.

[0104] In sub-step A32, comprehensive scores corresponding to the first number of historical black images are obtained based on the image information corresponding to the first number of historical black images, the image similarities corresponding to the first number of historical black images, the machine-reviewed labels of the input images, the text information of the input images, and the text information corresponding to the first number of historical black images. The comprehensive scores of the historical black images are used to measure the similarity between the historical black images and the input images.

[0105] In some embodiments, sub-step A32 includes at least one of sub-steps A321 to A325.

[0106] Sub-step A321, performing feature extraction on the text information of the input image to obtain the text features of the input image, and performing feature extraction on the text information corresponding to the first number of historical black images to obtain the text features corresponding to the first number of historical black images.

[0107] Inputting text information of an input image into a feature extraction model, and having the feature extraction model output text features of the input image. Furthermore, inputting text information corresponding to a first number of historical black images into the feature extraction model, and having the feature extraction model output text features corresponding to the first number of historical black images.

[0108] Sub-step A322 , calculating the similarity between the text features of the input image and the text features corresponding to the first number of historical black images, to obtain the text similarities corresponding to the first number of historical black images.

[0109] Optionally, a calculation method such as Euclidean distance or cosine similarity can be used to calculate the similarity between the text features of the input image and the text features corresponding to the first number of historical black images, thereby obtaining the text similarities corresponding to the first number of historical black images. The text similarity corresponding to each historical black image refers to the similarity between the text features of the input image and the text features of the historical black images.

[0110] Optionally, if there is no text information in the input image, the text similarities corresponding to the first number of historical black images are all 1; if there is text information in the input image and no text information in the historical black image, the text similarity corresponding to the historical black image is 1; if there is text information in both the input image and the historical black image, the similarity between the text features of the input image and the text features corresponding to the historical black image is calculated to obtain the text similarity corresponding to the historical black image.

[0111] For example, the text similarity corresponding to the historical black image can be expressed as S_ocr_text.

[0112] Sub-step A323 , calculating the similarity between the machine-reviewed label of the input image and the machine-reviewed labels corresponding to the first number of historical black images, to obtain the label similarities corresponding to the first number of historical black images.

[0113] The label similarity corresponding to each historical black image refers to the similarity between the machine-reviewed label of the input image and the machine-reviewed label of the historical black image.

[0114] Optionally, if the machine-reviewed label of the input image is the same as the machine-reviewed label of the historical black image, the label similarity corresponding to the historical black image is 1; if the machine-reviewed label of the input image is different from the machine-reviewed label of the historical black image, the label similarity corresponding to the historical black image is 0. That is, if the machine-reviewed label of the input image and the machine-reviewed label of the historical black image are both compliant labels, or if the machine-reviewed label of the input image and the machine-reviewed label of the historical black image are both illegal labels, the label similarity corresponding to the historical black image is 1. If one of the machine-reviewed label of the input image and the machine-reviewed label of the historical black image is a compliant label and the other is an illegal label, the label similarity corresponding to the historical black image is 0.

[0115] For example, the label similarity corresponding to the historical black image can be expressed as C_machine.

[0116] Sub-step A324, obtaining the time differences corresponding to the first number of historical black images according to the current time and the review times corresponding to the first number of historical black images.

[0117] Subtract the review time of the historical black image from the current time to get the time difference corresponding to the historical black image. For example, if the review time of the historical image is in days, subtract the review date of the historical black image from the current date to get the day difference corresponding to the historical black image.

[0118] Exemplarily, the time difference corresponding to the historical black image can be expressed as Δt.

[0119] Sub-step A325, obtaining the comprehensive scores corresponding to the first number of historical black images according to the image similarities corresponding to the first number of historical black images, the text similarities corresponding to the first number of historical black images, the time differences corresponding to the first number of historical black images, and the label similarities corresponding to the first number of historical black images.

[0120] For each historical black image in the first number of historical black images, the image similarity corresponding to the historical black image, the text similarity corresponding to the historical black image, the time difference corresponding to the historical black image, and the label similarity corresponding to the historical black image are weighted to obtain a comprehensive score corresponding to the historical black image.

[0121] For example, the comprehensive score corresponding to the historical black map can be expressed as:

[0122] Score=α*S visual +β*S ocr_text +γ*exp(-λΔt)+η*C machine

[0123] α+β+γ+η=1

[0124] This application does not limit the values ​​of the above α, β, γ, and η. For example, α=0.6, β=0.2, γ=0.15, and η=0.05.

[0125] By calculating the comprehensive score of each historical black image based on the image similarity, text similarity, label similarity, and time difference between the input image and the historical black image, the target historical black image finally screened out not only considers the image content of the historical black image, but also the timeliness of the historical black image, so that the target historical black image can be more focused on historical images that have been manually reviewed recently, reducing the problem of invalid human review labels of historical images due to changes in review rules over a long period of time, thereby helping to improve the accuracy of image review.

[0126] In sub-step A33, among the comprehensive scores corresponding to the first number of historical black images, the historical black images corresponding to the second number of comprehensive scores that are ranked higher in descending order are determined as at least one target historical black image.

[0127] The comprehensive scores corresponding to the first number of historical black images are sorted in descending order of comprehensive scores to obtain the sorted first number of comprehensive scores. The second number of historical black images corresponding to the second number of comprehensive scores that have the highest ranking among the sorted first number of comprehensive scores are obtained to obtain the second number of historical black images. The second number of historical black images is the at least one target historical black image.

[0128] The second number is smaller than the first number. The specific value of the second number is set by the technical staff according to the image review requirements and is not limited in this application.

[0129] The first condition includes the image similarity between the image features of the target historical image and the image features of the input image, which is the image similarities of the first number of which are ranked higher among the image similarities corresponding to at least one historical black image, and the comprehensive score corresponding to the target historical image, which is the comprehensive scores of the second number of which are ranked higher among the comprehensive scores corresponding to the first number of historical black images.

[0130] After obtaining a first number of historical black images, further considering the comprehensive scores corresponding to the historical black images, a second number of historical black images are screened out from the first number of historical black images, so that at least one target historical black image obtained is a historical black image whose similarity with the input image meets the first condition, the number of target historical black images is reduced, and the correlation between the target historical black images and the input image is improved, thereby reducing the amount of calculation that needs to be performed by the multimodal large model, improving the speed of image review, and improving the accuracy of the image review results.

[0131] In some embodiments, for the process of determining at least one target historical white image, sub-step 223 includes at least one step among sub-steps B1 to B3.

[0132] Sub-step B1, calculating the similarity between the image features of the input image and the image features corresponding to at least one historical white image, to obtain the image similarity corresponding to at least one historical white image.

[0133] Sub-step B2: obtaining the historical white images corresponding to the image similarities corresponding to at least one historical white image, and sorting the image similarities in descending order to obtain the historical white images corresponding to the first number of image similarities.

[0134] Sub-step B3: obtaining at least one target historical white image according to the image information corresponding to the first number of historical white images.

[0135] In some embodiments, sub-step B3 includes at least one of sub-steps B31 to B33.

[0136] Sub-step B31 , performing optical character recognition on the input image and the first number of historical white images, respectively, to obtain text information of the input image and text information corresponding to the first number of historical white images, respectively.

[0137] Sub-step B32, based on the image information corresponding to the first number of historical white images, the image similarities corresponding to the first number of historical white images, the machine-reviewed labels of the input images, the text information of the input images and the text information corresponding to the first number of historical white images, obtain the comprehensive scores corresponding to the first number of historical white images. The comprehensive scores of the historical white images are used to measure the similarity between the historical white images and the input images.

[0138] In some embodiments, sub-step B32 includes at least one of sub-steps B321 to B325.

[0139] Sub-step B321, performing feature extraction on the text information of the input image to obtain the text features of the input image, and performing feature extraction on the text information corresponding to the first number of historical white images to obtain the text features corresponding to the first number of historical white images.

[0140] Sub-step B322 , calculating the similarity between the text features of the input image and the text features corresponding to the first number of historical white images, to obtain the text similarities corresponding to the first number of historical white images.

[0141] Sub-step B323 , calculating the similarity between the machine-reviewed label of the input image and the machine-reviewed labels corresponding to the first number of historical white images, to obtain the label similarities corresponding to the first number of historical white images.

[0142] Sub-step B324, obtaining the time differences corresponding to the first number of historical white images according to the current time and the review times corresponding to the first number of historical white images.

[0143] Sub-step B325, obtains the comprehensive scores corresponding to the first number of historical white images according to the image similarities corresponding to the first number of historical white images, the text similarities corresponding to the first number of historical white images, the time differences corresponding to the first number of historical white images, and the label similarities corresponding to the first number of historical white images.

[0144] Sub-step B33, among the comprehensive scores corresponding to the first number of historical white pictures, the historical white pictures corresponding to the second number of comprehensive scores that are ranked in descending order are determined as at least one target historical white picture.

[0145] Figure 6 A schematic diagram of the generation process of the target historical black image and the target historical white image is shown. First, the input image is passed through the feature extraction model to obtain the image features of the input image. Then, based on the image features of the input image, TOPK1 historical black images (the first number of historical black images) are screened out from the black library, and TOPK1 historical white images (the first number of historical white images, K1 is the first number) are screened out from the white library. A comprehensive score is calculated for the TOPK1 historical black images and they are sorted, and TOPK2 historical black images (the second number of historical black images) are screened out, and a comprehensive score is calculated for the TOPK1 historical white images and they are sorted, and TOPK2 historical white images (the second number of historical white images) are screened out, and K2 is the second number. The TOPK2 historical black image and TOPK2 historical white image finally obtained are the at least two target historical images mentioned in this application.

[0146] In some embodiments, step 230 includes at least one of sub-steps 231 - 236 .

[0147] Sub-step 231 : obtaining an audit rule index library corresponding to the audit rule library. The audit rule index library includes at least one index feature, and the index feature is related to the audit rule.

[0148] The audit rule index library is pre-generated based on the audit rule base. It includes at least one index feature, generated based on at least one audit rule. Each audit rule can generate one or more index features. The at least one index feature corresponding to each audit rule is used to indicate the audit rule. Different index features corresponding to each audit rule represent different feature presentations of the audit rule.

[0149] The construction process of the audit rule index library can refer to the following embodiment and will not be introduced here.

[0150] Sub-step 232 , calculating the similarity between the image feature of the input image and at least one index feature, and obtaining the index similarity corresponding to the at least one index feature.

[0151] Optionally, a calculation method such as Euclidean distance or cosine similarity can be used to calculate the similarity between the image feature of the input image and at least one index feature, thereby obtaining the index similarity corresponding to each index feature. The index similarity corresponding to each index feature refers to the similarity between the image feature of the input image and the index feature.

[0152] Sub-step 233 : obtaining index features corresponding to a third number of index similarities ranked in descending order among the index similarities corresponding to at least one index feature.

[0153] The index similarities corresponding to the at least one index feature are sorted in descending order of index similarity to obtain the at least one sorted index similarity. The index features corresponding to the top third number of index similarities in the sorted at least one index similarity are obtained to obtain the third number of index features.

[0154] The specific value of the third quantity is set by the technical personnel according to the image review requirements and is not limited in this application.

[0155] Sub-step 234: determining the audit rules corresponding to the third number of index features according to the third number of index features.

[0156] According to the mapping relationship between the index features and the audit rules, the audit rules corresponding to the third number of index features are obtained.

[0157] Optionally, the third number of audit rules may be directly determined as at least one target audit rule.

[0158] Sub-step 235 , performing deduplication filtering on the audit rules corresponding to the third number of index features to obtain filtered audit rules.

[0159] Since each audit rule can correspond to at least one index feature, a mapping relationship can exist between one audit rule and multiple index features. Therefore, among the audit rules corresponding to the third number of index features, there may be duplicate audit rules. It is necessary to perform deduplication filtering on the duplicate audit rules to obtain filtered audit rules. The filtered audit rule is at least one audit rule that does not appear duplicated among the audit rules corresponding to the third number of index features.

[0160] In sub-step 236, the audit rules corresponding to the fourth number of index similarities ranked top in descending order among the filtered audit rules are determined as at least one target audit rule.

[0161] The similarity corresponding to the audit rule refers to the index similarity corresponding to the index feature corresponding to the audit rule. Since each audit rule can correspond to at least one index feature, each audit rule can correspond to at least one index similarity. The maximum index similarity among the at least one index similarity corresponding to the audit rule is determined as the index similarity corresponding to the audit rule.

[0162] Obtain the index similarities corresponding to the filtered audit rules, sort the index similarities corresponding to the filtered audit rules in descending order of index similarity, and obtain at least one re-sorted index similarity. Obtain the audit rules corresponding to the fourth number of index similarities that are ranked highest among the at least one re-sorted index similarity, thereby obtaining a fourth number of audit rules. The fourth number of audit rules is the at least one target audit rule.

[0163] The fourth number is smaller than the third number. The specific value of the fourth number is set by the technical personnel according to the image review requirements and is not limited in this application.

[0164] The second condition includes the index similarity between the index feature corresponding to the target review rule and the image feature of the input image, which is the third number of index similarities ranked higher among the index similarities corresponding to at least one index feature, and the index similarity corresponding to the target review rule, which is the fourth number of index similarities ranked higher among at least one review rule.

[0165] By extracting a fourth number of audit rules that meet the second condition based on the similarity between the audit rules and the input image, the multimodal large model can specifically analyze the image content of the input image based on the target audit rules, thereby enhancing the targeted nature of the image audit and improving the accuracy of the image audit results.

[0166] In some embodiments, the process of constructing the audit rule index library includes at least one of the following steps C1 to C4.

[0167] Step C1: For each audit rule in at least one audit rule, perform synonymous rewriting on the audit rule to obtain at least one rewriting rule of the audit rule.

[0168] Synonymous rewriting is performed on the audit rule using the multimodal large model to obtain at least one rewritten rule of the audit rule. The semantics of the at least one rewritten rule are the same as or similar to the semantics of the audit rule corresponding to the at least one rewritten rule.

[0169] Step C2: Obtain keywords of the audit rules according to the audit rules.

[0170] The multimodal large model is used to extract keywords from the audit rules to obtain keywords of the audit rules. Optionally, the keywords of the audit rules can be words directly extracted from the audit rules, or words generated based on the semantic summary of the audit rules, which is not limited in this application.

[0171] Step C3, performing feature extraction on the audit rule, the at least one rewriting rule and the keyword respectively, to obtain the semantic features of the audit rule, the semantic features corresponding to the at least one rewriting rule and the semantic features of the keyword.

[0172] The audit rules are input into the feature extraction model, and the feature extraction model outputs the semantic features of the audit rules. At least one rewriting rule is input into the feature extraction model, and the feature extraction model outputs the semantic features corresponding to at least one rewriting rule. The keywords of the audit rules are input into the feature extraction model, and the feature extraction model outputs the semantic features of the keywords.

[0173] Step C4: obtaining an audit rule index library based on the semantic features of the audit rule, the semantic features corresponding to at least one rewriting rule, and the semantic features of the keywords.

[0174] The audit rule index library includes at least one index feature corresponding to at least one audit rule, and the at least one index feature corresponding to each audit rule includes at least one of the following: the semantic feature of the audit rule, the semantic feature corresponding to at least one rewriting rule corresponding to the audit rule, and the semantic feature of the keyword.

[0175] By semantically rewriting each audit rule to obtain at least one rewriting rule and extracting keywords for each audit rule, the content of the audit rule index library is expanded, so that the audit rule index library can include different forms of semantic expressions of each audit rule. Therefore, when extracting target audit rules based on the audit rule index library, it can avoid missing some target audit rules due to the limitations of the semantic expression of the audit rules, which helps to improve the accuracy of image audit.

[0176] Figure 7 A schematic diagram of the construction process of the audit rule index library is shown. The audit rule library includes N audit rules. Keywords and at least one rewriting rule are generated for each audit rule. For example, based on audit rule 1, the keywords of audit rule 1 and at least one rewriting rule corresponding to audit rule 1 are obtained. The at least one rewriting rule corresponding to audit rule 1 includes rewriting rule 1-1, ..., rewriting rule 1-n. The N audit rules, the keywords corresponding to the N audit rules, and the at least one rewriting rule corresponding to the N audit rules are respectively input into the feature extraction model, and the audit rule index library is constructed based on the output semantic features. The TOPN1 audit rule (the third number of audit rules) is extracted from the audit rule index library, where N1 is the third number, and then the TOPN2 audit rule (the fourth number of audit rules) is extracted from the TOPN1 audit rule, where N2 is the fourth number.

[0177] In some embodiments, the process of constructing the historical image library includes at least one of the following steps D1 to D4.

[0178] Step D1: For each of the at least two historical images, perform feature extraction on the historical image to obtain image features of the historical image.

[0179] Step D2: Obtain the human review label of the historical image and the review time of the historical image.

[0180] Step D2: Obtain machine-reviewed labels of historical images based on the historical images using an image security algorithm.

[0181] In step D3, a label analysis text of the historical image is obtained through a multimodal large model based on the historical image, the human review label of the historical image, the machine review label of the historical image, and at least one review rule.

[0182] Step D4 , analyzing text based on image features of historical images, human review labels of historical images, machine review labels of historical images, review time of historical images, and labels of historical images to obtain a historical image library.

[0183] Figure 8 A schematic diagram of the construction process of a historical image library is shown. Historical black images and historical white images are input into a feature extraction model, respectively, to obtain image features of the historical black images and image features of the historical white images. The image features of the historical black images are added to the black library, and the image features of the historical white images are added to the white library. A multimodal large model is used to analyze the reasons for violations of the historical black images based on the historical black images, their human review labels, their machine review labels, and at least one audit rule, to obtain label analysis text for the historical black images. Furthermore, a multimodal large model is used to analyze the reasons for compliance of the historical white images based on the historical white images, their human review labels, their machine review labels, and at least one audit rule, to obtain label analysis text for the historical white images. The human review labels of the historical black images, their machine review labels, and the label analysis text for the historical black images are added to the black library, and the human review labels of the historical white images, their machine review labels, and the label analysis text for the historical white images are added to the white library.

[0184] In some embodiments, if the input image is not configured with a human review label of the input image, the above method further includes step 250, and step 250 includes at least one step of sub-steps 251 to 252.

[0185] Sub-step 251, when the audit result of the input image indicates that the input image is an illegal image, or when the audit result of the input image indicates that the input image is a compliant image, obtain a human review label of the input image according to the audit result of the input image.

[0186] If the audit result of the input image indicates that the input image is an illegal image, the human review label of the input image is the illegal label. If the audit result of the input image indicates that the input image is a compliant image, the human review label of the input image is the compliant label.

[0187] Sub-step 252 , when the review result of the input image indicates that the input image is an uncertain image, determining that the input image is an image that needs to be manually reviewed to obtain a human review label.

[0188] If the review result of the input image indicates that the input image is an uncertain image, the input image needs to be manually reviewed to determine the human review label of the input image.

[0189] By manually reviewing input images whose review results indicate uncertain images when the input images are not configured with human review labels, and directly using the review results as human review labels for input images whose review results indicate illegal images or compliant images, the number of images that need to be manually reviewed is greatly reduced, the cost of manual review is reduced, and the efficiency of determining human review labels is improved.

[0190] In some embodiments, if the input image is configured with a human review label of the input image, the above method further includes step 260, and step 260 includes at least one step of sub-steps 261-262.

[0191] Sub-step 261 : When the human review label of the input image does not match the review result of the input image, it is determined that the input image is an image that needs to be manually reviewed again.

[0192] For example, if the human review label of the input image is a violation label, then if the review result of the input image indicates that the input image is a compliant image or an uncertain image, the input image is determined to be an image that requires manual review again. If the human review label of the input image is a compliance label, then if the review result of the input image indicates that the input image is a violation image or an uncertain image, the input image is determined to be an image that requires manual review again.

[0193] If the input image is an image that needs to be manually reviewed again, the input image will be manually reviewed again to obtain an updated human review label, and the updated human review label will be updated to the historical image library.

[0194] Sub-step 262 , when the human review label of the input image matches the review result of the input image, determining that the input image is an image that does not need to be manually reviewed again.

[0195] For example, if the human review label for an input image is a violation label, then if the review result of the input image indicates that the input image is a violation image, the input image is determined to be an image that does not require further manual review. If the human review label for an input image is a compliance label, then if the review result of the input image indicates that the input image is a compliance image, the input image is determined to be an image that does not require further manual review. This means that there is no need to update the historical image library.

[0196] By manually reviewing the input images whose human review labels do not match the review results, and mining image samples that differ from the manual review, it is easier to discover potential missed and mistaken detection problems in the image samples, improve the efficiency of quality inspection and error correction, and thus improve the quality of manual review.

[0197] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0198] Please refer to Figure 9 , which shows a block diagram of an image audit device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned image audit method, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Figure 9 As shown, the apparatus 900 may include: a data acquisition module 910 , an image determination module 920 , a rule determination module 930 and a result generation module 940 .

[0199] Data acquisition module 910 is used to obtain input images, a historical image library and an audit rule library; the historical image library includes image information corresponding to at least two historical images, and the image information of the historical images includes human review labels of the historical images, and the human review labels are audit labels obtained through manual review, and the human review labels are compliance labels or violation labels; the audit rule library includes at least one audit rule, and the audit rule is used to describe the judgment conditions of the image violation type.

[0200] The image determination module 920 is configured to obtain at least two target historical images in the historical image library based on the input image and the historical image library, wherein the similarity between the target historical images and the input image satisfies a first condition.

[0201] The rule determination module 930 is configured to obtain at least one target audit rule in the audit rule library based on the input image and the audit rule library, wherein the similarity between the target audit rule and the input image satisfies a second condition.

[0202] The result generation module 940 is used to generate the audit result of the input image based on the input image, the at least two target historical images and the at least one target audit rule through a multimodal large model, and the audit result of the input image is used to indicate that the input image is one of an illegal image, a compliant image, and an uncertain image.

[0203] In some embodiments, the image information of the historical image further includes image features of the historical image, the at least two historical images include at least one historical black image and at least one historical white image, the human review label of the historical black image is the violation label, and the human review label of the historical white image is the compliance label; the image determination module 920 is configured to:

[0204] Performing feature extraction on the input image to obtain image features of the input image;

[0205] Obtaining at least one target historical black image from the at least one historical black image according to image features of the input image and image features corresponding to the at least one historical black image;

[0206] Obtaining at least one target historical white image from the at least one historical white image according to image features of the input image and image features corresponding to the at least one historical white image;

[0207] The at least two target history images are obtained according to the at least one target history black image and the at least one target history white image.

[0208] In some embodiments, the image determination module 920 is configured to:

[0209] Calculating similarities between image features of the input image and image features corresponding to the at least one historical black image, to obtain image similarities corresponding to the at least one historical black image;

[0210] Obtaining the image similarities corresponding to the at least one historical black image, wherein the historical black images corresponding to a first number of image similarities that are ranked in descending order of image similarities;

[0211] The at least one target historical black image is obtained according to the image information respectively corresponding to the first number of historical black images.

[0212] In some embodiments, the image information of the historical image further includes a machine review tag of the historical image and the review time of the historical image. The machine review tag is a review tag obtained through an image security algorithm, and the machine review tag is the violation tag or the compliance tag. The input image is configured with the machine review tag of the input image. The image determination module 920 is used to:

[0213] performing optical character recognition on the input image and the first number of historical black images respectively to obtain text information corresponding to the input image and the first number of historical black images respectively;

[0214] Obtaining comprehensive scores corresponding to the first number of historical black images based on image information corresponding to the first number of historical black images, image similarities corresponding to the first number of historical black images, machine-reviewed labels of the input image, text information of the input image, and text information corresponding to the first number of historical black images, wherein the comprehensive scores of the historical black images are used to measure the similarity between the historical black images and the input image;

[0215] Among the comprehensive scores corresponding to the first number of historical black images, the historical black images corresponding to the second number of comprehensive scores that are ranked higher in descending order are determined as the at least one target historical black image.

[0216] In some embodiments, the image determination module 920 is configured to:

[0217] performing feature extraction on the text information of the input image to obtain text features of the input image, and performing feature extraction on the text information corresponding to the first number of historical black images to obtain text features corresponding to the first number of historical black images;

[0218] Calculating similarities between text features of the input image and text features corresponding to the first number of historical black images, to obtain text similarities corresponding to the first number of historical black images;

[0219] Calculating the similarity between the machine-reviewed label of the input image and the machine-reviewed labels corresponding to the first number of historical black images, to obtain the label similarities corresponding to the first number of historical black images;

[0220] Obtaining time differences corresponding to the first number of historical black images according to the current time and the review times corresponding to the first number of historical black images;

[0221] According to the image similarities corresponding to the first number of historical black images, the text similarities corresponding to the first number of historical black images, the time differences corresponding to the first number of historical black images, and the label similarities corresponding to the first number of historical black images, the comprehensive scores corresponding to the first number of historical black images are obtained.

[0222] In some embodiments, the rule determination module 930 is configured to:

[0223] Acquire an audit rule index library corresponding to the audit rule library, wherein the audit rule index library includes at least one index feature, and the index feature is related to the audit rule;

[0224] Calculating the similarity between the image feature of the input image and the at least one index feature to obtain index similarities corresponding to the at least one index feature;

[0225] Obtaining, among the index similarities corresponding to the at least one index feature, index features corresponding to a third number of index similarities that are ranked top in descending order;

[0226] Determining, based on the third number of index features, audit rules corresponding to the third number of index features respectively;

[0227] Performing deduplication filtering on the audit rules corresponding to the third number of index features to obtain filtered audit rules;

[0228] Among the filtered audit rules, the audit rules corresponding to the fourth number of index similarities that are ranked top in descending order are determined as the at least one target audit rule.

[0229] In some embodiments, the rule determination module 930 is configured to:

[0230] For each audit rule in the at least one audit rule, performing synonymous rewriting on the audit rule to obtain at least one rewriting rule of the audit rule;

[0231] According to the audit rule, obtain the keyword of the audit rule;

[0232] Performing feature extraction on the audit rule, the at least one rewriting rule, and the keyword to obtain semantic features of the audit rule, semantic features corresponding to the at least one rewriting rule, and semantic features of the keyword;

[0233] The audit rule index library is obtained according to the semantic features of the audit rule, the semantic features corresponding to the at least one rewriting rule, and the semantic features of the keyword.

[0234] In some embodiments, the image information of the historical image further includes label analysis text of the historical image, where the label analysis text is used to indicate the analysis process of the human review label of the historical image; the result generation module 940 is used to:

[0235] Generate, by a prompt text generator, a prompt text according to the input image, the at least two target historical images, and the at least one target review rule, wherein the prompt text is used to instruct the multimodal large model to generate a review result of the input image according to a preset template;

[0236] Generating review information of the input image according to the prompt text through the multimodal large model;

[0237] The audit information of the input image includes at least one of the audit result of the input image, the violation type of the input image, and the audit analysis text of the input image.

[0238] In some embodiments, the input image is not configured with a human review label of the input image; the apparatus 900 further includes an image review module, the image review module being configured to:

[0239] When the review result of the input image indicates that the input image is an illegal image, or when the review result of the input image indicates that the input image is a compliant image, obtaining a human review label for the input image according to the review result of the input image;

[0240] When the review result of the input image indicates that the input image is an uncertain image, the input image is determined to be an image that needs to be manually reviewed to obtain a human review label.

[0241] In some embodiments, the input image is configured with a human review label of the input image; and the image review module is configured to:

[0242] If the human review label of the input image does not match the review result of the input image, determining that the input image is an image that needs to be manually reviewed again;

[0243] In a case where the human review label of the input image matches the review result of the input image, it is determined that the input image is an image that does not need to be manually reviewed again.

[0244] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0245] Please refer to Figure 10 , which shows a block diagram of a computer device 1000 provided in one embodiment of the present application. The computer device 1000 can be any electronic device with data computing, processing, and storage functions. The computer device 1000 can be used to implement the image review method provided in the above embodiment.

[0246] Typically, the computer device 1000 includes a processor 1001 and a memory 1002 .

[0247] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0248] Memory 1002 may include one or more computer-readable storage media, which may be non-transitory. Memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage media in memory 1002 is used to store a computer program configured to be executed by one or more processors to implement the above-described image review method.

[0249] Those skilled in the art will understand that Figure 10 The structure shown in the figure does not constitute a limitation on the computer device 1000, and the computer device 1000 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0250] In an exemplary embodiment, a computer-readable storage medium is also provided, storing a computer program that, when executed by a processor of a computer device, implements the image review method. Optionally, the computer-readable storage medium may be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0251] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, so that the computer device performs the above-mentioned image review method.

[0252] It should be noted that this application can display a prompt interface, pop-up window or output voice prompt information before collecting the user's relevant data and during the process of collecting the user's relevant data. The prompt interface, pop-up window or voice prompt information is used to remind the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are terminated, that is, the user's relevant data is not obtained. In other words, all user data collected by this application are processed strictly in accordance with the requirements of relevant national laws and regulations. The informed consent or separate consent of the personal information subject is obtained only when the user agrees and authorizes it to collect the data. Subsequent data use and processing are carried out within the scope of authorization of laws and regulations and the personal information subject, and the collection, use and processing of relevant user data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0253] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0254] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An image review method, characterized in that: The method comprises: Obtain an input image, a historical image library, and an audit rule library; the historical image library includes image information corresponding to at least two historical images, the image information of the historical images includes a human review label for the historical images, the human review label is an audit label obtained through manual review, and the human review label is a compliance label or a violation label; the audit rule library includes at least one audit rule, the audit rule being used to describe a judgment condition for an image violation type; Obtaining, based on the input image and the historical image library, at least two target historical images in the historical image library, wherein similarities between the target historical images and the input image satisfy a first condition; According to the input image and the audit rule library, at least one target audit rule in the audit rule library is obtained, wherein the similarity between the target audit rule and the input image satisfies a second condition; A multimodal large model is used to generate an audit result of the input image based on the input image, the at least two target historical images and the at least one target audit rule. The audit result of the input image is used to indicate that the input image is one of an illegal image, a compliant image, and an uncertain image.

2. The method according to claim 1, characterized in that The image information of the historical image further includes image features of the historical image, the at least two historical images include at least one historical black image and at least one historical white image, the human review label of the historical black image is the violation label, and the human review label of the historical white image is the compliance label; The step of obtaining at least two target historical images in the historical image library according to the input image and the historical image library includes: Performing feature extraction on the input image to obtain image features of the input image; Obtaining at least one target historical black image from the at least one historical black image according to image features of the input image and image features corresponding to the at least one historical black image; Obtaining at least one target historical white image from the at least one historical white image according to image features of the input image and image features corresponding to the at least one historical white image; The at least two target history images are obtained according to the at least one target history black image and the at least one target history white image.

3. The method according to claim 2, characterized in that The obtaining, based on the image features of the input image and the image features corresponding to the at least one historical black image, at least one target historical black image in the at least one historical black image includes: Calculating similarities between image features of the input image and image features corresponding to the at least one historical black image, to obtain image similarities corresponding to the at least one historical black image; Obtaining the image similarities corresponding to the at least one historical black image, wherein the historical black images corresponding to a first number of image similarities that are ranked in descending order of image similarities; The at least one target historical black image is obtained according to the image information respectively corresponding to the first number of historical black images.

4. The method according to claim 3, characterized in that The image information of the historical image also includes a machine review label of the historical image and the review time of the historical image, wherein the machine review label is a review label obtained through an image security algorithm, and the machine review label is the violation label or the compliance label; the input image is configured with the machine review label of the input image; The obtaining, according to the image information corresponding to the first number of historical black images, the at least one target historical black image includes: performing optical character recognition on the input image and the first number of historical black images respectively to obtain text information corresponding to the input image and the first number of historical black images respectively; Obtaining comprehensive scores corresponding to the first number of historical black images based on image information corresponding to the first number of historical black images, image similarities corresponding to the first number of historical black images, machine-reviewed labels of the input image, text information of the input image, and text information corresponding to the first number of historical black images, wherein the comprehensive scores of the historical black images are used to measure the similarity between the historical black images and the input image; Among the comprehensive scores corresponding to the first number of historical black images, the historical black images corresponding to the second number of comprehensive scores that are ranked higher in descending order are determined as the at least one target historical black image.

5. The method according to claim 4, characterized in that Obtaining comprehensive scores corresponding to the first number of historical black images based on the image information corresponding to the first number of historical black images, the image similarities corresponding to the first number of historical black images, the machine-reviewed label of the input image, the text information of the input image, and the text information corresponding to the first number of historical black images, includes: performing feature extraction on the text information of the input image to obtain text features of the input image, and performing feature extraction on the text information corresponding to the first number of historical black images to obtain text features corresponding to the first number of historical black images; Calculating similarities between text features of the input image and text features corresponding to the first number of historical black images, to obtain text similarities corresponding to the first number of historical black images; Calculating the similarity between the machine-reviewed label of the input image and the machine-reviewed labels corresponding to the first number of historical black images, to obtain the label similarities corresponding to the first number of historical black images; Obtaining time differences corresponding to the first number of historical black images according to the current time and the review times corresponding to the first number of historical black images; According to the image similarities corresponding to the first number of historical black images, the text similarities corresponding to the first number of historical black images, the time differences corresponding to the first number of historical black images, and the label similarities corresponding to the first number of historical black images, the comprehensive scores corresponding to the first number of historical black images are obtained.

6. The method according to any one of claims 1 to 5, characterized in that The step of obtaining at least one target audit rule in the audit rule library according to the input image and the audit rule library includes: Acquire an audit rule index library corresponding to the audit rule library, wherein the audit rule index library includes at least one index feature, and the index feature is related to the audit rule; Calculating the similarity between the image feature of the input image and the at least one index feature to obtain index similarities corresponding to the at least one index feature; Obtaining, among the index similarities corresponding to the at least one index feature, index features corresponding to a third number of index similarities that are ranked top in descending order; Determining, based on the third number of index features, audit rules corresponding to the third number of index features respectively; Performing deduplication filtering on the audit rules corresponding to the third number of index features to obtain filtered audit rules; Among the filtered audit rules, the audit rules corresponding to the fourth number of index similarities that are ranked top in descending order are determined as the at least one target audit rule.

7. The method according to claim 6, characterized in that The method further comprises: For each audit rule in the at least one audit rule, performing synonymous rewriting on the audit rule to obtain at least one rewriting rule of the audit rule; According to the audit rule, obtain the keyword of the audit rule; Performing feature extraction on the audit rule, the at least one rewriting rule, and the keyword to obtain semantic features of the audit rule, semantic features corresponding to the at least one rewriting rule, and semantic features of the keyword; The audit rule index library is obtained according to the semantic features of the audit rule, the semantic features corresponding to the at least one rewriting rule, and the semantic features of the keyword.

8. The method according to any one of claims 1 to 7, characterized in that The image information of the historical image also includes label analysis text of the historical image, and the label analysis text is used to indicate the analysis process of the human review label of the historical image; Generating the audit result of the input image by the multimodal large model according to the input image, the at least two target historical images, and the at least one target audit rule includes: Generate, by a prompt text generator, a prompt text according to the input image, the at least two target historical images, and the at least one target review rule, wherein the prompt text is used to instruct the multimodal large model to generate a review result of the input image according to a preset template; Generating review information of the input image according to the prompt text through the multimodal large model; The audit information of the input image includes at least one of the audit result of the input image, the violation type of the input image, and the audit analysis text of the input image.

9. The method according to any one of claims 1 to 8, characterized in that The input image is not configured with a human review label of the input image; and the method further includes: When the review result of the input image indicates that the input image is an illegal image, or when the review result of the input image indicates that the input image is a compliant image, obtaining a human review label for the input image according to the review result of the input image; When the review result of the input image indicates that the input image is an uncertain image, the input image is determined to be an image that needs to be manually reviewed to obtain a human review label.

10. The method according to any one of claims 1 to 8, characterized in that The input image is configured with a human review label of the input image; and the method further comprises: If the human review label of the input image does not match the review result of the input image, determining that the input image is an image that needs to be manually reviewed again; In a case where the human review label of the input image matches the review result of the input image, it is determined that the input image is an image that does not need to be manually reviewed again.

11. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the image review method according to any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the image review method according to any one of claims 1 to 10.

13. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the image review method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Image sensitive content auditing method and system

    CN121353653A

  • Image sensitive content review method and system

    CN121353653B