Image recognition method and device, storage medium, equipment and program product

By building and training large models in the field of vision, the cost and difficulty of custom sensitive image recognition is solved, efficient and accurate recognition results are achieved, and good migration and flexibility are provided.

CN120014312APending Publication Date: 2025-05-16HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411846875.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems such as high cost, high difficulty, and difficult to split up the detailed labels related to subjective cognition of objects in custom sensitive image recognition, which makes it difficult to effectively solve the growing demand for custom sexy image recognition of objects.

Method used

By building a large model of the visual field and training it with a custom sensitive image dataset, the target recognition model is obtained. This model can efficiently and accurately realize object custom sensitive image recognition, with good migration, and can be customized for custom sensitive image standards for different objects.

Benefits of technology

It realizes efficient and accurate custom sensitive image recognition, reduces development and maintenance costs, and improves the flexibility and accuracy of recognition effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014312A_ABST
    Figure CN120014312A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method and device, a storage medium, equipment and a program product, and the method comprises the steps: obtaining a to-be-recognized image; the to-be-recognized image is input into a target recognition model for image recognition, so that an output result of the target recognition model is obtained, and the output result comprises at least one prediction category and a probability value corresponding to each prediction category; determining the category of the to-be-recognized image according to the output result; wherein the target recognition model is obtained by training a visual field large model according to a user-defined sensitive image data set, and the visual field large model is obtained by training a general visual large model according to a label-free initial sensitive image data set. According to the scheme provided by the embodiment of the invention, self-defined sensitive image recognition of the object can be simply and accurately realized, good mobility is achieved, customization can be carried out according to self-defined sensitive image standards of different objects, and the self-defined sensitive image recognition effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image recognition method, apparatus, storage medium, device and program product. Background Art

[0002] In related technologies, fine-label recognition tasks, represented by sensitive image recognition, usually adopt a complex solution of "multi-model recognition + strategy rule combination". This method has effectively solved many business problems in the past, but with the continuous expansion and deepening of the business, its inherent limitations have become increasingly prominent.

[0003] First, as business needs continue to grow, the number of sensitive image labels that need to be identified and classified has increased dramatically, which has directly led to a significant increase in the number and complexity of the required models. At the same time, in order to deal with various special situations, the number and complexity of policy rules have also increased, significantly increasing the expansion and maintenance costs.

[0004] Secondly, there are significant differences between different subjects in the standards for sensitive image identification. Due to the subjectivity and diversity of sensitive standards, subjects often have different definitions of sensitive images. This difference becomes more obvious against the backdrop of the rapid development of artificial intelligence generated content (AIGC) technology, as the generation of image content accelerates and data boundaries become increasingly blurred, making it difficult to objectively and accurately describe some objects that fall below the standards for sensitive images in words.

[0005] Therefore, it is becoming increasingly difficult to perform customized sensitive image recognition. How to achieve a more efficient, flexible and customizable sensitive image recognition method has become a problem that needs to be solved urgently. Summary of the invention

[0006] The embodiments of the present application provide an image recognition method, storage medium, device and program product, which can easily and accurately realize the recognition of customized sensitive images of objects, have good portability, and can be customized according to the customized sensitive image standards of different objects, thereby improving the recognition effect of customized sensitive images.

[0007] On the one hand, an embodiment of the present application provides an image recognition method, the method comprising:

[0008] Obtain an image to be recognized;

[0009] Inputting the image to be recognized into a target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category;

[0010] Determine the category to which the image to be identified belongs according to the output result;

[0011] Among them, the target recognition model is obtained by training a visual field big model based on a custom sensitive image data set, and the visual field big model is obtained by training a general visual big model based on an unlabeled initial sensitive image data set.

[0012] On the other hand, an embodiment of the present application provides an image recognition device, the device comprising:

[0013] An acquisition unit, used for acquiring an image to be recognized;

[0014] an identification unit, configured to input the image to be identified into a target identification model for image identification, so as to obtain an output result of the target identification model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category;

[0015] A determination unit, configured to determine the category to which the image to be identified belongs according to the output result;

[0016] Among them, the target recognition model is obtained by training a visual field big model based on a custom sensitive image data set, and the visual field big model is obtained by training a general visual big model based on an unlabeled initial sensitive image data set.

[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the method described in any of the above embodiments.

[0018] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the image recognition method described in any of the above embodiments by calling the computer program stored in the memory.

[0019] On the other hand, an embodiment of the present application provides a computer program product, including computer instructions, which, when executed by a processor, implement the method described in any of the above embodiments.

[0020] The embodiments of the present application provide for obtaining an image to be identified; inputting the image to be identified into a target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category; determining the category to which the image to be identified belongs according to the output result; wherein the target recognition model is obtained by training a visual domain big model according to a custom sensitive image data set, and the visual domain big model is a solution obtained by training a general visual big model according to an unlabeled initial sensitive image data set, constructing a visual domain big model using the general visual big model, and then training the visual domain big model using the custom sensitive image data set to obtain a target recognition model, so that the target recognition model can efficiently and accurately realize object custom sensitive image recognition, has good portability, can be customized according to custom sensitive image standards of different objects, and improves the custom sensitive image recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flowchart of an image recognition method provided in an embodiment of the present application.

[0022] Figure 2 Schematic diagram of the training principle of the large model in the visual field provided in the embodiment of the present application.

[0023] Figure 3 A schematic diagram of the principle of obtaining a target recognition model based on training of a large model in the visual field provided in an embodiment of the present application.

[0024] Figure 4 A schematic diagram of the principle of obtaining a new target recognition model provided in an embodiment of the present application.

[0025] Figure 5 A schematic diagram of the structure of an image recognition device provided in an embodiment of the present application.

[0026] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application.

[0027] Figure 7 A schematic diagram of a storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The embodiments of the present application provide an image recognition method, storage medium, device and program product. Specifically, the image recognition method of the embodiments of the present application can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a smart car terminal and other devices. The terminal can also include a client, which can be a video client, a browser client, an instant messaging client or a small program. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0029] The embodiments of the present application can be applied to various scenarios such as customized sensitive image recognition.

[0030] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the game data involved in this application are all obtained with full authorization.

[0031] First, some nouns or terms that appear in the description of the embodiments of the present application are explained as follows:

[0032] OpenCV: Open Source Computer Vision Library, an open source computer vision library, is a cross-platform computer vision and machine learning software library released under the Apache 2.0 license (open source).

[0033] Self-supervised model for computer vision: A model that uses the structure and properties of the dataset itself for supervised learning. Unlike traditional supervised learning models that rely on a large amount of manually annotated data, self-supervised models establish the learning process by predicting future tasks in the dataset or the properties or structure of the data itself.

[0034] DINOv2 (Dense Image Object Navigation with Vision Transformers version2): DINOv2 is a self-supervised model based on the Vision Transformer (ViT) architecture, which is particularly suitable for tasks such as object navigation and image recognition. It learns rich visual features by capturing the associations between images and uses these features to perform various visual tasks.

[0035] MoCo, SimCLR, and BYOL are all contrastive learning algorithms in unsupervised learning and are part of self-supervised learning.

[0036] Data cleaning: refers to cleaning data with incorrect data labels;

[0037] Noisy learning: Data with incorrect labels is called noisy data. When it is difficult to further clean up the noisy data, the technical settings of the training process are used to perform training on a noisy data set, which is called noisy learning.

[0038] Semi-supervised: Semi-supervised learning is a machine learning paradigm that can be used to process data that is partially labeled and partially unlabeled. In semi-supervised learning, the model needs to use labeled data (supervised data) for supervised learning, and also needs to use unlabeled data (unsupervised data) to extract hidden structures and patterns in the data.

[0039] Self-supervised: Self-supervised learning is an unsupervised learning method in which the model is trained using the features of the data itself without externally provided labels.

[0040] In the related art, when solving the customization requirements related to sensitive image recognition of different objects (i.e., customized sensitive image requirements), it is mainly achieved by developing model fine labels and combining them with strategy combinations. As the number of model fine labels increases, the development cost of this solution becomes higher and higher, and the difficulty of strategy combination becomes greater and greater. At the same time, the fine labels related to the subjective cognition of the object become increasingly difficult to separate. Therefore, the customized sensitive image recognition solutions in the related art are increasingly difficult to solve the growing demand for customized sensitive image recognition of objects.

[0041] In view of the problems existing in the above-mentioned related technologies, the embodiments of the present application provide an image recognition method, apparatus, storage medium, device and program product, which are used to solve the technical problems of high cost and high difficulty of the object customized image recognition method in the related technology.

[0042] It should be noted that the order of description of the following embodiments is not intended to limit the priority order of the embodiments.

[0043] See also Figure 1 , Figure 1 This is a flow chart of an image recognition method provided in an embodiment of the present application. The method may include the following steps 110 to 130.

[0044] Step 110: Obtain an image to be recognized.

[0045] In some embodiments, the image to be recognized may be captured in real time by a camera device, read from a storage device, downloaded or received from a network, or obtained from an image processing library (eg, OpenCV).

[0046] Step 120: input the image to be recognized into the target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category.

[0047] Among them, the target recognition model is obtained by training the visual domain big model based on the customized sensitive image dataset, and the visual domain big model is obtained by training the general visual big model based on the unlabeled initial sensitive image dataset.

[0048] The images included in the customized sensitive image dataset are images with labels customized by the object as sensitive images, and the images included in the initial sensitive image dataset are images of related fields without labels.

[0049] In some embodiments, sensitive images refer to images containing content that is not suitable for public disclosure or may cause controversy. More specifically, these images may involve personal privacy, violence, pornography, abuse, malicious attacks or other content that is not suitable for public disclosure.

[0050] The general vision big model refers to a computer vision model that can process and understand a wide range of image content and perform a variety of image understanding tasks. It can be based on deep learning technology and can learn the distribution and characteristics of image data through training with a large amount of image data, thereby achieving efficient understanding and application of images.

[0051] The visual domain big model is a visual big model used in a specific field that is obtained by further training or fine-tuning the data set for the specific field on the basis of the general visual big model. For example, if the initial sensitive image data set is a violent image data set, then the visual domain big model is a violent image recognition big model; if the initial sensitive image data set is a pornographic image data set, then the visual domain big model is a pornographic image recognition big model.

[0052] In an embodiment of the present application, the image recognition method provided by the present application is introduced by taking the initial sensitive image data set as the initial pornographic image data set, the visual field large model as the pornographic image recognition large model, and the target recognition model as the custom pornographic image recognition model for identifying custom pornographic images as an example. The predicted categories in the output results of the target recognition model may include pornographic images and non-pornographic images.

[0053] Step 130: Determine the category of the image to be identified according to the output result.

[0054] In some embodiments, the image recognition method provided in the embodiments of the present application can also be migrated to custom tasks such as custom audio recognition tasks and custom text recognition tasks.

[0055] In some embodiments, the corresponding audio recognition method may include:

[0056] Obtain audio to be recognized; input the audio to be recognized into an audio recognition model for audio recognition to obtain the output result of the audio recognition model, wherein the audio recognition model is obtained by training a large auditory model based on a custom sensitive audio data set, and the large auditory model is obtained by training a general auditory large model based on an unlabeled initial sensitive audio data set; determine the category of the audio to be recognized based on the output result.

[0057] In some embodiments, the corresponding text recognition method may include:

[0058] Obtain the text to be recognized; input the text to be recognized into a text recognition model for text recognition to obtain the output result of the text recognition model, wherein the text recognition model is obtained by training a large model in the field of natural language processing based on a custom sensitive text data set, and the large model in the field of natural language processing is obtained by training a general natural language processing large model based on an unlabeled initial sensitive text data set; determine the category of the text to be recognized based on the output result.

[0059] Among them, the training methods of the audio recognition model and the text recognition model are consistent in principle with the training method of the target image recognition model provided in the embodiment of the present application, and the training methods of the auditory field large model and the natural language processing field large model are consistent in principle with the training method of the visual field large model provided in the embodiment of the present application, which will not be elaborated in this application.

[0060] In some embodiments, in step 130, determining the category of the image to be identified according to the output result includes the following steps 1301 to 1302:

[0061] Step 1301, sorting all prediction categories according to the probability values ​​in the output results, so as to select the prediction categories whose probability values ​​in the sorting results exceed a predefined threshold as candidate categories, and the predefined threshold is used to judge the credibility of the prediction category;

[0062] Predefined thresholds can be reasonably set by relevant personnel based on model performance or application scenario requirements.

[0063] By setting a predefined threshold, you can filter out prediction categories with higher credibility and avoid using prediction categories with low probability as candidate categories.

[0064] Step 1302: If there are multiple candidate categories, select the candidate category with the highest probability value as the category to which the image to be identified belongs.

[0065] When there are multiple candidate categories, selecting the candidate category with the highest probability value as the category to which the image to be identified belongs can maximize the accuracy of recognition.

[0066] In some embodiments, the method further includes the following steps 1303 to 1304:

[0067] Step 1303: if the highest probability value in the sorting result is lower than the predefined threshold, all the prediction categories in the output result and the probability value corresponding to each prediction category are displayed on the graphical user interface;

[0068] When the highest probability value is lower than the predefined threshold, it means that the target recognition model is not confident enough about the recognition result of the image to be recognized. At this time, displaying all predicted categories and their probability values ​​can help the subject understand the recognition status of the target recognition model for the image to be recognized and make further judgments or decisions.

[0069] Step 1304 , in response to the category selection information input based on the graphical user interface, determine the category to which the image to be identified belongs.

[0070] Specifically, an interface for an object to input category selection information is provided on the image user interface, and the category to which the image to be identified belongs is determined according to the category selection information input by the object.

[0071] When the target recognition model cannot determine the category of the image to be recognized (i.e., the highest probability value is lower than the predefined threshold), allowing the subject to enter the category selection information through the graphical user interface can be used as a remedial measure to ensure that the image to be recognized can be correctly classified. This method combines the automatic recognition ability of the target recognition model with the judgment ability of the user, improving the flexibility and accuracy of the image recognition method.

[0072] See also Figure 2 , Figure 2Schematic diagram of the training principle of the large model in the visual field provided in the embodiment of the present application.

[0073] In some embodiments, the training step of the large model in the visual field may include the following steps 210 to 220:

[0074] Step 210, by setting a scheduled task to reflow the online data of the trained large model and perform deduplication processing on the online data, an unlabeled initial sensitive image data set is obtained, and the trained large model is a model with the same application field as the target recognition model;

[0075] Specifically, Figure 2 As shown, for large-scale random data online, the large-scale random data can be timed to flow back at a specific time within each preset period (for example, every day), and the trained large model can be used to perform model filtering on the returned random data to obtain filtered sensitive image data, and the filtered sensitive image data can be further deduplicated to obtain an unlabeled initial sensitive image data set.

[0076] In some embodiments, after deduplication processing is performed on the filtered sensitive image data, balancing processing is also performed on the deduplication-processed sensitive image data, that is, the proportion of image data of different prediction categories is controlled to be roughly balanced.

[0077] Among them, the trained large models have the same application field as the target recognition model, that is, the target recognition models of the trained large models are all designed to solve the same problem or be applied to the same field. For example, if the target recognition model is used to identify pornographic images, the trained large model should also be a model trained for the pornographic image recognition task.

[0078] Step 220, based on the self-supervised pre-training framework of the general visual big model and the initial sensitive image data set, the general visual big model is domain migrated to obtain the visual domain big model.

[0079] Specifically, Figure 2 As shown, using the initial sensitive image dataset, the general visual large model is pre-trained under the self-supervised pre-training framework to adapt it to the image features of a specific field, and a visual field large model is obtained.

[0080] In the embodiment of the present application, the general vision big model can be any type of computer vision self-supervisory model. For example, the general vision big model can be a model such as DINOv2, MoCo, SimCLR, BYOL, etc.

[0081] See also Figure 3 , Figure 3 A schematic diagram of the principle of obtaining a target recognition model based on training of a large model in the visual field provided in an embodiment of the present application.

[0082] In some embodiments, the training step of the target recognition model may include the following steps 310 to 320:

[0083] Step 310, obtaining a custom sensitive image dataset, where the custom sensitive image dataset includes custom sensitive image samples and corresponding true labels;

[0084] The custom sensitive image dataset is a sensitive image dataset that is labeled by different objects themselves.

[0085] Step 320, performing a first parameter adjustment process on the large model in the visual domain according to the custom sensitive image data set to obtain a target recognition model.

[0086] In this embodiment, the size of the subjectively annotated custom sensitive image data set is much smaller than the size of the initial sensitive image data set, and the large model in the visual field can be customized and fine-tuned through a small amount of subjectively annotated data, which greatly reduces the customization cost and complexity of the strategy combination originally required for the development of fine labels to achieve customized sensitive image recognition, and can improve the customized recognition effect.

[0087] In some embodiments, in step 320, performing a first parameter adjustment process on the visual domain large model according to the custom sensitive image dataset to obtain a target recognition model may include the following steps 31 to 35:

[0088] Step 31, performing data cleaning on the custom sensitive image data set to obtain a cleaned data set;

[0089] By performing data cleaning on custom sensitive image datasets, errors in the custom sensitive image datasets can be discovered and corrected.

[0090] In some embodiments, in step 31, performing data cleaning on the custom sensitive image dataset to obtain a cleaned dataset may include the following steps 311 to 313:

[0091] Step 311, dividing the custom sensitive image data set into N sub-data, where N is a positive integer;

[0092] Wherein, N is a positive integer, and the specific value of N can be reasonably set by relevant personnel according to the application scenario.

[0093] Step 312, for each target sub-data in the N sub-data, use the remaining N-1 sub-data in the N sub-data as a training set, use the visual domain large model to predict the target sub-data, and mark the wrongly predicted data in the target sub-data as the first noise data;

[0094] Specifically, for each target sub-data in the N sub-data, the remaining N-1 sub-data in the N sub-data except the target sub-data are used as training sets to train the large model in the visual field, and then the target sub-data is predicted, and the wrongly predicted data in the target sub-data is marked as the first noise data.

[0095] Step 313, traverse and process the N data elements, and remove the first noise data marked in the custom sensitive image data set to obtain a cleaned data set.

[0096] Specifically, after repeating step 312N times, the traversal processing of N molecular data is completed, and the marking of all noise data in the custom sensitive image data set (that is, the first noise data marked in each traversal) is completed, and the marked first noise data is removed to obtain the cleaned data set.

[0097] Step 32, performing noise learning training on the large model in the visual field based on the cleaned data set to obtain an initial target recognition model;

[0098] In some embodiments, in step 32, performing noise learning training on the visual domain large model based on the cleaned data set to obtain an initial target recognition model may include the following steps 321 to 322:

[0099] Step 321, identifying second noise data in the cleaned data set;

[0100] After cleaning, there are still some noise data in the data set that have not been removed, and these data are called second noise data.

[0101] Step 322, perform noisy learning training on the cleaned data set through the large model of the visual domain until the training result reaches the expected performance index on the cleaned data set, and obtain the initial target recognition model, wherein the expected performance index is that the large model of the visual domain can distinguish between the clean data and the second noise data in the cleaned data set.

[0102] Specifically, the large model in the visual field is trained with noisy learning using the cleaned data set, and the training parameters of the large model in the visual field are efficiently fine-tuned using the loss function calculated using the training results until the training results reach the expected performance indicators on the cleaned data set, thereby obtaining the initial target recognition model.

[0103] Step 33, identifying the initial sensitive image data set based on the initial target recognition model, obtaining a sample prediction category of each initial sensitive image sample in the initial sensitive image data set, and using the sample prediction category as a pseudo label of each initial sensitive image sample;

[0104] Specifically, the initial sensitive image data set is input into the initial target recognition model for recognition, and the sample prediction category of each initial sensitive image sample in the initial sensitive image data set is obtained, and the sample prediction category of each initial sensitive image sample is used as a pseudo label of each initial sensitive image sample.

[0105] Step 34, merging the initial sensitive image dataset with pseudo labels with the custom sensitive image dataset to obtain a mixed training dataset;

[0106] Step 35, performing semi-supervised learning training on the initial target recognition model according to the mixed training data set to obtain the target recognition model.

[0107] In this embodiment, when the first parameter adjustment processing is performed on the large model in the visual field according to the custom sensitive image data set, data cleaning and noisy learning are introduced to ensure the high utilization rate of the labeled data. At the same time, the labeled data is expanded with the help of semi-supervised learning to further improve the fine-tuning effect.

[0108] In some embodiments, in step 35, semi-supervised learning training is performed on the initial target recognition model according to the mixed training data set to obtain the target recognition model, which may include the following steps 351 to 356:

[0109] Step 351, inputting the annotated data with real labels and the pseudo-labeled data with pseudo-labels in the mixed training data set into the initial target recognition model respectively, wherein the annotated data comes from the custom sensitive image data set, and the pseudo-labeled data comes from the initial sensitive image data set with pseudo-labels;

[0110] Step 352, performing supervised learning training on the labeled data, and calculating a supervised learning loss function, wherein the supervised learning loss function is determined based on a cross entropy loss between a true label of the labeled data and a first prediction result of the initial target recognition model on the labeled data;

[0111] In some embodiments, in addition to the cross entropy loss, a supervised learning loss function may also be calculated based on the true label of the labeled data and the first prediction result of the initial target recognition model for the labeled data using other loss calculation methods.

[0112] Step 353, performing unsupervised learning training on the pseudo-label data, and calculating an unsupervised learning loss function, wherein the unsupervised learning loss function is determined based on a consistency loss between a pseudo-label of the pseudo-label data and a second prediction result of the initial target recognition model on the pseudo-label data;

[0113] Step 354, optimizing the model parameters of the initial target recognition model according to the supervised learning loss function and the unsupervised learning loss function;

[0114] Step 355, in each iteration, dynamically adjusting the weight between the supervised learning loss function and the unsupervised learning loss function;

[0115] Step 356, after multiple iterations of training, a target recognition model is obtained.

[0116] In some embodiments, the method may further include step 330:

[0117] Step 330 , performing a second parameter adjustment process on the target recognition model according to different customized sensitive image data sets labeled with different objects, to obtain a new target recognition model suitable for different objects.

[0118] Specifically, different objects have different definitions for different sensitive images. Therefore, by using different customized sensitive image data sets labeled with different objects to adjust the second parameter of the target recognition model, a new target recognition model suitable for different objects can be obtained.

[0119] See also Figure 4 , Figure 4 A schematic diagram of the principle of obtaining a new target recognition model provided in an embodiment of the present application.

[0120] In some embodiments, according to different customized sensitive image data sets annotated with different objects, performing a second parameter adjustment process on the target recognition model to obtain a new target recognition model suitable for different objects may include the following steps 3301 to 3303:

[0121] Step 3301, receiving a custom sensitive image dataset annotated with different objects;

[0122] like Figure 4 As shown, the custom sensitive image datasets annotated with different objects may include a custom sensitive image dataset of object A, a custom sensitive image dataset of object B, and a custom sensitive image dataset of object C.

[0123] Step 3302, for each object-annotated custom sensitive image dataset, add a corresponding model layer in the object recognition model;

[0124] like Figure 4 As shown, the original model structure in the target recognition model is used as the model trunk, and the newly added model layer is used as the model branch. The parameters of the model trunk are frozen parameters, and the parameters of the model branch are trainable parameters.

[0125] Step 3303, when adjusting the parameters of the target recognition model according to different custom sensitive image data sets labeled with different objects, freeze the model backbone parameters of the target recognition model, and adjust the parameters of the model layers suitable for different custom sensitive image data sets to obtain a new target recognition model suitable for different objects.

[0126] Specifically, when adjusting the parameters of the target recognition model according to different custom sensitive image data sets labeled with different objects, the model backbone parameters of the original model backbone in the target recognition model are set to be non-trainable, that is, the model backbone parameters of the target recognition model are frozen, and the parameters of the newly added model layers suitable for different custom sensitive image data sets are adjusted to obtain new target recognition models suitable for different objects.

[0127] like Figure 4 As shown, the custom sensitive image dataset of object A, the custom sensitive image dataset of object B, and the custom sensitive image dataset of object C are used to adjust the parameters of the target recognition model, respectively, to obtain a new target recognition model suitable for object A, a new target recognition model suitable for object B, and a new target recognition model suitable for object C, respectively.

[0128] In this embodiment, by performing a second parameter adjustment on custom sensitive image data sets annotated with different objects, the new target recognition model can more accurately identify custom sensitive image data of specific objects, thereby improving the custom sensitive image recognition capability of the model. In addition, while keeping the backbone parameters of the model unchanged, only the parameters of the newly added model layer are adjusted, which can not only reduce training time and computing resource consumption, but also enable the new target recognition model to better adapt to custom sensitive image recognition tasks of different objects while maintaining the original general feature extraction capability.

[0129] All of the above technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.

[0130] The embodiments of the present application provide for obtaining an image to be identified; inputting the image to be identified into a target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category; determining the category to which the image to be identified belongs according to the output result; wherein the target recognition model is obtained by training a visual domain big model according to a custom sensitive image data set, and the visual domain big model is a solution obtained by training a general visual big model according to an unlabeled initial sensitive image data set, constructing a visual domain big model using the general visual big model, and then training the visual domain big model using the custom sensitive image data set to obtain a target recognition model, so that the target recognition model can efficiently and accurately realize the object custom sensitive image recognition, has good portability, can be customized according to the custom sensitive image standards of different objects, and improves the custom sensitive image recognition effect.

[0131] The present application also provides an image recognition device. Figure 5 , Figure 5 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of the present application. The image recognition device 500 may include:

[0132] An acquisition unit 510 is used to acquire an image to be recognized;

[0133] A recognition unit 520, configured to input the image to be recognized into a target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category;

[0134] A determination unit 530, configured to determine the category to which the image to be identified belongs according to the output result;

[0135] Among them, the target recognition model is obtained by training a visual field big model based on a custom sensitive image data set, and the visual field big model is obtained by training a general visual big model based on an unlabeled initial sensitive image data set.

[0136] Optionally, the image recognition device 500 further includes a training unit, which, when used to train the visual domain large model, is specifically used to:

[0137] By setting a scheduled task to reflow the online data of the trained large model and performing deduplication processing on the online data, an unlabeled initial sensitive image data set is obtained, wherein the trained large model is a model having the same application field as the target recognition model;

[0138] According to the self-supervised pre-training framework of the general visual big model and the initial sensitive image data set, the general visual big model is transferred to the domain to obtain the visual domain big model.

[0139] Optionally, when the training unit is used to train the target recognition model, it is specifically used to:

[0140] Obtain a custom sensitive image dataset, where the custom sensitive image dataset includes custom sensitive image samples and corresponding real labels;

[0141] The first parameter adjustment process is performed on the visual domain large model according to the custom sensitive image data set to obtain the target recognition model.

[0142] Optionally, when the training unit is used to perform a first parameter adjustment process on the visual domain large model according to the custom sensitive image data set to obtain the target recognition model, it is specifically used to:

[0143] Performing data cleaning on the custom sensitive image data set to obtain a cleaned data set;

[0144] Based on the cleaned data set, the large model in the visual field is trained with noise to obtain an initial target recognition model;

[0145] Identify the initial sensitive image data set based on the initial target recognition model, obtain a sample prediction category of each initial sensitive image sample in the initial sensitive image data set, and use the sample prediction category as a pseudo label of each initial sensitive image sample;

[0146] Merging the initial sensitive image dataset with the pseudo-labels with the custom sensitive image dataset to obtain a mixed training dataset;

[0147] According to the mixed training data set, semi-supervised learning training is performed on the initial target recognition model to obtain the target recognition model.

[0148] Optionally, when the training unit is used to perform data cleaning on the custom sensitive image data set to obtain a cleaned data set, it is specifically used to:

[0149] Divide the custom sensitive image data set into N sub-data, where N is a positive integer;

[0150] For each target sub-data in the N sub-data, the remaining N-1 sub-data in the N sub-data are used as training sets, the target sub-data are predicted using the visual domain large model, and the wrongly predicted data in the target sub-data are marked as first noise data;

[0151] The N data elements are traversed and processed, and the first noise data marked in the custom sensitive image data set is removed to obtain a cleaned data set.

[0152] Optionally, when the training unit is used to perform noise learning training on the large model in the visual field based on the cleaned data set to obtain an initial target recognition model, it is specifically used to:

[0153] identifying second noise data in the cleaned data set;

[0154] The cleaned data set is subjected to noisy learning training by the large visual domain model until the training result reaches an expected performance indicator on the cleaned data set, thereby obtaining an initial target recognition model, wherein the expected performance indicator is that the large visual domain model can distinguish between the clean data and the second noise data in the cleaned data set.

[0155] Optionally, when the training unit is used to perform semi-supervised learning training on the initial target recognition model according to the mixed training data set to obtain the target recognition model, it is specifically used to:

[0156] Inputting the annotated data with real labels and the pseudo-labeled data with pseudo-labels in the mixed training data set into the initial target recognition model respectively, wherein the annotated data comes from the custom sensitive image data set, and the pseudo-labeled data comes from the initial sensitive image data set with pseudo-labels;

[0157] Performing supervised learning training on the labeled data, and calculating a supervised learning loss function, wherein the supervised learning loss function is determined based on a cross entropy loss between a true label of the labeled data and a first prediction result of the initial target recognition model on the labeled data;

[0158] Performing unsupervised learning training on the pseudo-label data, and calculating an unsupervised learning loss function, wherein the unsupervised learning loss function is determined based on a consistency loss between a pseudo-label of the pseudo-label data and a second prediction result of the initial target recognition model on the pseudo-label data;

[0159] Optimizing the model parameters of the initial target recognition model according to the supervised learning loss function and the unsupervised learning loss function;

[0160] In each iteration, dynamically adjusting the weight between the supervised learning loss function and the unsupervised learning loss function;

[0161] After multiple iterations of training, the target recognition model is obtained.

[0162] Optionally, the training unit is further used for:

[0163] According to different customized sensitive image data sets annotated with different objects, a second parameter adjustment process is performed on the target recognition model to obtain a new target recognition model suitable for different objects.

[0164] Optionally, when the training unit is used to perform a second parameter adjustment process on the target recognition model according to different customized sensitive image data sets labeled with different objects to obtain a new target recognition model suitable for different objects, it is specifically used to:

[0165] Receive custom sensitive image datasets with different object annotations;

[0166] For each object-annotated custom sensitive image dataset, a corresponding model layer is added to the object recognition model;

[0167] When adjusting the parameters of the target recognition model according to different custom sensitive image data sets labeled with different objects, the model backbone parameters of the target recognition model are frozen, and the parameters of the model layers suitable for different custom sensitive image data sets are adjusted to obtain new target recognition models suitable for different objects.

[0168] Optionally, when the determining unit 530 is used to determine the category to which the image to be identified belongs according to the output result, it is specifically used to:

[0169] Sorting all the prediction categories according to the probability values ​​in the output results, so as to select the prediction categories whose probability values ​​in the sorted results exceed a predefined threshold as candidate categories, wherein the predefined threshold is used to judge the credibility of the prediction categories;

[0170] If there are multiple candidate categories, the candidate category with the highest probability value is selected as the category to which the image to be identified belongs.

[0171] Optionally, the determining unit 530 is further configured to:

[0172] If the highest probability value in the sorting result is lower than the predefined threshold, displaying all the prediction categories in the output result and the probability value corresponding to each prediction category on a graphical user interface;

[0173] In response to category selection information input based on the graphical user interface, the category to which the image to be identified belongs is determined.

[0174] Each unit in the above-mentioned image recognition device 500 can be implemented in whole or in part by software, hardware or a combination thereof. Each unit can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each unit.

[0175] The image recognition device 500 may be integrated in a terminal or a server that has a storage device and a processor and has computing capabilities, or the above device may be a terminal or a server.

[0176] Optionally, the present application further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0177] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a terminal or a server. Figure 6 As shown, the computer device 600 may include: a communication interface 601, a memory 602, a processor 603 and a communication bus 604. The communication interface 601, the memory 602, and the processor 603 communicate with each other through the communication bus 604. The communication interface 601 is used for the computer device 600 to communicate data with external devices. The memory 602 can be used to store software programs and modules, and the processor 603 runs the software programs and modules stored in the memory 602, such as the software programs of the corresponding operations in the aforementioned method embodiments.

[0178] Optionally, the processor 603 can call the software program and modules stored in the memory 602 to perform the following operations: obtain the image to be identified; input the image to be identified into the target recognition model for image recognition to obtain the output result of the target recognition model, the output result includes at least one prediction category and a probability value corresponding to each prediction category; determine the category to which the image to be identified belongs based on the output result; wherein the target recognition model is obtained by training the visual domain big model based on the custom sensitive image data set, and the visual domain big model is obtained by training the general visual big model based on the unlabeled initial sensitive image data set.

[0179] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program enables the computer device to execute the corresponding process in the image recognition method in the embodiment of the present application, which will not be described in detail for the sake of brevity.

[0180] A schematic diagram of a storage medium provided in an embodiment of the present application, such as Figure 7As shown, a program product 700 for implementing the above method according to an exemplary embodiment of the present application is described, which can adopt a portable compact disk read-only memory (CDROM) and include program code, and can be run on a computer device, such as a mobile phone. However, the program product of the present application is not limited thereto. In the present application, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0181] The present application also provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the image recognition method in the embodiment of the present application, which will not be described here for the sake of brevity.

[0182] The present application also provides a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the image recognition method in the present application, which will not be described here for the sake of brevity.

[0183] It should be understood that the processor of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and performed. The software module may be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0184] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0185] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0186] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0187] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0188] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0189] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An image recognition method, characterized in that: The method comprises: Obtain an image to be recognized; Inputting the image to be recognized into a target recognition model for image recognition to obtain an output result of the target recognition model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category; Determine the category to which the image to be identified belongs according to the output result; Among them, the target recognition model is obtained by training a visual field big model based on a custom sensitive image data set, and the visual field big model is obtained by training a general visual big model based on an unlabeled initial sensitive image data set.

2. The image recognition method according to claim 1, characterized in that: The training steps of the large model in the visual field include: By setting a scheduled task to reflow the online data of the trained large model and performing deduplication processing on the online data, an unlabeled initial sensitive image data set is obtained, wherein the trained large model is a model having the same application field as the target recognition model; According to the self-supervised pre-training framework of the general visual big model and the initial sensitive image data set, the general visual big model is transferred to the domain to obtain the visual domain big model.

3. The image recognition method according to claim 1 or 2, characterized in that: The training steps of the target recognition model include: Obtain a custom sensitive image dataset, where the custom sensitive image dataset includes custom sensitive image samples and corresponding real labels; The first parameter adjustment process is performed on the visual domain large model according to the custom sensitive image data set to obtain the target recognition model.

4. The image recognition method according to claim 3, characterized in that: The step of performing a first parameter adjustment process on the visual domain large model according to the custom sensitive image data set to obtain the target recognition model includes: Performing data cleaning on the custom sensitive image data set to obtain a cleaned data set; Based on the cleaned data set, the large model in the visual field is trained with noise to obtain an initial target recognition model; Identify the initial sensitive image data set based on the initial target recognition model, obtain a sample prediction category of each initial sensitive image sample in the initial sensitive image data set, and use the sample prediction category as a pseudo label of each initial sensitive image sample; Merging the initial sensitive image dataset with the pseudo-labels with the custom sensitive image dataset to obtain a mixed training dataset; According to the mixed training data set, semi-supervised learning training is performed on the initial target recognition model to obtain the target recognition model.

5. The image recognition method according to claim 3, characterized in that: The method further comprises: According to different customized sensitive image data sets annotated with different objects, a second parameter adjustment process is performed on the target recognition model to obtain a new target recognition model suitable for different objects.

6. The image recognition method according to claim 5, characterized in that: The second parameter adjustment process is performed on the target recognition model according to different customized sensitive image data sets annotated with different objects to obtain a new target recognition model suitable for different objects, including: Receive custom sensitive image datasets with different object annotations; For each object-annotated custom sensitive image dataset, a corresponding model layer is added to the object recognition model; When adjusting the parameters of the target recognition model according to different custom sensitive image data sets labeled with different objects, the model backbone parameters of the target recognition model are frozen, and the parameters of the model layers applicable to different custom sensitive image data sets are adjusted to obtain new target recognition models applicable to different objects.

7. An image recognition device, characterized in that: The device comprises: An acquisition unit, used for acquiring an image to be recognized; an identification unit, configured to input the image to be identified into a target identification model for image identification, so as to obtain an output result of the target identification model, wherein the output result includes at least one predicted category and a probability value corresponding to each predicted category; A determination unit, configured to determine the category to which the image to be identified belongs according to the output result; Among them, the target recognition model is obtained by training a visual field big model based on a custom sensitive image data set, and the visual field big model is obtained by training a general visual big model based on an unlabeled initial sensitive image data set.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the image recognition method according to any one of claims 1 to 6.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the image recognition method according to any one of claims 1 to 6 by calling the computer program stored in the memory.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the image recognition method according to any one of claims 1 to 6 is implemented.