Image classification method and device, equipment and storage medium
By combining the characteristics and classification results of the images to be reviewed and standard images in the image processing model, the model is dynamically adjusted to adapt to changes in the fields and standards, and the problems of model flexibility and low cost efficiency in the prior art are solved, and efficient and low-cost image classification is achieved.
Patent Information
- Application Number
- CN202410108872.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, image processing models need to be retrained after changes in the field of the image to be reviewed or the audit standards change, resulting in poor flexibility, high cost and low efficiency.
By inputting the images to be reviewed into the pre-trained image processing model, and combining the features and classification results of multiple standard images, the image processing model is dynamically adjusted to determine the target classification results, reducing the need for retraining.
It realizes the flexibility and classification efficiency of image processing models under low cost conditions, adapts to changes in audit standards, and ensures accurate classification.
Smart Images

Figure CN120375029A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing technologies, and in particular, to an image classification method, apparatus, device, and storage medium. Background Art
[0002] With the development and progress of science and technology, content security has gradually become the main content of Internet ecological governance. For the huge amount of Internet image data, relying solely on manual review not only has low recognition efficiency but also consumes a great deal of human resources. Therefore, in risk control work such as image content review, the mainstream review algorithms have gradually changed to artificial intelligence review based on computer vision.
[0003] In the prior art, in a scenario where image content is recognized to perform multi-label classification on an image, the image can be input into a pre-trained classification model capable of multi-label determination, and the classification result corresponding to the image can be determined based on its output result. Therefore, before performing multi-label classification on an image, it is necessary to construct a classification model capable of multi-label determination and perform model training based on a large number of training data sets to obtain a classification model with determined parameters. The construction cost and training cost of the classification model are relatively high, and moreover, the universality is poor.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following technical problems in the prior art:
[0005] When the field to which the image to be reviewed belongs or the review criteria change, it is necessary to retrain the image processing model to correctly classify, resulting in poor flexibility of the image processing model and high cost and low efficiency of image classification. Summary of the Invention
[0006] The present invention provides an image classification method, apparatus, device, and storage medium to achieve low-cost image classification.
[0007] In a first aspect, an embodiment of the present invention provides an image classification method, including:
[0008] Inputting the image to be reviewed into a pre-trained image processing model, so that the image processing model determines the review features and review classification result of the image to be reviewed;
[0009] Inputting a plurality of standard images corresponding to the image to be reviewed into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image;
[0010] Determining the target classification result of the image to be reviewed according to the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
[0011] In a second aspect, an embodiment of the present invention further provides an image classification device, including:
[0012] A first execution module, configured to input an image to be reviewed into a pre-trained image processing model, so that the image processing model determines the review features and review classification results of the image to be reviewed;
[0013] A second execution module, configured to input a plurality of standard images corresponding to the image to be reviewed into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image;
[0014] A determination module, configured to determine the target classification result of the image to be reviewed according to the review features and review classification results of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
[0015] In a third aspect, an embodiment of the present invention further provides a computer device, where the computer device includes:
[0016] One or more processors;
[0017] A storage device, configured to store one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement any one of the image classification methods in the present invention.
[0019] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, characterized in that the computer-executable instructions are used to execute any one of the image classification methods in the present invention when executed by a computer processor.
[0020] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program or instruction, where the computer program or instruction implements any one of the image classification methods in the present invention when executed by a processor.
[0021] The embodiments in the above invention have the following advantages or beneficial effects:
[0022] An embodiment of the present invention provides an image classification method, including: inputting an image to be audited into a pre-trained image processing model, so that the image processing model determines the audit features and audit classification results of the image to be audited; inputting a plurality of standard images corresponding to the image to be audited into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image; and determining the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each standard image, and the true classification results of each standard image. Through the adaptive adjustment of a plurality of standard images corresponding to the image to be audited, the above technical solution enables the image processing model to better adapt to the adjustment of audit criteria, increasing the flexibility of the image processing model. When the audit criteria or the data of the image to be audited changes, the existing trained image processing model can also accurately classify images, without the need for additional training or only requiring a small amount of additional training of the image processing model, achieving a more accurate target classification result for the image to be audited at a lower cost and higher efficiency, that is, realizing high-efficiency classification of the image to be audited while reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flowchart of an image classification method provided by an embodiment of the present invention;
[0024] Figure 2 is an implementation flowchart of an image classification method provided by an embodiment of the present invention;
[0025] Figure 3 is a flowchart of another image classification method provided by an embodiment of the present invention;
[0026] Figure 4 is a schematic structural diagram of an initial processing model in another image classification method provided by an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of determining the target classification result of an image to be audited provided by an embodiment of the present invention;
[0028] Figure 6 is a schematic diagram of another determination of the target classification result of an image to be audited provided by an embodiment of the present invention;
[0029] Figure 7 is a schematic structural diagram of an image classification device provided by an embodiment of the present invention;
[0030] Figure 8 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that, for the sake of convenience of description, only the parts related to the present invention rather than all the structures are shown in the drawings.
[0032] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc. In addition, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0033] In a scenario where it is necessary to identify the image content to perform multi-label classification on the image, in one case, a proprietary image classification model can be adopted. That is, after converting the actual business requirements into a solution that can be implemented by an algorithm, a common classification network (such as the resnet network of CNN) and a detection network (such as the YOLO series network) are used to construct an initial classification network, and then a large amount of training data is collected to train the initial classification network to obtain a proprietary image classification model that can be used for image classification. This process requires algorithm engineers to convert the actual business requirements into algorithms, with relatively high personnel costs. Moreover, a large amount of data is required for model training, and the data cost and time cost of model training are relatively high. In another case, a general large model can be adopted. Large models are usually used in scenarios such as chatting, searching, and recommendation, and currently, functions for image classification scenarios have not been developed. Moreover, the training cost of large models is even higher.
[0034] Therefore, the present application proposes an image classification method to achieve low-cost classification of images.
[0035] The image classification method proposed in the present application will be described in detail below with reference to the drawings and embodiments.
[0036] Figure 1 It is a flowchart of an image classification method provided by an embodiment of the present invention. Figure 2 It is an implementation flowchart of an image classification method provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation where it is necessary to perform high-efficiency classification on images at a reduced cost. This method can be executed by an image classification device, and the device can be implemented in a software and / or hardware manner. As Figure 1 and Figure 2 described, the method specifically includes the following steps:
[0037] Step 110: Input the image to be reviewed into a pre-trained image processing model, so that the image processing model determines the review features and review classification results of the image to be reviewed.
[0038] Specifically, a pre-trained image processing model can be used to determine the image features and classification results of an image. As Figure 2 shown, after inputting the image to be reviewed into a pre-trained image processing model, the image processing model can determine the review features and review classification results of the image to be reviewed. Among them, the review features can be understood as the image features of the image to be reviewed determined by the image processing model, and the review classification results can be understood as the classification results of the image to be reviewed determined by the image processing model.
[0039] The image processing model can be composed of a feature extraction module, a feature optimization module, and a classification module. In one embodiment, the feature extraction module includes a first feature extraction network and a second feature extraction network.
[0040] Therefore, the initial image features of the image to be reviewed can be determined based on the feature extraction module, and then the initial image features of the image to be reviewed can be optimized based on the feature optimization module to obtain the review features of the image to be reviewed. The classification module determines the review classification results of the image to be reviewed based on the initial image features of the image to be reviewed. In one embodiment, the initial image features are extracted through the first feature extraction network.
[0041] It should be noted that the feature optimization module can be used to optimize the dimensionality of the image features of the image to be reviewed extracted by the feature extraction module to obtain more accurate image features of the image to be reviewed, that is, to obtain the review features of the image to be reviewed.
[0042] In the embodiment of the present invention, based on a pre-trained image processing model, the review features and review classification results of the image to be reviewed are determined at low cost.
[0043] Step 120: Input multiple standard images corresponding to the image to be reviewed into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image.
[0044] The standard image can be understood as a standard image that is preferably selected according to experience and experiments and is consistent with the field of the image to be reviewed, including at least a normal image and a violation image. In one embodiment, by calculating the similarity between the image to be reviewed and the standard image, it can be determined whether the image to be reviewed is a normal image or a violation image.
[0045] Specifically, after inputting multiple standard images corresponding to the image to be reviewed into a pre-trained image processing model, as Figure 2As shown, the image processing model can respectively determine the standard features and standard classification results of each standard image. Among them, the standard features can be understood as the image features of the standard image determined by the image processing model, and the standard classification result can be understood as the classification result of the standard image determined by the image processing model. In one embodiment, Figure 2 the image processing model into which the image to be reviewed is input and the image processing models into which multiple standard images corresponding to the image to be reviewed are input are the same model. In another embodiment, Figure 2 the image processing model into which the image to be reviewed is input and the image processing models into which multiple standard images corresponding to the image to be reviewed are input are different models.
[0046] In addition, the image features of the image to be reviewed and the image features of the standard image are both composed of a preset number of floating-point numbers. The number of floating-point numbers that make up the image features of the image to be reviewed and the image features of the standard image is not specifically limited here and can be determined according to actual needs. Images can generally be divided into normal images and illegal images. The illegal categories of illegal images can be pornographic violations, violent violations, terrorist-related violations, political-related violations, etc. Therefore, the review classification results of the image to be reviewed and the standard classification results of the standard image can include normal images, pornographic violation images, violent violation images, terrorist-related violation images, political-related violation images, etc. The review classification results and the standard classification results can also include various other types of classification results, which will not be enumerated here. In one embodiment, the type of the review classification result is included in the standard classification result, that is, the type of the review classification result is a subset of the type of the standard classification result. In another embodiment, the type of the review classification result and the type of the standard classification result can be different.
[0047] Of course, it is also possible to determine the initial image features of each standard image based on the feature extraction module, and then optimize the initial image features of each standard image based on the feature optimization module to obtain the standard features of each standard image. The classification module determines the standard classification results of each standard image based on the initial image features of each standard image.
[0048] As mentioned above, the feature optimization module can be used to perform dimensionality optimization on the image features of each standard image extracted by the feature extraction module to obtain more accurate image features of each standard image, that is, to obtain the standard features of each standard image.
[0049] In the embodiments of the present invention, based on a pre-trained image processing model, the standard features and standard classification results of each standard image corresponding to the image to be reviewed are determined at low cost.
[0050] In practical applications, step 110 and step 120 can be executed simultaneously, or step 110 can be executed first and then step 120, or step 120 can be executed first and then step 110. This is not specifically limited here.
[0051] Step 130: Determine the target classification result of the image to be reviewed based on the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
[0052] In one embodiment, the review classification result of the image to be reviewed can be directly used as the classification result of the image to be reviewed. In some cases, such as changes in the field of the image to be reviewed or changes in the review criteria, the accuracy of this classification result may not be high. In another embodiment, the classification result of the image to be reviewed can be determined based on the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image, thereby improving the accuracy of the classification result.
[0053] Specifically, when determining the classification result of the image to be reviewed based on the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image, first, the similarity between the image to be reviewed and each standard image can be determined based on the review features of the image to be reviewed and the standard features of each standard image. The standard features corresponding to the largest M similarities or similarities greater than a preset threshold are determined as the target standard features, and then the classification result of the image to be reviewed is determined based on the standard classification results and true classification results of the standard images corresponding to each target standard feature. Among them, the specific values of M and the preset threshold can be set according to actual needs and are not specifically limited here.
[0054] In practical applications, in one embodiment, when classifying multiple images to be reviewed in the same field, the multiple standard images corresponding to the images to be reviewed form a standard image set. It is only necessary to determine the standard features and standard classification results of the standard image set once, and then store the standard features and standard classification results of the standard image set for use in subsequent image classification steps. That is, it is only necessary to determine the standard features and standard classification results of each standard image and store them when classifying the first or multiple images to be reviewed. When classifying other images to be reviewed, only the standard features and standard classification results of the standard image set need to be called. In another embodiment, the standard image set corresponding to the image to be reviewed can be adjusted, and the standard features and standard classification results of the adjusted standard image set are determined and then stored for use in subsequent classification steps. Optionally, the standard features and standard classification results of different standard image sets stored can be stored and marked separately for the user to select when reviewing and classifying the images to be reviewed, without having to input repeatedly. Optionally, the standard features and standard classification results of the adjusted standard image set can also be used to update the standard features and standard classification results of the pre-adjustment standard image set stored.
[0055] Of course, when it is necessary to classify the images to be reviewed in other fields, the standard images can be updated to the standard images of that field. After determining the review features and review classification results of the images to be reviewed and the standard features and standard classification results of each standard image based on the image processing model, according to the review features and review classification results of the images to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image, the target classification result of the images to be reviewed can be determined, so as to achieve accurate classification of the images to be reviewed at a low cost and high efficiency.
[0056] In the embodiments of the present invention, by combining the review features and review classification results of the images to be reviewed determined by the foregoing image processing model, the standard features and standard classification results of each standard image, and the true classification results of each standard image, a more accurate target classification result of the images to be reviewed is determined at a low cost and high efficiency.
[0057] The image classification method provided by the embodiments of the present invention includes: inputting the image to be reviewed into a pre-trained image processing model, so that the image processing model determines the review features and review classification results of the image to be reviewed, and inputting a plurality of standard images corresponding to the image to be reviewed into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image; determining the target classification result of the image to be reviewed according to the review features and review classification results of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image. According to the above technical solution, through the adaptive adjustment of a plurality of standard images corresponding to the image to be reviewed, the image processing model better adapts to the adjustment of the review standard, increasing the flexibility of the image processing model. When the review standard or the data of the image to be reviewed changes, the existing trained image processing model can also accurately classify the images, without the need for additional training or only a small amount of additional training of the image processing model, so as to determine a more accurate target classification result of the image to be reviewed at a low cost and high efficiency, that is, to achieve high-efficiency classification of the image to be reviewed while reducing costs.
[0058] Figure 3The flowchart of another image classification method provided by an embodiment of the present invention is applicable to the situation where images need to be classified efficiently at a reduced cost. On the basis of the above embodiments, before inputting the image to be reviewed into the pre-trained image processing model, the embodiment of the present invention adds "constructing an initial processing model based on a feature extraction module, a feature optimization module, and a classification module, where the feature extraction module includes a first feature extraction network and a second feature extraction network, and the second feature extraction network is the feature extraction part of a pre-trained large classification model". The explanations of the same or corresponding terms in the above embodiments are not repeated here. Refer to Figure 3 , the image classification method provided by the embodiment of the present invention includes:
[0059] Step 310, construct an initial processing model based on a feature extraction module, a feature optimization module, and a classification module.
[0060] Wherein, the feature extraction module includes a first feature extraction network and a second feature extraction network, and the second feature extraction network is the feature extraction part of a pre-trained large model.
[0061] Figure 4 The structural schematic diagram of the initial processing model in another image classification method provided by the embodiment of the present invention, as Figure 4As shown, the initial processing model consists of a feature extraction module, a feature optimization module, and a classification module. The feature extraction module consists of a first feature extraction network and a second feature extraction network. In one embodiment, the first feature extraction network can be a small model with relatively small parameters and scale. Preferably, it can be transformers (a graph neural network based on the attention mechanism) or CNN (convolutional neural network). The second feature extraction network can be a large model with relatively large parameters and scale. Preferably, it is the feature extraction part of large models such as CLIP, dino, dinoV2, BLIP, BLIP2, Instruct BLIP, etc. that can extract image features. For example, when selecting the latest and relatively stable BLIP2 as the large model, the feature extraction part of BLIP2 can be used as the large model. That is, the output of Q-former in BLIP2 can be used as the output of the second feature extraction network. The feature optimization module can consist of at least one fully connected layer, and the classification module can consist of at least one fully connected layer. In one embodiment, an initial processing model can be constructed based on the first feature extraction network, the second feature extraction network, the feature optimization module, and the classification module. Since the second feature extraction network can pre-determine or real-time determine the initial image features of the image, and when the second feature extraction network pre-determines the initial features of the image, during the training of the initial processing model, there is no need to train the second feature extraction network. That is, only the sub-model composed of the first feature extraction network, the feature optimization module, and the classification module in the initial processing model needs to be trained. At this time, the training efficiency of the initial processing model is relatively high. When the second feature extraction network real-time determines the initial features of the image, during the training of the initial processing model, the second feature extraction network needs to be trained. That is, the initial processing model composed of the first feature extraction network, the second feature extraction network, the feature optimization module, and the classification module needs to be trained. At this time, the training efficiency of the initial processing model is relatively low. However, the accuracy of the trained image processing model is higher. Therefore, after constructing the initial processing model based on the first feature extraction network, the second feature extraction network, the feature optimization module, and the classification module, a sub-model of the initial processing model can be constructed based on the first feature extraction network, the feature optimization module, and the classification module in the initial processing model.
[0062] In the embodiment of the present invention, an initial processing model is implemented based on the first feature extraction network, the second feature extraction network, the feature optimization module, and the classification module.
[0063] Step 320: Train the initial processing model to obtain the image processing model.
[0064] Before training the initial processing model, a large number of unlabeled training images and initial labeled training images need to be prepared, and the initial labeled training images also need to be augmented to obtain augmented labeled training images.
[0065] The data sources of the untagged training images can be: augmented open-source images (such as ImageNet, COCO, LAION, etc.), private images, and illegal images crawled by keyword and other technologies (for example, bloody images can be easily recognized as ketchup, and "ketchup" can be crawled to obtain illegal images with the illegal category of violence), and difficult images that are difficult to learn.
[0066] In practical applications, sampling can be performed among the above data sources of untagged training images according to the training cost (such as machine resources and training cycle) to obtain a large number of untagged training images. The sampling method can be uniform sampling, proportional sampling, or weighted sampling customized according to value. After determining a large number of untagged training images, data deduplication operations can be performed, and unique names can be given to the selected untagged training images. For example, the untagged training images can be named based on the MD5 value of the images.
[0067] The initial labeled training images include normal images and illegal images. The initial labeled training images can be sampled from the above-mentioned untagged training images, and the sampling method can also be uniform sampling, proportional sampling, or weighted sampling customized according to value. The data ratio of normal images to illegal images in the initial labeled training images obtained by sampling here is 1:1 - 10:1. After determining the initial labeled training images, data deduplication operations can also be performed, and unique names can be given to the selected initial labeled training images. For example, the initial labeled training images can be named based on the MD5 value of the images.
[0068] When training the feature output function of the image processing model, labeled training data and unlabeled training data can be used. When training the classification function of the image processing model, labeled training data is used to better fit the feature output function and classification function of the image processing model.
[0069] Optionally, the labeled training images in the training images used to train the initial processing model include initial labeled training images and augmented labeled training images. Correspondingly, the process of determining the augmented labeled training images includes:
[0070] Obtain the initial labeled training images and the untagged training images; for each of the initial labeled training images, determine the similarity between the initial labeled training image and each of the untagged training images according to the image features of the initial labeled training image and the image features of each of the untagged training images; determine the untagged training images corresponding to the similarities that meet the first preset condition as the augmented labeled training images corresponding to the initial labeled training images, where the category label of the augmented labeled training images corresponding to the initial labeled training images is the same as the category label of the initial labeled training images.
[0071] Specifically, after obtaining the initial labeled training images and unlabeled training images as described in the foregoing steps, the image features of the initial labeled training images and the image features of the unlabeled training images can be determined based on the second feature extraction network, and the image features of the initial labeled training images and the image features of the unlabeled training images can be stored in the form of key-value. The key is the MD5 value of the image, and the value is the image feature of the image. The image feature is generally composed of a preset number of floating-point numbers. For example, an image feature is composed of 768 floating-point numbers: 0.001, 0.350,..., 0.007. Among them, the image feature composed of 768 floating-point numbers is a vector with 768 elements, which can be obtained by encoding the image through a neural network. The neural network can be, for example, CNN, transformer, etc. The number of floating-point numbers in the vector can be 1600 or other real numbers, and the present invention does not limit this.
[0072] For each initial labeled training image, the similarity between the initial labeled training image and each unlabeled training image can be determined according to the image features of the initial labeled training image and the image features of each unlabeled training image, that is, the similarity between the values of the initial labeled training image and each unlabeled training image can be determined. Preferably, the cosine similarity can be determined. Furthermore, the unlabeled training image corresponding to the similarity that meets the first preset condition can be determined as the augmented labeled training image corresponding to the initial labeled training image.
[0073] The first preset condition can be the largest N similarities or similarities greater than the first preset threshold. That is, the unlabeled training image corresponding to the largest N similarities or similarities greater than the first preset threshold can be determined as the augmented labeled training image corresponding to the initial labeled training image. At this time, the class label of the augmented labeled training image corresponding to the initial labeled training image is the same as the class label of the initial labeled training image. Among them, N is a positive integer, and the specific values of N and the first preset threshold can be set according to actual needs and are not specifically limited here.
[0074] After determining the augmented labeled training images corresponding to each initial labeled training image, the augmented labeled training images can also be manually verified to screen out the images that do not meet the requirements.
[0075] In one implementation, step 320 may specifically include:
[0076] Obtain training images and the image features of each of the training images, where the training images include labeled training images and unlabeled training images, and the image features of each of the training images are pre-determined by the second feature extraction network; use the training images, the image features of each of the training images, and the class labels of each of the labeled training images as training data to perform network training on the sub-model composed of the first feature extraction network, the feature optimization module, and the classification module in the initial processing model, and calculate the loss function; perform network optimization based on the backpropagation algorithm until the loss function converges to obtain the image processing model.
[0077] Further, the loss function includes a feature loss function and a classification loss function. Correspondingly, calculating the loss function includes:
[0078] Determine the feature loss function according to the image features and training features of each of the training images, where the training features of each of the training images are determined by the initial processing model; determine the classification loss function according to the class labels and training classes of each of the labeled training images, where the training classes of each of the training images are determined by the initial processing model.
[0079] Specifically, obtain unlabeled training images and labeled training images as training images, where the labeled training images include initial labeled training images and augmented labeled training images. After obtaining the training images, the image features of each training image, that is, the image labels of each training image, can be pre-determined based on the second feature extraction network. After determining the training images and the image features of each training image, the training images, the image features of each training image, and the class labels of each labeled training image can be used as training data to perform network training on the sub-model composed of the first feature extraction network, the feature optimization module, and the classification module in the initial processing model, and calculate the loss function. After inputting the training images into the initial processing model, the initial processing model can determine the initial training features of the training images based on the first feature extraction network. The initial training features are the image features of the training images determined by the first feature extraction network, and the initial training features are optimized based on the feature optimization module to obtain the training features.
[0080] Generally, the dimension of the image features of the training images determined by the second feature extraction network is relatively large, and the dimension of the initial image features of the training images determined by the first feature extraction network is relatively small. It is impossible to directly determine the feature loss function based on the image features with a large difference in dimensions. Therefore, after determining the initial training features of the training images based on the first feature extraction network, the initial training features can be optimized based on the feature optimization module to obtain the training features, and the dimension of the training features of the training images is consistent with the dimension of the image features of the training images.
[0081] Furthermore, a feature loss function can be determined based on the image features and training features of each training image, that is, the image features of each training image can be used as feature labels, and the feature loss function can be determined according to the feature labels and training features of each training image. After the training category of the training image is obtained based on the initial training features of the training image processed by the classification module, a classification loss function can be determined according to the category label and training category of each labeled training image. Based on the backpropagation algorithm for network optimization until both the feature loss function and the classification loss function converge, an image processing model can be obtained.
[0082] In practical applications, the feature loss function can be one of MSE, KL divergence, and cosine similarity loss, and the classification loss function can be one or a combination of several of softmax, sigmoid, focal loss, arcface, binary loss, and triplet loss, etc.
[0083] In this embodiment, during the process of training the initial processing model to obtain the image processing model, there is no need to train the second feature extraction network, the model training cost is relatively low, and furthermore, the training process does not require too much labeled training data, further reducing the training cost.
[0084] In the embodiment of the present invention, after obtaining the training images and the image features corresponding to each training image, the initial processing model is trained based on the training images, the image features corresponding to each training image, and the category labels of each labeled training image to obtain an image processing model. The feature optimization module in the image processing model can ensure that the image processing model has seen enough samples, maintain the zero-shot learning ability of the image processing model, and reduce the cost of the image processing model for image processing.
[0085] In another embodiment, step 320 may specifically include:
[0086] Obtain training images, and determine the image features of each training image based on the second feature extraction network, where the training images include labeled training images and unlabeled training images; use the training images, the image features of each training image, and the category labels of each labeled training image as training data to perform network training on the initial processing model, and calculate the loss function; based on the backpropagation algorithm for network optimization until the loss function converges to obtain the image processing model.
[0087] Specifically, unlabeled training images and labeled training images are obtained as training images. Among them, the labeled training images include initial labeled training images and augmented labeled training images. The training images, the image features of each training image, and the class labels of each labeled training image are used as training data to perform network training on the initial processing model and calculate the loss function. After inputting the training images into the initial processing model, the image features of the training images, that is, the image labels of the training images, are determined in real time based on the second feature extraction network. The initial training features of the training images are determined based on the first feature extraction network. The initial training features are the image features of the training images determined by the first feature extraction network. The initial training features are optimized by the feature optimization module to obtain training features. The dimension of the training features of the training images is consistent with the dimension of the image features of the training images.
[0088] Furthermore, the feature loss function can be determined according to the image features and training features of each training image. That is, the image features of each training image can be used as feature labels, and the feature loss function is determined according to the feature labels and training features of each training image. After the initial training features of the training images are processed by the classification module to obtain the training classes of the training images, the classification loss function can be determined according to the class labels and training classes of each labeled training image. Based on the backpropagation algorithm for network optimization until both the feature loss function and the classification loss function converge, an image processing model can be obtained.
[0089] Similarly, the feature loss function can be one of MSE, KL divergence, and cosine similarity loss, and the classification loss function can be one or a combination of several of softmax, sigmoid, focal loss, arcface, binary loss, and triplet loss.
[0090] It should be noted that after the first feature extraction network and the second feature extraction network determine the image features and the feature optimization module optimizes the image features, the image features are normalized.
[0091] In this embodiment, during the process of training the initial processing model to obtain the image processing model, the second feature extraction network is trained, enabling parameter fine-tuning of the second feature extraction network during the model training process. Moreover, the more accurate image features of the training images, that is, the more accurate feature labels, can be determined in real time, improving the model accuracy. Additionally, the learning rate of the second feature extraction network can be set to 0, and only the image features of the training images are determined in real time based on the second feature extraction network, further improving the model accuracy.
[0092] In an image processing model, the first feature extraction network can be selected as transformers or CNN according to actual needs, and the second feature extraction network can be selected as the feature extraction part of large models such as CLIP, dino, dinoV2, BLIP, BLIP2, Instruct BLIP, etc. The invariant feature optimization module can ensure that the image processing model has always seen enough samples, maintain the zero-shot learning ability of the image processing model, and reduce the cost of the image processing model for image processing.
[0093] In an embodiment of the present invention, after obtaining a training image, the initial processing model is trained based on the training image, the image features of each training image, and the category label of each labeled training image to obtain an image processing model, and the image processing model can determine a more accurate image classification result.
[0094] Step 330: Input the image to be reviewed into a pre-trained image processing model, so that the image processing model determines the review features and review classification result of the image to be reviewed.
[0095] Step 340: Input multiple standard images corresponding to the image to be reviewed into the image processing model, so that the image processing model determines the standard features and standard classification results of each of the standard images.
[0096] Specifically, after obtaining the image to be reviewed and multiple standard images corresponding to the image to be reviewed, image scaling processing and normalization processing can be performed on the image to be reviewed and each standard image corresponding to the image to be reviewed. Then, the image to be reviewed can be input into a pre-trained image processing model, and the image processing model can determine the review features and review classification result of the image to be reviewed. Multiple standard images corresponding to the image to be reviewed can also be input into the image processing model, and the image processing model can respectively determine the standard features and standard classification results of each standard image. Furthermore, the review features and review classification result of the image to be reviewed, as well as the standard features and standard classification results of each standard image, can be stored in the form of key-value to implement the corresponding storage of the image features and classification results of the images. In addition, in key-value, the key can be an image feature or a classification result, and the value can be the classification result of the image or the image feature, which is not specifically limited herein.
[0097] In an embodiment of the present invention, based on the image processing model, the review features and review classification result of the image to be reviewed and the standard features and standard classification results of each standard image corresponding to the image to be reviewed are determined at low cost.
[0098] As described above, step 330 and step 340 can be executed simultaneously, or step 330 can be executed first and then step 340, or step 340 can be executed first and then step 330. No specific limitation is made here.
[0099] Step 350: Determine the target classification result of the image to be audited according to the audit features and the audit classification result of the image to be audited, the standard features and the standard classification results of each standard image, and the true classification results of each standard image.
[0100] In one implementation, step 350 may specifically include:
[0101] Determine the similarity between the image to be audited and each standard image according to the audit features of the image to be audited and the standard features of each standard image; determine the standard features corresponding to the similarity that meets the second preset condition as the target standard features, and determine the intermediate classification result of the image to be audited based on the standard classification result and the true classification result of the standard image corresponding to the target standard features; determine the target classification result of the image to be audited from the intermediate classification result of the image to be audited according to the audit classification result of the image to be audited.
[0102] Specifically, Figure 5 is a schematic diagram for determining the target classification result of the image to be audited provided by an embodiment of the present invention. As Figure 5As shown, taking the image to be reviewed, multiple standard images corresponding to the image to be reviewed, and the true classification results of the multiple standard images corresponding to the image to be reviewed as input data, after determining the review features and review classification results of the image to be reviewed, the standard features and standard classification results of each standard image based on the image processing model, according to the review features of the image to be reviewed and the standard features of each standard image, determine the similarity between the image to be reviewed and each standard image, and determine the target standard features as the standard features corresponding to the similarities that meet the second preset condition. Then, according to the standard classification results and true classification results of the target standard images corresponding to each target standard feature, determine the intermediate classification result of the image to be reviewed. Specifically, on the one hand, the standard classification result and / or true classification result with the highest frequency can be determined as the intermediate classification result of the image to be reviewed. On the other hand, the standard classification results with a frequency greater than the preset frequency can be determined as the intermediate classification result of the image to be reviewed. On the other hand, the standard classification result and / or true classification result of the standard image corresponding to the target standard feature with the maximum similarity can be determined as the intermediate classification result of the image to be reviewed. Furthermore, the target classification result of the image to be reviewed can be determined by combining the review classification result and the intermediate classification result of the image to be reviewed. Specifically, the overlapping classification results in the review classification result and the intermediate classification result can be determined as the target classification result, that is, the target classification result of the image to be reviewed is the intersection of the review classification result of the image to be reviewed and the intermediate classification result of the image to be reviewed, so as to realize the output data of the target classification result of the image to be reviewed.
[0103] In this embodiment, the second preset condition can be the largest K similarities or similarities greater than the second preset threshold. That is, the standard features corresponding to the largest K similarities or similarities greater than the second preset threshold can be determined as the target standard features. Among them, K is a positive integer, and the specific values of K, the second preset threshold, and the preset frequency can be set according to actual needs and are not specifically limited here. For example, K can be 10, the second preset threshold can be 0.5, and the preset frequency can be 3 times.
[0104] In another embodiment, step 350 can specifically include:
[0105] Determine the target standard image in the standard images according to the review classification result of the image to be reviewed; according to the review features of the image to be reviewed and the standard features of each target standard image, determine the similarity between the image to be reviewed and each target standard image; determine the standard features corresponding to the similarities that meet the third preset condition as the target standard features, and determine the target classification result of the image to be reviewed according to the standard classification result and the true classification result of the target standard image corresponding to each target standard feature.
[0106] Specifically,Figure 6 Another schematic diagram for determining the target classification result of the image to be reviewed provided by the embodiment of the present invention is shown as follows Figure 6 As shown, taking the image to be reviewed, multiple standard images corresponding to the image to be reviewed, and the true classification results of the multiple standard images corresponding to the image to be reviewed as input data, after determining the review features and review classification results of the image to be reviewed, the standard classification results of the standard features of each standard image based on the image processing model, first, determine the target standard image among the standard images according to the review classification result of the image to be reviewed, that is, the standard image with the same review classification result as the image to be reviewed can be determined as the target standard image, reducing the data volume of the standard images for which the similarity needs to be compared; secondly, determine the similarity between the review features of the image to be reviewed and the standard features of each target standard image, determine the standard features corresponding to the similarity that meets the third preset condition as the target standard features, and then determine the intermediate classification result of the image to be reviewed according to the standard classification results and true classification results of the target standard images corresponding to each target standard feature. As described above, the standard classification result with the most frequent occurrence, the standard classification result with the frequency greater than the preset frequency, or the standard classification result and / or true classification result of the standard image corresponding to the target standard feature with the largest similarity can be determined as the intermediate classification result of the image to be reviewed. Furthermore, the target classification result of the image to be reviewed can be determined by combining the review classification result and the intermediate classification result of the image to be reviewed. Specifically, the overlapping classification results in the review classification result and the intermediate classification result can be determined as the target classification result. Similarly, the target classification result of the image to be reviewed is the intersection part of the review classification result and the intermediate classification result of the image to be reviewed, realizing the output data of the target classification result of the image to be reviewed.
[0107] In this embodiment, the third preset condition may be the same as the second preset condition in the foregoing embodiment, or may be different from the second preset condition in the foregoing embodiment, and no specific limitation is made here. In the embodiment of the present invention, by combining the review features and review classification results of the image to be reviewed determined by the foregoing image processing model, the standard features and standard classification results of each standard image, and the true classification results of each standard image, the more accurate target classification result of the image to be reviewed is determined at a lower cost and higher efficiency.
[0108] The image classification method provided by the embodiments of the present invention includes: constructing an initial processing model based on a feature extraction module, a feature optimization module, and a classification module; training the initial processing model to obtain the image processing model; inputting the image to be audited into the pre-trained image processing model so that the image processing model determines the audit features and audit classification results of the image to be audited; inputting a plurality of standard images corresponding to the image to be audited into the image processing model so that the image processing model determines the standard features and standard classification results of each of the standard images; and determining the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each of the standard images, and the true classification results of each of the standard images.Based on the first feature extraction network, the second feature extraction network, the feature optimization module, and the classification module, an initial processing model is constructed. Based on the second feature extraction network, the initial features of each training image are determined in advance or in real time. After obtaining the training images and the image features corresponding to each training image, based on the training images, the image features corresponding to each training image determined in advance, and the category labels of each labeled training image in the training images, the sub-models constructed by the first feature extraction network, the feature optimization module, and the classification module in the initial processing model are trained to obtain an image processing model for image classification according to the pre-determined image features. This image processing model has a relatively high training efficiency. It is also possible to, after obtaining the training images, train the initial processing model based on the training images, the image features corresponding to each training image determined in real time, and the category labels of each labeled training image in the training images to obtain an image processing model for image classification according to the image features determined in real time. This image processing model can obtain more accurate image classification results. In addition, the feature optimization module in the image processing model can ensure that the image processing model has seen enough samples, maintain the zero-shot learning ability of the image processing model, and reduce the cost of the image processing model for image processing. After using the image to be reviewed as input data and inputting it into the image processing model, the image processing model can determine the review features and review classification results of the image to be reviewed. After using each standard image as input data and inputting it into the image processing model, the image processing model can respectively determine the standard features and standard classification results of each standard image, achieving the determination of the review features and review classification results of the image to be reviewed and the standard features and standard classification results of each standard image corresponding to the image to be reviewed at low cost. After determining the review features and review classification results of the image to be reviewed, the standard features and standard classification results of each standard image based on the image processing model, the target classification result of the image to be reviewed can be determined according to the review features and review classification results of the image to be reviewed, the standard features and standard classification results of each standard image, and the true classification results of each standard image, achieving the determination of a more accurate target classification result of the image to be reviewed at lower cost and higher efficiency. That is, by combining the small models that make up the first feature extraction network and the large models that make up the second feature extraction network, high-efficiency classification of the image to be reviewed is achieved while reducing costs.
[0109] Figure 7 FIG. is a schematic structural diagram of an image classification device provided by an embodiment of the present invention. This device and the image classification methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the image classification device, reference can be made to the embodiments of the above image classification methods.
[0110] The specific structure of this image classification device is as Figure 7 shown and includes:
[0111] The first execution module 710 is configured to input the image to be audited into a pre-trained image processing model, so that the image processing model determines the audit features and audit classification results of the image to be audited;
[0112] The second execution module 720 is configured to input a plurality of standard images corresponding to the image to be audited into the image processing model, so that the image processing model determines the standard features and standard classification results of each of the standard images;
[0113] The determination module 730 is configured to determine the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each of the standard images, and the true classification results of each of the standard images.
[0114] In one embodiment, the first execution module 710 and the second execution module 720 may be the same execution module. In another embodiment, the first execution module 710 and the second execution module 720 may be different execution modules.
[0115] Based on the above embodiments, the apparatus further includes:
[0116] The construction module is configured to construct an initial processing model based on a feature extraction module, a feature optimization module, and a classification module, wherein the feature extraction module includes a first feature extraction network and a second feature extraction network, and the second feature extraction network is a feature extraction part of a pre-trained large model.
[0117] Based on the above embodiments, the apparatus further includes:
[0118] The first training module is configured to obtain training images and the image features of each of the training images, wherein the training images include labeled training images and unlabeled training images, and the image features of each of the training images are pre-determined by the second feature extraction network; use the training images, the image features of each of the training images, and the class labels of each of the labeled training images as training data to perform network training on the sub-model composed of the first feature extraction network, the feature optimization module, and the classification module in the initial processing model, and calculate the loss function; perform network optimization based on the backpropagation algorithm until the loss function converges to obtain the image processing model.
[0119] Based on the above embodiments, the apparatus further includes:
[0120] The second training module is used to obtain training images, determine the image features of each of the training images based on the second feature extraction network, where the training images include labeled training images and unlabeled training images; use the training images, the image features of each of the training images, and the class labels of each of the labeled training images as training data to perform network training on the initial processing model, and calculate the loss function; perform network optimization based on the backpropagation algorithm until the loss function converges to obtain the image processing model.
[0121] In one implementation, the loss function includes a feature loss function and a classification loss function. Correspondingly, calculating the loss function includes:
[0122] Determine the feature loss function according to the image features and training features of each of the training images, where the training features of each of the training images are determined by the initial processing model; determine the classification loss function according to the class labels and training classes of each of the labeled training images, where the training classes of each of the training images are determined by the initial processing model.
[0123] In one implementation, the labeled training images in the training images include initial labeled training images and augmented labeled training images. Correspondingly, the process of determining the augmented labeled training images includes:
[0124] Obtain the initial labeled training images and the unlabeled training images; for each of the initial labeled training images, determine the similarity between the initial labeled training image and each of the unlabeled training images according to the image features of the initial labeled training image and the image features of each of the unlabeled training images; determine the unlabeled training images corresponding to the similarities that meet the first preset condition as the augmented labeled training images corresponding to the initial labeled training image, where the class labels of the augmented labeled training images corresponding to the initial labeled training image are the same as the class labels of the initial labeled training image.
[0125] Based on the above embodiments, the determining module 730 is specifically configured to:
[0126] Determine the similarity between the image to be audited and each of the standard images according to the audit features of the image to be audited and the standard features of each of the standard images; determine the target standard features corresponding to the similarities that meet the second preset condition, and determine the intermediate classification result of the image to be audited according to the standard classification result and the true classification result of the standard image corresponding to the target standard features; determine the target classification result of the image to be audited from the intermediate classification results of the image to be audited according to the audit classification result of the image to be audited.
[0127] Based on the above embodiments, a determination module 730 is specifically configured to:
[0128] Determine a target standard image from the standard images according to the review classification result of the image to be reviewed; determine the similarity between the image to be reviewed and each target standard image according to the review features of the image to be reviewed and the standard features of each target standard image; determine the standard features corresponding to the similarity that meets the third preset condition as the target standard features, and determine the target classification result of the image to be reviewed according to the standard classification result and the true classification result of the target standard image corresponding to each target standard feature.
[0129] The image classification device provided by the embodiments of the present invention can execute the image classification method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the image classification method.
[0130] It should be noted that in the embodiments of the above image classification device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0131] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Figure 8 It shows a block diagram of an exemplary computer device 8 suitable for implementing the embodiments of the present invention. Figure 8 The displayed computer device 8 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0132] As Figure 8 shown, the computer device 8 is presented in the form of a general-purpose computing electronic device. The components of the computer device 8 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0133] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0134] Computer device 8 typically includes a variety of computer system readable media. These media can be any available media accessible by computer device 8, including volatile and non-volatile media, removable and non-removable media.
[0135] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 8 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 8 not shown, typically referred to as a "hard disk drive"). Although Figure 8 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. System memory 28 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0136] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.
[0137] Computer device 8 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the computer device 8, and / or communicate with any device that enables the computer device 8 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. And, computer device 8 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 20. As Figure 8 shown, network adapter 20 communicates with other modules of computer device 8 through bus 18. It should be understood that although Figure 8is not shown in the figure, and other hardware and / or software modules can be used in combination with the computer device 8, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0138] The processing unit 16 executes various functional applications and page displays by running the programs stored in the system memory 28. For example, the image classification method provided by the present embodiment is implemented. The method includes:
[0139] Input the image to be audited into a pre-trained image processing model, so that the image processing model determines the audit features and audit classification results of the image to be audited;
[0140] Input multiple standard images corresponding to the image to be audited into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image;
[0141] Determine the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
[0142] Of course, those skilled in the art can understand that the processor can also implement the technical solutions of the image classification method provided in any embodiment of the present invention.
[0143] The embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the image classification method provided by the present embodiment is implemented. The method includes:
[0144] Input the image to be audited into a pre-trained image processing model, so that the image processing model determines the audit features and audit classification results of the image to be audited;
[0145] Input multiple standard images corresponding to the image to be audited into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image;
[0146] Determine the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
[0147] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component.
[0148] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable program codes. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component.
[0149] The program codes contained on the computer-readable media may be transmitted by any appropriate media, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0150] The computer program codes for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program codes may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0151] Those of ordinary skill in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be centralized on a single computing device or distributed across a network composed of multiple computing devices. Optionally, they can be implemented with program codes executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0152] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An image classification method, characterized in that, Including: Input the image to be audited into a pre-trained image processing model, so that the image processing model determines the audit features and audit classification results of the image to be audited; Input multiple standard images corresponding to the image to be audited into the image processing model, so that the image processing model determines the standard features and standard classification results of each standard image; Determine the target classification result of the image to be audited according to the audit features and audit classification results of the image to be audited, the standard features and standard classification results of each standard image, and the true classification results of each standard image.
2. The image classification method according to claim 1, wherein Before inputting the image to be audited into the pre-trained image processing model, it further includes: Construct an initial processing model based on a feature extraction module, a feature optimization module, and a classification module. Among them, the feature extraction module includes a first feature extraction network and a second feature extraction network, and the second feature extraction network is the feature extraction part of a pre-trained large model.
3. The image classification method according to claim 2, wherein The image processing model is trained through the following process: Obtain training images and the image features of each training image. Among them, the training images include labeled training images and unlabeled training images, and the image features of each training image are pre-determined by the second feature extraction network; Use the training images, the image features of each training image, and the class labels of each labeled training image as training data to perform network training on the sub-model composed of the first feature extraction network, the feature optimization module, and the classification module in the initial processing model, and calculate the loss function; Perform network optimization based on the backpropagation algorithm until the loss function converges to obtain the image processing model.
4. The image classification method according to claim 2, wherein The image processing model is trained through the following process: Obtain training images, and determine the image features of each training image based on the second feature extraction network. Among them, the training images include labeled training images and unlabeled training images; Use the training images, the image features of each training image, and the class labels of each labeled training image as training data to perform network training on the initial processing model, and calculate the loss function; Perform network optimization based on the backpropagation algorithm until the loss function converges to obtain the image processing model.
5. The image classification method according to claim 3 or 4, characterized in that The loss function includes a feature loss function and a classification loss function. Correspondingly, calculating the loss function includes: Determine the feature loss function according to the image features and training features of each training image, where the training features of each training image are determined by the initial processing model; Determine the classification loss function according to the class labels and training classes of each labeled training image, where the training classes of each labeled training image are determined by the initial processing model.
6. The image classification method according to claim 3 or 4, characterized in that, The labeled training images in the training images include initial labeled training images and augmented labeled training images. Correspondingly, the process of determining the augmented labeled training images includes: Obtain the initial labeled training images and the unlabeled training images; For each of the initial labeled training images, determine the similarity between the initial labeled training image and each of the unlabeled training images according to the image features of the initial labeled training image and the image features of each of the unlabeled training images; Determine the augmented labeled training image corresponding to the initial labeled training image for the unlabeled training image corresponding to the similarity that meets the first preset condition, where the class label of the augmented labeled training image corresponding to the initial labeled training image is the same as the class label of the initial labeled training image.
7. The image classification method according to claim 1, characterized in that, Determine the target classification result of the image to be reviewed according to the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each of the standard images, and the true classification results of each of the standard images, including: Determine the similarity between the image to be reviewed and each of the standard images according to the review features of the image to be reviewed and the standard features of each of the standard images; Determine the target standard features for the similarity that meets the second preset condition, and determine the intermediate classification result of the image to be reviewed according to the standard classification result and the true classification result of the standard image corresponding to the target standard features; Determine the target classification result of the image to be reviewed from the intermediate classification result of the image to be reviewed according to the review classification result of the image to be reviewed.
8. The image classification method according to claim 1, wherein Determine the target classification result of the image to be reviewed according to the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each of the standard images, and the true classification results of each of the standard images, including: Determine the target standard image from the standard images according to the review classification result of the image to be reviewed; Determine the similarity between the image to be reviewed and each of the target standard images according to the review features of the image to be reviewed and the standard features of each of the target standard images; Determine the target standard features for the similarity that meets the third preset condition, and determine the target classification result of the image to be reviewed according to the standard classification result and the true classification result of the target standard image corresponding to each of the target standard features.
9. An image classification device, characterized in that, Including: A first execution module for inputting the image to be reviewed into a pre-trained image processing model so that the image processing model determines the review features and review classification result of the image to be reviewed; A second execution module for inputting multiple standard images corresponding to the image to be reviewed into the image processing model so that the image processing model determines the standard features and standard classification results of each of the standard images; A determination module for determining the target classification result of the image to be reviewed according to the review features and review classification result of the image to be reviewed, the standard features and standard classification results of each of the standard images, and the true classification results of each of the standard images.
10. A computer device, characterized in that, The computer device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image classification method according to any one of claims 1-8.
11. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the image classification method according to any one of claims 1-8 when executed by a computer processor.
12. A computer program product, characterized in that, Comprising a computer program or instructions which, when executed by a processor, implement the image classification method according to any one of claims 1-8.