Image classification method, device and equipment and readable storage medium

By splicing text description information and image classification prompt word vectors in the image classification model and adjusting the text feature extraction module, the problem of difficulty in extracting text description features in the image data is solved, and the accurate classification of image risk categories is achieved.

CN119939318APending Publication Date: 2025-05-06CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100213.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, in the classification of image content risk information, it is difficult to effectively extract the text description content features of risk information in image data, resulting in insufficient classification accuracy.

Method used

By obtaining the image to be classified and its corresponding text description data, the word vector corresponding to the text description information and the image classification prompt word vector are stitched, and the text feature extraction module of the image classification model is adjusted, thereby improving the text feature extraction ability.

Benefits of technology

The accurate classification of risk categories for classified images is realized, and the image classification model's ability to extract the content characteristics of risk information text description is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939318A_ABST
    Figure CN119939318A_ABST
Patent Text Reader

Abstract

The invention provides an image classification method, apparatus and device, and a readable storage medium. The method comprises the steps of obtaining text description data corresponding to a to-be-classified image; splicing the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector; and determining a risk category of the to-be-classified image based on the first text feature vector corresponding to the first word vector. Therefore, the text feature extraction capability of the text feature module can be improved by adjusting the image classification prompt word vector; the word vector corresponding to the text description information of the to-be-classified image and the image classification prompt word vector are spliced, so that more description contents of the risk information of the to-be-classified image can be obtained; the text feature extraction module can obtain a first text feature vector containing rich text feature representation based on the first word vector obtained after splicing processing, so that the risk category of the to-be-classified image can be classified more accurately based on the first text feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and specifically to an image classification method, device, equipment and readable storage medium. Background Art

[0002] In the field of image content security processing, it is necessary to identify the risk information in the image, such as personal privacy, commercial secrets and other target information contained in the image, and process such information accordingly. How to effectively protect the target content in the image data and achieve appropriate processing of the image data has become a key issue for individuals or institutions.

[0003] In the related technology, there is a solution that combines image and text to train the model to achieve understanding of image content. The risk information of the image content is screened and classified through specific text descriptions, and then the images with risky targets are processed. It has high practical value. However, accurate classification of risk information depends on the acquisition of text description content features of risk information in image data. How to improve the image classification model's ability to extract text description content features of risk information in image data is a key issue in the accurate classification of risk information in image data. Summary of the invention

[0004] The present application provides an image classification method, apparatus, device and readable storage medium, which can obtain text descriptions rich in risk information, thereby improving the feature extraction capability of the text feature extraction module of the image classification model and achieving accurate classification of the risk category of the image to be classified.

[0005] The technical solution of this application is implemented as follows:

[0006] An embodiment of the present application provides an image classification method, including: obtaining an image to be classified, and text description data corresponding to the image to be classified; the text description data includes description content of risk information for the image to be classified; concatenating the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of a text feature extraction module of a first image classification model; the first word vector is input into the text feature extraction module of the trained image classification model to obtain a first text feature vector; based on the first text feature vector, determining the risk category corresponding to the image to be classified.

[0007] According to the above-mentioned technical means, the text feature extraction capability of the text feature module can be improved by adjusting the image classification prompt word vector during the training process of the text feature extraction module of the first image classification model. The word vector corresponding to the text description information of the image to be classified and the image classification prompt word vector are further spliced ​​to obtain more descriptive content of the risk information of the image to be classified. The text feature extraction module can obtain a first text feature vector containing rich text feature representation based on the first word vector obtained after the splicing process, so that the risk category of the image to be classified can be more accurately classified based on the first text feature vector.

[0008] Furthermore, based on the first text feature vector, the risk category corresponding to the image to be classified is determined, including: inputting the image to be classified into an image feature extraction module of a trained image classification model to obtain a first image feature vector; fusing the first image feature vector and the first text feature vector based on a feature weight parameter to obtain a target feature vector; the feature weight parameter is obtained by tuning during the training of the classification module of the first image classification model; based on the target feature vector, the risk category of the image to be classified is determined.

[0009] According to the above technical means, the first image feature vector and the first text feature vector are fused based on the feature weight parameter, so that the target feature vector contains both image and text features, so that in the subsequent classification process based on the first target feature vector, a more accurate classification result can be obtained.

[0010] Furthermore, a method for determining an image classification prompt word vector includes: obtaining an image sample data set; the image sample data set includes multiple image samples, and text description samples corresponding to each image sample; concatenating the word vectors corresponding to the multiple text description samples and the initial image classification prompt word vector to obtain a first reference text matrix; inputting the first reference text matrix into a text feature extraction module of a first image classification model to obtain a first reference text feature matrix; and inputting the multiple image samples into an image feature extraction module of the first image classification model to obtain a first reference image feature matrix; based on the similarity between the first reference text feature matrix and the first reference image feature matrix, updating the initial image classification prompt word vector to obtain an image classification prompt word vector.

[0011] According to the above technical means, a first reference text matrix is ​​obtained by concatenating the word vector corresponding to the text description sample of the image sample and the initial image classification prompt word vector, and obtaining the first reference text feature matrix corresponding to the first reference text matrix and the first reference image feature matrix corresponding to the image sample. The initial image classification prompt word vector is updated according to the similarity between the first reference text feature matrix and the first reference image feature matrix. The initial image classification prompt word vector can be changed in a direction that is more conducive to the text feature extraction module of the image classification model to extract text features, thereby obtaining an image classification prompt word vector that can enhance the feature extraction capability of the text feature extraction module.

[0012] Furthermore, the word vectors corresponding to the multiple text description samples and the initial image classification prompt word vectors are concatenated to obtain a first reference text matrix, including: concatenating the initial image classification prompt word vector and the word vector corresponding to each text description sample to obtain multiple first reference word vectors; and determining the first reference text matrix based on the multiple first reference word vectors.

[0013] According to the above technical means, the initial image classification prompt word vector and the word vector corresponding to each text description sample are concatenated to obtain a first reference text matrix including richer text representations, so that in the process of training the text feature extraction module, the text feature extraction capability of the text feature extraction module can be improved.

[0014] Furthermore, the image sample data set also includes category labels corresponding to each image sample; the first reference text feature matrix includes multiple reference text feature vectors, and the first reference image feature matrix includes multiple reference image feature vectors; based on the similarity between the first reference text feature matrix and the first reference image feature matrix, the initial image classification prompt word vector is updated to obtain the image classification prompt word vector, including: determining the similarity between the candidate text feature vector in the first reference text feature matrix and each reference image feature vector; the candidate image feature vector is any one of the multiple reference text feature vectors; based on the similarity, determining the predicted category corresponding to each image sample; based on the predicted category and the corresponding category label, continuously adjusting the initial image classification prompt word vector to perform back-propagation training on the text feature extraction module of the first image classification model until the first convergence condition is met to obtain the image classification prompt word vector.

[0015] According to the above-mentioned technical means, by determining the similarity between the candidate text feature vector and each reference image feature vector in the first reference text feature matrix, determining the predicted category corresponding to the image sample based on the similarity, and continuously adjusting the initial image classification hint word vector based on the predicted category and the corresponding category label, the feature extraction capability of the text feature extraction module of the image classification model can be continuously improved. After the training of the text feature extraction module is completed, an image classification hint word vector can be obtained that enables the text feature extraction module to have good feature extraction performance.

[0016] Furthermore, the method also includes: after obtaining the image classification prompt vector, concatenating the word vectors corresponding to multiple text description samples and the image classification prompt word vector to obtain a second reference text matrix; inputting the second reference text matrix into the text feature extraction module of the first image classification model to obtain a second reference text feature matrix; based on the initial feature weight parameters, fusing the second reference text feature matrix and the first reference image feature matrix to obtain a reference feature matrix; based on the reference feature matrix, updating the initial feature weight parameters to obtain feature weight parameters.

[0017] According to the above technical means, a second reference text matrix is ​​obtained by concatenating the word vectors corresponding to the text description samples and the image classification prompt word vectors, and a second reference text feature matrix and a first reference image feature matrix corresponding to the second reference text matrix are fused based on the initial feature weight parameters. A reference feature matrix containing multiple text features and image features can be obtained, so that the initial feature weight parameters can be subsequently updated based on the reference feature matrix, and feature weight parameters that can increase the classification accuracy of the image classification model can be obtained.

[0018] Furthermore, based on the reference feature matrix, the initial feature weight parameters are updated to obtain the feature weight parameters, including: based on the reference feature matrix, determining the category probabilities corresponding to each of the multiple image samples; based on the category probabilities, continuously adjusting the initial feature weight parameters and performing back propagation training on the classification module of the first image classification model until the second convergence condition is met to obtain the feature weight parameters.

[0019] According to the above technical means, the initial feature weight parameters are adjusted by referring to the category probabilities corresponding to the multiple image samples corresponding to the feature matrix, so that the initial feature weight parameters change in the direction of higher classification accuracy of the first image classification model, so that the final feature weight parameters are more conducive to the classification module to perform correct classification.

[0020] Furthermore, the initial feature weight parameters include initial image feature weights, initial text feature weights and normalization coefficients; based on the initial feature weight parameters, the second reference text feature matrix and the first reference image feature matrix are fused to obtain a reference feature matrix, including: determining the first feature matrix based on the initial image feature weights, the initial text feature weights, the second reference text feature matrix and the first reference image feature matrix; determining the reference feature matrix based on the first feature matrix and the normalization coefficients.

[0021] According to the above technical means, the second reference text feature matrix and the first reference image feature matrix can be fused through the initial image feature weights and the initial text feature weights to obtain the first feature matrix. The first feature matrix is ​​further processed based on the normalization coefficient to improve the processing efficiency of the image classification model.

[0022] Furthermore, the method also includes: if the reference image sample includes multiple risk information with different risk levels, the risk level corresponding to the risk information with the highest risk level among the multiple risk information is determined as the category label corresponding to the reference image sample; the reference image sample is any one of the multiple image samples.

[0023] According to the above-mentioned technical means, by determining the risk level corresponding to the risk information with the highest risk level as the category label of the image sample, the recognition ability of high-risk information can be continuously improved during the training of the image classification model based on image samples, thereby improving the classification accuracy of the image classification model.

[0024] Furthermore, the trained image classification model also includes multiple image processing modules, and different image processing modules have different image processing methods; the method also includes: determining the target processing method for the image to be classified according to the risk category corresponding to the image to be classified, and the preset correspondence between the risk category and the image processing method; inputting the image to be classified into the image processing module corresponding to the target processing method to obtain the image processing result.

[0025] According to the above-mentioned technical means, based on the preset correspondence between risk categories and image processing methods, the target processing method corresponding to the image to be classified is determined, and the image to be classified is input into the image processing module corresponding to the target processing method. This can achieve different processing for images of different risk categories, solve the problem of singleness of multiple image content processing methods, and meet the needs of diversified post-processing solutions.

[0026] The present application provides an image classification device, including:

[0027] A first acquisition module is used to acquire an image to be classified and text description data corresponding to the image to be classified; the text description data includes description content of risk information for the image to be classified;

[0028] A first splicing processing module is used to splice the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model;

[0029] A first feature extraction module, used for inputting the first word vector into a text feature extraction module of a trained image classification model to obtain a first text feature vector;

[0030] The first determination module is used to determine the risk category corresponding to the image to be classified based on the first text feature vector.

[0031] The present application provides an image classification device, including:

[0032] A memory for storing computer programs that can be executed on the processor;

[0033] The processor is used to execute the image classification method provided in the embodiment of the present application when running the computer program.

[0034] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are configured to execute the above-mentioned image classification method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A flowchart of an image classification method provided in an embodiment of the present application;

[0036] Figure 2 A flowchart of an image classification model training method provided in an embodiment of the present application;

[0037] Figure 3 A schematic diagram of the classification of a sample image provided in an embodiment of the present application;

[0038] Figure 4 A schematic diagram of a processing process of a CLIP pre-training model provided in an embodiment of the present application;

[0039] Figure 5 A schematic diagram of a fine-tuning process of a CLIP pre-training model provided in an embodiment of the present application;

[0040] Figure 6 A schematic diagram of an improved CLIP pre-training model provided in an embodiment of the present application;

[0041] Figure 7 A schematic diagram of different processing methods after image classification provided in an embodiment of the present application;

[0042] Figure 8 A schematic diagram of the structure of an image classification device provided in an embodiment of the present application;

[0043] Fig. 9 A schematic diagram of the composition structure of an image classification device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0046] In the following description, reference is made to “some embodiments\other embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments\other embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0047] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0049] In the related art, an image processing method includes obtaining an image to be subjected to risk detection, inputting the image into a risk detection model, extracting image features of the image through the risk detection model, the image features are features used to represent the correlation between the image content of the image and the risk information to be detected, the risk information to be detected includes multiple categories of risk information, and determining the risk information of the image according to the image features and the semantic features corresponding to the risk information to be detected obtained by pre-training the risk detection model. The first image feature extraction module in the method is a feature extraction module pre-trained using a contrastive language-image pre-training model (CLIP). The specific source of this module is not specified in the article. If the module is obtained by self-training, the method will consume a lot of resources, and it is difficult for non-professional teams to complete the training of the module, which is not suitable for general users; if an open source trained model is used, since the image text data used for the training of the open source model is mostly public on the Internet, it is not compatible with confidential or privacy or niche image sets, and cannot meet the accuracy of feature extraction; the risk level of the single target in the image content and the different processing methods for different levels are not taken into account, and only the risk information of the image can be confirmed.

[0050] In the related art, an image processing method includes acquiring a first number of target recognition images, which are images of target objects collected in a predetermined area; determining image features of each target recognition image in the first number of target recognition images, which include foreground features and background features; fusing and desensitizing the foreground features and background features of the first number of target recognition images to generate corresponding foreground fused desensitized images and background fused desensitized images; fusing the foreground fused desensitized images and the background fused desensitized images to generate fused desensitized images corresponding to the first number of target recognition images. This method does not take into account the risk level of each target in the image content, and the different processing methods for different levels, and only uses the same processing method.

[0051] In the related art, an image data processing method for text-based image search, by marking the image as text data in the form of keywords and storing the information in a database, can conveniently classify various images stored in the device according to the keywords contained in the image tag, and the image data processing method for inputting search words through text involves an image data processing method for text-based image search. This method does not take into account the hierarchical distinction of various targets in the image content, and the different processing methods for different levels, and only uses text as a label for image classification, rather than using a text description to explain the content that the image may contain, the content of the image and the text cannot interact effectively, and the classification efficiency is low.

[0052] Based on the problems existing in the related art, the embodiment of the present application provides an image classification method, which can obtain text descriptions rich in risk information, thereby improving the feature extraction capability of the text feature extraction module of the image classification model and realizing accurate classification of the risk category of the image to be classified. Figure 1 FIG. 1 is a flow chart of an image classification method provided in an embodiment of the present application, and the method comprises the following steps:

[0053] S101, obtaining an image to be classified and text description data corresponding to the image to be classified.

[0054] It should be noted that the text description data includes description content for the risk information of the image to be classified. The risk information may be privacy or sensitive information such as personal privacy information, commercial confidential information, etc. contained in the image to be classified. Among them, personal privacy information may be faces, names, user license plate numbers, etc. that do not want to be made public, and commercial confidential information includes product design methods, sales strategies, etc.

[0055] In some embodiments, the text description data corresponding to the image to be classified may include description content for the target object containing risk information in the image to be classified. For example, if the image to be classified includes a face, the corresponding text description data may be: the image to be classified containing a face.

[0056] S102: Concatenate the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector.

[0057] In some embodiments, the image classification hint word vector may be a vector including one or more hint words, and the image classification hint words may include descriptive content used to classify the image. For example, for an image containing risk information, the hint words that may be included in the image classification hint word vector include classification, risk, etc.

[0058] In some embodiments, the image classification prompt word vector can be obtained after tuning during the training process of the text feature extraction module of the first image classification model, that is, during the training process of the text feature extraction module of the first image classification model, the initial image classification prompt word vector can be continuously adjusted until the text feature extraction module training is completed, and the trained image classification prompt word can be obtained. Among them, the initial image classification prompt word vector can be an image classification prompt word vector randomly generated during the initialization process, or a pre-set image classification prompt word vector.

[0059] In some embodiments, during the training of the text feature extraction model of the first image model, the initial image classification prompt word vector is adjusted to improve the text feature extraction effect of the text feature extraction module. After the training of the text feature extraction module is completed, the tuned image classification prompt word vector can be obtained, and the text feature extraction model can obtain richer text feature representation in the process of extracting text features based on the image classification prompt words.

[0060] In some embodiments, the first image classification model may be a pre-trained image classification model, which may be a CLIP model or other models that can perform image classification based on text-image data.

[0061] In some embodiments, the word vector corresponding to the text description information may be obtained by performing a word embedding operation using a word embedding module (Embedding) in a trained image classification model. The number of prompt words in the image classification prompt word vector and the number of word vectors corresponding to the text description information may be the same or different. In the process of splicing the word vector corresponding to the text description information and the image classification prompt word vector, the image classification prompt word vector may be spliced ​​before the word vector corresponding to the text description information, and the image classification prompt word vector may be spliced ​​after the word vector corresponding to the text description information, thereby obtaining the first spliced ​​word vector.

[0062] S103: Input the first word vector into a text feature extraction module of a trained image classification model to obtain a first text feature vector.

[0063] In some real-time, the trained image classification model may be an image classification model obtained after training a first image classification model using image samples, and the text feature extraction module of the trained image classification model may include a text encoder, which may perform feature encoding on the input first word vector to obtain a first text feature vector.

[0064] In some embodiments, since the first word vector includes the word vector corresponding to the text description information and the image classification prompt word vector, feature extraction is performed through the text feature extraction module, and the first text feature vector can include richer image content description features.

[0065] S104: Determine the risk category corresponding to the image to be classified based on the first text feature vector.

[0066] In some embodiments, the risk category may include multiple ones, and different risk categories may correspond to different risk levels, and different risk levels may represent different risk levels. The correspondence between risk levels and risk categories may be pre-established. For example, if the risk levels include high risk, medium risk, and low risk, the high risk level may be classified into one risk category, the medium risk level may be classified into one risk category, and the low risk level may be classified into one risk category.

[0067] In some embodiments, a higher risk may indicate a higher degree of risk, and risk information of different risk levels may be divided into different risk levels. For example, a person's name, license plate number, etc. may be classified as a low risk level, facial information may be classified as a medium risk level, and a person's contact information, ID number, etc. may be classified as a high risk level.

[0068] In some embodiments, the risk category of the image to be classified can be determined by the risk information in the image to be classified, and the trained image classification model can analyze and identify the risk information in the image to be classified based on the first text feature vector through its own image classification module, thereby determining the classification category of the image to be classified. In the case where the image to be classified includes target objects of multiple risk levels, the category corresponding to the highest risk level can be determined as the risk category of the image to be classified.

[0069] In an embodiment of the present application, an image to be classified and text description data corresponding to the image to be classified are obtained; the text description data includes description content of risk information for the image to be classified; the word vector corresponding to the text description information and the image classification prompt word vector are spliced ​​to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model; the first word vector is input into the text feature extraction module of the trained image classification model to obtain a first text feature vector; based on the first text feature vector, the risk category corresponding to the image to be classified is determined. In this way, by adjusting the image classification prompt word vector during the training process of the text feature extraction module of the first image classification model, the text feature extraction capability of the text feature module can be improved, and further the word vector corresponding to the text description information of the image to be classified and the image classification prompt word vector are spliced ​​to obtain more description content of the risk information of the image to be classified. The text feature extraction module can obtain a first text feature vector containing rich text feature representation based on the first word vector obtained after the splicing process, so that the risk category of the image to be classified can be more accurately classified based on the first text feature vector.

[0070] In some embodiments of the present application, based on the first text feature vector, the risk category corresponding to the image to be classified is determined, that is, the above step S104 can be implemented by the following steps S1041 to S1043, and each step is described separately below.

[0071] S1041. Input the image to be classified into an image feature extraction module of a trained image classification model to obtain a first image feature vector.

[0072] In some embodiments, the image feature extraction module of the trained image classification model can extract image features of the image to be classified input thereto, such as image color features, texture features, etc., to obtain a first image feature vector. In implementation, this step can be performed simultaneously with the above step S103, or before the above step S103.

[0073] S1042: Fusing the first image feature vector and the first text feature vector based on a feature weight parameter to obtain a target feature vector.

[0074] It should be noted that the feature weight parameter is obtained by tuning during the training process of the classification module of the first image classification model, and the feature weight parameter may be a feature adjustment parameter during the fusion process of the first image feature vector and the first text feature vector.

[0075] In some embodiments, during the training of the classification module of the first image classification model, the initial feature weight parameter can be continuously modified according to the classification effect of the classification module until the classification effect of the classification module meets the preset conditions (for example, the classification accuracy is greater than the preset threshold), and the adjusted feature weight parameter can be obtained. The initial feature weight parameter can be a randomly generated value during the initialization process, or it can be any preset value.

[0076] In some embodiments, if the fusion processing method of the first image feature vector and the first text feature vector is feature combination, the feature weight parameter can be used as a weighting parameter, and the two feature vectors can be combined using the weighting parameter to obtain a weighted fusion feature vector, that is, a target feature vector.

[0077] Exemplarily, the first image feature vector can be weighted by the feature weight parameters corresponding to the first image feature vector to obtain a weighted first image feature vector, and the first text feature vector can be weighted by the feature weight parameters corresponding to the first text feature vector to obtain a weighted first text feature vector, and then the weighted first image feature vector and the weighted first text feature vector can be added together to obtain the target feature vector.

[0078] S1043. Determine the risk category of the image to be classified based on the target feature vector.

[0079] In some embodiments, after obtaining a target feature vector including a first text feature vector and a first image vector, a classification module of a trained image classification model can be used to determine the category probabilities of the image to be classified corresponding to different risk categories based on the two features in the target feature vector, and determine the risk category with the largest category probability as the risk category of the image to be classified.

[0080] According to the above technical means, the first image feature vector and the first text feature vector are fused based on the feature weight parameter, so that the target feature vector contains both image and text features, so that in the subsequent classification process based on the first target feature vector, a more accurate classification result can be obtained.

[0081] In some embodiments of the present application, a method for determining image classification prompt words includes: obtaining an image sample data set; the image sample data set includes multiple image samples, and text description samples corresponding to each image sample; concatenating the word vectors corresponding to the multiple text description samples and the initial image classification prompt word vector to obtain a first reference text matrix; inputting the first reference text matrix into a text feature extraction module of a first image classification model to obtain a first reference text feature matrix; and inputting multiple image samples into an image feature extraction module of the first image classification model to obtain a first reference image feature matrix; based on the similarity between the first reference text feature matrix and the first reference image feature matrix, updating the initial image classification prompt word vector to obtain the image classification prompt word vector.

[0082] In some embodiments, multiple image samples in an image sample data set may include risk information, the text description sample corresponding to the image sample may be a description content of the risk information in the image sample, and the word vector corresponding to the text description sample may be a word vector obtained through a word embedding operation.

[0083] In some embodiments, the word vector corresponding to each text description sample and the initial image classification prompt word vector can be spliced ​​to obtain multiple spliced ​​word vectors, each spliced ​​word vector corresponds to a text description sample, and the first reference text matrix can be obtained by combining the spliced ​​word vectors.

[0084] In some embodiments, the text feature extraction module of the first image classification model can perform feature extraction on the input first reference text matrix to obtain a first reference text feature matrix. At the same time, the image feature extraction module of the first image classification model can also perform image feature extraction on multiple image samples to obtain a first reference image feature matrix corresponding to each image sample.

[0085] In some embodiments, the similarity between the first reference text feature matrix and the first reference image feature matrix can be cosine similarity, and the cosine similarity between each text feature vector in the first reference text feature matrix and each image feature vector in the first reference image feature matrix can be determined, and then the image classification prompt word vector is adjusted according to each cosine similarity to obtain an image classification prompt word vector that can enable the text feature extraction module of the image classification model to have good feature extraction performance.

[0086] According to the above technical means, a first reference text matrix is ​​obtained by concatenating the word vector corresponding to the text description sample of the image sample and the initial image classification prompt word vector, and obtaining the first reference text feature matrix corresponding to the first reference text matrix and the first reference image feature matrix corresponding to the image sample. The initial image classification prompt word vector is updated according to the similarity between the first reference text feature matrix and the first reference image feature matrix. The initial image classification prompt word vector can be changed in a direction that is more conducive to the text feature extraction module of the image classification model to extract text features, thereby obtaining an image classification prompt word vector that can enhance the feature extraction capability of the text feature extraction module.

[0087] In some embodiments of the present application, in the process of splicing the word vectors corresponding to multiple text description samples and the initial image classification prompt word vector, the initial image classification prompt word vector and the word vector corresponding to each text description sample can be spliced ​​to obtain multiple first reference word vectors; based on the multiple first reference word vectors, the first reference text matrix is ​​determined.

[0088] In some embodiments, the initial image classification prompt word vector can be spliced ​​before or after the word vector corresponding to each text description sample to obtain multiple first reference word vectors, and the vector length of the first reference word vector is the sum of the length of the initial image classification prompt word vector and the length of the word vector corresponding to the text description sample.

[0089] In some embodiments, multiple first reference word vectors can be combined, each first reference word vector corresponds to a word vector of a text description sample, and the first reference word vector is used as a row of the first reference text matrix, so that the first reference text matrix can be obtained. The number of rows of the first reference text matrix can be the number of text description samples, and the number of columns of the first reference text matrix can be the length of the first reference word vector, which can be the number of word vectors in the first reference word vector.

[0090] According to the above technical means, the initial image classification prompt word vector and the word vector corresponding to each text description sample are concatenated to obtain a first reference text matrix including richer text representations, so that in the process of training the text feature extraction module, the text feature extraction capability of the text feature extraction module can be improved.

[0091] In some embodiments of the present application, the image sample data set also includes a category label corresponding to each image sample; the first reference text feature matrix includes multiple reference text feature vectors, and the first reference image feature matrix includes multiple reference image feature vectors. The category label corresponding to the image sample can be a real category determined based on the risk information in the image sample.

[0092] In some embodiments, if the reference image sample includes multiple risk information with different risk levels, the risk level corresponding to the risk information with the highest risk level among the multiple risk information is determined as the category label corresponding to the reference image sample; the reference image sample is any one of the multiple image samples. By determining the risk level corresponding to the risk information with the highest risk level as the category label of the image sample, the recognition capability for high-risk information can be continuously improved during the training of the image classification model based on the image samples, thereby improving the classification accuracy of the image classification model.

[0093] In some embodiments, in the process of updating the initial image classification hint word vector based on the similarity between the first reference text feature matrix and the first reference image feature matrix, the similarity between the candidate text feature vector in the first reference text feature matrix and each reference image feature vector can be determined; based on the similarity, the predicted category corresponding to each image sample is determined; based on the predicted category and the corresponding category label, the initial image classification hint word vector is continuously adjusted to perform back-propagation training on a text feature extraction module of an image classification model until the first convergence condition is met to obtain the image classification hint word vector.

[0094] In some embodiments, each reference text feature vector in the first reference text feature matrix corresponds to a text feature vector of a text description sample, and each reference image feature vector in the first reference image feature matrix corresponds to an image feature vector of an image sample. By determining the similarity between the candidate text feature vector and each reference image feature vector in the first reference image feature matrix, the degree of similarity between the candidate text feature vector and each reference image feature vector can be determined.

[0095] In some embodiments, the candidate image feature vector is any one of a plurality of reference text feature vectors, and the similarity between the candidate text feature vector and each reference image feature vector may be cosine similarity. After obtaining the similarities between the candidate text feature vector and each reference image feature vector, the predicted category corresponding to the image sample corresponding to the candidate text feature vector may be determined based on each similarity. According to the difference between the preset category and the category label, and the first preset loss function, the corresponding first loss function value is determined. When the first convergence condition is not met, for example, the first loss function value is greater than the first preset threshold, the initial image classification hint word vector is adjusted to train the text feature extraction module of the first image classification model.

[0096] In some embodiments, during the process of adjusting the initial image classification prompt word vector, the model parameters of the first image classification model remain unchanged, that is, only the initial image classification prompt word vector is adjusted, and the first loss function value under different initial image classification prompt word vectors is determined until the first convergence condition is met, so that the trained image classification prompt word vector can be obtained.

[0097] According to the above-mentioned technical means, by determining the similarity between the candidate text feature vector and each reference image feature vector in the first reference text feature matrix, determining the predicted category corresponding to the image sample based on the similarity, and continuously adjusting the initial image classification hint word vector based on the predicted category and the corresponding category label, the feature extraction capability of the text feature extraction module of the image classification model can be continuously improved. After the training of the text feature extraction module is completed, an image classification hint word vector can be obtained that enables the text feature extraction module to have good feature extraction performance.

[0098] In some embodiments of the present application, after obtaining the image classification hint vector, the word vectors corresponding to multiple text description samples and the image classification hint word vector can be concatenated to obtain a second reference text matrix; the second reference text matrix is ​​input into the text feature extraction module of the first image classification model to obtain a second reference text feature matrix; based on the initial feature weight parameters, the second reference text feature matrix and the first reference image feature matrix are fused to obtain a reference feature matrix; based on the reference feature matrix, the initial feature weight parameters are updated to obtain feature weight parameters.

[0099] In some embodiments, since the image classification hint word vector and the initial image classification hint word vector may be different, the second reference text matrix and the first reference text matrix may also be different. The image classification hint word vector may be spliced ​​before the word vector corresponding to each text description sample, or the image classification hint word vector may be spliced ​​after the word vector of each text description sample, thereby obtaining the second reference text matrix.

[0100] In some embodiments, the text features in the first reference text matrix can be extracted by the text feature extraction module of the first image classification model to obtain the corresponding second reference text feature matrix. The second reference text feature matrix can include text feature vectors corresponding to multiple text description samples.

[0101] In some embodiments, the initial feature weight parameter may be any pre-set value, and the initial feature weight may be used to implement weighted fusion of the second reference text feature matrix and the first reference image feature matrix, thereby obtaining a fused reference feature matrix.

[0102] In some embodiments, the prediction category corresponding to each image sample can be determined based on the reference feature matrix obtained after fusion, the classification accuracy can be determined based on the difference between the prediction category and the corresponding label category, and the initial feature weight parameters can be updated based on the classification accuracy until the classification accuracy meets the accuracy threshold, and the feature weight parameters can be obtained.

[0103] According to the above technical means, a second reference text matrix is ​​obtained by concatenating the word vectors corresponding to the text description samples and the image classification prompt word vectors, and a second reference text feature matrix and a first reference image feature matrix corresponding to the second reference text matrix are fused based on the initial feature weight parameters. A reference feature matrix containing multiple text features and image features can be obtained, so that the initial feature weight parameters can be subsequently updated based on the reference feature matrix, and feature weight parameters that can increase the classification accuracy of the image classification model can be obtained.

[0104] In some embodiments of the present application, in the process of updating the initial feature weight parameters based on the reference feature matrix, the category probabilities corresponding to each of the multiple image samples can be determined based on the reference feature matrix; based on the category probabilities, the initial feature weight parameters are continuously adjusted and the classification module of the first image classification model is back-propagation trained until the second convergence condition is met to obtain the feature weight parameters.

[0105] In some embodiments, the second convergence condition may be a condition for determining that the training of the classification module of the first image classification model is completed. The second convergence condition may be the same as or different from the first convergence condition.

[0106] In some embodiments, the category probability of each image sample belonging to a different risk category can be calculated by referring to each eigenvalue in the feature matrix, and the risk category with the largest probability of each category can be determined as the reference category corresponding to the image sample. According to the difference between the reference category and the label category, and the second preset loss function, the corresponding second loss function value is determined. When it is determined that the second convergence condition is not met, such as when the second loss function value is greater than the second preset threshold, the initial feature weight parameters are adjusted and the classification module of the first image classification model is trained until the second convergence condition is met, such as when the second loss function value is less than or equal to the second preset threshold, the trained feature weight parameters can be obtained.

[0107] In some embodiments, during the training of the classification module of the first image classification model, the parameters of other modules of the first image classification model, such as the text feature extraction module, the image feature extraction module, etc., can remain unchanged, that is, only the initial feature weight parameters and the parameters of the classification module can be adjusted.

[0108] According to the above technical means, the initial feature weight parameters are adjusted by referring to the category probabilities corresponding to the multiple image samples corresponding to the feature matrix, so that the initial feature weight parameters change in the direction of higher classification accuracy of the first image classification model, so that the final feature weight parameters are more conducive to the classification module to perform correct classification.

[0109] In some embodiments of the present application, the initial feature weight parameters include initial image feature weights, initial text feature weights and normalization coefficients; in the process of fusing the second reference text feature matrix and the first reference image feature matrix based on the initial feature weight parameters, the first feature matrix can be determined based on the initial image feature weights, the initial text feature weights, the second reference text feature matrix and the first reference image feature matrix; and the reference feature matrix can be determined based on the first feature matrix and the normalization coefficient.

[0110] In some embodiments, the initial image feature weights may be the weights corresponding to the first reference image feature matrix, and the initial text feature weights may be the weights corresponding to the second reference text feature matrix. The first reference image feature matrix and the second reference text feature matrix may be weightedly summed based on the initial image feature weights and the initial text feature weights to obtain the first feature matrix.

[0111] In some embodiments, the normalization coefficient and the first characteristic matrix can be multiplied to obtain the first characteristic matrix. The normalization coefficient can transform large values ​​in the first characteristic matrix into small values. For example, values ​​of hundreds or thousands can be converted into values ​​within the range of 0 to 1, or into values ​​within the range of 0 to 10.

[0112] According to the above technical means, the second reference text feature matrix and the first reference image feature matrix can be fused through the initial image feature weights and the initial text feature weights to obtain the first feature matrix. The first feature matrix is ​​further processed based on the normalization coefficient to improve the processing efficiency of the image classification model.

[0113] In some embodiments of the present application, the trained image classification model also includes multiple image processing modules, and different image processing modules have different image processing methods. After obtaining the risk category of the image to be classified, the target processing method for the image to be classified can be determined based on the risk category corresponding to the image to be classified, and the preset correspondence between the risk category and the image processing method; the image to be classified is input into the image processing module corresponding to the target processing method to obtain the image processing result.

[0114] In some embodiments, depending on the image processing method, the image processing module may include a deletion module and a desensitization module. The deletion module may delete the image to be classified, and the desensitization module may replace the risky content in the image to be classified.

[0115] In some embodiments, the preset correspondence between the image risk category and the image processing method can be predetermined, and the processing methods for different risk categories may be different or the same. For example, if the risk categories include high-risk category, relatively high-risk category, medium-risk category, low-risk category and no-risk category, the processing method for images corresponding to the high-risk category may be deletion, the processing method for images corresponding to the relatively high-risk category and the medium-risk category may be desensitization, and the images corresponding to the low-risk category and the no-risk category may not be processed. It should be noted that the processing method for the risk category, as well as the risk category and the corresponding image processing method are merely exemplary descriptions, and this application does not limit this.

[0116] In some embodiments, the corresponding target processing method can be determined from the preset correspondence between the risk category and the image processing method according to the risk category of the image to be classified, and the corresponding image processing result can be obtained by inputting the image to be classified into the image processing module corresponding to the target processing method.

[0117] According to the above-mentioned technical means, based on the preset correspondence between risk categories and image processing methods, the target processing method corresponding to the image to be classified is determined, and the image to be classified is input into the image processing module corresponding to the target processing method. This can achieve different processing for images of different risk categories, solve the problem of singleness of multiple image content processing methods, and meet the needs of diversified post-processing solutions.

[0118] In an embodiment of the present application, an image to be classified and text description data corresponding to the image to be classified are obtained; the text description data includes description content of risk information for the image to be classified; the word vector corresponding to the text description information and the image classification prompt word vector are spliced ​​to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model; the first word vector is input into the text feature extraction module of the trained image classification model to obtain a first text feature vector; based on the first text feature vector, the risk category corresponding to the image to be classified is determined. In this way, by adjusting the image classification prompt word vector during the training process of the text feature extraction module of the first image classification model, the text feature extraction capability of the text feature module can be improved, and further the word vector corresponding to the text description information of the image to be classified and the image classification prompt word vector are spliced ​​to obtain more description content of the risk information of the image to be classified. The text feature extraction module can obtain a first text feature vector containing rich text description features based on the first word vector obtained after the splicing process, so that the risk category of the image to be classified can be more accurately classified based on the first text feature vector.

[0119] Next, the implementation process of the application embodiment in the actual application scenario is introduced.

[0120] Figure 2 A flowchart of an image classification model training method provided in an embodiment of the present application is provided. The method can be implemented through the following steps S201 to S203, and each step is described separately below.

[0121] S201, obtaining image-text annotation data.

[0122] The first step is image collection. After determining that the target category to be processed is T (a positive integer greater than 0), image data for these T categories of targets must be collected separately. During the process, it is ensured that the image information of each category of data set collected contains the target corresponding to that category, that is, it is ensured that the images of this category contain the same specified features.

[0123] The more similar the target features in a class of images are, the easier it is for the model to extract the feature information, and the better the training effect. It should be noted that according to actual needs, the concept of risk level can be set for these T types of targets, such as Figure 3 When preparing data, it is necessary to prepare it according to the level. A high-level class can contain targets other than the target of that class. Figure 3 It can be seen that the D-type target is higher than the A, B, C, ... T-type targets, so according to actual needs, the A-type data set 301A can only contain the sample image 301B of target A; the sample image of the B-type data set 302A can include target A, or include the sample image 302B of targets A and B at the same time; the C-type data set 303A can include the sample image of target C, or include the sample image of targets C and B at the same time, or include the sample image 303B of targets C and A at the same time; the D-type data set 304A can include the sample image of targets C and D at the same time, or include the sample image of targets D and B at the same time, or include the sample image 304B of targets D, C, and B at the same time. In addition to collecting the target data of the T-type that needs to be processed, it is also necessary to separately collect a type of data that does not need to be processed, that is, the image of this type does not contain any of the T-type targets (called the T+1 class).

[0124] Secondly, after the image collection is completed, it is necessary to perform corresponding statistical analysis on each category (including the T+1 category) of the images to be processed (to count which specific targets may be included in the category) and then perform text description. The text description must include the targets that may appear in the images of this category and need to specify their categories. The corresponding text description and category are used as their labels. The actual annotation depends on the specific level classification requirements. For example, assuming that the data image of category A only contains target A, then a label similar to "DESCRIBE_A: images containing target A, CLASS_A category" can be used as its label, where the text after DESCRIBE_A is its text description, and CLASS_A is the category; and if the data image of category D may only contain sample targets B, C and D, since the level of category D is higher than that of categories C and B, a label similar to "DESCRIBE_D: images containing target D and may contain targets B and C, CLASS_D category" can be used, where the text after DESCRIBE_D is its text description, and CLASS_D is its category. For the T+1th category, OTHER can be used to represent its category. The text descriptions of all samples in the same category conform to the same theme. The number of images in the same category cannot be too small. Too few images may affect the classification performance of this type of data. You can also add more images based on subsequent results.

[0125] S202. Fine-tune the CLIP pre-trained model based on the image-text annotation data.

[0126] like Figure 4 As shown, the CLIP pre-trained model includes an image encoder 401 and a text encoder 402. An image 403 is input into the image encoder 401 to obtain an image feature vector 404. The corresponding text category is subjected to a word embedding operation (Embdding) to obtain a word vector 405, which is then input into the image-text encoder 402 to obtain a text feature vector 406. The classification probability 408 of the image is obtained by calculating the cosine similarity 407 between the image feature vector 404 and the text feature vector 404.

[0127] First, you need to select a model that has been pre-trained by CLIP. You can choose it according to the actual task and requirements without any restrictions. In order to make the pre-trained model have better feature extraction capabilities for the text description category we are processing, it is necessary to partially fine-tune the pre-trained basic model. Figure 5 As shown, during fine-tuning, we add the prefix vector group E before the text word vector P 501,E P 501 can be passed [V1][V2][V3] … [Vm] To represent, m represents the total number of vectors in the prefix word vector group.

[0128] The combined word vector 502 of the text with the prefix word vector group 501 added is the input of the text encoder 402. The word vector 405 can be obtained by mapping the text 503 through the word embedding module (Embdding) 504. Add E in front of the word vector 405 of the text P 501, where Vt (t = 1, 2, ... m) is an n-dimensional vector, whose length needs to be equal to the word vector generated by the subsequent text description). During training, only the prefix of the input word vector is learned, and all other parameters remain unchanged. The fine-tuning effect is achieved by learning to adapt these values. This is equivalent to adding some adaptive prompts to the text input, which allows the model to learn the particularity of the input text and make the model more adaptable to the task requirements. This method does not change the original value of the feature extractor in the model, and retains the original excellent feature extraction capability.

[0129] It can be understood that when fine-tuning the model, adding the word vector to be trained before the input of the text encoder, that is, the word vector of the text, is equivalent to adding an adaptive prompt word, and then training and optimizing it by annotating the image to make its text feature extraction effect better. After fine-tuning, a classifier is added after the image feature extractor and the text feature extractor, and the overall classification and recognition effect of the model is optimized by reusing the annotated data for training, which solves the problem that the pre-trained model has weak feature extraction ability and overall classification effect for a small number of image samples or special image samples.

[0130] In S201, we collected and labeled a total of T+1 (T target samples and 1 non-target sample) training samples. During training, we only need to use these T target samples for learning. Through back propagation, the model updates the corresponding E P In actual training, Figure 5 As shown, the word embedding module 504 generates a word vector 405 after the word embedding operation of the T-th image text description (DESCRIBE_T) 505, and then concatenates it with the prefix vector group 501 to obtain the combined word vectors [V1][V2][V3]…[V m ][Embedding_T] 502, and finally input it into the text encoder 402 of the model to obtain the text feature vector 506 of the T-th type of text description: V t =Text encoding ([V1][V2][V3]…[V m ][Embedding_T]), where Text encoding represents the encoding operation of the text encoder 402, and Embedding_T represents the word vector of the text description of the T-th class label. Because our actual label has T classes, a vector matrix V can be generated according to the above rules N507, where N = 1, 2, … T, that is, V N is a matrix of size T×M, where M is the length of the text feature vector.

[0131] In addition, the image is input into the image encoder to obtain the feature vector I of the image t , the process can be expressed as: I t =Image encoding ([Image t ]), Image encoding Represents the encoding operation of the image encoder, Image t Represents the single image corresponding to the above t-th class label. After obtaining the complete image feature vector and text feature matrix, the cross entropy of each text feature vector in the image feature vector and text feature matrix is ​​calculated respectively to obtain the cosine similarity between the current image feature vector and each text feature vector. All image samples are input into the model in turn to obtain the corresponding similarity value, and the value of each vector is optimized by back propagation so that the similarity between the image feature vector and the text feature vector obtained by the image corresponding label is maximized, and the smaller the similarity with other text feature vectors, the better. The cosine similarity calculation is shown in the following formula (1):

[0132]

[0133] Among them, t represents the t-th class label, x represents the current input image, and τ is a hyperparameter used to prevent the parameter from exceeding the reasonable range. After multiple learning, E P The value will gradually stabilize, and the model has a good recognition ability for T-type image targets. Save the E with the best overall effect P The value is used as the weight of the model after fine-tuning, and the fine-tuning stage is completed.

[0134] S203, connecting a classifier to the fine-tuned CLIP pre-trained model, and training the image classifier based on the image-text annotation data to obtain a trained image classification model.

[0135] Although the model has been fine-tuned and has good classification and recognition capabilities, it is still necessary to further improve the overall accuracy of the model and increase the stability of the task. Figure 6 As shown in the figure, in order to meet this requirement, a classifier 601 is added after the two encoders of the model to combine the text features and image features to complete the classification. The type and specific implementation of the classifier 601 here can be set according to the actual classification effect and requirements, and are not limited. It should be noted that the number of output neurons in the last layer of the classifier 601 in the figure is the total number of classes to be identified, that is, T+1. By freezing the original model parameters and E PThe 501 parameters are then trained to further enhance the recognition ability of the model and meet the task requirements.

[0136] The training set collected in step S201 can be reused at this time. The difference from step S202 is that all image samples of the dataset (including the T+1th non-target class) and the CLASS_T attribute in the label need to be used at this time. According to the introduction in S202, when the target class of the text is T, the text feature vector matrix V N 507. During training, the parameters such as the text description of the T-type image are fixed, and the classifier weight parameters are initialized. t The image feature vector I is obtained by inputting the image feature encoder 401 of the model. t 602, the vector can be regarded as a 1×H feature matrix. Then the image feature vector I t 602 and a single vector V in the text feature vector matrix 507 t Fusion, get the fusion vector S t 603,S t It can be expressed by formula (2):

[0137] S t =a*I t +b*V t (2);

[0138] Since the values ​​and lengths of the image feature vector and the text feature vector are inconsistent, a direct concatenation method is used to achieve fusion, which can be represented by the symbol "+", and the image feature vector uses two hyperparameters a and b as the hyperparameters of the feature vector to facilitate subsequent training and optimization. The resulting fusion vector S t The size of 603 is 1×(M+H).

[0139] Similarly, by concatenating the image features with each vector of each text feature vector matrix in the above manner, we can obtain the fusion vector matrix S N =[S1,S2,…S T ]604, whose matrix size is T×(M+H), represents the combination of the input image features and each target text feature. To further link the feature relationships together, the row vectors of the matrix are weighted to obtain the relationship vector V rela , that is, the value of each row vector is multiplied by 1 / T and then added to obtain the final feature fusion vector 605, the relationship vector V rela It can be expressed by formula (3):

[0140]

[0141] The vector is input into the preset classifier 601, and the classification result V of T+1 is obtained through forward propagation. out , and then the classification probability 607 is obtained through the SOFTMAX function 606. This process can be expressed by formula (4):

[0142]

[0143] Among them, y t (t=1,2, … T+1) represents the network prediction result corresponding to the t-th class label, the T+1-th class is the label class that does not need to be processed, x represents the current input image, and f FC represents the forward process of the classifier, f softmax It represents the calculation process of the SOFTMAX function 606. The loss is calculated based on the final actual label CLASS_T and the obtained classification result, and the error is back-propagated to each layer of the classifier to optimize the weight parameters and hyperparameters a and b of the classifier. After multiple rounds of training, the classification effect of the model will gradually increase, and finally achieve a better effect. At this time, the training of the classifier is completed.

[0144] An image processing module can also be added after the classifier. This module is a preset module that uses the preset image processing method to process the corresponding image in the form of model prediction probability triggering. This module is at the end of the entire image processing model. Figure 7 As shown. This module includes multiple submodules: a first processing submodule 701, a second processing submodule 702 and a first processing submodule 703. Different submodules have different processing methods for images to meet different processing requirements. After the image is input to the model, if the image contains a T-type image target and meets the processing requirements, the image is passed to the corresponding post-processing module for processing. In this module, you can customize the image processing method and subsequent requirements. It can be a desensitization module to desensitize specific content and replace the original image; it can also be an image deletion module to directly delete the original file and other submodules to meet various different processing requirements.

[0145] The training method of the image classification model provided in this application uses the trained image-text model as the basic feature model, and uses the data labeled according to the actual target level. After fine-tuning the basic model with the labeled data, it is connected to the classifier for training, and finally a model with excellent text-image classification ability can be obtained. Using the trained model avoids the resource consumption of training from scratch, and the basic model used will also have good feature extraction ability and strong generalization ability; fine-tuning the model with a specific type of data will make it sensitive to this type of data, improve the model's feature extraction ability for this type of data, and solve the problem of the model's feature extraction ability; in order to further improve the classification ability and monitorability of the model, the classifier is added to control the final output result of the model, improve the classification ability and reduce the uncontrollable risk of the model, and connect the preset and diverse image processing methods behind the classifier. Various processing methods are triggered by the scores of the model classification results to meet the diversity of processing requirements. Through the above method, the targets of different levels in the image can be classified and the specified image processing can be performed, and the method can be applied to the classification and processing of various images, including but not limited to the image desensitization and encryption of intelligent driving, and the classification and processing of self-built data sets.

[0146] The present application embodiment provides an image classification device, Figure 8 A schematic diagram of the structure of an image classification device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the image classification device 800 includes:

[0147] The first acquisition module 801 is used to acquire an image to be classified and text description data corresponding to the image to be classified; the text description data includes description content of risk information for the image to be classified;

[0148] A first splicing processing module 802 is used to splice the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model;

[0149] A first feature extraction module 803, used for inputting the first word vector into a text feature extraction module of a trained image classification model to obtain a first text feature vector;

[0150] The first determination module 804 is configured to determine the risk category corresponding to the image to be classified based on the first text feature vector.

[0151] In some embodiments, the first determining module 804 includes:

[0152] A first feature extraction submodule, used for inputting the image to be classified into the image feature extraction module of the trained image classification model to obtain a first image feature vector;

[0153] A first feature fusion module, configured to fuse the first image feature vector and the first text feature vector based on a feature weight parameter to obtain a target feature vector; the feature weight parameter is obtained by tuning during the training of the classification module of the first image classification model;

[0154] The first determination submodule is used to determine the risk category of the image to be classified based on the target feature vector.

[0155] In some embodiments, the image classification device 800 further includes:

[0156] A second acquisition module is used to acquire an image sample data set; the image sample data set includes a plurality of image samples and text description samples corresponding to each image sample;

[0157] A first splicing processing module is used to splice word vectors corresponding to multiple text description samples and initial image classification prompt word vectors to obtain a first reference text matrix;

[0158] a second feature extraction module, configured to input the first reference text matrix into a text feature extraction module of the first image classification model to obtain a first reference text feature matrix; and input the plurality of image samples into an image feature extraction module of the first image classification model to obtain a first reference image feature matrix;

[0159] The first updating module is used to update the initial image classification prompt word vector based on the similarity between the first reference text feature matrix and the first reference image feature matrix to obtain the image classification prompt word vector.

[0160] In some embodiments, the first splicing processing module includes:

[0161] A first splicing processing submodule is used to splice the initial image classification prompt word vector and the word vector corresponding to each text description sample to obtain multiple first reference word vectors;

[0162] The second determination submodule is used to determine the first reference text matrix based on the multiple first reference word vectors.

[0163] In some embodiments, the image sample data set further includes category labels corresponding to the respective image samples; the first reference text feature matrix includes a plurality of reference text feature vectors, and the first reference image feature matrix includes a plurality of reference image feature vectors; the first update module includes:

[0164] A third determination submodule is used to determine the similarity between the candidate text feature vector in the first reference text feature matrix and each reference image feature vector; the candidate image feature vector is any one of the multiple reference text feature vectors;

[0165] A fourth determination submodule, used to determine the prediction category corresponding to each image sample according to the similarity;

[0166] The first training submodule is used to continuously adjust the initial image classification prompt word vector based on the predicted category and the corresponding category label to perform back-propagation training on the text feature extraction module of the first image classification model until the first convergence condition is met to obtain the image classification prompt word vector.

[0167] In some embodiments, the image classification device 800 further includes:

[0168] A second splicing processing module is used to splice the word vectors corresponding to the multiple text description samples and the image classification prompt word vector after obtaining the image classification prompt vector to obtain a second reference text matrix;

[0169] A third feature extraction module, used for inputting the second reference text matrix into the text feature extraction module of the first image classification model to obtain a second reference text feature matrix;

[0170] A second feature fusion module, configured to fuse the second reference text feature matrix and the first reference image feature matrix based on an initial feature weight parameter to obtain a reference feature matrix;

[0171] The second updating module is used to update the initial feature weight parameter based on the reference feature matrix to obtain the feature weight parameter.

[0172] In some embodiments, the second update module includes:

[0173] A fourth determination submodule, configured to determine, based on the reference feature matrix, the category probabilities corresponding to each of the plurality of image samples;

[0174] The second training submodule is used to continuously adjust the initial feature weight parameters based on the category probability and perform back propagation training on the classification module of the first image classification model until a second convergence condition is met to obtain the feature weight parameters.

[0175] In some embodiments, the initial feature weight parameters include initial image feature weights, initial text feature weights and normalization coefficients; and the second feature fusion module includes:

[0176] a fifth determination submodule, configured to determine a first feature matrix based on the initial image feature weight, the initial text feature weight, the second reference text feature matrix, and the first reference image feature matrix;

[0177] A sixth determination submodule is configured to determine the reference feature matrix based on the first feature matrix and the normalization coefficient.

[0178] In some embodiments, if the reference image sample includes multiple risk information with different risk levels, the risk level corresponding to the risk information with the highest risk level among the multiple risk information is determined as the category label corresponding to the reference image sample; the reference image sample is any one of the multiple image samples.

[0179] In some embodiments, the trained image classification model further includes multiple image processing modules, and different image processing modules have different image processing methods; the image classification device 800 further includes:

[0180] A second determination module is used to determine a target processing method for the image to be classified according to the risk category corresponding to the image to be classified and a preset correspondence between the risk category and the image processing method;

[0181] The input module is used to input the image to be classified into the image processing module corresponding to the target processing method to obtain the image processing result.

[0182] It should be noted that the description of the image classification device in the embodiment of the present application is similar to the description of the corresponding method embodiment above, and has similar beneficial effects as the method embodiment, so it is not repeated. For technical details not disclosed in this embodiment, please refer to the description of the method embodiment of the present application for understanding.

[0183] It should be noted that in the embodiment of the present application, if the above-mentioned image classification method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or the part that contributes to the relevant solution can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0184] Accordingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the image classification method provided in the above embodiment is implemented.

[0185] Fig. 9 A schematic diagram of the structure of an image classification device provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the image classification device 900 includes: a memory 901, a processor 902, a communication interface 903 and a communication bus 904. The memory 901 is used to store instructions for downloading executable vehicle upgrade files; the processor 902 is used to execute the executable image classification instructions stored in the memory 901 to implement the image classification method provided in the above embodiment.

[0186] The description of the above image classification device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the image classification device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0187] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises at least one ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0188] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0189] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0190] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0191] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, disks or optical disks.

[0192] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a product to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0193] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technical object familiar with the technical field can be easily thought of within the technical scope disclosed in the present application. Changes or substitutions should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. An image classification method, characterized in that: include: Acquire an image to be classified and text description data corresponding to the image to be classified; the text description data includes description content of risk information for the image to be classified; The word vector corresponding to the text description information and the image classification prompt word vector are concatenated to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model; Inputting the first word vector into a text feature extraction module of a trained image classification model to obtain a first text feature vector; Based on the first text feature vector, a risk category corresponding to the image to be classified is determined.

2. The method according to claim 1, characterized in that The step of determining the risk category corresponding to the image to be classified based on the first text feature vector includes: Inputting the image to be classified into the image feature extraction module of the trained image classification model to obtain a first image feature vector; The first image feature vector and the first text feature vector are fused based on a feature weight parameter to obtain a target feature vector; the feature weight parameter is obtained by tuning during the training of the classification module of the first image classification model; Based on the target feature vector, a risk category of the image to be classified is determined.

3. The method according to claim 1, characterized in that The method for determining the image classification prompt word vector includes: Acquire an image sample data set; the image sample data set includes a plurality of image samples and text description samples corresponding to each image sample; The word vectors corresponding to the multiple text description samples and the initial image classification prompt word vectors are concatenated to obtain a first reference text matrix; Inputting the first reference text matrix into a text feature extraction module of the first image classification model to obtain a first reference text feature matrix; and inputting the plurality of image samples into an image feature extraction module of the first image classification model to obtain a first reference image feature matrix; Based on the similarity between the first reference text feature matrix and the first reference image feature matrix, the initial image classification prompt word vector is updated to obtain the image classification prompt word vector.

4. The method according to claim 3, characterized in that The word vectors corresponding to the multiple text description samples and the initial image classification prompt word vectors are concatenated to obtain a first reference text matrix, including: The initial image classification prompt word vector and the word vector corresponding to each text description sample are concatenated to obtain a plurality of first reference word vectors; Based on the multiple first reference word vectors, determine the first reference text matrix.

5. The method according to claim 3 or 4, characterized in that: The image sample data set also includes the category labels corresponding to the respective image samples; the first reference text feature matrix includes a plurality of reference text feature vectors, and the first reference image feature matrix includes a plurality of reference image feature vectors; The updating of the initial image classification prompt word vector based on the similarity between the first reference text feature matrix and the first reference image feature matrix to obtain the image classification prompt word vector includes: Determine the similarity between the candidate text feature vector in the first reference text feature matrix and each reference image feature vector; the candidate image feature vector is any one of the multiple reference text feature vectors; Determining a prediction category corresponding to each image sample according to the similarity; Based on the predicted category and the corresponding category label, the initial image classification prompt word vector is continuously adjusted to perform back-propagation training on the text feature extraction module of the first image classification model until a first convergence condition is met to obtain the image classification prompt word vector.

6. The method according to claim 3 or 4, characterized in that: The method further comprises: After obtaining the image classification prompt vector, concatenating the word vectors corresponding to the multiple text description samples and the image classification prompt word vector to obtain a second reference text matrix; Inputting the second reference text matrix into the text feature extraction module of the first image classification model to obtain a second reference text feature matrix; Based on the initial feature weight parameter, the second reference text feature matrix and the first reference image feature matrix are fused to obtain a reference feature matrix; Based on the reference feature matrix, the initial feature weight parameters are updated to obtain feature weight parameters.

7. The method according to claim 6, characterized in that The updating of the initial feature weight parameter based on the reference feature matrix to obtain the feature weight parameter includes: Based on the reference feature matrix, determining the category probabilities corresponding to each of the plurality of image samples; Based on the category probability, the initial feature weight parameters are continuously adjusted and back-propagation training is performed on the classification module of the first image classification model until a second convergence condition is met to obtain the feature weight parameters.

8. The method according to claim 6, characterized in that The initial feature weight parameters include initial image feature weight, initial text feature weight and normalization coefficient; The step of fusing the second reference text feature matrix and the first reference image feature matrix based on the initial feature weight parameter to obtain a reference feature matrix includes: Determine a first feature matrix based on the initial image feature weight, the initial text feature weight, the second reference text feature matrix and the first reference image feature matrix; The reference feature matrix is ​​determined based on the first feature matrix and the normalization coefficients.

9. The method according to claim 3 or 4, characterized in that: The method further comprises: If the reference image sample includes multiple risk information with different risk levels, the risk level corresponding to the risk information with the highest risk level among the multiple risk information is determined as the category label corresponding to the reference image sample; the reference image sample is any one of the multiple image samples.

10. The method according to any one of claims 1 to 4, characterized in that: The trained image classification model further includes a plurality of image processing modules, and different image processing modules have different image processing modes; the method further includes: Determining a target processing method for the image to be classified according to the risk category corresponding to the image to be classified and a preset correspondence between the risk category and the image processing method; The image to be classified is input into an image processing module corresponding to the target processing method to obtain an image processing result.

11. An image classification device, characterized in that: include: A first acquisition module is used to acquire an image to be classified and text description data corresponding to the image to be classified; the text description data includes description content of risk information for the image to be classified; A first splicing processing module is used to splice the word vector corresponding to the text description information and the image classification prompt word vector to obtain a first word vector; the image classification prompt word vector is obtained after tuning during the training process of the text feature extraction module of the first image classification model; A first feature extraction module, used for inputting the first word vector into a text feature extraction module of a trained image classification model to obtain a first text feature vector; The first determination module is used to determine the risk category corresponding to the image to be classified based on the first text feature vector.

12. An image classification device, characterized in that: include: processor and memory; wherein, The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the method according to any one of claims 1 to 10 when running the computer program.

13. A computer-readable storage medium, characterized in that: The storage medium stores computer program code, and when the computer program code is executed by a computer, the method according to any one of claims 1 to 10 is executed.