Image recognition model training method, image recognition method and device

By performing significant object detection and adversarial classification network training on multiple training samples, the problem that traditional image recognition models cannot recognize similar images is solved, and efficient recognition of target images and similar images of structures is achieved, which improves recall effect and review efficiency.

CN114638304BActive Publication Date: 2025-05-23BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210270415.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-05-23
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

Traditional image recognition models can only recognize images that are exactly the same as the target image, and cannot recognize similar images. Especially in content review scenarios, it is difficult to identify images with similar structures to sensitive images.

Method used

By performing significant object detection on multiple training samples, the target feature data of each training sample is obtained, the picture structure information is characterized, and the adversarial classification network is used to train these feature data to obtain an image recognition model that can recognize the target image and its structure similar to the image.

Benefits of technology

It can not only identify the target image, but also identify images with similar structure to the target image, greatly improving the recall effect of the target image, reducing false detection of normal images, and reducing manual review costs and business violation risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638304B_ABST
    Figure CN114638304B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a training method for an image recognition model, an image recognition method and a device, and relates to the field of image recognition technology. The training method comprises: obtaining a training sample data set, wherein the training sample data set includes a plurality of training samples; performing significant target detection on each of the training samples, and obtaining target feature data corresponding to each of the training samples based on the obtained significant target detection results; the target feature data is used to characterize the picture structure information of the training sample; and the image recognition model is obtained by training according to the target feature data of the training sample. The image recognition model obtained by the training method can not only recognize the target image, but also recognize images with structures similar to the target image, which greatly improves the recall effect of the target image and can reduce the false detection of non-target images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a training method for an image recognition model, an image recognition method and an image recognition device. Background Art

[0002] With the development of computer technology, image recognition can be applied to a variety of scenarios, such as face recognition, vehicle recognition, medical recognition, content review, etc. In the content review scenario, when performing some target image recognition tasks, it is found that there are many similar images that are highly related to specific targets in the target image. For example, when reviewing some sensitive images, such as images that do not comply with national laws and regulations, industry norms, or social order and morality, and images with negative values, vulgarity, and indecentity, many malicious cos images are encountered, and these cos images are not allowed to be exposed. However, traditional image recognition models can only recognize images that are exactly the same as the target image, and cannot recognize similar images. Summary of the invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, an embodiment of the present invention provides a training method and device for an image recognition model, an image recognition method and device, an electronic device and a computer-readable storage medium.

[0004] In the first aspect of the implementation of the present invention, a method for training an image recognition model is first provided, comprising: obtaining a training sample data set, the training sample data set including a plurality of training samples; performing salient target detection on each of the training samples respectively, and based on the obtained salient target detection results, obtaining target feature data corresponding to each of the training samples, the target feature data being used to characterize the picture structure information of the training sample; and training the image recognition model according to the target feature data of the plurality of training samples.

[0005] Optionally, performing significant target detection on each of the training samples respectively and obtaining target feature data corresponding to each of the training samples based on the obtained significant target detection results include: for each training sample, performing significant target detection on the training sample using a pre-constructed significant target detection model, determining a significant area of ​​the training sample, and saving the significant area as a significant image; inputting the training sample into a pre-constructed feature extraction model to obtain an output result of the pre-constructed feature extraction model, and using the output result as the first feature data of the training sample; inputting the significant image into the pre-constructed feature extraction model to obtain an output result of the pre-constructed feature extraction model, and using the output result as the second feature data of the significant image; and fusing the first feature data and the second feature data to obtain the target feature data of the training sample.

[0006] Optionally, the training to obtain the image recognition model according to the target feature data of the multiple training samples comprises: training a preset adversarial classification network according to the target feature data of the multiple training samples to obtain the image recognition model; the adversarial classification network comprises an autoencoder and a classifier, and the autoencoder comprises an encoder and a decoder;

[0007] The process of training a preset adversarial classification network based on the target feature data of the multiple training samples includes: using a preset sample reconstruction loss function to train the target feature data of the multiple training samples, determining the first network parameters of the encoder and the first network parameters of the decoder, and obtaining hidden layer feature data obtained after the encoder encodes the target feature data of the multiple training samples based on its first network parameters; using a preset adversarial loss function to train the hidden layer feature data, determine the second network parameters of the classifier, determine the second network parameters of the encoder, and determine the second network parameters of the decoder.

[0008] Optionally, the multiple training samples include positive samples and negative samples; the training of the image recognition model according to the target feature data of the multiple training samples includes: when the proportion of negative samples in the training sample data set is greater than the proportion of positive samples, in the current iteration round of training the image recognition model, sampling the negative samples in the training sample data set to obtain multiple sampled negative samples, and the number of the sampled negative samples is the same as the number of the positive samples; performing training for the current iteration round according to the target feature data of the positive samples and the target feature data of the sampled negative samples; when training the image recognition model for the next iteration round, sampling the remaining negative samples in the training sample data set except the sampled negative samples to obtain multiple new sampled negative samples, and the number of the new sampled negative samples is the same as the number of the positive samples; performing training for the next iteration round according to the target feature data of the positive samples and the target feature data of the new sampled negative samples.

[0009] In a second aspect of the implementation of the present invention, an image recognition method is provided, comprising: acquiring an image to be recognized; performing significant target detection on the image to be recognized, and based on the obtained significant target detection result, acquiring target feature data of the image to be recognized, wherein the target feature data of the image to be recognized is used to characterize the picture structure information of the image to be recognized; and recognizing the image to be recognized based on the target feature data of the image to be recognized and a preset image recognition model, and determining the category of the image to be recognized.

[0010] Optionally, the preset image recognition model includes an autoencoder and a classifier; the autoencoder includes an encoder and a decoder;

[0011] According to the target feature data of the image to be identified and a preset image recognition model, the image to be identified is identified, and determining the category of the image to be identified includes: inputting the target feature data of the image to be identified into the autoencoder, and obtaining hidden layer feature data obtained after the encoder of the autoencoder encodes the target feature data; inputting the hidden layer feature data into the classifier to determine the category of the image to be identified.

[0012] Optionally, performing salient target detection on the image to be identified, and obtaining target feature data of the image to be identified based on the obtained salient target detection result includes: performing salient target detection on the image to be identified using a pre-constructed salient target detection model, determining a salient area of ​​the image to be identified, and saving the salient area as a salient image; inputting the image to be identified into a pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as third feature data of the image to be identified; inputting the salient image into the pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as fourth feature data of the salient image; and fusing the third feature data and the fourth feature data to obtain target feature data of the image to be identified.

[0013] In a third aspect of the implementation of the present invention, a training device for an image recognition model is provided, comprising: a sample acquisition module, used to acquire a training sample data set, wherein the training sample data set includes multiple training samples; a feature engineering module, used to perform significant target detection on each of the training samples respectively, and based on the obtained significant target detection results, acquire corresponding target feature data, wherein the target feature data is used to characterize the picture structure information of the training sample; and a model training module, used to train the image recognition model according to the target feature data of the multiple training samples.

[0014] In a fourth aspect of the implementation of the present invention, an image recognition device is provided, comprising: an image acquisition module, used to acquire an image to be recognized; a feature determination module, used to perform significant target detection on the image to be recognized, and based on the obtained significant target detection result, obtain target feature data of the image to be recognized, wherein the target feature data of the image to be recognized is used to characterize the picture structure information of the image to be recognized; an image recognition module, used to recognize the image to be recognized according to the target feature data of the image to be recognized and a preset image recognition model, and determine the category of the image to be recognized.

[0015] In the fifth aspect of the implementation of the present invention, an electronic device is provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement the image recognition model training method or image recognition method provided by an embodiment of the present invention when executing the program stored in the memory.

[0016] In a sixth aspect of the implementation of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the training method of the image recognition model or the image recognition method provided by an embodiment of the present invention is implemented.

[0017] The training method of the image recognition model provided by the embodiment of the present invention detects the salient areas of multiple training samples by performing salient target detection on them, and obtains corresponding target feature data based on the obtained salient target detection results, so that the target feature data can characterize the picture structure information of the training samples, and then the target feature data of the training samples are learned and trained to obtain an image recognition model, so that the image recognition model can not only recognize the target image, but also recognize images with a structure similar to the target image, thereby greatly improving the recall effect of the target image, wherein the images with a structure similar to the target image include images with a picture structure, picture composition and mutual positional relationship similar to the target image.

[0018] The image recognition method provided by the embodiment of the present invention can not only recognize the target image, but also recognize images with similar structures to the target image, which greatly improves the recall effect of the target image and reduces the false detection of normal images (i.e., non-target images). Among them, the target image can be an image containing a specific target, such as an image that does not comply with national laws and regulations, industry norms, or social order and morality, and an image with negative values, vulgarity, and indecentness. Exemplarily, this method can be applied to image content review scenarios, and can analyze and identify whether the image content has a specific target, thereby reducing the cost of manual review and the risk of business violations. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0020] Figure 1 A schematic diagram schematically shows the main process of a method for training an image recognition model according to an embodiment of the present invention;

[0021] Figure 2 A schematic diagram schematically shows a sub-process of a training method for an image recognition model according to an embodiment of the present invention;

[0022] Figure 3 A schematic diagram schematically shows the results of salient object detection in the training method of the image recognition model according to an embodiment of the present invention;

[0023] Figure 4 The structure diagram of the image recognition model according to the embodiment of the present invention is schematically shown;

[0024] Figure 5 The following is a schematic diagram showing a flow chart of an image recognition method according to an embodiment of the present invention;

[0025] Figure 6 A schematic diagram schematically shows the structure of a training device for an image recognition model according to an embodiment of the present invention;

[0026] Figure 7 A schematic diagram of the structure of an image recognition device according to an embodiment of the present invention is shown;

[0027] Figure 8 The structure diagram of an electronic device applicable to the training method of the image recognition model or the image recognition method according to the embodiment of the present invention is schematically shown. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0029] Figure 1 A schematic diagram schematically shows the main process of the training method of the image recognition model according to an embodiment of the present invention, as shown in FIG. Figure 1 As shown, the method includes:

[0030] Step 101: Acquire a training sample data set, where the training sample data set includes a plurality of training samples.

[0031] The multiple training samples include positive samples and negative samples, the positive sample is an image including the target object in the picture, and the negative sample is an image not including the target object in the picture. Among them, the target object can be flexibly selected according to the application scenario, and the present invention is not limited here. Exemplarily, the target object can be a sensitive object, and the positive sample can also be called a sensitive image. A sensitive object can be an object that does not comply with national laws and regulations, industry norms or social order and morality, and has negative values, vulgarity, and indecent. A negative sample is an image that does not include the target object in the picture, and a negative sample can also be called a normal image.

[0032] Step 102: Perform salient target detection on each of the training samples respectively, and obtain corresponding target feature data based on the obtained salient target detection results.

[0033] In order to highlight the target object in the positive sample and highlight the structural information of the positive sample screen, this embodiment performs salient target detection on the positive sample and the negative sample respectively to detect their salient regions or salient objects, thereby extracting the salient regions or salient objects in the positive sample and the negative sample, and constructing the target feature data of the positive sample and the negative sample based on the salient target detection results. The target feature data can characterize the screen structural information of the positive and negative samples. Among them, the screen structural information of the positive and negative samples is used to illustrate the structure of the screen, the components of the screen, and the relative position relationship.

[0034] In this step, the pre-built salient target detection model can be used to perform salient target detection on positive samples and negative samples respectively. The salient target detection model can be trained by a deep learning algorithm, such as a traditional convolutional neural network (CNN) or a fully convolutional neural network (FCN).

[0035] Step 103: training the image recognition model according to the target feature data of the plurality of training samples.

[0036] In this step, the traditional neural network model, such as VGGNets network and ResNets network, can be trained according to the target feature data of the positive sample and the target feature data of the negative sample to obtain the image recognition model. Since the target feature data of the training image recognition model can represent the picture structure information of the positive and negative samples, the image recognition model obtained by training the target feature data can not only accurately identify the target image, but also identify images with similar picture structure information to the target image, which greatly improves the recall effect of the target image.

[0037] The training method of the image recognition model provided by the embodiment of the present invention detects the salient areas of multiple training samples by performing salient target detection on them, and obtains corresponding target feature data based on the obtained salient target detection results, so that the target feature data can characterize the picture structure information of the training samples, and then the target feature data of the training samples are learned and trained to obtain an image recognition model, so that the image recognition model can not only recognize the target image, but also recognize images with a structure similar to the target image, thereby greatly improving the recall effect of the target image, wherein the images with a structure similar to the target image include images with a picture structure, picture composition and mutual positional relationship similar to the target image.

[0038] The process of performing salient target detection on each training sample and obtaining target feature data of each training sample based on the obtained salient target detection results is as follows: Figure 2 As shown, the process includes:

[0039] Step 201: for each training sample, perform salient object detection on the training sample using a pre-built salient object detection model, determine a salient region of the training sample, and save the salient region as a salient image.

[0040] The salient target detection model in this step can be obtained by training with a deep learning algorithm, for example, by training with a traditional convolutional neural network (CNN) or a fully convolutional neural network (FCN). The salient target detection model in this embodiment can segment the salient target in the training sample from the image background, and detect its skeleton, edge and other information, so as to determine the boundary of the salient target, and the area surrounded by the boundary of the salient target is used as the salient area, and the salient area is saved as an image, and the image is used as the salient image corresponding to the training sample. Figure 3 As shown, the salient object detection model in this embodiment can segment the salient objects of the training sample, such as tanks and people, from the image background, detect their skeleton and edge information, and determine the boundaries of each salient object. Then, the area surrounded by the boundaries of the salient objects is used as a salient area, and the salient area is saved as an image, which is used as a salient image corresponding to the training sample.

[0041] Step 202: Input the training sample into a pre-constructed feature extraction model, obtain an output result of the pre-constructed feature extraction model, and use the output result as the first feature data of the training sample.

[0042] The feature extraction model in this step can be obtained by training a convolutional neural network (CNN). The first feature data of the training sample can represent the picture structure information of the original image of the training sample, and the picture structure information can be used to illustrate the structure of the picture, the components of the picture, and the relative position relationship.

[0043] Step 203: Input the saliency image into the pre-built feature extraction model, obtain the output result of the pre-built feature extraction model, and use the output result as the second feature data of the saliency image. The second feature data of the saliency image corresponding to the training sample can represent the structural information of the salient object in the training sample.

[0044] Step 204: Fusing the first feature data and the second feature data to obtain target feature data of the training sample.

[0045] In this embodiment, the target feature data of the training sample can be determined by fusing the first feature data of the training sample and the second feature data of the saliency image corresponding to the training sample, for example, the first feature data and the second feature data are concatenated to obtain the target feature data of the training sample. The first feature data and the second feature data can also be fused by other feature fusion algorithms, such as multiplying, calculating the Cartesian product, or dividing the first feature data and the second feature data to determine the target feature data.

[0046] This step determines the target feature data of the training sample through the first feature data of the training sample and the second feature data of the salient image corresponding to the training sample, so that the target feature data can represent both the picture structure information of the original image of the training sample and the structural information of the salient target in the training sample, thereby enabling the image recognition model trained by the target feature data to not only recognize specific images, but also recognize images with structures similar to the specific images.

[0047] In an optional embodiment, training the image recognition model according to the target feature data of the plurality of training samples includes:

[0048] According to the target feature data of the multiple training samples, a preset adversarial classification network is trained to obtain the image recognition model; wherein the network parameters of the preset adversarial classification network are updated by adversarial learning.

[0049] This embodiment trains the network parameters of a preset adversarial classification network by adversarial learning. In the process of training the preset adversarial classification network, adversarial data is generated for the current network through adversarial learning, and then the adversarial data is learned by updating the network parameters of the current network. This cycle is repeated until the model converges or other stopping conditions are met (for example, the maximum number of iterations is reached), thereby obtaining an image recognition model that can not only recognize images with similar structures to sensitive images, but also reduce false detection of normal images.

[0050] Optional, such as Figure 4As shown, the preset adversarial classification network includes an autoencoder and a classifier. Among them, the autoencoder is an unsupervised neural network model, which can learn the implicit features of the input data, which is called coding, and at the same time, the original input data can be reconstructed with the learned new features, which is called decoding. The network structure of the autoencoder is divided into an encoder E and a decoder G, wherein the input of the encoder is called the input layer, the output is called the hidden layer, and the input of the decoder is called the hidden layer, and the output is called the reconstruction layer. The classifier can be a classifier of a multi-layer perceptron structure, wherein the multi-layer perceptron (MLP, Multilayer Perceptron) is a feedforward artificial neural network model that maps multiple input data sets to a single output data set. The autoencoder will input the feature data of the hidden layer into the classifier to learn a specific distribution P(y).

[0051] In this embodiment, the network parameters of the autoencoder and the classifier are updated by adversarial learning, that is, the target feature data of the positive samples and the negative samples are passed through the encoder of the autoencoder to obtain the hidden layer feature data, and then the hidden layer feature data is input into the classifier for adversarial learning.

[0052] In an optional embodiment, the process of updating the network parameters of the autoencoder and the classifier through adversarial learning includes:

[0053] Using a preset sample reconstruction loss function, training the target feature data of the multiple training samples, determining a first network parameter of the encoder and a first network parameter of the decoder, and obtaining hidden layer feature data obtained after the encoder encodes the target feature data of the multiple training samples based on its first network parameter;

[0054] The hidden layer feature data is trained using a preset adversarial loss function to determine the second network parameters of the classifier, the second network parameters of the encoder, and the second network parameters of the decoder.

[0055] The above training update process includes two stages:

[0056] Sample reconstruction stage: Update the network parameters of the encoder and decoder to minimize the preset sample reconstruction loss function. The network parameters of the encoder E and decoder G can be updated by the gradient descent method, and the preset sample reconstruction loss function can adopt the mean square error loss function MSE(X, G(z)).

[0057] Distribution constraint stage: Update the network parameters of the classifier D and the encoder E by minimizing the preset adversarial loss function to improve the ability of the adversarial classification network. The preset adversarial loss function can be a cross entropy loss function, as shown in the following formula:

[0058]

[0059] Among them, loss(o, t) represents the value of the preset adversarial loss function, n represents the total number of positive samples and negative samples, t represents the sample label, the label of the positive sample is 0, the label of the negative sample is 1, and o represents the output of the classifier.

[0060] The training method of the image recognition model of an embodiment of the present invention adopts the network structure of an autoencoder and the parameter updating method of adversarial learning when learning the target feature data of the training samples. It can accurately identify images similar to specific image structures while reducing false detection of normal images.

[0061] In an optional embodiment, the training method of the image recognition model provided by the embodiment of the present invention further includes the following steps:

[0062] When the proportion of negative samples in the training sample data set is greater than the proportion of positive samples, in the current iteration round of training the image recognition model, the negative samples in the training sample data set are sampled to obtain a plurality of sampled negative samples, and the number of the sampled negative samples is the same as the number of the positive samples;

[0063] Performing training for the current iteration round according to the target feature data of the positive sample and the target feature data of the sampled negative sample;

[0064] In the next iteration round of training the preset adversarial classification network, sampling the remaining negative samples in the training sample data set except the sampled negative samples to obtain a plurality of new sampled negative samples, wherein the number of the new sampled negative samples is the same as the number of the positive samples;

[0065] The next iteration round of training is performed according to the target feature data of the positive sample and the target feature data of the newly sampled negative sample.

[0066] In actual application scenarios, compared with a large number of normal images, the frequency of occurrence of sensitive images and images similar to sensitive images is low. Therefore, the number of collected positive samples is less than or even much less than the number of negative samples, that is, the proportion of negative samples in the training sample data set is greater than or even much greater than the proportion of positive samples. For example, there are 3,000 positive samples and 100,000 negative samples in the training sample set, and the proportion of positive samples is much smaller than the proportion of negative samples. In order to avoid overfitting and improve the accuracy of image recognition, when iteratively training the image recognition model, the embodiment of the present invention needs to sample negative samples (for example, uniform sampling without replacement) to obtain multiple sampled negative samples with the same number as the positive samples, and then train the target feature data of the sampled negative samples and the positive samples until the model converges or reaches other stop conditions (for example, reaching the maximum number of iterations). For example, in the first iteration round of training the image recognition model, 3,000 sampled negative samples are uniformly sampled from 100,000 negative samples, and the current iteration round of training is performed based on the target feature data of the 3,000 sampled negative samples and the 3,000 positive samples. In the second iteration of training the adversarial classification network, 3,000 new negative samples are uniformly sampled from the remaining 97,000 negative samples, and the current iteration round of training is performed based on the target feature data of the 3,000 new negative samples and the 3,000 positive samples. The above iterative training process is repeated until the model converges or reaches other stopping conditions (such as reaching the maximum number of iterations), thereby obtaining an image recognition model.

[0067] Figure 5 The following is a schematic diagram showing a flow chart of an image recognition method according to an embodiment of the present invention. Figure 5 As shown, the method includes:

[0068] Step 501: Obtain an image to be recognized.

[0069] Step 502: Perform salient target detection on the image to be identified, and based on the obtained salient target detection result, obtain target feature data of the image to be identified, wherein the target feature data of the image to be identified is used to characterize the picture structure information of the image to be identified.

[0070] The target feature data can be used to perform salient target detection on the image to be identified through a pre-built salient target detection model. The salient target detection model can be trained by a deep learning algorithm, such as a traditional convolutional neural network (CNN) or a fully convolutional neural network (FCN).

[0071] Step 503: Identify the image to be identified based on the target feature data of the image to be identified and a preset image recognition model, and determine the category of the image to be identified.

[0072] Among them, the preset image recognition model is obtained according to the training method of the image recognition model in the above embodiment, and the image recognition model can accurately identify whether the image to be recognized is a target image or an image with a structure similar to the target image.

[0073] The image recognition method provided by the embodiment of the present invention can not only recognize the target image, but also recognize images with structures similar to the target image, which greatly improves the recall effect of the target image and reduces the false detection of normal images. The method can be applied to image content review scenarios, and can analyze whether the image to be recognized is an image containing a sensitive object or whether the image to be recognized is an image with a similar structure to an image containing a sensitive object, thereby reducing the cost of manual review and the risk of business violations.

[0074] In an optional embodiment, the process of performing salient target detection on the image to be identified and acquiring target feature data of the image to be identified based on the obtained salient target detection result includes:

[0075] Performing salient object detection on the image to be identified using a pre-built salient object detection model, determining a salient region of the image to be identified, and saving the salient region as a salient image;

[0076] Inputting the image to be identified into a pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as third feature data of the image to be identified;

[0077] Inputting the saliency image into the pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as fourth feature data of the saliency image;

[0078] The third feature data and the fourth feature data are integrated to obtain target feature data of the image to be identified.

[0079] Among them, the salient target detection model can be obtained by training a deep learning algorithm, for example, by training a traditional convolutional neural network (CNN) or by training a fully convolutional neural network (FCN). The salient target detection model in this embodiment can segment the salient target in the image to be identified from the image background, and detect its skeleton, edge and other information, so as to determine the boundary of the salient target, and the area surrounded by the boundary of the salient target is used as the salient area, and the salient area is saved as an image, and the image is used as the salient image corresponding to the image to be identified. .

[0080] The feature extraction model can be obtained by training a convolutional neural network (CNN). The third feature data extracted by the feature extraction model can represent the picture structure information of the original image to be identified, and the fourth feature data extracted by the feature extraction model can represent the structural information of the salient target in the image to be identified.

[0081] After the third feature data and the fourth feature data are extracted, the third feature data and the fourth feature data are fused to determine the target feature data of the image to be identified. For example, the third feature data and the fourth feature data may be directly concatenated to obtain the target feature data of the image to be identified, or the third feature data and the fourth feature data may be fused by other feature fusion algorithms, such as multiplying or calculating the Cartesian product or dividing the third feature data and the fourth feature data to determine the target feature data of the image to be identified.

[0082] In an optional embodiment, the image recognition model includes an autoencoder and a classifier; the autoencoder includes an encoder and a decoder. The network parameters of the image recognition model are determined by adversarial learning. The structure of the image recognition model is as follows: Figure 4 As shown in Figure 2, the process of building and training the image recognition model is as follows: Figure 4 The embodiment shown in the figure is not described in detail in the present invention. According to the image recognition model, the process of identifying the image to be identified and determining the category of the image to be identified may include:

[0083] Inputting the target feature data of the image to be recognized into the autoencoder, and obtaining hidden layer feature data obtained after the encoder of the autoencoder encodes the target feature data;

[0084] The hidden layer feature data is input into the classifier to determine the category of the image to be identified.

[0085] The image recognition model in this embodiment adopts the network structure of an autoencoder and a parameter update method of adversarial learning, which can accurately identify images with similar structures to the target image while reducing false detection of normal images.

[0086] Figure 6 The structure diagram of the training device 600 of the image recognition model according to the embodiment of the present invention is schematically shown. Figure 6 As shown, the training device 600 includes:

[0087] The sample acquisition module 601 is used to acquire a training sample data set, where the training sample data set includes a plurality of training samples;

[0088] A feature engineering module 602 is used to perform salient target detection on each of the training samples, and obtain corresponding target feature data based on the obtained salient target detection results, wherein the target feature data is used to characterize the picture structure information of the training sample;

[0089] The model training module 603 is used to train the image recognition model according to the target feature data of the multiple training samples.

[0090] Optionally, the feature engineering module is also used to: for each training sample, use a pre-constructed significant target detection model to perform significant target detection on the training sample, determine the significant area of ​​the training sample, and save the significant area as a significant image; input the training sample into a pre-constructed feature extraction model to obtain the output result of the pre-constructed feature extraction model, and use the output result as the first feature data of the training sample; input the significant image into the pre-constructed feature extraction model to obtain the output result of the pre-constructed feature extraction model, and use the output result as the second feature data of the significant image; and fuse the first feature data and the second feature data to obtain the target feature data of the training sample.

[0091] Optionally, the model training module is further used to: train a preset adversarial classification network according to the target feature data of the multiple training samples to obtain the image recognition model; the adversarial classification network includes an autoencoder and a classifier, and the autoencoder includes an encoder and a decoder;

[0092] The model training module is also used to: use a preset sample reconstruction loss function to train the target feature data of the multiple training samples, determine the first network parameters of the encoder and the first network parameters of the decoder, and obtain the hidden layer feature data obtained after the encoder encodes the target feature data of the multiple training samples based on its first network parameters; use a preset adversarial loss function to train the hidden layer feature data, determine the second network parameters of the classifier, determine the second network parameters of the encoder, and determine the second network parameters of the decoder.

[0093] Optionally, the multiple training samples include positive samples and negative samples; the model training module is also used to: when the proportion of negative samples in the training sample data set is greater than the proportion of positive samples, in the current iteration round of training the image recognition model, sample the negative samples in the training sample data set to obtain multiple sampled negative samples, and the number of the sampled negative samples is the same as the number of the positive samples; perform training for the current iteration round based on the target feature data of the positive samples and the target feature data of the sampled negative samples; when training the next iteration round of the image recognition model, sample the remaining negative samples in the training sample data set except the sampled negative samples to obtain multiple new sampled negative samples, and the number of the new sampled negative samples is the same as the number of the positive samples; perform training for the next iteration round based on the target feature data of the positive samples and the target feature data of the new sampled negative samples.

[0094] The training device of the image recognition model provided in the embodiment of the present invention detects the salient areas of a plurality of training samples by performing salient target detection on them, and obtains the corresponding target feature data based on the obtained salient target detection results, so that the target feature data can characterize the picture structure information of the training samples, and then the target feature data of the training samples are trained to obtain the image recognition model, so that the image recognition model can not only recognize the target image, but also recognize images with structures similar to the target image, which greatly improves the recall effect of the target image, wherein the images with structures similar to the target image include images with picture structures, picture compositions, and mutual positional relationships similar to the target image. The above-mentioned device can execute the method provided in the embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, please refer to the method provided in the embodiment of the present invention.

[0095] Figure 7 The structure diagram of the image recognition device 700 according to the embodiment of the present invention is schematically shown. Figure 7 As shown, the image recognition device 700 includes:

[0096] An image acquisition module 701 is used to acquire an image to be recognized;

[0097] The feature determination module 702 is used to perform salient target detection on the image to be identified, and obtain target feature data of the image to be identified based on the obtained salient target detection result, wherein the target feature data of the image to be identified is used to characterize the picture structure information of the image to be identified;

[0098] The image recognition module 703 is used to recognize the image to be recognized based on the target feature data of the image to be recognized and a preset image recognition model, and determine the category of the image to be recognized.

[0099] Optionally, the preset image recognition model includes an autoencoder and a classifier; the autoencoder includes an encoder; the image recognition module is also used to: input the target feature data of the image to be recognized into the autoencoder, and obtain the hidden layer feature data obtained after the encoder of the autoencoder encodes the target feature data; input the hidden layer feature data into the classifier to determine the category of the image to be recognized.

[0100] Optionally, the feature determination module is also used to: perform significant target detection on the image to be identified using a pre-constructed significant target detection model, determine the significant area of ​​the image to be identified, and save the significant area as a significant image; input the image to be identified into a pre-constructed feature extraction model, obtain the output result of the pre-constructed feature extraction model, and use the output result as the third feature data of the image to be identified; input the significant image into the pre-constructed feature extraction model, obtain the output result of the pre-constructed feature extraction model, and use the output result as the fourth feature data of the significant image; and fuse the third feature data and the fourth feature data to obtain the target feature data of the image to be identified.

[0101] The image recognition device provided by the embodiment of the present invention can not only recognize the target image, but also recognize images with similar structures to the target image, which greatly improves the recall effect of the target image and reduces the false detection of normal images (i.e., non-target images). This method can be applied to image content review scenarios, and can analyze and identify whether the image content contains sensitive content, thereby reducing the cost of manual review and the risk of business violations. The above-mentioned device can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided by the embodiment of the present invention.

[0102] Figure 8 The structure diagram of the electronic device according to the embodiment of the present invention is schematically shown. Figure 8As shown, the electronic device includes a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804.

[0103] Memory 803, used for storing computer programs;

[0104] The processor 801 is used to implement the image recognition model training method described in any of the above embodiments or the image recognition method described in any of the above embodiments when executing the program stored in the memory 803.

[0105] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0106] The communication interface is used for communication between the above terminal and other devices.

[0107] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0108] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0109] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the training method of the image recognition model described in any of the above embodiments or the image recognition method described in any of the above embodiments.

[0110] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When the computer is run on a computer, the computer executes the training method of the image recognition model described in any of the above embodiments or the image recognition method described in any of the above embodiments.

[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0112] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements limited by the sentence "comprise a..." do not exclude the existence of other identical elements in the process, method, article or equipment including the elements.

[0113] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A training method for an image recognition model, It is characterized in that include: Acquire a training sample data set, wherein the training sample data set includes a plurality of training samples; Performing salient target detection on each of the training samples respectively, and acquiring target feature data of each of the training samples based on the obtained salient target detection results, wherein the target feature data is used to characterize the picture structure information of the training sample; The image recognition model is obtained by training according to the target feature data of the plurality of training samples; The performing salient target detection on each of the training samples respectively, and acquiring target feature data of each of the training samples based on the obtained salient target detection results comprises: For each training sample, using a pre-built salient object detection model to perform salient object detection on the training sample, determine a salient region of the training sample, and save the salient region as a salient image; Inputting the training sample into a pre-built feature extraction model, obtaining an output result of the pre-built feature extraction model, and using the output result as first feature data of the training sample; Inputting the saliency image into the pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as second feature data of the saliency image; The first feature data and the second feature data are fused to obtain target feature data of the training sample.

2. The method according to claim 1, It is characterized in that The training of the image recognition model according to the target feature data of the plurality of training samples comprises: According to the target feature data of the plurality of training samples, a preset adversarial classification network is trained to obtain the image recognition model; the adversarial classification network includes an autoencoder and a classifier, and the autoencoder includes an encoder and a decoder; The process of training a preset adversarial classification network according to the target feature data of the plurality of training samples includes: Using a preset sample reconstruction loss function, training the target feature data of the multiple training samples, determining a first network parameter of the encoder and a first network parameter of the decoder, and obtaining hidden layer feature data obtained after the encoder encodes the target feature data of the multiple training samples based on its first network parameter; The hidden layer feature data is trained using a preset adversarial loss function to determine the second network parameters of the classifier, the second network parameters of the encoder, and the second network parameters of the decoder.

3. The method according to claim 1, It is characterized in that The multiple training samples include positive samples and negative samples; The training of the image recognition model according to the target feature data of the plurality of training samples comprises: When the proportion of negative samples in the training sample data set is greater than the proportion of positive samples, in the current iteration round of training the image recognition model, the negative samples in the training sample data set are sampled to obtain a plurality of sampled negative samples, and the number of the sampled negative samples is the same as the number of the positive samples; Performing training for the current iteration round according to the target feature data of the positive sample and the target feature data of the sampled negative sample; In the next iteration round of training the image recognition model, sampling the remaining negative samples in the training sample data set except the sampled negative samples to obtain a plurality of new sampled negative samples, wherein the number of the new sampled negative samples is the same as the number of the positive samples; The next iteration round of training is performed according to the target feature data of the positive sample and the target feature data of the newly sampled negative sample.

4. An image recognition method, It is characterized in that include: Obtain an image to be recognized; Performing salient target detection on the image to be identified, and acquiring target feature data of the image to be identified based on the obtained salient target detection result, wherein the target feature data of the image to be identified is used to characterize the picture structure information of the image to be identified; Identify the image to be identified based on the target feature data of the image to be identified and a preset image recognition model, and determine the category of the image to be identified; The performing salient target detection on the image to be identified and acquiring target feature data of the image to be identified based on the obtained salient target detection result includes: Using a pre-built salient object detection model to perform salient object detection on the image to be identified, determining a salient region of the image to be identified, and saving the salient region as a salient image; Inputting the image to be identified into a pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as third feature data of the image to be identified; Inputting the saliency image into the pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as fourth feature data of the saliency image; The third feature data and the fourth feature data are integrated to obtain target feature data of the image to be identified.

5. The method according to claim 4, It is characterized in that The preset image recognition model includes an autoencoder and a classifier; the autoencoder includes an encoder and a decoder; According to the target feature data of the image to be identified and a preset image recognition model, the image to be identified is identified, and determining the category of the image to be identified includes: Inputting the target feature data of the image to be recognized into the autoencoder, and obtaining hidden layer feature data obtained after the encoder of the autoencoder encodes the target feature data; The hidden layer feature data is input into the classifier to determine the category of the image to be identified.

6. A training device for an image recognition model, It is characterized in that include: A sample acquisition module is used to acquire a training sample data set, wherein the training sample data set includes a plurality of training samples; A feature engineering module, used to perform salient target detection on each of the training samples, and based on the obtained salient target detection results, obtain target feature data corresponding to each of the training samples, wherein the target feature data is used to characterize the picture structure information of the training sample; A model training module, used for training the image recognition model according to the target feature data of the plurality of training samples; The feature engineering module is also used to: for each training sample, use a pre-built salient target detection model to perform salient target detection on the training sample, determine a salient region of the training sample, and save the salient region as a salient image; Inputting the training sample into a pre-built feature extraction model, obtaining an output result of the pre-built feature extraction model, and using the output result as first feature data of the training sample; The saliency image is input into the pre-constructed feature extraction model to obtain an output result of the pre-constructed feature extraction model, and the output result is used as the second feature data of the saliency image; the first feature data and the second feature data are fused to obtain the target feature data of the training sample.

7. An image recognition device, It is characterized in that include: An image acquisition module, used for acquiring an image to be identified; A feature determination module, used to perform significant target detection on the image to be identified, and based on the significant target detection result obtained, obtain target feature data of the image to be identified, wherein the target feature data of the image to be identified is used to characterize the picture structure information of the image to be identified; An image recognition module, used to recognize the image to be recognized and determine the category of the image to be recognized based on the target feature data of the image to be recognized and a preset image recognition model; The feature determination module is also used to: perform salient target detection on the image to be identified using a pre-built salient target detection model, determine a salient region of the image to be identified, and save the salient region as a salient image; Inputting the image to be identified into a pre-constructed feature extraction model, obtaining an output result of the pre-constructed feature extraction model, and using the output result as third feature data of the image to be identified; The saliency image is input into the pre-constructed feature extraction model to obtain an output result of the pre-constructed feature extraction model, and the output result is used as the fourth feature data of the saliency image; and the third feature data and the fourth feature data are fused to obtain target feature data of the image to be identified.

8. An electronic device, It is characterized in that It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1-3 or 4-5 when executing a program stored in a memory.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method described in any one of claims 1-3 or 4-5 is implemented.

Citation Information

Patent Citations

  • Sensitive image recognition model training method, training device and electronic equipment

    CN113936195A

  • Content check model training method and apparatus, video content check method and apparatus, computer device, and storage medium

    WO2021082589A1