Image recognition method and device, processing equipment and storage medium
By using a self-step learning algorithm to train the AUC model in image recognition, and combining the stochastic gradient parameter update strategy, the problems of unreliability and inaccuracy of image recognition in the prior art are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411845500.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems of unreliability and inaccuracy in the process of identifying pictures, especially when dealing with unbalanced data sets and noise samples.
The self-step learning algorithm is used to train the area under the curve (AUC) model. By adding a full connection layer to the neural network structure, and combining stochastic gradient parameters to update the image sample weight, AUC model parameters and step size parameters, the model identification accuracy and robustness of the model are improved.
Through self-step learning algorithm and stochastic gradient parameter update strategy, the adverse effects of noise samples on model training are reduced, and the reliability and accuracy of image recognition are improved.
Smart Images

Figure CN120125870A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to and are not limited to the field of information processing, and particularly relate to an image recognition method, device, processing equipment and storage medium. Background Art
[0002] Pictures are important carriers of information, and the compliance of pictures is crucial. The spread of illegal pictures will bring serious harm to society and individuals. In related technologies, illegal pictures can be identified based on technical means. For example, illegal pictures can be identified based on intelligent models, and after the illegal pictures are identified, they will be processed in time to block the spread of illegal pictures. However, in the process of picture recognition, there are still problems such as unreliable and inaccurate picture recognition. Summary of the Invention
[0003] In view of this, the present invention discloses an image recognition method, device, processing equipment and storage medium.
[0004] According to a first aspect of the embodiments of the present application, an image recognition method is provided. The method includes: obtaining a to-be-trained image sample, where the to-be-trained image sample includes: a first type of image and a second type of image, the first type of image is an image marked as not containing illegal content, and the second type of image is an image marked as containing illegal content; adding a fully connected layer to a selected neural network structure to construct a to-be-trained Area Under the Curve (AUC) model, where the trained AUC model is used to identify whether a to-be-identified image is a first type of image or a second type of image; determining an objective function for training the to-be-trained AUC model based on a self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; determining a training strategy for training the to-be-trained AUC model based on the objective function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameter using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter; training the to-be-trained AUC model based on the training strategy and the to-be-trained image sample until the objective function converges to obtain a trained AUC model.
[0005] In some embodiments, the objective function includes at least one of the following: a first term, which is a regularization term for reducing model overfitting; a second term for assisting the model in identifying noise samples based on reference image samples; a third term for weighting the AUC loss using image sample weights; a fourth term, which is a self-paced regularization term; a fifth term, which is a binary classification balance regularization term for balancing the weights of positive and negative samples; a sixth term, which is a consistency regularization term for reducing the impact of augmentation operations on the self-paced learning algorithm, and the augmentation operations include at least one of the following: image scaling; image rotation; image shearing; image flipping.
[0006] In some embodiments, the objective function includes the second term, and the method further includes: screening out the reference image samples from the to-be-trained image samples; wherein, the reference image samples include the first type of image samples and the second type of image samples for assisting the model in identifying noise samples.
[0007] In some embodiments, updating the image sample weights using the stochastic gradient parameter includes: randomly selecting a first image sample from the to-be-trained image samples; randomly selecting a second image sample from the reference image samples; determining the partial derivative of the objective function with respect to the image sample weights based on the first image sample and the second image sample to obtain a first stochastic gradient parameter; updating the image sample weights based on the first stochastic gradient parameter.
[0008] In some embodiments, updating the image sample weights based on the first stochastic gradient parameter includes: updating the image sample weights based on the first stochastic gradient parameter until the image sample weights converge.
[0009] In some embodiments, updating the AUC model parameters using the stochastic gradient parameter includes: randomly selecting a third image sample from the to-be-trained image samples; randomly selecting a fourth image sample from the reference image samples; determining the partial derivative of the objective function with respect to the AUC model parameters based on the third image sample and the fourth image sample to obtain a second stochastic gradient parameter; updating the AUC model parameters based on the second stochastic gradient parameter.
[0010] In some embodiments, updating the AUC model parameters based on the second stochastic gradient parameter includes: updating the AUC model parameters based on the second stochastic gradient parameter until the AUC model parameters converge.
[0011] According to a second aspect of the embodiments of the present application, there is provided an image recognition device, and the image recognition device includes:
[0012] An acquisition module, configured to: acquire image samples to be trained, where the image samples to be trained include: first-class images and second-class images, the first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content;
[0013] A construction module, configured to: add a fully-connected layer to a selected neural network structure to construct an area under the curve (AUC) model to be trained, where the trained AUC model is used to identify whether a to-be-identified image is a first-class image or a second-class image;
[0014] A determination module, configured to: determine an objective function for training the AUC model to be trained based on a self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or image sample weights, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; determine a training strategy for training the AUC model to be trained based on the objective function, where the training strategy includes at least one of the following: updating image sample weights using a stochastic gradient parameter; updating AUC model parameters using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter;
[0015] A training module, configured to: train the AUC model to be trained based on the training strategy and the image samples to be trained until the objective function converges, to obtain a trained AUC model.
[0016] In some embodiments, the image processing apparatus further includes a processing module, and the processing module is further configured to: screen out the reference image samples from the image samples to be trained; where the reference image samples include the first-class image samples and the second-class image samples for assisting the model in identifying noise samples.
[0017] In some embodiments, the processing module is further configured to: randomly select a first image sample from the image samples to be trained; randomly select a second image sample from the reference image samples; determine a partial derivative of the objective function with respect to the image sample weights based on the first image sample and the second image sample to obtain a first stochastic gradient parameter; update the image sample weights based on the first stochastic gradient parameter.
[0018] In some embodiments, the processing module is further configured to: update the image sample weights based on the first stochastic gradient parameter until the image sample weights converge.
[0019] In some embodiments, the processing module is further configured to: randomly select a third image sample from the to-be-trained image samples; randomly select a fourth image sample from the reference image samples; determine the partial derivative of the objective function with respect to the AUC model parameters based on the third image sample and the fourth image sample to obtain a second stochastic gradient parameter; and update the AUC model parameters based on the second stochastic gradient parameter.
[0020] In some embodiments, the processing module is further configured to: update the AUC model parameters based on the second stochastic gradient parameter until the AUC model parameters converge.
[0021] According to a third aspect of the embodiments of the present application, there is provided a processing device, which is used to execute the method described in the first aspect or the second aspect.
[0022] According to a fourth aspect of the embodiments of the present application, there is provided a computer storage medium storing an executable program, which when executed by a processor, implements the method described in any one of the embodiments of the present application.
[0023] According to a fifth aspect of the embodiments of the present application, there is provided a computer program product including a computer program or instruction, which when executed by a processor, implements the method described in any one of the embodiments of the present application.
[0024] In the embodiments of the application, an objective function for training the to-be-trained AUC model is determined based on the self-paced learning algorithm. Since the objective function is determined based on the self-paced learning algorithm and the objective function is determined based on the step size parameter and / or the image sample weight, the adverse effects brought by noise samples can be reduced. The to-be-trained AUC model is trained based on the training strategy and the to-be-trained image samples to obtain a trained AUC model. Since the training strategy includes at least one of the following: updating the image sample weight using the stochastic gradient parameter; updating the AUC model parameters using the stochastic gradient parameter; updating the step size parameter using the stochastic gradient parameter; during the process based on the training strategy, the image sample weight, the AUC model parameters, and / or the step size parameter will be continuously updated, so that the result of the trained AUC model for identifying images will be more reliable and accurate. Description of the Drawings
[0025] Figure 1 It is a schematic flowchart of an image recognition method shown according to the first embodiment;
[0026] Figure 2 It is a schematic flowchart of an image recognition method shown according to the second embodiment;
[0027] Figure 3Schematic flowchart of an image recognition method shown according to the third embodiment;
[0028] Figure 4 Schematic flowchart of an image recognition method shown according to the fourth embodiment;
[0029] Figure 5 Schematic diagram of an image recognition device shown according to the fifth embodiment. Detailed implementation manners
[0030] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0031] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0032] In the following description, the terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0033] In the following description, reference is made to "greater than" and "less than". It should be noted that in the present application, "greater than" can be used to indicate "greater than" or "equal to"; "less than" can be used to indicate "less than" or "equal to".
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0035] To better understand the embodiments of the present application, relevant examples are first described as follows:
[0036] In some embodiments, in an actual picture dataset, illegal pictures often account for only a very small part, while normal pictures account for the vast majority. However, if an artificial intelligence model that is not applicable to an imbalanced dataset is used. At this time, it will tend to predict most pictures as normal pictures, which may result in a situation where the accuracy rate is falsely high.
[0037] In some embodiments, since the labels of the image dataset are often obtained through manual annotation, some images may be mislabeled. These noisy images will have an adverse effect on the training of the artificial intelligence model, resulting in unstable performance of the model in actual applications. However, without a dedicated mechanism to handle these noisy images, the robustness and generalization ability will be limited.
[0038] In some embodiments, clean samples (including typical normal images and typical violation images) can usually be easily obtained and can guide the model to better identify noisy samples like a teacher.
[0039] In some embodiments, common image enhancement means (such as scaling, rotation, shearing, flipping, etc.) during the model training process will affect the self-paced learning technique, thereby further affecting the robustness and generalization ability of the model.
[0040] As Figure 1 shown, an image method is provided in an embodiment of the present application. The method includes:
[0041] Step S101, obtaining an image sample to be trained.
[0042] Step S102, adding a fully connected layer to the selected neural network structure to construct an area under the curve (AUC) model to be trained.
[0043] Step S103, determining an objective function for training the AUC model to be trained based on the self-paced learning algorithm.
[0044] Step S104, determining a training strategy for training the AUC model to be trained based on the objective function.
[0045] Step S105: Training the AUC model to be trained based on the training strategy and the image sample to be trained until the objective function converges, obtaining a trained AUC model.
[0046] The image recognition method of the embodiment of the present application can be applied to an electronic device. The electronic device involved in the embodiment of the present application can be, but is not limited to, a computer, a mobile phone, a wearable device, a vehicle-mounted terminal, a roadside unit (RSU), a smart home terminal, an industrial sensing device, and / or a medical device, etc.
[0047] In some embodiments, the AUC model may be an AUC model constructed based on the ResNet or DenseNet neural network structure. There are no special requirements for the neural network structure, and different neural network structures (such as ResNet or DenseNet) can be used according to the actual situation. For a specific neural network structure, only a fully connected layer with an output dimension of one needs to be added to construct the AUC model used in this solution. It should be noted that in the disclosure, θ represents the parameters of the AUC model, and f θ : R d ->R represents the AUC model.
[0048] In some embodiments, the above method corresponds to the training process of the AUC model, and the trained AUC model can be executed on an electronic device that needs to apply the method. Before applying the AUC model, the AUC model needs to be trained. The training process of the AUC model can be executed on a dedicated device for model training or on the electronic device that applies the method, which is not limited herein. When the training process of the AUC model is executed on the dedicated device, after the AUC model is trained, the trained AUC model can be transplanted from the dedicated device to the electronic device that applies the method for execution.
[0049] In some embodiments, meeting the convergence condition may include that the convergence function for training the AUC model meets the convergence condition, or may include determining that the convergence condition is met when the number of training times reaches a predetermined number, or other situations, which are not limited herein.
[0050] In some embodiments, the image samples to be trained include: the first type of images and the second type of images. The first type of images are images marked as not containing illegal content, and the second type of images are images marked as containing illegal content. The content of the illegal content is not limited.
[0051] In some embodiments, the image samples to be trained may be included in the training set, and the training set includes the first type of images and the second type of images.
[0052] Exemplarily, a training set is obtained, and these pictures are manually classified and marked as normal pictures (corresponding to the first type of images) and illegal pictures (corresponding to the second type of images). The normal pictures are represented as positive samples X + ∈R d , and the illegal pictures are represented as negative samples X - ∈R d , d represents the sample dimension, n+ and n- respectively represent the numbers of positive and negative samples, and n = n + +n - represents the number of all samples.
[0053] In some embodiments, the trained AUC model is used to identify the image to be recognized as a first type of image or a second type of image.
[0054] In some embodiments, the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm.
[0055] In some embodiments, the objective function includes at least one of the following:
[0056] The first term, which is a regularization term for reducing model overfitting;
[0057] The second term, which is a term for assisting the model to identify noise samples based on reference image samples;
[0058] The third term, which is a term for weighting the AUC loss using the image sample weight;
[0059] The fourth term, which is a self-paced regularization term;
[0060] The fifth term, which is a binary classification balance regularization term for balancing the weights of positive and negative samples;
[0061] The sixth term, which is a consistency regularization term for reducing the impact of augmentation operations on the self-paced learning algorithm, and the augmentation operations include at least one of the following: image scaling; image rotation; image shearing; image flipping.
[0062] In some embodiments, the reference image samples can be referred to as clean samples. Exemplarily, clean samples (including typical normal pictures and typical violation pictures) can guide the model to better identify noise samples like a teacher, thereby improving the robustness and generalization ability of the model. It should be noted that typical normal pictures (corresponding to the first sample image or the third sample image) are denoted as clean positive samples X + ∈R d , typical violation pictures (corresponding to the second sample image or the fourth sample image) are denoted as clean negative samples X - ∈R d , d represents the sample dimension, m + and m - respectively represent the numbers of clean positive and negative samples, and m = m + + m - represents the number of all clean samples.
[0063] In some embodiments, the training strategy includes at least one of the following: updating the image sample weights using stochastic gradient parameters; updating the AUC model parameters using stochastic gradient parameters; updating the step size parameters using stochastic gradient parameters.
[0064] It should be noted that the self-paced learning technique is a training strategy that simulates the learning processes of humans and animals, and its characteristic is to gradually train samples in the order from simple to complex. For this purpose, the step size parameter λ and the image sample weight w for controlling the training difficulty are introduced. In the embodiments of the present application, w + ∈[0; 1] n+ and w - ∈[0; 1] n- respectively represent the weight vectors of positive and negative image samples. It should be noted that in the embodiments of the present application, represents the data augmentation operation performed on the sample x, including scaling, rotation, shearing, flipping, etc.
[0065] In some embodiments, the objective function is:
[0066] Formula One:
[0067]
[0068] Among them, the first term is a model regularization term for avoiding model overfitting; the second and third terms make full use of clean samples, enabling clean samples to guide the model to better identify noisy samples like a teacher; the fourth term is the result of weighting the AUC loss using the image sample weights; the fifth term is a common self-paced regularization term in the self-paced learning technique; the sixth term is a binary classification balance regularization term, which can ensure the balance of the average weights of positive and negative image samples; the last term is a consistency regularization term, which can alleviate the adverse effects of the image augmentation means on the self-paced learning technique.
[0069] In some embodiments, obtain training image samples, where the training image samples include: first-class images and second-class images, the first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content; add a fully connected layer to the selected neural network structure to construct a training Area Under the Curve (AUC) model, where the trained AUC model is used to identify whether the image to be identified is a first-class image or a second-class image; determine an objective function for training the training AUC model based on the self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; the objective function includes at least one of the following: a first term, the first term is a regularization term, and the regularization term is used to reduce model overfitting; a second term, the second term is a term for assisting the model to identify noise samples based on reference image samples; a third term, the third term is a term for weighting the AUC loss using the image sample weight; a fourth term, the fourth term is a self-paced regularization term; a fifth term, the fifth term is a binary classification balance regularization term, and the binary classification balance regularization term is used to balance the weights of positive and negative samples; a sixth term, the sixth term is a consistency regularization term, and the consistency regularization term is used to reduce the impact of the augmentation operation on the self-paced learning algorithm, and the augmentation operation includes at least one of the following: image scaling; image rotation; image shearing; image flipping; determine a training strategy for training the training AUC model based on the objective function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameter using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter; train the training AUC model based on the training strategy and the training image samples until the objective function converges to obtain the trained AUC model.
[0070] In some embodiments, obtain an image sample to be trained, where the image sample to be trained includes: a first type of image and a second type of image, the first type of image is an image marked as not containing any illegal content, and the second type of image is an image marked as containing illegal content; add a fully connected layer to the selected neural network structure to construct an area under the curve (AUC) model to be trained, where the trained AUC model is used to identify whether the image to be identified is a first type of image or a second type of image; screen out the reference image samples from the image samples to be trained; where the reference image samples include the first type of image samples and the second type of image samples for assisting the model in identifying noise samples; determine an objective function for training the AUC model to be trained based on the self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; the objective function includes: a second term, and the second term is a term for assisting the model in identifying noise samples based on the reference image samples; determine a training strategy for training the AUC model to be trained based on the objective function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameter using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter; train the AUC model to be trained based on the training strategy and the image samples to be trained until the objective function converges to obtain a trained AUC model.
[0071] In some embodiments, obtain image samples to be trained, where the image samples to be trained include: first-class images and second-class images. The first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content; add a fully connected layer to the selected neural network structure to construct an area under the curve (AUC) model to be trained, where the trained AUC model is used to identify whether the image to be identified is a first-class image or a second-class image; determine an objective function for training the AUC model to be trained based on the self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter used to control the training difficulty of the self-paced learning algorithm; determine a training strategy for training the AUC model to be trained based on the objective function, where the training strategy includes at least one of the following: using a stochastic gradient parameter to update the image sample weight; using a stochastic gradient parameter to update the AUC model parameter; using a stochastic gradient parameter to update the step size parameter; the using a stochastic gradient parameter to update the image sample weight includes: randomly selecting a first image sample from the image samples to be trained; randomly selecting a second image sample from the reference image samples; determining the partial derivative of the objective function with respect to the image sample weight based on the first image sample and the second image sample to obtain a first stochastic gradient parameter; updating the image sample weight based on the first stochastic gradient parameter; training the AUC model to be trained based on the training strategy and the image samples to be trained until the objective function converges to obtain a trained AUC model.
[0072] In some embodiments, updating the image sample weight based on the first stochastic gradient parameter may be updating the image sample weight based on the first stochastic gradient parameter until the image sample weight converges.
[0073] In some embodiments, training image samples are obtained, where the training image samples include: first-class images and second-class images. The first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content; a fully connected layer is added to the selected neural network structure to construct a training Area Under the Curve (AUC) model, where the trained AUC model is used to identify whether a to-be-identified image is a first-class image or a second-class image; a target function for training the training AUC model is determined based on the self-paced learning algorithm, where the target function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; a training strategy for training the training AUC model is determined based on the target function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameters using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter; the updating the AUC model parameters using a stochastic gradient parameter includes: randomly selecting a third image sample from the training image samples; randomly selecting a fourth image sample from the reference image samples; determining the partial derivative of the target function with respect to the AUC model parameters based on the third image sample and the fourth image sample to obtain a second stochastic gradient parameter; updating the AUC model parameters based on the second stochastic gradient parameter; training the training AUC model based on the training strategy and the training image samples until the target function converges to obtain a trained AUC model.
[0074] In some embodiments, updating the AUC model parameters based on the second stochastic gradient parameter may be updating the AUC model parameters based on the second stochastic gradient parameter until the AUC model parameters converge.
[0075] In some embodiments, updating the step size parameter using a stochastic gradient parameter may be increasing the step size parameter based on the stochastic gradient parameter.
[0076] For a better understanding of the present invention, please refer to Figure 2 , an image recognition method provided by an embodiment of the present application includes:
[0077] Step S201: Collect a picture data set, including normal pictures (first-class pictures) and illegal pictures (second-class pictures).
[0078] In some embodiments, this step requires collecting pictures in the project and manually classifying these pictures, marking them as normal pictures (corresponding to the first-class images) and illegal pictures (corresponding to the second-class images). The normal pictures are represented as positive samples
[0079] X + ∈Rd The illegal pictures are represented as negative samples X - ∈R d , where d represents the sample dimension, n+ and n- represent the numbers of positive and negative samples respectively, and n = n + + n - represents the number of all samples.
[0080] Step S202: Further screen the data in step S201 to obtain a clean data set, including typical normal pictures (corresponding to the first image sample or the third image sample) and typical illegal pictures (corresponding to the second image sample or the fourth image sample).
[0081] In some embodiments, the clean samples (including typical normal pictures and typical illegal pictures) can guide the model to better identify noise samples like a teacher, thereby improving the robustness and generalization ability of the model. It should be noted that the typical normal pictures (corresponding to the first sample image or the third sample image) are represented as clean positive samples X + ∈R d , and the typical illegal pictures (corresponding to the second sample image or the fourth sample image) are represented as clean negative samples X - ∈R d , where d represents the sample dimension, m + and m - represent the numbers of clean positive and negative samples respectively, and m = m + + m - represents the number of all clean samples.
[0082] Step S203: Select a suitable neural network structure (such as ResNet or DenseNet), and construct a suitable AUC model therewith.
[0083] In some embodiments, there is no special requirement for the neural network structure, and different neural network structures (such as ResNet or DenseNet) can be used according to the actual situation. For a specific neural network structure, only a fully connected layer with an output dimension of one needs to be added to construct the AUC model used in this solution. It should be noted that in the disclosure, θ represents the parameters of the AUC model, and f θ : R d ->R represents the AUC model.
[0084] Step S204: Combine the self-paced learning technique to formulate a suitable objective function.
[0085] In some embodiments, the self-paced learning technique is a training strategy that mimics the learning processes of humans and animals, characterized by gradually training samples in the order from simple to complex. To this end, a step size parameter λ for controlling the training difficulty and an image sample weight w are introduced. In the embodiments of the present application, w + ∈[0; 1] n+ and w - ∈[0; 1] n- represent the weight vectors of positive and negative image samples respectively. It should be noted that in the embodiments of the present application, T(x) represents the data augmentation operation performed on the sample x, including scaling, rotation, shearing, flipping, etc. Exemplarily, the objective function can be seen in Formula 1.
[0086] Step S205, design a reasonable and effective model training strategy based on the objective function.
[0087] In some embodiments, referring to Formula 1 again, there are two coordinate blocks in the objective function used in the present invention: the sample weight coordinate block and the model parameter coordinate block. In view of this, the present invention fully follows the idea of stochastic optimization and uses stochastic gradients to alternately update the sample weights and model parameters. The specific steps are as Figure 3 shown and will be described in detail in the next part.
[0088] Step S206, use the datasets in Steps S201 and S202 to train the AUC model constructed in Step S203 according to the training strategy in Step S205.
[0089] In some embodiments, this step is to train the artificial intelligence model constructed in the present solution so that it can correctly distinguish normal pictures from illegal pictures.
[0090] Step S207, deploy the trained AUC model online to perform real-time recognition on the pictures uploaded by users.
[0091] In some embodiments, this step is to deploy the artificial intelligence model constructed in the present invention to an actual project to help the administrator identify the illegal pictures uploaded by users.
[0092] In some embodiments, for the above Step S205, please refer to Figure 3 , the embodiments of the present application provide a training strategy, including:
[0093] Step S301, randomly select samples from the dataset in Step S201.
[0094] In some embodiments, this step is to perform random sampling on the dataset. Among them, the number of positive samples obtained by sampling is the number of negative samples is And
[0095] Step S302: Randomly select samples from the dataset in Step S202.
[0096] This step performs random sampling on the dataset. Among them, the number of positive samples obtained by sampling is and the number of negative samples is And
[0097] Step S303: Update the sample weights through a formula.
[0098] In some embodiments, the formula is:
[0099] Formula Two:
[0100]
[0101] This step uses the samples obtained by random sampling in Step S301 and Step S302 to calculate the stochastic gradient and updates the sample weights using the stochastic gradient. Specifically, when calculating the stochastic gradient G i (w w ) with respect to the sample weight w i , the partial derivative of Formula One with respect to w i needs to be calculated.
[0102] Step S304: Update the model parameters through a formula.
[0103] In some embodiments, the formula is:
[0104] Formula Three:
[0105]
[0106] This step uses the samples obtained by random sampling in Step S501 and Step S502 to calculate the stochastic gradient and updates the model parameters using the stochastic gradient. Specifically, when calculating the stochastic gradient Gθ(θ) with respect to the model parameter θ, the partial derivative of Formula One with respect to θ needs to be calculated.
[0107] Step S305: Repeat Steps S301 to S304 until the sample weights and model parameters converge.
[0108] Step S306: Update the step size parameter through a formula:
[0109] Formula Four:
[0110] λ k = min(cλ k-1 , λ ∞ )
[0111] In some embodiments, this step is a conventional step in self-paced learning technology. By gradually increasing the step size parameter to the maximum value λ ∞ , the training difficulty is gradually increased.
[0112] Step S307, repeat steps S301 to S306 until the sample weights and model parameters converge.
[0113] In this application, as shown in Formula 1, the guiding role of clean samples is emphasized, further improving the robustness and generalization ability of the model. The adverse effects of common image enhancement means on self-paced learning technology are alleviated by the consistency regularization term. As steps S301 to S307, the random optimization idea is followed, reducing the time complexity of the algorithm.
[0114] The above steps S301 to S307 use a random optimization strategy to optimize the objective function in Formula 1.
[0115] In the embodiments of this application, an objective function for training the to-be-trained AUC model is determined based on the self-paced learning algorithm. Since the objective function is determined based on the self-paced learning algorithm and the objective function is determined based on the step size parameter and / or the image sample weights, the adverse effects brought by noise samples can be reduced. Training the to-be-trained AUC model based on the training strategy and the to-be-trained image samples to obtain a trained AUC model. Since the training strategy includes at least one of the following: using random gradient parameters to update the image sample weights; using random gradient parameters to update the AUC model parameters; using random gradient parameters to update the step size parameter; during the process based on the training strategy, the image sample weights, the AUC model parameters, and / or the step size parameter will be continuously updated, so that the result of the trained AUC model for identifying images will be more reliable and accurate.
[0116] As Figure 4 shown, in the embodiments of this application, an image recognition method is provided, and the method includes:
[0117] Step S401, obtain the training image samples to be recognized.
[0118] Step S402, recognize the training image samples to be recognized based on the trained AUC model.
[0119] In some embodiments, obtain training image samples to be trained, where the training image samples to be trained include: first-class images and second-class images, the first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content; add a fully connected layer to the selected neural network structure to construct an area under the curve (AUC) model to be trained, where the trained AUC model is used to identify whether the image to be identified is a first-class image or a second-class image; determine an objective function for training the AUC model to be trained based on the self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; determine a training strategy for training the AUC model to be trained based on the objective function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameter using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter; train the AUC model to be trained based on the training strategy and the training image samples to be trained until the objective function converges to obtain a trained AUC model. Obtain the training image samples to be identified. Identify the training image samples to be identified based on the trained AUC model.
[0120] As Figure 5 shown, an embodiment of the present application provides an image recognition device, and the image recognition device includes:
[0121] An acquisition module 51, configured to: acquire training image samples to be trained, where the training image samples to be trained include: first-class images and second-class images, the first-class images are images marked as not containing illegal content, and the second-class images are images marked as containing illegal content;
[0122] A construction module 52, configured to: add a fully connected layer to the selected neural network structure to construct an area under the curve (AUC) model to be trained, where the trained AUC model is used to identify whether the image to be identified is a first-class image or a second-class image;
[0123] A determination module 53, configured to: determine an objective function for training the AUC model to be trained based on the self-paced learning algorithm, where the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the training difficulty of the self-paced learning algorithm; determine a training strategy for training the AUC model to be trained based on the objective function, where the training strategy includes at least one of the following: updating the image sample weight using a stochastic gradient parameter; updating the AUC model parameter using a stochastic gradient parameter; updating the step size parameter using a stochastic gradient parameter;
[0124] A training module 54, configured to: train the AUC model to be trained based on the training strategy and the image samples to be trained until the objective function converges, so as to obtain the trained AUC model.
[0125] An embodiment of the present application provides a processing device, which includes: implementing any of the methods described in the embodiments of the application.
[0126] An embodiment of the present application provides a processing device, which includes:
[0127] A memory, configured to store an executable program;
[0128] A processor, configured to implement any of the methods described in the embodiments of the present application when executing the executable program stored in the memory.
[0129] It can be understood that the memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.
[0130] Among them, the method for determining the topology structure disclosed in the embodiments of this application can be applied to or implemented by the processor. The processor can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the method for determining the topology structure can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the method for determining the topology structure provided in the embodiments of this application.
[0131] The embodiments of this application also provide a computer storage medium. The computer storage medium stores an executable program, and when the executable program is executed by a processor, it implements the method for image recognition as described in any one of the embodiments of this application. Specifically, it can be a computer-readable storage medium, such as a memory storing a computer program. The above-mentioned computer program can be executed by the processor of the processing device to complete the steps of the method described in the embodiments of this application. The computer-readable storage medium can be a ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0132] The embodiments of this application provide a computer program product, which includes: a computer program or executable instructions, and the computer program or executable instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer program or executable instructions from the computer-readable storage medium, and the processor executes the computer program or executable instructions, so that the computer device executes any one of the above-mentioned methods for image recognition in the embodiments of this application.
[0133] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An image recognition method, characterized in that: The method comprises: Acquire a sample of images to be trained, wherein the sample of images to be trained includes: a first type of images and a second type of images, wherein the first type of images are images marked as not containing illegal content, and the second type of images are images marked as containing illegal content; A fully connected layer is added to the selected neural network structure to construct an area under the curve AUC model to be trained, wherein the trained AUC model is used to identify whether the image to be identified is a first-category image or a second-category image; Determine an objective function for training the AUC model to be trained based on a self-paced learning algorithm, wherein the objective function is determined based on a step size parameter and / or an image sample weight, and the step size parameter is a parameter for controlling the difficulty of training the self-paced learning algorithm; Determining a training strategy for training the AUC model to be trained based on the objective function, wherein the training strategy includes at least one of the following: updating image sample weights using random gradient parameters; updating AUC model parameters using random gradient parameters; updating step size parameters using random gradient parameters; The AUC model to be trained is trained based on the training strategy and the image samples to be trained until the objective function converges to obtain a trained AUC model.
2. The method according to claim 1, characterized in that The objective function includes at least one of the following: The first item, the first item is a regularization item, and the regularization item is used to reduce the regularization item of the model overfitting; A second item, the second item being an item for identifying noise samples based on the reference image sample auxiliary model; A third item, the third item is a item for weighting the AUC loss using the image sample weight; The fourth term, the fourth term is a self-stepping regularization term; The fifth item is a binary balanced regularization item, and the binary balanced regularization item is used to balance the weights of positive samples and negative samples; The sixth item is a consistency regularization item, and the consistency regularization item is used to reduce the impact of the enhancement operation on the self-paced learning algorithm, and the enhancement operation includes at least one of the following: image scaling; image rotation; image shearing; image flipping.
3. The method according to claim 2, characterized in that The objective function includes the second term, and the method further includes: Filtering the reference image samples from the image samples to be trained; The reference image samples include the first category image samples and the second category image samples used to assist the model in identifying noise samples.
4. The method according to claim 3, characterized in that The method of updating the image sample weights using the stochastic gradient parameters comprises: Randomly selecting a first image sample from the image samples to be trained; randomly selecting a second image sample from the reference image samples; Determine a partial derivative of an objective function with respect to an image sample weight based on the first image sample and the second image sample to obtain a first stochastic gradient parameter; The image sample weights are updated based on the first stochastic gradient parameters.
5. The method according to claim 4, characterized in that The updating of the image sample weight based on the first stochastic gradient parameter comprises: The image sample weights are updated based on the first stochastic gradient parameters until the image sample weights converge.
6. The method according to any one of claims 3 to 5, characterized in that The method of updating the AUC model parameters using stochastic gradient parameters includes: Randomly selecting a third image sample from the image samples to be trained; randomly selecting a fourth image sample from the reference image samples; Determine a partial derivative of an objective function with respect to an AUC model parameter based on the third image sample and the fourth image sample to obtain a second stochastic gradient parameter; The AUC model parameters are updated based on the second stochastic gradient parameters.
7. The method according to claim 6, characterized in that The updating of the AUC model parameters based on the second stochastic gradient parameters comprises: The AUC model parameters are updated based on the second stochastic gradient parameters until the AUC model parameters converge.
8. A processing device, characterized in that: The processing device is configured to execute the method according to any one of claims 1 to 7.
9. A computer storage medium, characterized in that The computer storage medium stores an executable program, and when the executable program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program or instructions, characterized in that: When the computer program or instruction is executed by a processor, the method according to any one of claims 1 to 7 is implemented.