Model training method, long-tail image recognition method and electronic equipment

By combining the weighted loss processing of category frequency and sample gradient during model training, the problem of low recognition accuracy of tail classes in long-tail image recognition is solved, balanced recognition of head and tail classes is achieved, and the overall recognition accuracy is improved.

CN120823441APending Publication Date: 2025-10-21AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957029.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies cannot guarantee the accuracy of long-tail image recognition, mainly because the model is biased towards the head class during training, resulting in low accuracy in tail class recognition.

Method used

By obtaining a dataset that obeys a long-tail distribution, using the frequency of occurrence of sample categories for weighted processing, combining the sample gradient to adjust the loss function, and fusing the weighted loss based on category frequency and sample gradient to train the target image recognition model.

Benefits of technology

While improving the recognition accuracy of tail classes, the recognition accuracy of head classes is maintained, ensuring the overall accuracy of long-tail image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823441A_ABST
    Figure CN120823441A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method, a long-tail image recognition method and electronic equipment. The method comprises the following steps: acquiring a data set obeying long-tail distribution, and inputting the data set into an image recognition model so as to obtain sample loss corresponding to each sample image according to an output result of the image recognition model and a sample category corresponding to each sample image in the data set; determining an occurrence frequency corresponding to each sample category in the data set, and weighting each sample loss based on each occurrence frequency to obtain a first loss; determining a sample gradient corresponding to each sample image, and weighting each sample loss based on each sample gradient to obtain a second loss; and obtaining target loss according to the first loss and the second loss, and adjusting parameters in the image recognition model according to the target loss to train a target image recognition model. According to the technical scheme of the embodiment of the invention, the long-tail image can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a model training method, a long-tail image recognition method, and an electronic device. Background Art

[0002] In the process of visual image research, in order to verify the robustness and reliability of the research results, they are usually fully verified on multiple large-scale datasets. Most of these datasets are carefully designed to be balanced datasets, that is, the number of images in each category in the dataset is equal.

[0003] However, datasets collected from nature often follow a long-tail distribution, such as Figure 1 As shown in the figure, images of the minority class (i.e., the head class) occupy the vast majority of the dataset, while the majority class (i.e., the tail class) is underrepresented in the dataset. Using the above dataset to train a model and then perform image recognition based on it is called long-tail image recognition.

[0004] In the process of realizing the present invention, the inventors discovered that the following technical problems exist in the prior art: the accuracy of long-tail image recognition cannot be guaranteed, which needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present invention provide a model training method, a long-tail image recognition method, and an electronic device, which solve the problem of being unable to guarantee the accuracy of long-tail image recognition.

[0006] According to one aspect of the present invention, a model training method is provided, which may include:

[0007] Obtain a dataset that follows a long-tail distribution and input the dataset into an image recognition model to obtain the sample loss corresponding to each sample image based on the output of the image recognition model and the sample category corresponding to each sample image in the dataset;

[0008] Determine the frequency of occurrence of each sample category in the data set, and weight the loss of each sample based on the frequency of occurrence to obtain the first loss;

[0009] Determine the sample gradient corresponding to each sample image, and weight each sample loss based on the sample gradient to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image;

[0010] A target loss is obtained according to the first loss and the second loss, and parameters in the image recognition model are adjusted according to the target loss to train and obtain a target image recognition model.

[0011] According to another aspect of the present invention, a long-tail image recognition method is provided, which may include:

[0012] Obtaining a target image to be recognized and a target image recognition model trained according to the model training method described in any embodiment of the present invention, wherein each sample category used in the training process of the target image recognition model includes the target category to which the target image belongs;

[0013] The target image is input into the target image recognition model, and the target category is obtained according to the output result of the target image recognition model.

[0014] According to another aspect of the present invention, a model training device is provided, which may include:

[0015] The sample loss obtaining module is used to obtain a dataset that follows a long-tail distribution and input the dataset into the image recognition model to obtain the sample loss corresponding to each sample image based on the output results of the image recognition model and the sample category corresponding to each sample image in the dataset;

[0016] A first loss obtaining module is used to determine the occurrence frequency of each sample category in the data set, and perform weighted processing on each sample loss based on each occurrence frequency to obtain the first loss;

[0017] A second loss obtaining module is used to determine the sample gradient corresponding to each sample image, and to perform weighted processing on each sample loss based on each sample gradient to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image;

[0018] The model training module can be used to obtain a target loss based on the first loss and the second loss, and adjust the parameters in the image recognition model based on the target loss to train and obtain the target image recognition model.

[0019] According to another aspect of the present invention, a long-tail image recognition device is provided, which may include:

[0020] a model acquisition module, configured to acquire a target image to be recognized and a target image recognition model trained according to the model training method described in any embodiment of the present invention, wherein each sample category used in the training process of the target image recognition model includes the target category to which the target image belongs;

[0021] The image recognition module is used to input the target image into the target image recognition model and obtain the target category based on the output result of the target image recognition model.

[0022] According to another aspect of the present invention, there is provided an electronic device, which may include:

[0023] at least one processor; and

[0024] a memory communicatively connected to at least one processor; wherein,

[0025] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor implements the model training method or long-tail image recognition method provided by any embodiment of the present invention when executing.

[0026] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer instructions are stored. The computer instructions are used to enable a processor to implement the model training method or long-tail image recognition method provided by any embodiment of the present invention when executed.

[0027] According to another aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the model training method or the long-tail image recognition method provided by any embodiment of the present invention is implemented.

[0028] The technical solution of an embodiment of the present invention obtains a data set that obeys a long-tail distribution and inputs the data set into an image recognition model to obtain a sample loss corresponding to each sample image based on the output result of the image recognition model and the sample category corresponding to each sample image in the data set; on this basis, the occurrence frequency corresponding to each sample category in the data set is determined, and each sample loss is weighted based on each occurrence frequency to obtain a first loss, which can be considered as a weighted loss based on the category frequency; and the sample gradient corresponding to each sample image is determined, which can characterize the difficulty of the image recognition model learning the corresponding sample image, and each sample loss is weighted based on each sample gradient to obtain a second loss, which can be considered as a weighted loss based on the sample gradient; further, based on the first loss and the second loss, a target loss is obtained, and the parameters in the image recognition model are adjusted based on the target loss to train a target image recognition model. During the model training process, the above technical solution can reduce the negative impact of inter-class imbalance and intra-class imbalance on the model by fusing the weighted loss based on category frequency and the weighted loss based on sample gradient. This can improve the accuracy of tail class recognition while maintaining the accuracy of head class recognition, thereby ensuring the accuracy of long-tail image recognition.

[0029] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0031] Figure 1 It is a schematic diagram of the long-tail distribution of data sets in nature;

[0032] Figure 2 is a flowchart of a model training method provided according to an embodiment of the present invention;

[0033] Figure 3 is a flowchart of another model training method provided according to an embodiment of the present invention;

[0034] Figure 4 is a flowchart of an optional example of another model training method provided according to an embodiment of the present invention;

[0035] Figure 5 is a flowchart of a long-tail image recognition method provided according to an embodiment of the present invention;

[0036] Figure 6 is a structural block diagram of a model training device provided according to an embodiment of the present invention;

[0037] Figure 7 2 is a structural block diagram of a long-tail image recognition device provided according to an embodiment of the present invention;

[0038] Figure 8 It is a structural diagram of an electronic device for implementing the model training method or long-tail image recognition method of an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0040] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The situations of "target", "original", etc. are similar and will not be repeated here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0041] Before introducing the embodiments of the present invention, an example is given to explain why the accuracy of long-tail image recognition cannot be guaranteed at present, so as to better understand how the embodiments of the present invention ensure this accuracy.

[0042] For example, as explained above, images of the head class occupy the vast majority of the dataset, while images of the tail class are under-represented in the dataset. This may cause the model to be biased toward the head class when training the model on this dataset. That is, the model has a higher recognition accuracy on the head class, but a lower recognition accuracy on the tail class, which makes it impossible to guarantee the accuracy of long-tail image recognition.

[0043] Therefore, to ensure the accuracy of long-tail image recognition, the key is to bias the model toward the tail classes during training to improve recognition accuracy on these classes, while also ensuring that the model is not overly biased toward these classes, thereby compromising recognition accuracy on the head classes. Guided by this principle, the model training methods described in the following examples are proposed.

[0044] Figure 2 This is a flow chart of a model training method provided by an embodiment of the present invention. This embodiment is applicable to image recognition model training, and is particularly applicable to image recognition model training adapted for long-tail images. This method can be performed by the model training device provided by an embodiment of the present invention. This device can be implemented in software and / or hardware and can be integrated into an electronic device, such as various user terminals or servers.

[0045] See also Figure 2 The method of the embodiment of the present invention specifically includes the following steps:

[0046] S110. Obtain a data set that obeys a long-tail distribution, and input the data set into an image recognition model to obtain a sample loss corresponding to each sample image based on an output result of the image recognition model and a sample category corresponding to each sample image in the data set.

[0047] The data set includes multiple sample images, some of which belong to sample category 1, some belong to sample category 2, and some belong to sample category C, where C is a positive number greater than 1. It should be noted that the data set obeys a long-tail distribution, that is, the multiple sample images obey a long-tail distribution, which is Zipf's law. For the sake of illustration, for example, assume that the data set is a data set with N training samples (i.e., sample images) and C sample categories. The data set can be represented as D train ={(x i ,y i )|i∈{1,2,...,N},y i ∈{1, 2, ..., C}}, where x i is the i-th sample image, y i is its corresponding label (i.e., sample category).

[0048] The image recognition model can be understood as a deep learning model with image recognition capabilities, wherein a parameter can be denoted as θ. By adjusting the parameter θ (i.e., training the image recognition model), the image recognition model can minimize the classification loss by effectively processing the imbalanced nature of the data set. In an embodiment of the present invention, optionally, the image recognition model may include a feature backbone and a classifier, wherein the feature backbone may be responsible for extracting discriminative features from the input sample image, and the classifier may map the learned features to a logical space to identify (i.e., classify) the sample image.

[0049] During the model training process, the dataset is input into the image recognition model to target each sample image in the dataset, here the sample image x i For example, we can use the image recognition model to identify the sample image x i Perform feature extraction and classification to obtain the logical vector z i , the logical vector z i Including C logical values, here through z i,j Represents the logical vector z i The jth logical value of , j∈{1, 2, ..., C}; then, based on the logical vector z i And the sample image x i The sample category y i Perform loss calculation to obtain the sample image x iFor example, here we take cross entropy loss as an example, assuming that the sample category y i is c, then the sample loss calculated based on the cross entropy loss can be expressed as follows:

[0050]

[0051] S120 , determining the occurrence frequency of each sample category in the data set, and performing weighted processing on each sample loss based on each occurrence frequency to obtain a first loss.

[0052] Among them, in the data set that obeys the long-tail distribution, there is an inter-class imbalance phenomenon. This phenomenon refers to the fact that there are large differences in the number of sample images under each sample category in the data set, that is, the number of sample images between classes is unbalanced, which leads to the model optimization often being dominated by the head class, and then the model is seriously biased towards the head class, which is manifested in the obvious difference in the model's recognition accuracy on the head class and the tail class.

[0053] On this basis, to improve the accuracy of tail class recognition, we can consider reweighting the class distribution. The core idea of ​​this reweighting is to adjust the weights of sample images in the head and tail classes, allowing the model to give more attention to the tail class while reducing the model's attention to the head class. The frequency of occurrence of sample classes in the dataset, as a priori knowledge, directly reflects the degree of imbalance in the dataset. Therefore, the frequency of occurrence can be used for reweighting. In the embodiments of the present invention, this frequency of occurrence can also be referred to as the category frequency.

[0054] Specifically, determine the frequency of occurrence of each sample category in the data set. For example, the number of samples corresponding to each sample category in the data set can be determined. The number of samples can represent the number of sample images belonging to the corresponding sample category in the data set. Then, for each sample category, based on its corresponding sample images and the number of samples corresponding to each sample category, the frequency of occurrence of the sample category in the data set can be determined. On this basis, combined with the above example about the data set, here we take sample category c as an example, its frequency of occurrence π c It can be expressed by the following formula:

[0055]

[0056] Among them, n c is the number of samples of sample category c.

[0057] Furthermore, based on the frequency of occurrence of each sample category, each sample loss can be weighted to obtain the first loss. Specifically, if sample category c has a small number of sample images, its contribution to model optimization is minimal. In this case, the loss weight of sample images from sample category c can be increased to encourage greater participation in the model optimization process. In other words, for sample categories with higher occurrence frequencies, the sample loss of their corresponding sample images is assigned a higher weight, while for samples with lower occurrence frequencies, a lower weight is assigned. In this way, the sample losses corresponding to each sample image are weighted to obtain the first loss.

[0058] On this basis, optionally, the first loss can be obtained in the following manner: for each sample category, the category weight of the sample category is determined according to the frequency of occurrence corresponding to the sample category, wherein the category weight is negatively correlated with the corresponding frequency of occurrence, that is, the higher the frequency of occurrence, the smaller the category weight; according to the sample categories to which each sample image belongs and the category weights corresponding to each sample category, the sample losses corresponding to each sample image are weighted to obtain the first loss.

[0059] For example, here we take sample category c as an example, and its category weight w c It can be expressed by the following formula:

[0060]

[0061] According to the above formula, if the sample category c is a tail category, that is, its frequency of occurrence is small, then the category weight w calculated based on the above formula is c The larger the , the greater the loss contribution of the tail class.

[0062] In addition, the "1+" term in the above formula is to ensure that the calculated class weights are all positive numbers. The second-order derivative of the logarithmic function is less than 0. Therefore, when the input value of the above formula (i.e., the frequency of occurrence) increases, its output value (i.e., the class weight) also increases slowly. Therefore, even if the difference between the frequencies of occurrence is large, the class weight calculated by the above formula changes slowly. This avoids the risk of the model overemphasizing the tail class and ignoring the head class, and ensures that the recognition accuracy of both the head and tail classes is balanced.

[0063] Further, based on the above example, here we take the sample image x i For example, the first loss Loss1 can be expressed as follows:

[0064]

[0065] The above step (S120) performs loss weighting by the frequency of occurrence (i.e., class frequency), which can solve the model bias problem caused by the imbalance between classes to a certain extent.

[0066] S130. Determine the sample gradient corresponding to each sample image, and perform weighted processing on each sample loss based on each sample gradient to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image.

[0067] Among them, in the data set that obeys the long-tail distribution, there is an intra-class imbalance phenomenon. This phenomenon refers to the fact that there are differences in the difficulty of the model learning sample images under the same sample category, that is, the difficulty of learning sample images within the class is unbalanced, which causes the model to be able to learn sample images with lower difficulty but unable to learn sample images with higher difficulty, and thus cannot guarantee the accuracy of long-tail image recognition.

[0068] On this basis, in order to reduce the difficulty of learning difficult samples (i.e., sample images with high learning difficulty), we can consider difficult sample discovery, locate difficult samples, and increase their loss weight to encourage the model to better learn these difficult samples. In this regard, considering that the sample gradient of difficult samples is greater than the sample gradient of simple samples, the sample gradient can be used to enhance the model's learning of difficult samples.

[0069] Specifically, the sample gradient corresponding to each sample image is determined. The sample gradient represents the difficulty of the image recognition model learning the corresponding sample image. For example, for each sample image, the parameter gradient corresponding to each parameter in the image recognition model can be determined based on its sample loss. For example, the parameter gradient can be determined by taking the partial derivative. Further, the sample gradient of the sample image is obtained based on each parameter gradient. Here is an example. The parameter gradient and sample gradient can be calculated by the following formula:

[0070]

[0071] Among them, g represents the sample gradient, g i,j Represents the sample image x i The parameter gradient at the jth parameter θ among all parameters θ, abs(·) represents the absolute value, g i Represents the sample image x i The corresponding combination of parameter gradients.

[0072] Furthermore, since the sample gradient can directly reflect the difficulty of the model learning the corresponding sample image, each sample gradient can be used as a basis to perform weighted processing on each sample loss to obtain the second loss. For example, for each sample image, its sample weight is obtained according to its sample gradient. The sample gradient is positively correlated with the corresponding difficulty and the corresponding sample weight. For example, the sample gradient can be directly used as the sample weight, that is, w i =g, which helps to quickly obtain the sample weight. Of course, the sample weight can also be obtained based on other methods. According to the sample weight corresponding to each sample image, the sample loss corresponding to each sample image is weighted to obtain the second loss.

[0073] On this basis, in order to ensure the stability of model training and avoid gradient explosion, each sample weight can be normalized or smoothed, and the corresponding sample weight can be updated according to the obtained processing results. The smoothing process can be achieved by logarithmic function or other methods. Here, the sample image x i Taking normalization as an example, the sample weight after normalization can be expressed by the following formula:

[0074]

[0075] On this basis, the second loss can be expressed by the following formula:

[0076]

[0077] The above step (S130) performs loss weighting through sample gradients, which enhances the model's learning of difficult samples and promotes balanced optimization of the model.

[0078] S140. Obtain a target loss based on the first loss and the second loss, and adjust parameters in the image recognition model based on the target loss to train and obtain a target image recognition model.

[0079] Among them, S120 can improve the accuracy of tail class recognition, but at the expense of the accuracy of head class recognition. This is because while increasing the importance of the tail class, it may ignore some valuable information of the head class, resulting in underfitting of difficult samples in the head class.

[0080] S130 balances the contribution of difficult samples and simple samples to sample loss based on the difficulty of learning the sample images. However, since the number of difficult samples in the head class is far greater than that in the tail class, this may cause the model to be more biased towards the head class, further exacerbating the degree of model bias.

[0081] On this basis, in order to avoid sacrificing the accuracy of head class recognition or making the model more inclined to the head class, the first loss and the second loss can be integrated for model training to reduce the negative impact of inter-class imbalance and intra-class imbalance on the model. While improving the accuracy of tail class recognition, the accuracy of head class recognition is maintained, thereby ensuring the accuracy of long-tail image recognition.

[0082] Specifically, a target loss is obtained based on the first and second losses. For example, the target loss can be obtained by summing the first and second losses. Another example is the weighted sum of the first and second losses. The weights corresponding to the two losses can be preset or dynamically set based on actual conditions. These can be set based on actual needs and are not specifically limited here. The parameters of the image recognition model are then adjusted based on the target loss to train the target image recognition model.

[0083] On this basis, optionally, in order to verify the recognition accuracy of the target image recognition model obtained by the above training, the test set D test ={(x i ,y i )|i∈{1,2,...,T},y i ∈{1, 2, ..., C}}, and test it to ensure that the target image recognition model of the final application has a good recognition accuracy.

[0084] The technical solution of an embodiment of the present invention obtains a data set that obeys a long-tail distribution and inputs the data set into an image recognition model to obtain a sample loss corresponding to each sample image based on the output result of the image recognition model and the sample category corresponding to each sample image in the data set; on this basis, the occurrence frequency corresponding to each sample category in the data set is determined, and each sample loss is weighted based on each occurrence frequency to obtain a first loss, which can be considered as a weighted loss based on the category frequency; and the sample gradient corresponding to each sample image is determined, which can characterize the difficulty of the image recognition model learning the corresponding sample image, and each sample loss is weighted based on each sample gradient to obtain a second loss, which can be considered as a weighted loss based on the sample gradient; further, based on the first loss and the second loss, a target loss is obtained, and the parameters in the image recognition model are adjusted based on the target loss to train a target image recognition model. During the model training process, the above technical solution can reduce the negative impact of inter-class imbalance and intra-class imbalance on the model by fusing the weighted loss based on category frequency and the weighted loss based on sample gradient. This can improve the accuracy of tail class recognition while maintaining the accuracy of head class recognition, thereby ensuring the accuracy of long-tail image recognition.

[0085] Figure 3 It is a flowchart of another model training method provided in an embodiment of the present invention. This embodiment is optimized based on the above-mentioned technical solutions. In this embodiment, optionally, the target image recognition model is obtained by performing multiple rounds of training using a data set, and a target loss is obtained according to a first loss and a second loss, including: obtaining the current iteration round and the total iteration rounds, and determining a loss balance factor according to the current iteration round and the total iteration rounds; and weighting the first loss and the second loss according to the loss balance factor to obtain the target loss. Among them, the explanations of the terms that are the same as or corresponding to the above-mentioned embodiments are not repeated here.

[0086] See also Figure 3 , and combined with Figure 4 The method of this embodiment may specifically include the following steps:

[0087] S210. Obtain a data set that obeys a long-tail distribution, and input the data set into an image recognition model to obtain a sample loss corresponding to each sample image based on an output result of the image recognition model and a sample category corresponding to each sample image in the data set.

[0088] S220: Determine the occurrence frequency of each sample category in the data set, and perform weighted processing on each sample loss based on each occurrence frequency to obtain a first loss.

[0089] S230. Determine the sample gradient corresponding to each sample image, and perform weighted processing on each sample loss based on each sample gradient to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image.

[0090] S240: Obtain the current iteration round and the total iteration rounds for the image recognition model, and determine a loss balance factor based on the current iteration round and the total iteration rounds.

[0091] Among them, according to the above explanation, it can be seen that the first loss pays more attention to the sample images of the tail class, while the second loss pays more attention to the difficult samples, that is, compared with the first loss, the second loss is more difficult to learn. On this basis, in order to make the model learn and generalize better, inspired by the human concept of learning from easy to difficult, the first loss and the second loss can be integrated through course learning, that is, first learn simple samples, and then gradually transition to the learning of difficult samples. Combined with the embodiments of the present invention, this can be understood as gradually shifting the focus of the target loss in the model training process from the first loss (i.e., weighted loss based on category frequency) to the second loss (i.e., weighted loss based on sample gradient). During the transfer process, the regularization of the model for difficult samples is gradually enhanced.

[0092] Specifically, the target image recognition model is obtained by performing multiple rounds of iterative training using a data set. On this basis, the current iteration round can represent the number of iterative training rounds currently being performed, and the total iteration rounds can represent the total number of iterative training rounds required. Therefore, the loss balance factor is determined based on the current iteration round and the total iteration rounds. The loss balance factor can represent whether the first loss or the second loss should be used as the main learning method. For example, the loss balance factor can be determined by the following formula:

[0093]

[0094] Among them, α represents the loss balance factor, t represents the current iteration round, and T represents the total iteration round.

[0095] S250. Perform weighted processing on the first loss and the second loss according to the loss balance factor to obtain a target loss, and adjust parameters in the image recognition model according to the target loss to update the image recognition model.

[0096] Among them, according to the loss balance factor, the first loss and the second loss can be weighted to obtain the target loss. Exemplarily, the loss balance factor is negatively correlated with the current iteration round, then the first weight corresponding to the first loss and the second weight corresponding to the second loss can be determined according to the loss balance factor, and the loss balance factor is positively correlated with the first weight and negatively correlated with the second weight; then, according to the first weight and the second weight, the first loss and the second loss are weighted to obtain the target loss. Here, still taking the above example as an example, the objective function Loss can be expressed by the following formula:

[0097] Loss = αLoss1 + (1-α) Loss2;

[0098] Among them, α is the first weight and 1-α is the second weight.

[0099] S260. If the current iteration round is less than the total iteration round, the current iteration round is incremented and updated, and the process returns to execute S210; otherwise, the image recognition model is used as the target image recognition model.

[0100] The technical solution of the embodiment of the present invention integrates the weighted loss based on category frequency and the weighted loss based on sample gradient through curriculum learning, so that the model can learn and generalize better.

[0101] Figure 5This is a flow chart of a long-tail image recognition method provided in an embodiment of the present invention. This embodiment is applicable to image recognition, and particularly to long-tail image recognition. This method can be performed by a long-tail image recognition device provided in an embodiment of the present invention. This device can be implemented in software and / or hardware and can be integrated into an electronic device, such as various user terminals or servers.

[0102] See also Figure 5 The method of the embodiment of the present invention specifically includes the following steps:

[0103] S310. Obtain a target image to be identified and a target image recognition model trained according to the model training method described in any embodiment of the present invention, wherein the sample categories used in the training process of the target image recognition model include the target category to which the target image belongs.

[0104] The target image can be understood as the image of the category to be identified (i.e., the target category). It should be noted that the target category is one of the sample categories involved in the model training process. Whether it is the head class or the tail class depends on the actual situation and is not specifically limited here.

[0105] S320: Input the target image into the target image recognition model, and obtain the target category according to the output result of the target image recognition model.

[0106] Among them, according to the above explanation, it can be seen that the target image recognition model can accurately identify both the head class and the tail class. Therefore, the target image recognition model can be used to identify the target image, thereby achieving accurate recognition of the target image, that is, achieving the effect of accurate recognition of the long-tail image.

[0107] Figure 6 This is a structural block diagram of a model training device provided in an embodiment of the present invention, which is used to execute the model training method provided in any of the above embodiments. This device and the model training method of each of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the model training device, please refer to the embodiment of the above model training method. Figure 6 The device may specifically include: a sample loss obtaining module 410, a first loss obtaining module 420, a second loss obtaining module 430 and a model training module 440.

[0108] The sample loss obtaining module 410 may be used to obtain a dataset that follows a long-tail distribution and input the dataset into an image recognition model to obtain a sample loss corresponding to each sample image based on the output of the image recognition model and the sample category corresponding to each sample image in the dataset.

[0109] A first loss obtaining module 420 is configured to determine the occurrence frequency of each sample category in the data set, and perform weighted processing on each sample loss based on the occurrence frequency to obtain a first loss;

[0110] A second loss obtaining module 430 may be configured to determine a sample gradient corresponding to each sample image, and perform weighted processing on each sample loss based on each sample gradient to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image;

[0111] The model training module 440 can be used to obtain a target loss based on the first loss and the second loss, and adjust the parameters in the image recognition model based on the target loss to train and obtain the target image recognition model.

[0112] Optionally, the first loss obtaining module 420 may include:

[0113] A sample quantity determination unit may be used to determine the number of samples corresponding to each sample category in the data set, wherein the sample quantity represents the number of sample images belonging to the corresponding sample category in the data set;

[0114] The occurrence frequency determination unit may be used to determine the occurrence frequency of each sample category in the data set according to the sample images corresponding to the sample category and the number of samples corresponding to each sample category.

[0115] Optionally, the first loss obtaining module 420 may include:

[0116] a category weight determination unit, configured to determine, for each sample category, a category weight of the sample category according to an occurrence frequency corresponding to the sample category, wherein the category weight is negatively correlated with the corresponding occurrence frequency;

[0117] The first loss obtaining unit is used to perform weighted processing on each sample loss according to the sample category to which each sample image belongs and the category weight corresponding to each sample category to obtain the first loss.

[0118] Optionally, the second loss obtaining module 430 may include:

[0119] a parameter gradient determining unit, configured to determine, for each sample image, the parameter gradient corresponding to each parameter of the sample image according to the sample loss of the sample image;

[0120] The sample gradient obtaining unit is used to obtain the sample gradient of the sample image according to the gradients of each parameter.

[0121] Optionally, the second loss obtaining module 430 may include:

[0122] a sample weight determination unit, configured to obtain, for each sample image, a sample weight of the sample image according to a sample gradient of the sample image, wherein the sample gradient is positively correlated with the corresponding difficulty level, and the sample gradient is positively correlated with the corresponding sample weight;

[0123] The second loss obtaining unit is used to perform weighted processing on the sample losses corresponding to each sample image according to the sample weights corresponding to each sample image to obtain the second loss.

[0124] On this basis, optionally, the second loss obtaining module 430 may further include:

[0125] The sample weight updating unit is used to perform normalization or smoothing processing on the sample weight after obtaining the sample weight of the sample image, and update the sample weight according to the obtained processing result.

[0126] Optionally, the target image recognition model is obtained by performing multiple rounds of training using a data set, and the model training module 440 may include:

[0127] A loss balance factor determination unit, configured to obtain a current iteration round and a total iteration round, and determine a loss balance factor based on the current iteration round and the total iteration round;

[0128] The target loss obtaining unit is used to perform weighted processing on the first loss and the second loss according to the loss balance factor to obtain the target loss.

[0129] On this basis, optionally, the loss balance factor is negatively correlated with the current iteration round, and the target loss is obtained as a unit, which is specifically used for:

[0130] Determining, according to a loss balancing factor, a first weight corresponding to the first loss and a second weight corresponding to the second loss, wherein the loss balancing factor is positively correlated with the first weight and negatively correlated with the second weight;

[0131] The first loss and the second loss are weighted according to the first weight and the second weight to obtain the target loss.

[0132] The model training device provided by an embodiment of the present invention can obtain a data set that obeys a long-tail distribution through a sample loss acquisition module, and input the data set into an image recognition model to obtain the sample loss corresponding to each sample image based on the output result of the image recognition model and the sample category corresponding to each sample image in the data set; on this basis, through the first loss acquisition module, the occurrence frequency corresponding to each sample category in the data set can be determined, and based on each occurrence frequency, each sample loss is weighted to obtain a first loss, which can be considered as a weighted loss based on the category frequency; and through the second loss acquisition module, the sample gradient corresponding to each sample image can be determined, and the sample gradient can characterize the difficulty of the image recognition model learning the corresponding sample image, and based on each sample gradient, each sample loss is weighted to obtain a second loss, which can be considered as a weighted loss based on the sample gradient; further, through the model training module, the target loss is obtained based on the first loss and the second loss, and the parameters in the image recognition model are adjusted according to the target loss to train the target image recognition model. During the model training process, the above-mentioned device can reduce the negative impact of inter-class imbalance and intra-class imbalance on the model by fusing the weighted loss based on category frequency and the weighted loss based on sample gradient. This can improve the accuracy of tail class recognition while maintaining the accuracy of head class recognition, thereby ensuring the accuracy of long-tail image recognition.

[0133] The model training device provided in the embodiment of the present invention can execute the model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0134] It is worth noting that in the embodiment of the above-mentioned model training device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0135] Figure 7 This is a structural block diagram of a long-tail image recognition device provided in an embodiment of the present invention. The device is used to execute the long-tail image recognition method provided in any of the above embodiments. The device and the long-tail image recognition method of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the long-tail image recognition device, please refer to the embodiment of the long-tail image recognition method. Figure 7 , the device specifically includes: a model acquisition module 510 and an image recognition module 520. Among them,

[0136] a model acquisition module 510 for acquiring a target image and a target image recognition model trained according to the model training method described in any embodiment of the present invention, wherein each sample category used in the training process of the target image recognition model includes the target category to which the target image belongs;

[0137] The image recognition module 520 is used to input the target image into the target image recognition model and obtain the target category according to the output result of the target image recognition model.

[0138] The long-tail image recognition device provided by the embodiment of the present invention achieves the effect of accurate recognition of long-tail images through the mutual cooperation of the model acquisition module and the image recognition module.

[0139] The long-tail image recognition device provided by the embodiment of the present invention can execute the long-tail image recognition method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0140] It is worth noting that in the embodiment of the above-mentioned long-tail image recognition device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0141] Figure 8 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0142] like Figure 8As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0143] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0144] The processor 11 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method or the long-tail image recognition method.

[0145] In some embodiments, the model training method or the long-tail image recognition method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method or the long-tail image recognition method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the model training method or the long-tail image recognition method in any other appropriate manner (e.g., by means of firmware).

[0146] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips or systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0151] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0152] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0153] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0154] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A model training method, characterized in that: include: Obtaining a data set that obeys a long-tail distribution, and inputting the data set into an image recognition model to obtain a sample loss corresponding to each sample image based on an output result of the image recognition model and a sample category corresponding to each sample image in the data set; Determine the occurrence frequency of each sample category in the data set, and perform weighted processing on each sample loss based on the occurrence frequency to obtain a first loss; Determining a sample gradient corresponding to each of the sample images, and performing weighted processing on each of the sample losses based on each of the sample gradients to obtain a second loss, wherein the sample gradient represents the difficulty of the image recognition model learning the corresponding sample image; A target loss is obtained according to the first loss and the second loss, and parameters in the image recognition model are adjusted according to the target loss to train a target image recognition model.

2. The method according to claim 1, characterized in that Determining the occurrence frequency of each sample category in the data set includes: Determine the number of samples corresponding to each sample category in the data set, wherein the number of samples represents the number of sample images belonging to the corresponding sample category in the data set; For each of the sample categories, the occurrence frequency of the sample category in the data set is determined according to the sample images corresponding to the sample category and the number of samples corresponding to each of the sample categories.

3. The method according to claim 1 or 2, characterized in that The step of performing weighted processing on each of the sample losses based on the occurrence frequencies to obtain a first loss includes: For each of the sample categories, determining a category weight of the sample category according to the occurrence frequency corresponding to the sample category, wherein the category weight is negatively correlated with the corresponding occurrence frequency; The sample losses are weighted according to the sample categories to which the sample images belong and the category weights corresponding to the sample categories to obtain a first loss.

4. The method according to claim 1, wherein Determining the sample gradient corresponding to each of the sample images includes: For each of the sample images, determining a parameter gradient corresponding to each of the parameters of the sample image according to the sample loss of the sample image; A sample gradient of the sample image is obtained according to each of the parameter gradients.

5. The method according to claim 1 or 4, characterized in that The step of performing weighted processing on the sample losses based on the gradients of the samples to obtain the second loss includes: For each of the sample images, obtaining a sample weight of the sample image according to the sample gradient of the sample image, wherein the sample gradient is positively correlated with the corresponding difficulty level, and the sample gradient is positively correlated with the corresponding sample weight; The sample losses corresponding to each of the sample images are weighted according to the sample weights corresponding to each of the sample images to obtain a second loss.

6. The method according to claim 5, characterized in that After obtaining the sample weight of the sample image, the method further includes: The sample weights are normalized or smoothed, and the sample weights are updated according to the obtained processing results.

7. The method according to claim 1, characterized in that The target image recognition model is obtained by performing multiple rounds of training using the data set, and the target loss is obtained according to the first loss and the second loss, including: Obtaining a current iteration round and a total iteration round, and determining a loss balance factor based on the current iteration round and the total iteration round; The first loss and the second loss are weighted according to the loss balance factor to obtain a target loss.

8. The method according to claim 7, characterized in that The loss balance factor is negatively correlated with the current iteration round, and the first loss and the second loss are weighted according to the loss balance factor to obtain the target loss, including: determining, according to the loss balancing factor, a first weight corresponding to the first loss and a second weight corresponding to the second loss, wherein the loss balancing factor is positively correlated with the first weight and negatively correlated with the second weight; The first loss and the second loss are weighted according to the first weight and the second weight to obtain a target loss.

9. A long-tail image recognition method, characterized in that: include: Obtaining a target image to be identified and a target image recognition model trained according to the model training method of any one of claims 1 to 8, wherein each sample category used in the training process of the target image recognition model includes the target category to which the target image belongs; The target image is input into the target image recognition model, and the target category is obtained according to the output result of the target image recognition model.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the model training method according to any one of claims 1 to 8, or the long-tail image recognition method according to claim 9.