Training Method, Device, Electronic Device and Storage Medium of Image Processing Model
By classifying and data enhancing the image sample set, the parameters of the image processing model are optimized, and the missed detection and false detection problems of the object detection model in the recognition of small targets and similar targets are solved, and the detection performance of the model is improved.
Patent Information
- Application Number
- CN202111057705.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-09-09
AI Technical Summary
When existing target detection models deal with small targets, seriously occluded targets and similar targets, there are problems of missed detection and missed detection, especially the high rate of missed detection and similar targets in small targets.
By classifying the image sample set, cross-training is performed using multiple classified data sets, combining data augmentation technology and a decreasing learning rate method, the parameters of the image processing model are optimized, including performing different data augmentation processing on each classified data set, and using the optimal model in each loop until the predetermined conditions are met.
The error detection rate of the target detection model is reduced, the recall, accuracy and mAP of the model are improved, and the robustness and visualization effect of the model are improved.
Smart Images

Figure CN113869376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, electronic device, and storage medium for training an image processing model applied in computer vision technology. Background Art
[0002] In recent years, various data augmentation methods have been proposed to continuously refresh the various indicators of the object detection model. However, for some special objects, there is still no good solution at present. There are still problems such as missed detection of small objects, missed detection of objects with severe occlusion, and misdetection of similar objects.
[0003] In view of the above problems, the present invention proposes a new method for training an image processing model based on misdetection of similar objects, and excellent results have been achieved through actual application tests. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for training an image processing model, which can reduce the misdetection rate of the object detection model and slightly improve the indicators of the model.
[0005] The purpose of this application is achieved by the following technical solutions:
[0006] In a first aspect, this application provides a method for training an image processing model, including: obtaining an image sample set matching the image processing model, determining the categories of the image samples in the image sample set, and classifying the image samples in the image sample set into multiple classification data sets of different categories according to the categories of the image samples; training the multiple classification data sets respectively based on the initial setup model of the image processing model, where training each of the classification data sets with a first predetermined number of times as a training cycle is regarded as one loop, and repeating the loop until a predetermined condition is met to continuously optimize the model parameters of the initial setup model, and obtaining an intermediate optimized model of the image processing model after meeting the predetermined condition; and training the entire image sample set based on the intermediate optimized model for a second predetermined number of times in a loop to further optimize the model parameters and obtain the final model of the image processing model.
[0007] The beneficial effect of this solution is that it can reduce the misdetection rate of the object detection model and slightly improve the indicators of the model, such as recall rate, precision, mAP, etc. Moreover, the final model obtained through final fine-tuning has been greatly improved in terms of visualization effect and indicators compared with ordinary training.
[0008] In some alternative embodiments, the continuously optimizing the model parameters of the initial setup model and / or the further optimizing the model parameters includes: in one cycle training process of the image processing model, at the end of the Nth cycle, selecting the optimal model from the model obtained in the Nth cycle and the models retained in the previous N - 1 cycles as the model retained in the Nth cycle and the model to be used in the (N + 1)th cycle, and wherein, N is a positive integer greater than 1. Selecting the optimal model from the model obtained in the Nth cycle and the models retained in the previous N - 1 cycles at the end of the Nth cycle includes: using the validation set that remains unchanged in the image sample set to determine the optimal model from the model obtained in the Nth cycle and the models retained in the previous N - 1 cycles.
[0009] The beneficial effect of this solution is that it can continuously optimize the model parameters of the initial setup model, and realizes using the optimal model in all previous cycles in each cycle.
[0010] In some alternative embodiments, the separately training the multiple classification data sets includes: in repeatedly executing the cycle, for each of the classification data sets, using different sub - data sets included in each of the classification data sets in different cycles for training to separately train for the first predetermined number of times.
[0011] The beneficial effect of this solution is that it can improve the diversity of data and further optimize the model parameters.
[0012] In some alternative embodiments, during the entire training process of the image processing model, performing data augmentation processing on the image sample set to increase the number of image samples in the image sample set. Performing the data augmentation processing on the image sample set includes performing different data augmentation processing on each of the classification data sets respectively to obtain multiple sub - data sets of each of the classification data sets respectively.
[0013] The beneficial effect of this solution is that it can further increase the sample capacity and optimize the model parameters.
[0014] In some alternative embodiments, the method further includes: keeping the hyperparameters used in the data augmentation processing unchanged; and configuring the learning rate of the image processing model to decrease as the number of cycles increases.
[0015] The beneficial effect of this solution is that it increases the robustness of the model to the detection target, reduces the false detection rate, and can slightly improve various indicators of the model, such as recall rate, precision, mAP, etc.
[0016] In some alternative embodiments, determining the categories of the image samples in the image sample set includes determining the category corresponding to each image sample according to the label information of each image sample in the image sample set. Determining the category corresponding to each image sample according to the label information of each image sample in the image sample set includes: when a certain image sample contains multiple pieces of label information, classifying the certain image sample into the category corresponding to the label information with the largest number among the multiple pieces of label information.
[0017] The beneficial effect of this solution is that by classifying the image samples and then training them by category, various indicators of the model are improved, and the model parameters are further optimized.
[0018] In a second aspect, the present application provides an image processing method, including: in response to an image processing request, obtaining a target image collected by a terminal; using an image processing model to identify a detection target in the target image, where the image processing model is trained based on the method described in the first aspect above.
[0019] The beneficial effect of this solution is that it can reduce the false detection rate of the target detection model and slightly improve the indicators of the model, such as recall rate, precision, mAP, etc. Moreover, the final model obtained after final fine-tuning has been greatly improved in terms of visualization effect and indicators compared with ordinary training.
[0020] In a third aspect, the present application provides a training device for an image processing model, including: an image sample set acquisition module for acquiring an image sample set matching the image processing model; a classification data set determination module for determining the categories of the image samples in the image sample set and classifying the image samples in the image sample set into different classification data sets according to the categories of the image samples; a classification data set training module for training the multiple classification data sets respectively based on an initial setting model of the image processing model, where training each of the classification data sets with a first predetermined number of times as a training cycle and repeating the cycle until a predetermined condition is met to continuously optimize the model parameters of the initial setting model and obtaining an intermediate optimized model of the image processing model after meeting the predetermined condition; and a final model training module for training the entire image sample set based on the intermediate optimized model for a second predetermined number of times in a loop, so as to further optimize the model parameters and obtain a final model of the image processing model.
[0021] The beneficial effects of this solution are as follows: it can reduce the false detection rate of the target detection model and slightly improve the model's metrics, such as recall rate, precision, mAP, etc. Moreover, the final model obtained through the final fine-tuning has achieved significant improvements in terms of visualization effects and metrics compared to ordinary training.
[0022] In some alternative embodiments, the classification dataset training module and / or the final model training module includes: a model selection unit, which, during one cycle of training of the image processing model, selects the optimal model among the models obtained in the Nth cycle and the models retained in the previous N - 1 cycles at the end of the Nth cycle as the model retained in the Nth cycle and the model to be used in the (N + 1)th cycle, where N is a positive integer greater than 1.
[0023] The beneficial effects of this solution are as follows: it can continuously optimize the model parameters of the initial setup model, and realizes using the optimal model in all previous cycles in each cycle.
[0024] In some alternative embodiments, the classification dataset training module includes: a sub-dataset training unit, which is used to, during the repeated execution of the cycle, for each classification dataset, use different sub-datasets among the multiple sub-datasets included in each classification dataset in different cycles for training to train respectively for the first predetermined number of times.
[0025] The beneficial effects of this solution are as follows: it can improve the diversity of data and further optimize the model parameters.
[0026] In some alternative embodiments, the device further includes: a data augmentation module, which is used to perform data augmentation processing on the image sample set during the entire training process of the image processing model to increase the number of image samples in the image sample set. The data augmentation module includes: a multi-augmentation processing unit, which is used to perform different data augmentation processing on each classification dataset respectively to obtain the multiple sub-datasets of each classification dataset respectively.
[0027] The beneficial effects of this solution are as follows: it can further increase the sample capacity and optimize the model parameters.
[0028] In some alternative embodiments, the device further includes: a hyperparameter setting module, which is used to keep the hyperparameters used in the data augmentation processing unchanged; and configure the learning rate of the image processing model to decrease as the number of cycles increases.
[0029] The beneficial effects of this solution are as follows: it increases the robustness of the model to the detection target, reduces the false detection rate, and can slightly improve various indicators of the model, such as recall rate, precision, mAP, etc.
[0030] Fourthly, the present application provides an image processing device, including: a target image acquisition module for acquiring a target image collected by a terminal; and an identification module for identifying a detection target in the target image by using an image processing model, wherein the image processing model is trained based on the device described in the third aspect.
[0031] Fifthly, the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above methods are implemented.
[0032] Sixthly, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any one of the above methods are implemented. Description of the Drawings
[0033] The present application will be further described below with reference to the drawings and embodiments.
[0034] Figure 1 is a schematic flowchart of a method for training an image processing model provided by an embodiment of the present application;
[0035] Figure 2 is a schematic flowchart of a method for training an image processing model provided by an embodiment of the present application;
[0036] Figure 3 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0037] Figure 4 is a schematic structural diagram of a device for training an image processing model provided by an embodiment of the present application;
[0038] Figure 5 is a schematic structural diagram of a device for training an image processing model provided by an embodiment of the present application;
[0039] Figure 6 is a schematic structural diagram of an image processing device provided by an embodiment of the present application;
[0040] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0041] Figure 8It is a schematic structural diagram of a program product for implementing a training method of an image processing model provided by an embodiment of the present application. Detailed implementation manners
[0042] Next, in combination with the accompanying drawings and specific implementation manners, the present application will be further described. It should be noted that, on the premise of no conflict, the following-described embodiments or technical features can be arbitrarily combined with each other to form new embodiments.
[0043] The training method of the image processing model according to an embodiment of the present invention includes steps S101 - S104. The following will refer to the attached Figure 1 Describe each step of the above training method of the image processing model of the embodiment of the present invention.
[0044] S101: Image sample set acquisition step.
[0045] Acquire an image sample set that matches the image processing model.
[0046] The image samples included in the image sample set can be, for example, three-object detection samples of no wearing, wearing a safety helmet, and wearing a helmet; or can be, for example, two-object detection samples of a person wearing a reflective vest and a person wearing other clothes.
[0047] S102: Classification data set determination step.
[0048] Determine the categories of each image sample in the image sample set, and classify each image sample in the image sample set into classification data sets of different categories according to the categories of the image samples.
[0049] Specifically, determine the category of the sample according to the label information of each image sample in the image sample set, and form a new set of all samples under each sample category to form a classification data set, so as to divide all the image samples in the image sample set into multiple classification data sets according to the categories of the samples.
[0050] In addition, there is also a situation where a certain image sample contains multiple sample information. At this time, when a certain image sample in the image sample set contains multiple sample information, set the category of the sample corresponding to the sample information with the largest number among the multiple sample information as the category of the certain image sample.
[0051] The image sample set includes the training set and the validation set described above. The above classification is only performed on the training set, and the validation set remains unchanged.
[0052] Step S103: Classification data set training step.
[0053] The initial setting model based on the image processing model trains the multiple classification data sets respectively. Taking the first predetermined number of times as the training cycle, each of the classification data sets is trained as a cycle, and the cycle is repeatedly executed until a predetermined condition is met, so as to continuously optimize the model parameters of the initial setting model, and obtain an intermediate optimized model of the image processing model after the predetermined condition is met.
[0054] Specifically, for example, when the multiple classification data sets are the first classification data set, the second classification data set, and the third classification data set respectively, first, train the first classification data set, for example, 20 times, which is the first predetermined number of times, then train the third classification data set for the first predetermined number of times, and finally train the second classification data set for the first predetermined number of times, taking this as a cycle. Among them, preferably, the training of the first, second, and third classification data sets does not follow a specific order.
[0055] Preferably, the first predetermined number of times, which is the number of cycles for each cycle, can be an integer greater than 10 and less than 50. Further preferably, for example, it can be set to 20.
[0056] In addition, it should be noted that the above-mentioned cycle of the predetermined number of cycles only applies to the multiple classification data sets into which the training set is divided, rather than the validation set. By keeping the validation set unchanged, the model can continue to be fine-tuned towards other targets after fitting a certain type of sample target. After multiple cycles, the model has the ability to recognize multiple targets, and can effectively grasp the characteristics of each category, preventing overfitting and misdetection.
[0057] Further preferably, the classification data set training step S103 may further include: a sub-data set training step.
[0058] In the sub-data set training step, each classification data set can be divided into multiple sub-data sets, and in the repeated execution of the cycle, for each classification data set, different sub-data sets included in each classification data set are used for training in different cycles, so as to train the first predetermined number of times respectively, thereby improving the diversity of data and further optimizing the model parameters.
[0059] Further preferably, the classification data set training step S103 may further include: a model selection step.
[0060] In the model selection step, in the cycle of the above-mentioned predetermined number of cycles of the image processing model, at the end of the Nth cycle, the model obtained in the Nth cycle is selected and compared with the optimal model among the models retained in the previous N - 1 cycles, as the model retained in the Nth cycle and the model to be used in the N + 1th cycle, and among them, N is a positive integer greater than 1.
[0061] Further, since the validation set remains unchanged, in the model selection step, the validation set can be used to determine the model obtained in N cycles and the optimal model among the models retained in the previous N-1 cycles.
[0062] Thus, it is possible to continuously optimize the model parameters of the initial setup model, and the optimal model in all previous cycles is used in each loop.
[0063] Further preferably, a predetermined condition can be set according to system settings or specific training situations. The predetermined condition can be, for example, a predetermined number of loops, or it can be, for example, until all sub-datasets of each classification dataset have been trained, or the validation metrics obtained based on the validation set no longer update, etc. Among them, the predetermined number of loops is, for example, 10-20, and 30-50 training cycles are trained in each loop. The predetermined number of loops and the number of training cycles in each loop can be determined according to the specific needs of the user. After meeting the above predetermined conditions, the obtained image processing model is called an intermediate optimized model. The reason why this model cannot be used as the final model is that during the loop, the feature extraction ability of the model will shift towards the dataset used in each loop. Therefore, the model in the last loop (intermediate optimized model) in the above predetermined number of loops cannot be used as the final model.
[0064] Step S104: Final model training step.
[0065] Based on the intermediate optimized model, train the entire image sample set for a second predetermined number of loops to further optimize the model parameters and obtain the final model of the image processing model.
[0066] Specifically, in the present invention, further, fine-tune the total dataset, that is, the entire image sample set, for a predetermined number of cycles, for example, 100 cycles, so as to complete the final fine-tuning of the model to obtain the final model. The above 100 cycle number is only exemplary. Preferably, by observing the loss curve, when the loss remains unchanged, or the optimal model no longer updates, the loop can be manually terminated. Generally, 100 cycles can meet the requirements.
[0067] This final model training step S104 may also include the above model selection step to further optimize the model parameters using the above model selection step.
[0068] The above-mentioned overall image sample set as the training object is only a specific example, and the present invention is not limited thereto. For example, it can be a set mixed with various sample images including detection targets of all categories. In view of having obtained the mixed image sample set including the above various detection targets in step S101, therefore, preferably, in step S104, the overall image sample set obtained in step S101 is used as the training object.
[0069] The final model obtained through the above final fine-tuning has achieved a significant improvement in visualization effects and metrics compared to ordinary training.
[0070] According to the training method of the image processing model of the above embodiment of the present invention, the false detection rate of the target detection model can be reduced, and the metrics of the model, such as recall rate, precision, mAP, etc., can be slightly improved.
[0071] Furthermore, as Figure 2 shown, the training method of the image processing model according to the embodiment of the present invention includes step S201. The following will refer to the attached Figure 2 to describe steps S201-S202 of the embodiment of the present invention.
[0072] Step S201: Data augmentation step.
[0073] During the entire training process of the image processing model, data augmentation processing is performed on the image sample set to increase the number of image samples in the image sample set.
[0074] Further preferably, this data augmentation step includes: multi-augmentation processing step S501.
[0075] In this multi-augmentation processing step S501, different data augmentation processing can be performed on each classification data set respectively to obtain the above-mentioned multiple sub-data sets of each classification data set respectively.
[0076] Examples of the above data augmentation processing include but are not limited to mosaic, mixup, random flipping, random cropping, random changing of brightness and saturation, etc.
[0077] After obtaining different multiple sub-data sets through the above data augmentation processing, different sub-data sets can be used for training in each loop of step S103 to further increase the sample capacity and optimize the model parameters.
[0078] Step S202: Hyperparameter setting step.
[0079] The hyperparameters in the data augmentation processing are set to remain unchanged throughout the loop process, which can increase the diversity of images and the robustness of the model.
[0080] In addition, as an important hyperparameter, the learning rate is configured to decrease as the number of cycles increases. For example, the initial value of the learning rate can be set to 0.0032, and the learning rate is decreased by 0.00015 after each cycle. After, for example, 20 cycles, the learning rate becomes 0.0002.
[0081] And for example, during each cycle, the learning rate can be kept continuously decaying, such as cosine decay.
[0082] According to the above embodiments of the present invention, the robustness of the model to the detection target is increased, the false detection rate is reduced, and the various indicators of the model can be slightly improved, such as recall rate, precision, mAP, etc.
[0083] The above describes the exemplary embodiments of the present invention. Below, specific examples will be given to better understand the training method of the image processing model according to the embodiments of the present invention.
[0084] [Example 1]
[0085] Below, a safety helmet detection model applying the training method of the image processing model of the present invention will be used as an example for illustration.
[0086] During the training process of the safety helmet detection model, first, in step S101, a safety helmet image sample set including each detection target as a detection sample is obtained, where the detection targets include not wearing, wearing a safety helmet, and wearing a helmet.
[0087] Then, in step S102, the category of each image sample in the safety helmet image sample set is determined, and the safety helmet image sample set is divided according to the label information (the label information indicating not wearing, the label information indicating wearing a safety helmet, and the label information indicating wearing a helmet) to finally form a not wearing data set, a wearing a safety helmet data set, and a wearing a helmet data set. Among them, the safety helmet image sample set includes a safety helmet training set and a safety helmet validation set. The above classification is only performed on the safety helmet training set, and no classification processing is performed on the safety helmet validation set.
[0088] In special cases, for example, in the recognition image of a detection target, there are both people wearing safety helmets and people not wearing safety helmets at the same time. At this time, the detection target includes multiple label information, that is, several label information of people wearing safety helmets and several label information of people not wearing safety helmets. At this time, the detection target is classified according to the label information with a larger quantity. For example, if the detection target contains 3 people not wearing safety helmets, 2 people wearing safety helmets, and 1 person wearing a helmet, then the detection target is classified into the not wearing data set.
[0089] Subsequently, in step S103, using, for example, a loop training script, the initial setup model of the image processing model is used to train the non-wearing data set, the safety helmet wearing data set, and the safety helmet wearing data set respectively. For example, the non-wearing data set is trained with 20 times as a training cycle, then the safety helmet wearing data set is trained with 20 times as a training cycle, and then the safety helmet wearing data set is trained with 20 times as a training cycle as the first loop. The above loop is repeatedly executed for a predetermined number of times, for example, 10 - 50 times. Among them, in each loop, taking the non-wearing data set as an example, different sub-data sets of the non-wearing data set can be used for training (this sub-data set can be obtained through the data augmentation processing described above). In addition, the training of the classification data sets in each loop is not in a specific order. For example, in the second loop, the better model among the model obtained in the first loop and the initial training model can be selected first as the model used in this loop (using, for example, a validation set to determine which one is the better model). Then, the safety helmet wearing data set is trained with 20 times as a training cycle first, then the non-wearing data set is trained with 20 times as a training cycle, and then the safety helmet wearing data set is trained with 20 times as a training cycle. And during this training process, sub-data sets different from those in the first loop of each classification data set are used.
[0090] In addition, during the above-mentioned loop of the predetermined number of times, various hyperparameters are reasonably configured, such as the hyperparameters in the data augmentation method described above and the learning rate, etc. Preferably, as described above, the hyperparameters in the data augmentation processing are set to remain unchanged, and the learning rate is set to continuously decay.
[0091] After completing the above-mentioned loop of the predetermined number of times, an intermediate optimized model is obtained. Subsequently, step S104 is entered. In step S104, the intermediate optimized model obtained through the loop of the predetermined number of times is used to train the entire safety helmet training set, and the loop is, for example, 100 cycles for fine-tuning, so as to obtain the final model of the image processing model.
[0092] The entire above-mentioned safety helmet training set being used as the training object is only a specific example, and the present invention is not limited thereto. For example, it can be a set mixed with various sample images with detection targets of non-wearing, wearing a safety helmet, and wearing a safety helmet. In view of the fact that a mixed image sample set including the above various detection targets has been obtained in step S101, therefore, preferably, in step S104, the entire image sample set obtained in step S101 is used as the training object.
[0093] Traditionally, in the training method using existing image processing models, for example, a safety helmet hanging on the wall may be recognized as being worn, resulting in a relatively high false detection rate. However, the final model obtained through the above process of the training method of the image processing model of the present invention, due to the use of cross-training on the data set divided by category, can reduce the false detection rate and slightly improve various indicators of the model.
[0094] [Example 2]
[0095] The following will be described by taking the reflective vest detection model applying the training method of the image processing model of the present invention as an example.
[0096] In the reflective vest detection model, the detection targets include people wearing reflective vests and people not wearing reflective vests (people wearing other clothes).
[0097] First, in step S101, two-object detection samples of people wearing reflective vests and people not wearing reflective vests are obtained as the reflective vest image sample set.
[0098] Then, in step S102, the reflective vest image sample set is divided according to the label information (label information indicating reflective vests, label information indicating other clothes) to finally form a reflective vest data set and an other clothes data set. Among them, the reflective vest image sample set includes a reflective vest training set and a reflective vest validation set.
[0099] Among them, the above classification is only performed on the reflective vest training set, and no classification processing is performed on the reflective vest validation set.
[0100] In special cases, for example, in the recognition image of a detection target, there are both people wearing reflective vests and people wearing other clothes at the same time. At this time, the detection target includes multiple label information, that is, several reflective vest label information and several other clothes label information. At this time, the detection target is classified based on the label information with a larger quantity. For example, if the detection target contains 3 people wearing reflective vests and 1 person wearing other clothes, then the detection target is classified into the reflective vest data set.
[0101] Subsequently, in step S103, for example, using a loop training script, the initial setting model of the image processing model is used to train the reflective vest data set and the other clothes data set respectively. Specifically, for example, taking 20 times as the training cycle for the reflective vest data set, and then taking 20 times as the training cycle for the other clothes data set as the first loop, and then repeating the loop for a predetermined number of times. Among them, in each loop, taking the non-wearing data set as an example, different sub-data sets of the non-wearing data set can be used for training (this sub-data set can be obtained through the data augmentation processing described above).
[0102] In addition, the training of the classified datasets in each loop is not in any particular order. For example, in the second loop, the better model among the model obtained in the first loop and the initial training model can be selected first as the model to be used in this loop (judging which one is the better model by using, for example, the validation set). Subsequently, first train other clothing datasets with a training cycle of 20 times, and then train the reflective clothing dataset with a training cycle of 20 times. And during this training process, sub-datasets different from those in the first loop are adopted for each classified dataset.
[0103] In addition, during the process of the above-mentioned loop for a predetermined number of times, various hyperparameters are reasonably configured, such as the hyperparameters in the data augmentation method described above and the learning rate, etc. Preferably, as described above, the hyperparameters in the data augmentation process are set to remain unchanged, and the learning rate is set to continuously decay.
[0104] After completing the above-mentioned loop for a predetermined number of times, an intermediate optimized model is obtained. Subsequently, enter step S104. In step S104, the entire safety helmet training set is trained using the intermediate optimized model obtained through the loop for a predetermined number of times. When it is found that the optimal model no longer updates during the loop process, the loop can be manually terminated.
[0105] Traditionally, using the training method of existing image processing models, for example, it is easy to identify a person wearing bright clothes as a person wearing a reflective vest, with a relatively high false detection rate. However, the final model obtained through the above process of the training method of the image processing model of the present invention, due to the use of cross-training on the datasets divided by category, can reduce the false detection rate and can slightly improve various indicators of the model.
[0106] The above describes the embodiments and specific application examples of the training method of the image processing model of the embodiments of the present invention. The above embodiments and specific application examples are only used for listing to facilitate understanding of the present invention, and are not intended to limit the scope of the present invention in any way. And it is obvious to those skilled in the art that various changes and deformations can be made within the scope of the present invention.
[0107] According to the training method of the image processing model of the embodiments of the present invention, due to the use of cross-training on the datasets divided by category, the robustness of the model to the detection target is increased, the false detection rate is reduced, and various indicators of the model, such as recall rate, precision, mAP, etc., can be slightly improved.
[0108] See Figure 3, an embodiment of the present application further provides an image processing method, including: Step S601, in response to an image processing request, obtain a target image collected by a terminal; Step S602, use an image processing model to identify a detection target in the target image, where the image processing model is trained based on the training method of the image processing model described above.
[0109] The specific implementation manner is the same as the implementation manner and the achieved technical effects described in the embodiment of the above method, and some contents will not be elaborated.
[0110] Among them, the above terminal may be a camera device, a communication device, etc. that can collect images.
[0111] See Figure 4 , an embodiment of the present application further provides a training device for an image processing model. The specific implementation manner is the same as the implementation manner and the achieved technical effects described in the embodiment of the above method, and some contents will not be elaborated.
[0112] The training device for the image processing model according to the embodiment of the present invention includes modules 101-104. The following will refer to the attached Figure 4 Describe each module of the above training device for the image processing model according to the embodiment of the present invention.
[0113] Module 101: Image sample set acquisition module.
[0114] The image sample set acquisition module 101 is used to acquire an image sample set matching the image processing model.
[0115] The image samples included in the image sample set may be, for example, three-target detection samples of no wearing, wearing a safety helmet, and wearing a helmet; or may be, for example, two-target detection samples of a person wearing a reflective vest and a person wearing other clothes.
[0116] Module 102: Classification data set determination module.
[0117] The classification data set determination module 102 is used to determine the category of each image sample in the image sample set, and classify each image sample in the image sample set into different classification data sets according to the category of the image sample.
[0118] Specifically, the classification data set determination module 102 is used to determine the category of the sample according to the label information of each image sample in the image sample set, and form a new set of all samples under each sample category to form a classification data set, so as to divide all the image samples in the image sample set into multiple classification data sets according to the category of the sample.
[0119] In addition, there is a situation where a certain image sample contains multiple sample information. At this time, when a certain image sample in the image sample set contains multiple sample information, the classification data set determination module 102 is used to set the category of the sample corresponding to the sample information with the largest number among the multiple sample information as the category of the certain image sample.
[0120] The image sample set includes the training set and the validation set described above. The above classification is only performed on the training set, and the validation set remains unchanged.
[0121] Module 103: Classification data set training module.
[0122] The classification data set training module 103 is used to train the multiple classification data sets respectively based on the initial setting model of the image processing model. Taking the first predetermined number of times as the training cycle, each of the classification data sets is trained as one cycle, and the cycle is repeated for a predetermined number of cycles to continuously optimize the model parameters of the initial setting model, and after passing through the predetermined number of cycles, an intermediate optimized model of the image processing model is obtained.
[0123] Specifically, for example, when the multiple classification data sets are the first classification data set, the second classification data set, and the third classification data set respectively, first, the classification data set training module 103 trains the first classification data set, for example, 20 times, then trains the third classification data set, for example, the first predetermined number of times, and finally trains the second classification data set the first predetermined number of times, and this is used as one cycle. Among them, preferably, the training of the first, second, and third classification data sets does not follow a specific order.
[0124] Preferably, the first predetermined number of times as the number of cycles for each cycle can be an integer greater than 10 and less than 50. Further preferably, for example, it can be set to 20.
[0125] In addition, it should be pointed out that the above cycle of the predetermined number of cycles only applies to the multiple classification data sets divided from the training set, and not to the validation set. By keeping the validation set unchanged, the model can continue to be fine-tuned towards other targets after fitting to a certain type of sample target. After multiple cycles, the model has the ability to recognize multiple targets, and can effectively grasp the characteristics of each category, preventing overfitting and false detection.
[0126] Further preferably, the classification data set training module 103 may further include a sub-data set training unit 301.
[0127] The sub - dataset training unit 301 is used to divide each classification dataset into multiple sub - datasets, and in the repeated execution of the loop, for each classification dataset, in different loops, different sub - datasets among the multiple sub - datasets included in each classification dataset are used for training, so as to train the first predetermined number of times respectively, thereby improving the diversity of data and further optimizing the model parameters.
[0128] Further preferably, the classification dataset training module 103 may further include: a model selection unit. The model selection unit is used to, in the loop of the above - mentioned predetermined number of cycles of the image processing model, select, at the end of the Nth cycle, the model obtained in the Nth cycle and the optimal model among the models retained in the previous N - 1 cycles as the model retained in the Nth cycle and the model to be used in the (N + 1)th cycle, where N is a positive integer greater than 1.
[0129] Furthermore, since the validation set remains unchanged, the model selection unit can use the validation set to judge the model obtained in N cycles and the optimal model among the models retained in the previous N - 1 cycles.
[0130] Thus, it is possible to continuously optimize the model parameters of the initially set model, and achieve using the optimal model in all previous cycles in each loop.
[0131] Further preferably, a predetermined condition can be set according to system settings or specific training situations. The predetermined condition can be, for example, a predetermined number of cycles, or it can be, for example, until all sub - datasets of each classification dataset have been trained, etc. After meeting the above - mentioned predetermined condition, the obtained image processing model is called an intermediate optimized model. The reason why this model cannot be used as the final model is that during the loop, the feature extraction ability of the model will shift towards the dataset used in each loop. Therefore, the model in the last cycle of the above - mentioned predetermined number of cycles (intermediate optimized model) cannot be used as the final model.
[0132] Module 104: Final model training module.
[0133] The final model training module 104 is used to train the entire image sample set based on the intermediate optimized model for a second predetermined number of cycles, so as to further optimize the model parameters and obtain the final model of the image processing model.
[0134] Specifically, in the present invention, further, the entire general dataset, i.e., the overall image sample set, is fine-tuned for a predetermined number of cycles, for example, 100 cycles, so as to complete the final fine-tuning of the model and obtain the final model. The above-mentioned 100 cycles are only exemplary. Preferably, by observing the loss curve, when the loss remains unchanged or the optimal model no longer updates, the loop can be manually terminated. Generally, 100 cycles can meet the requirements.
[0135] The final model training module 104 may further include the above-mentioned model selection unit to further optimize the model parameters by using the model selection unit.
[0136] Taking the overall above-mentioned image sample set as the training object is only a specific example, and the present invention is not limited thereto. For example, it may be a set mixed with various sample images including detection targets of all categories. In view of having obtained the mixed image sample set including the above various detection targets in module 101, therefore, preferably, module 104 takes the overall image sample set obtained by module 101 as the training object.
[0137] The final model obtained through the above-mentioned final fine-tuning has made significant improvements in terms of visualization effect and metrics compared with ordinary training, etc.
[0138] According to the training method of the image processing model of the above-mentioned embodiment of the present invention, the false detection rate of the target detection model can be reduced, and the metrics of the model, such as recall rate, precision, mAP, etc., can be slightly improved.
[0139] Further, as Figure 5 shown, the training device of the image processing model according to the embodiment of the present invention further includes modules 201-202. The following will refer to the attached Figure 5 description of modules 201-202 of the embodiment of the present invention.
[0140] Module 201: Data augmentation module.
[0141] The data augmentation module 201 is used to perform data augmentation processing on the image sample set during the entire training process of the image processing model to increase the number of the image samples in the image sample set.
[0142] Further preferably, the data augmentation step includes: a multi-augmentation processing unit 501.
[0143] The multiple enhancement processing units 501 are used to perform different data enhancement processes on each classification data set respectively, so as to obtain the multiple sub-data sets described above for each of the classification data sets. Examples of the above data enhancement processes include but are not limited to mosaic, mixup, random flipping, random cropping, random changing of brightness and saturation, etc.
[0144] After obtaining different multiple sub-data sets through the above data enhancement processes, different sub-data sets can be used for training in each loop of module 103, so as to further increase the sample capacity and optimize the model parameters.
[0145] Module 202: Hyperparameter setting module.
[0146] The hyperparameter setting module 202 is used to set the hyperparameters in the data enhancement process to remain unchanged throughout the loop process, which can increase the diversity of images and the robustness of the model.
[0147] In addition, as an important hyperparameter, the hyperparameter setting module 202 configures the learning rate to decrease as the number of loops increases. For example, the initial value of the learning rate can be set to 0.0032, and the learning rate is decreased by 0.00015 after each loop ends. After experiencing, for example, 20 loops, the learning rate becomes 0.0002.
[0148] And for example, the hyperparameter setting module 202 can keep the learning rate continuously decaying during each loop process, such as cosine decay.
[0149] According to the above embodiments of the present invention, the robustness of the model for the detection target is increased, the false detection rate is reduced, and the various indicators of the model can be slightly improved, such as recall rate, precision, mAP, etc.
[0150] See Figure 6 , the embodiments of the present application also provide an image processing device 600, and its specific implementation manner is the same as the implementation manner and the achieved technical effects recorded in the embodiments of the above image processing method, and some contents will not be elaborated.
[0151] The image processing device 600 includes: a target image acquisition module 601, configured to acquire a target image collected by a terminal in response to an image processing request; an identification module 602, configured to identify a detection target in the target image by using an image processing model, wherein the image processing model is trained based on the training device of the image processing model described above.
[0152] See Figure 7, An embodiment of the present application also provides an electronic device 200, which includes at least one memory 210, at least one processor 220, and a bus 230 connecting different platform systems.
[0153] The memory 210 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 211 and / or a cache memory 212, and may further include a read-only memory (ROM) 213.
[0154] Among them, the memory 210 also stores a computer program, which can be executed by the processor 220, so that the processor 220 executes the steps of the real-time video processing method in the embodiment of the present application. The specific implementation manner is consistent with the implementation manner and the achieved technical effects recorded in the embodiment of the above real-time video processing method, and some contents will not be elaborated here.
[0155] The memory 210 may further include a utility 214 having at least one program module 215. Such program modules 215 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0156] Correspondingly, the processor 220 can execute the above computer program and can also execute the utility 214.
[0157] The bus 230 may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures.
[0158] The electronic device 200 can also communicate with one or more external devices 240, such as a keyboard, a pointing device, a Bluetooth device, etc., and can also communicate with one or more devices capable of interacting with the electronic device 200, and / or communicate with any device (such as a router, a modem, etc.) that enables the electronic device 200 to communicate with one or more other computing devices. This communication can be carried out through an input / output interface 250. And, the electronic device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 260. The network adapter 260 can communicate with other modules of the electronic device 200 through the bus 230. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 200, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0159] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, which, when executed, implements the steps of the real-time video processing method in the embodiment of the present application. The specific implementation manner is consistent with the implementation manner and the achieved technical effects described in the embodiment of the real-time video processing method above, and some content will not be repeated here.
[0160] Figure 8 Fig. 4 shows a program product 300 provided in this embodiment for implementing the above real-time video processing method. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product 300 of the present invention is not limited to this. In the present application, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. The program product 300 can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0161] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0162] This application is described from the viewpoints of usage purpose, efficacy, advancement, and novelty, and has met the requirements of functionality enhancement and usage emphasized by the patent law. The above description and the accompanying drawings of this application are only preferred embodiments of this application and do not limit this application thereby. Therefore, all those that are approximate or identical to the structure, device, features, etc. of this application, that is, all equivalent substitutions or modifications made according to the scope of the patent application of this application, shall fall within the scope of protection of the patent application of this application.
Claims
1. A training method for an image processing model, comprising: Obtaining an image sample set that matches the image processing model; Determining the categories of the respective image samples in the image sample set, and classifying the respective image samples in the image sample set into a plurality of classification data sets of different categories according to the categories of the image samples; Respectively training the plurality of classification data sets based on an initial setting model of the image processing model, wherein each of the classification data sets is trained for a first predetermined number of times as one cycle, and the cycle is repeatedly executed until a predetermined condition is satisfied, so as to continuously optimize the model parameters of the initial setting model, and obtaining an intermediate optimized model of the image processing model after the predetermined condition is satisfied; and Based on the intermediate optimized model, training the entire image sample set for a second predetermined number of times in a cycle, so as to further optimize the model parameters and obtain a final model of the image processing model; The continuously optimizing the model parameters of the initial setting model and / or the further optimizing the model parameters includes: During a cyclic training process of the image processing model, at the end of the Nth cycle, selecting the model obtained in the Nth cycle and the optimal model among the models retained in the previous N - 1 cycles as the model retained in the Nth cycle and the model to be used in the (N + 1)th cycle, and wherein, N is a positive integer greater than 1.
2. The method according to claim 1, characterized in that, The selecting, at the end of the Nth cycle, the model obtained in the Nth cycle and the optimal model among the models retained in the previous N - 1 cycles includes: Using a validation set that remains unchanged in the image sample set to judge the model obtained in the Nth cycle and the optimal model among the models retained in the previous N - 1 cycles.
3. The method according to any one of claims 1 to 2, characterized in that, The respectively training the plurality of classification data sets includes: During the repeated execution of the cycle, for each of the classification data sets, different sub-data sets among the plurality of sub-data sets included in each of the classification data sets are used for training in different cycles, so as to respectively train for the first predetermined number of times.
4. The method according to claim 3, wherein The method further includes: During the entire training process of the image processing model, performing data augmentation processing on the image sample set to increase the number of the image samples in the image sample set.
5. The method according to claim 4, wherein The performing the data augmentation processing on the image sample set includes: Performing different data augmentation processing on each of the classification data sets respectively to obtain the plurality of sub-data sets of each of the classification data sets.
6. The method according to claim 5, characterized in that, The method further includes: Keeping the hyperparameters used in the data augmentation processing unchanged; and configuring the learning rate of the image processing model to decrease as the number of cycles increases.
7. The method according to claim 6, characterized in that, The determining the categories of the respective image samples in the image sample set includes: Determining the category corresponding to each image sample according to the label information of the respective image samples in the image sample set.
8. The method according to claim 7, wherein The determining the category corresponding to each image sample according to the label information of the respective image samples in the image sample set includes: When a certain image sample contains multiple pieces of the tag information, classify the certain image sample into the category corresponding to the tag information with the largest number among the multiple pieces of the tag information.
9. An image processing method, comprising: In response to an image processing request, obtain a target image collected by a terminal; Use an image processing model to identify a detection target in the target image, where the image processing model is trained based on the method according to any one of claims 1 to 8.
10. A training device for an image processing model, comprising: An image sample set acquisition module, which is configured to acquire an image sample set matching the image processing model; A classification data set determination module, which determines the category of each image sample in the image sample set, and classifies each image sample in the image sample set into different classification data sets of different categories according to the category of the image sample; A classification data set training module, which is configured to respectively train a plurality of classification data sets based on an initial setting model of the image processing model, where, taking a first predetermined number of times as a training cycle, respectively training each of the classification data sets as one cycle, and repeating the execution of the cycle until a predetermined condition is satisfied, to continuously optimize the model parameters of the initial setting model, and obtain an intermediate optimized model of the image processing model after the predetermined condition is satisfied; and A final model training module, which is configured to train the entire image sample set based on the intermediate optimized model for a second predetermined number of times in a cycle, so as to further optimize the model parameters and obtain a final model of the image processing model; The classification data set training module and / or the final model training module includes: A model selection unit, during a cycle training process of the image processing model, at the end of the Nth cycle, select the model obtained in the Nth cycle and the optimal model among the models retained in the previous N - 1 cycles as the model retained in the Nth cycle and the model to be used in the N + 1th cycle, and where N is a positive integer greater than 1.
11. The device according to claim 10, characterized in that, The classification data set training module includes: A sub-data set training unit, which is configured to, during the repeated execution of the cycle, for each classification data set, use different sub-data sets included in each classification data set for training in different cycles to respectively train the first predetermined number of times.
12. The device according to claim 11, characterized in that, The device further includes: A data augmentation module, which is configured to perform data augmentation processing on the image sample set during the entire training process of the image processing model to increase the number of image samples in the image sample set.
13. The device according to claim 12, characterized in that, The data augmentation module includes: A multi-augmentation processing unit, which is configured to respectively perform different data augmentation processing on each classification data set to respectively obtain multiple sub-data sets of each classification data set.
14. The device according to claim 13, wherein The device further includes: A hyperparameter setting module, which is used to keep the hyperparameters used in the data augmentation process unchanged; and configure the learning rate of the image processing model to decrease as the number of cycles increases.
15. An image processing device, comprising: A target image acquisition module, which is used to acquire a target image collected by a terminal; And An identification module, which is used to identify a detection target in the target image by using an image processing model, wherein The image processing model is trained based on the device according to any one of claims 10 to 14.
16. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, the steps of the method according to any one of claims 1-8 are implemented.
17. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Image recognition model training method and device, electronic equipment and storage medium
CN111242217A