Image Recognition Method and Device Based on Data Augmentation Strategy Selection
By selecting data enhancement strategies from the training samples selected from the training image data, the problem of high calculation cost and high real-time difficulty in the prior art is solved, and the effect of quickly determining the optimal enhancement strategy is achieved, and the efficiency of the image data enhancement method is improved.
Patent Information
- Application Number
- CN202111273249.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the prior art, the calculation cost of choosing the optimal data enhancement strategy is high and the real-time difficulty is high, especially when the amount of data is large.
By selecting training samples in preset proportions from the training image data, data augmentation is performed using at least two data augmentation strategies, and the optimal data augmentation strategy is selected based on the augmentation results, so as to quickly determine the data augmentation strategy.
This method can quickly determine the data enhancement strategy, improve the efficiency of the image data enhancement method, reduce the calculation cost, and is suitable for training image data with large amounts of data.
Smart Images

Figure CN113971644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image recognition method and device based on data augmentation strategy selection. Background Art
[0002] Deep learning tasks usually require a large amount of data. In order to obtain such a large amount of data, some computer vision methods can be used to perform some transformations on the existing picture data to obtain equivalent new data that is highly correlated with the original data. This method is called data augmentation.
[0003] In related technologies, when selecting an augmentation strategy for picture data, the reinforcement learning method is mostly used to search for the optimal strategy. However, it usually requires tens of thousands of iterations to converge to obtain the optimal result, resulting in a relatively high implementation difficulty and computational cost for this method. Summary of the Invention
[0004] The present invention provides an image recognition method and device based on data augmentation strategy selection to solve the defect in the prior art that the computational cost of the optimal data augmentation strategy is relatively high and the real-time difficulty is relatively large due to a large amount of data, so as to quickly determine the data augmentation strategy and improve the usage efficiency of the image data augmentation method.
[0005] The present invention provides an image recognition method based on data augmentation strategy selection, including: inputting the acquired picture data into a target recognition model to obtain a target prediction result output by the target recognition model; wherein, the target recognition model is trained based on the data augmentation result obtained by performing data augmentation on the training image data according to a pre-selected data augmentation strategy and the target recognition result corresponding to the training image data, and the pre-selected data augmentation strategy is obtained based on training samples selected from the training image data according to a preset ratio.
[0006] According to an image recognition method based on data augmentation strategy selection provided by the present invention, training the target recognition model includes: selecting training samples from the acquired training image data according to a preset ratio; performing data augmentation on the training samples respectively by using at least two data augmentation strategies to obtain data augmentation results corresponding to the respective data augmentation strategies; selecting a data augmentation strategy from the at least two data augmentation strategies according to the respective data augmentation results, and performing data augmentation on the training image data by using the selected data augmentation strategy; using the training image data after the data augmentation as input data for training, and using the target recognition result corresponding to the training image as a label to train a network to be trained, so as to obtain a target recognition model for generating a target prediction result of an image to be recognized.
[0007] An image recognition method based on data augmentation strategy selection provided by the present invention, the selecting of the data augmentation strategy from the at least two data augmentation strategies according to each of the data augmentation results includes: obtaining loss values corresponding to the brightness, contrast and texture information according to each of the data augmentation results; scoring according to the loss values corresponding to the data augmentation strategies to obtain the highest scoring result; and selecting the corresponding data augmentation strategy according to the highest scoring result.
[0008] An image recognition method based on data augmentation strategy selection provided by the present invention, the scoring according to the loss values corresponding to the data augmentation strategies to obtain the highest scoring result includes: obtaining the iteration numbers corresponding to each of the data augmentation results when the minimum loss value is less than a preset threshold based on the minimum loss value and the preset threshold; comparing all the iteration numbers and using the minimum iteration number as the highest scoring result; or comparing the loss values corresponding to each of the data augmentation strategies to obtain the minimum loss value as the highest scoring result.
[0009] An image recognition method based on data augmentation strategy selection provided by the present invention, after selecting the data augmentation strategy with the highest corresponding scoring result according to the scoring result, further includes: using the data augmentation result corresponding to the data augmentation strategy with the highest corresponding scoring result as the input data for training, and training the network to be trained.
[0010] An image recognition method based on data augmentation strategy selection provided by the present invention, after comparing the loss values corresponding to each of the data augmentation strategies, further includes: using the training samples before data augmentation corresponding to the remaining loss values as the input data for training based on the minimum loss value, and training the network to be trained.
[0011] An image recognition method based on data augmentation strategy selection provided by the present invention, the data augmentation strategy includes at least one of cropping, rotation, translation, flipping, sharpening, illumination and occlusion.
[0012] The present invention also provides an image recognition device based on data augmentation strategy selection, including: a target recognition module, inputting the acquired picture data into a target recognition model to obtain a target prediction result output by the target recognition model; wherein, the target recognition model is trained based on the data augmentation result obtained by performing data augmentation on the training image data according to a pre-selected data augmentation strategy and the target recognition result corresponding to the training image data, and the pre-selected data augmentation strategy is obtained based on the training samples selected from the training image data according to a preset ratio.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the image recognition method selected based on the data augmentation strategy as described above are implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image recognition method selected based on the data augmentation strategy as described above are implemented.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image recognition method selected based on the data augmentation strategy as described above are implemented.
[0016] The image recognition method and device selected based on the data augmentation strategy provided by the present invention select a data augmentation strategy through training samples selected from training image data, so as to quickly determine the data augmentation strategy with fewer data samples, make the obtained optimal augmentation strategy applicable to training image data with a large amount of data, improve the selection time of the data augmentation strategy, reduce the calculation amount. In addition, the selected data augmentation strategy can be applied to image data of the same type of target objects for data augmentation, so as to improve the usage efficiency of the image data augmentation method. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic architecture diagram of the image recognition method selected based on the data augmentation strategy provided by the present invention;
[0019] Figure 2 is a schematic training process diagram of the target recognition model provided by the present invention;
[0020] Figure 3 is a schematic structural diagram of the image recognition device selected based on the data augmentation strategy provided by the present invention;
[0021] Figure 4 is a schematic structural diagram of the training module provided by the present invention;
[0022] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the protection scope of the present invention.
[0024] Figure 1 The architecture diagram of an image recognition method based on data augmentation strategy selection according to the present invention is shown. The method includes:
[0025] Input the obtained picture data into the target recognition model to obtain the target prediction result output by the target recognition model;
[0026] Among them, the target recognition model is trained based on the data augmentation result obtained by performing data augmentation on the training image data based on the pre-selected data augmentation strategy and the target recognition result corresponding to the training image data. The pre-selected data augmentation strategy is based on the training samples selected from the training image data at a preset ratio.
[0027] It should be noted that the pre-selected data augmentation strategy is selected by performing data augmentation on the training samples selected from the training image data using at least two data augmentation strategies and according to the data augmentation results.
[0028] Figure 2 The training process diagram of the target recognition model according to the present invention is shown. The training method includes:
[0029] S11, select training samples from the obtained training image data at a preset ratio;
[0030] S12, perform data augmentation on the training samples respectively using at least two data augmentation strategies to obtain the data augmentation results corresponding to the respective data augmentation strategies;
[0031] S13, select a data augmentation strategy from at least two data augmentation strategies according to the respective data augmentation results, and perform data augmentation on the training image data using the selected data augmentation strategy;
[0032] S14, use the training image data after data augmentation as the input data for training, use the target recognition result corresponding to the training image as the label, and train the network to be trained to obtain a target recognition model for generating the target prediction result of the image to be recognized.
[0033] It should be noted that S1N in this specification does not represent the sequence of the target recognition model training method. The following specifically describes the training method of the target recognition model of the present invention.
[0034] Step S11: Select training samples from the obtained training image data according to a preset ratio.
[0035] It should be noted that when selecting the optimal data augmentation strategy, if all the training image data is used, the computational complexity is large and the time consumption is long, which is not applicable to lightweight networks and some general application scenarios. Therefore, some training samples can be selected first according to a preset ratio to reduce the data volume and improve the efficiency of obtaining the data augmentation strategy, so as to be applicable to any network and task scenario, including classification, regression, detection, segmentation, etc.
[0036] In addition, the training image data can be understood as a set of picture data. The picture data can be for the same object, and the pictures of the object in different angles, different environmental conditions, different pixel colors and other states. The picture data stream is usually relatively large, reaching the level of millions of pictures. The above training image data is used to train the network to be trained to complete the construction of the model.
[0037] It should be noted that the data augmentation method selected based on this embodiment is applicable to perform data augmentation on the above training image data. That is, if data augmentation needs to be performed on specific image training data, it is necessary to first select training samples that meet the preset ratio from the image training data to facilitate the subsequent determination of the optimal data augmentation strategy. It can also be understood that the data augmentation strategy generated through this embodiment is only the optimal data augmentation strategy for the image training data obtained in the same application scenario, and it may not be applicable to the image training data in different scenarios.
[0038] Step S12: Use at least two data augmentation strategies to perform data augmentation on the training samples respectively to obtain data augmentation results corresponding to each data augmentation strategy.
[0039] It should be noted that data augmentation includes using some computer vision methods to perform some transformations on the existing picture data to obtain equivalent new data that is highly correlated with the original data.
[0040] In this embodiment, the data augmentation strategy includes at least one of cropping, rotation, translation, flipping, sharpening, lighting, and occlusion. For example, the data augmentation strategy includes cropping, rotation, translation, flipping, sharpening, lighting, or occlusion. For another example, the data augmentation strategy includes at least two of cropping, rotation, translation, flipping, sharpening, lighting, and occlusion. By using at least two data augmentation strategies to perform data augmentation on the obtained training samples respectively, it is convenient to obtain the optimal data augmentation strategy according to the data augmentation results corresponding to various data augmentation strategies.
[0041] Step S13: According to each data augmentation result, select a data augmentation strategy from at least two data augmentation strategies, and use the selected data augmentation strategy to perform data augmentation on the training image data.
[0042] In this embodiment, selecting a data augmentation strategy from at least two data augmentation strategies according to each data augmentation result includes: obtaining the loss values corresponding to the brightness, contrast, and texture information according to each data augmentation result; scoring according to the loss values corresponding to each data augmentation strategy to obtain the highest scoring result; and selecting the corresponding data augmentation strategy according to the highest scoring result. Specifically:
[0043] First, obtain the loss values corresponding to the brightness, contrast, and texture information according to each data augmentation result.
[0044] Specifically, first calculate the brightness, contrast, and texture information.
[0045] The calculation formula for brightness B is expressed as:
[0046]
[0047] where w represents the width of the image, h represents the height of the image, (i, j) represents the coordinates of the pixel point, and x(i, j) represents the gray value of the pixel point.
[0048] The calculation formula for contrast C is expressed as:
[0049]
[0050] where δ(i, j) = |i - j|, that is, the gray difference between adjacent pixels, and P δ (i, j) represents the pixel distribution probability where the gray difference between adjacent pixels is δ.
[0051] The calculation formula for texture information H is expressed as:
[0052]
[0053] where, P represents the feature quantity of the feature of the gray distribution space, f(i, j) represents the frequency of occurrence of the feature binary group (i, j), and N is the scale of the image.
[0054] Then, use the brightness, contrast, and texture information as the loss values.
[0055] Secondly, score according to the loss values corresponding to each data augmentation strategy to obtain the highest scoring result.
[0056] In this embodiment, scores are obtained based on the loss values corresponding to each data augmentation strategy, and the scoring results include: based on the minimum loss value and a preset threshold, obtaining the number of iterations corresponding to each data augmentation result when the minimum loss value is less than the preset threshold; comparing all the numbers of iterations, and using the minimum number of iterations as the highest scoring result.
[0057] In an alternative embodiment, when obtaining the highest scoring result by scoring based on the loss values corresponding to each data augmentation strategy, it further includes: comparing the loss values corresponding to each data augmentation strategy to obtain the minimum loss value as the highest scoring result.
[0058] It should be noted that this method is applicable to various tasks such as classification, regression, detection, and segmentation, and will not be further limited here. For ease of understanding, taking the binary classification task as an example, if its loss function L is expressed as:
[0059]
[0060] where y i represents the label of sample i, with the positive class being 1 and the negative class being 0, and P i represents the probability that sample i is predicted as the positive class.
[0061] It should be noted that when measuring the training speed, comparison is made through the number of iterations required when L < the preset threshold. For example, if there are m data augmentation strategies corresponding to the minimum loss value, and the corresponding iterations are N1,…N m , then find the minimum N i , use it as the highest score, and obtain the data augmentation strategy corresponding to N i as the optimal data augmentation strategy. Similarly, when measuring the accuracy, based on the comparison of each loss value, if the loss value is the smallest, the corresponding score is the highest, and the data augmentation strategy with the highest score is used as the optimal data augmentation strategy.
[0062] Finally, according to the highest scoring result, select the corresponding data augmentation strategy. It should be noted that after scoring based on each loss value, the scoring results are sorted to facilitate the selection of the highest scoring result, and then the corresponding data augmentation strategy is determined as the optimal data augmentation strategy based on the highest scoring result, so as to subsequently use the selected data augmentation strategy to perform data augmentation on all the data (i.e., the training image data).
[0063] In an alternative embodiment, after selecting the data augmentation strategy with the highest corresponding scoring result according to the scoring result, it further includes: using the data augmentation result corresponding to the data augmentation strategy with the highest corresponding scoring result as the input data for training, and training the network to be trained.
[0064] In an alternative embodiment, after comparing the loss values corresponding to the respective data augmentation strategies, it further includes: based on the minimum loss value, using the training samples before data augmentation corresponding to the remaining loss values as the input data for training, and training the network to be trained. It should be noted that when comparing the respective loss values, the minimum loss value is used for scoring to determine the optimal data augmentation strategy and its corresponding data augmentation result for training the network to be trained, and the original data (i.e., the training samples without data augmentation) corresponding to the remaining loss values other than the minimum loss value is used for training the network to be trained.
[0065] In addition, in this embodiment, after selecting a data augmentation strategy from the at least two data augmentation strategies, the selected data augmentation strategy is used to perform data augmentation on the training image data. It should be noted that after determining the data augmentation strategy, the training image data is augmented. For example, if the data augmentation strategy includes flipping by k degrees, the training image data is flipped by k degrees to obtain an image flipped by k degrees; if the data augmentation strategy includes flipping by x degrees and sharpening by y, the training image data is flipped by x degrees and sharpened by y to obtain a training image flipped by x degrees and sharpened by y. Here, k, x, and y can be set according to actual needs and are not further limited herein.
[0066] Step S14: Using the training image data after data augmentation as the input data for training, and using the target recognition result corresponding to the training image as the label, training the network to be trained to obtain a target recognition model for generating a target prediction result of the image to be recognized.
[0067] It should be noted that the network to be trained can be an existing network built into the training device, and this existing network usually includes a network structure, or it can be other networks specified by the user, such as various neural networks like CNN. The network to be trained usually includes a convolutional layer, a fully connected layer, and a loss function, and the loss function can specifically be a cross-entropy function; according to a preset iteration rule, the above-mentioned augmented training data is input into the model to be trained for training to obtain a trained target recognition model.
[0068] In an alternative embodiment, after obtaining the trained target recognition model, it further includes: using verification data to verify the target recognition model to obtain a verification result. It should be noted that the verification result can use classification accuracy, recognition accuracy, etc. as the main judgment criteria. For a network for recognizing picture categories, classification accuracy can be selected as the main judgment criteria, and for a network for picture targets, recognition accuracy can be selected as the main judgment criteria.
[0069] In summary, the present invention selects a data augmentation strategy from the training image data through the selected training samples, so as to quickly determine the data augmentation strategy with less data samples, make the obtained optimal augmentation strategy applicable to the training image data with a large amount of data, improve the selection time of the data augmentation strategy, reduce the calculation amount. In addition, the selected data augmentation strategy can be applied to the image data of the same type of target object for data augmentation, so as to improve the usage efficiency of the image data augmentation method.
[0070] The image recognition device based on data augmentation strategy selection provided by the present invention will be described below. The image recognition device based on data augmentation strategy selection described below can be mutually referred to the image recognition method based on data augmentation strategy selection described above.
[0071] Figure 3 The structural schematic diagram of an image recognition device based on data augmentation strategy selection of the present invention is shown. The device includes:
[0072] A target recognition module 31, which inputs the acquired picture data into the target recognition model to obtain the target prediction result output by the target recognition model;
[0073] Among them, the target recognition model is trained based on the data augmentation result obtained by performing data augmentation on the training image data based on the pre-selected data augmentation strategy and the target recognition result corresponding to the training image data. The pre-selected data augmentation strategy is based on the training samples selected from the training image data according to a preset ratio.
[0074] In an optional embodiment, in order to facilitate the training of the target recognition model, the device further includes a training module 32. Figure 4 The structural schematic diagram of the training module 32 of the present invention is shown. The training module includes:
[0075] A data acquisition unit 41, which selects training samples from the acquired training image data according to a preset ratio;
[0076] A data augmentation unit 42, which uses at least two data augmentation strategies to perform data augmentation on the training samples respectively to obtain the data augmentation results corresponding to each data augmentation strategy;
[0077] A strategy selection unit 43, which selects a data augmentation strategy from at least two data augmentation strategies according to each data augmentation result, and uses the selected data augmentation strategy to perform data augmentation on the training image data;
[0078] A training unit 44, which uses the training image data after data augmentation as the input data for training, uses the target recognition result corresponding to the training image as the label, and trains the network to be trained to obtain a target recognition model for generating the target prediction result of the image to be recognized.
[0079] In this embodiment, the data acquisition unit 41 selects some training samples according to a preset ratio to reduce the data volume and improve the efficiency of obtaining the data augmentation strategy.
[0080] The data augmentation unit 42 includes: a selection subunit that selects at least two data augmentation strategies; a first data augmentation subunit that performs data augmentation on the training samples respectively using the data augmentation strategies selected by the selection subunit to obtain data augmentation results corresponding to the respective data augmentation strategies. It should be noted that data augmentation includes using some computer vision methods to perform some transformations on the existing image data to obtain equivalent new data that is highly correlated with the original data. In addition, the data augmentation strategies include at least one of cropping, rotation, translation, flipping, sharpening, lighting, and occlusion. For example, the data augmentation strategies include cropping, rotation, translation, flipping, sharpening, lighting, or occlusion. For another example, the data augmentation strategies include at least two of cropping, rotation, translation, flipping, sharpening, lighting, and occlusion.
[0081] The strategy selection unit 43 includes: a loss value calculation subunit that obtains loss values corresponding to brightness, contrast, and texture information according to the respective data augmentation results; a scoring subunit that scores according to the loss values corresponding to the respective data augmentation strategies to obtain the highest scoring result; a strategy selection subunit that selects the corresponding data augmentation strategy according to the highest scoring result; a second data augmentation subunit that performs data augmentation on the training image data using the selected data augmentation strategy.
[0082] Furthermore, the scoring subunit includes: a comparison grandson subunit that, based on the minimum loss value and a preset threshold, obtains the iteration times corresponding to the respective data augmentation results when the minimum loss value is less than the preset threshold; a first scoring grandson subunit that compares all the iteration times and uses the minimum iteration time as the highest scoring result.
[0083] In an alternative embodiment, the scoring subunit further includes: a second scoring grandson subunit that compares based on the loss values corresponding to the respective data augmentation strategies to obtain the minimum loss value as the highest scoring result.
[0084] In an alternative embodiment, the strategy selection unit 43 further includes: a first data transmission subunit that uses the data augmentation result corresponding to the data augmentation strategy with the highest corresponding scoring result as the input data for training, and trains the network to be trained.
[0085] In an alternative embodiment, the policy selection unit 43 further includes: a second data transmission subunit, which, based on the minimum loss value, uses the training samples before data augmentation corresponding to the remaining loss values as the input data for training, and trains the network to be trained. It should be noted that the first data transmission subunit and the second data transmission subunit can be the same transmission subunit, that is, when the comparison subunit compares each loss value, the scoring subunit uses the minimum loss value for scoring to determine the optimal data augmentation policy and its corresponding data augmentation result to train the network to be trained, and uses the transmission subunit to input the data augmentation result corresponding to the selected data augmentation policy and the original data (i.e., the training samples without data augmentation) corresponding to the remaining loss values except the minimum loss value into the network to be trained for training.
[0086] Finally, the training unit 44 uses the training image data after data augmentation as the input data for training, and uses the target recognition result corresponding to the training image as the label to train the network to be trained, and obtains a target recognition model for generating the target prediction result of the image to be recognized. It should be noted that the network to be trained can be an existing network built into the training device, and this existing network usually includes a network structure, or it can be other networks specified by the user, such as various neural networks CNN, etc. The network to be trained usually includes a convolutional layer, a fully connected layer, and a loss function, and the loss function can specifically be a cross-entropy function; according to the preset iteration rule, the above-mentioned enhanced training data is input into the model to be trained for training to obtain the trained target recognition model.
[0087] In an alternative embodiment, the training module further includes: a verification unit, which uses the verification data to verify the target recognition model to obtain a verification result. It should be noted that the verification result can use the classification accuracy rate, recognition accuracy rate, etc. as the main judgment criteria. For a network that identifies picture categories, the classification accuracy rate can be selected as the main judgment criteria, and for a network of picture targets, the recognition accuracy rate can be selected as the main judgment criteria.
[0088] Figure 5 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 5As shown in the figure, the electronic device may include: a processor 51, a communications interface 52, a memory 53, and a communication bus 54. Among them, the processor 51, the communication interface 52, and the memory 53 complete mutual communication through the communication bus 54. The processor 51 can call the logical instructions in the memory 53 to execute an image recognition method selected based on a data enhancement strategy. The method includes: inputting the acquired picture data into a target recognition model to obtain a target prediction result output by the target recognition model; wherein, the target recognition model is trained based on the data enhancement result obtained by performing data enhancement on the training image data based on a pre-selected data enhancement strategy and the target recognition result corresponding to the training image data, and the pre-selected data enhancement strategy is obtained based on the training samples selected from the training image data according to a preset ratio.
[0089] In addition, when the logical instructions in the above-mentioned memory 53 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0090] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image recognition method selected based on the data enhancement strategy provided by the above-mentioned various methods. The method includes: inputting the acquired picture data into a target recognition model to obtain a target prediction result output by the target recognition model; wherein, the target recognition model is trained based on the data enhancement result obtained by performing data enhancement on the training image data based on a pre-selected data enhancement strategy and the target recognition result corresponding to the training image data, and the pre-selected data enhancement strategy is obtained based on the training samples selected from the training image data according to a preset ratio.
[0091] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an image recognition method based on the selection of a data augmentation strategy provided by the above-mentioned various methods. The method includes: inputting the acquired picture data into a target recognition model to obtain a target prediction result output by the target recognition model; wherein, the target recognition model is trained based on the data augmentation result obtained by performing data augmentation on the training image data based on a pre-selected data augmentation strategy and the target recognition result corresponding to the training image data, and the pre-selected data augmentation strategy is obtained based on training samples selected from the training image data according to a preset ratio.
[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image recognition method based on data augmentation strategy selection, characterized in that, Including: Input the acquired image data into the target recognition model to obtain the target prediction result output by the target recognition model; Among them, the target recognition model is trained based on the data augmentation result obtained by augmenting the training image data using a pre-selected data augmentation strategy and the target recognition result corresponding to the training image data, and the pre-selected data augmentation strategy is obtained based on the training samples selected from the training image data at a preset ratio; Training the target recognition model includes: Select training samples from the acquired training image data at a preset ratio; Use at least two data augmentation strategies to augment the training samples respectively to obtain data augmentation results corresponding to each data augmentation strategy; According to each of the data augmentation results, select a data augmentation strategy from the at least two data augmentation strategies, and use the selected data augmentation strategy to augment the training image data; Use the training image data after the data augmentation as the input data for training, and use the target recognition result corresponding to the training image as the label to train the network to be trained, so as to obtain a target recognition model for generating the target prediction result of the image to be recognized; The step of selecting a data augmentation strategy from the at least two data augmentation strategies according to each of the data augmentation results includes: According to each of the data augmentation results, obtain the loss values corresponding to the brightness, contrast and texture information; Score according to the loss values corresponding to the data augmentation strategies to obtain the highest scoring result; According to the highest scoring result, select the corresponding data augmentation strategy; The step of scoring according to the loss values corresponding to the data augmentation strategies to obtain the highest scoring result includes: Based on the loss value and a preset threshold, obtain the iteration times corresponding to each data augmentation result with the loss value less than the preset threshold; Compare all the iteration times and use the minimum iteration time as the highest scoring result.
2. The image recognition method selected based on the data augmentation strategy according to claim 1, wherein After selecting the data augmentation strategy with the highest corresponding scoring result according to the scoring result, it further includes: using the data augmentation result corresponding to the data augmentation strategy with the highest corresponding scoring result as the input data for training to train the network to be trained.
3. An image recognition device based on data augmentation strategy selection, characterized in that, Including: A target recognition module that inputs the acquired image data into the target recognition model to obtain the target prediction result output by the target recognition model; Among them, the target recognition model is trained based on the data augmentation result obtained by augmenting the training image data using a pre-selected data augmentation strategy and the target recognition result corresponding to the training image data, and the pre-selected data augmentation strategy is obtained based on the training samples selected from the training image data at a preset ratio; The device further includes a training module, and the training module includes: A data acquisition unit that selects training samples from the acquired training image data at a preset ratio; A data augmentation unit that uses at least two data augmentation strategies to augment the training samples respectively to obtain data augmentation results corresponding to each data augmentation strategy; A strategy selection unit, which selects a data augmentation strategy from the at least two data augmentation strategies according to each of the data augmentation results, and uses the selected data augmentation strategy to perform data augmentation on the training image data; A training unit, which uses the training image data after the data augmentation as the input data for training, uses the target recognition result corresponding to the training image as a label, trains the network to be trained, and obtains a target recognition model for generating a target prediction result of the image to be recognized; The strategy selection unit includes: A loss value calculation sub-unit, which obtains loss values corresponding to brightness, contrast, and texture information according to each of the data augmentation results; A scoring sub-unit, which scores according to the loss values corresponding to the data augmentation strategies and obtains the highest scoring result; A strategy selection sub-unit, which selects the corresponding data augmentation strategy according to the highest scoring result; The scoring sub-unit includes: A comparison sub-subunit, which, based on the loss value and a preset threshold, obtains the number of iterations corresponding to each data augmentation result for which the loss value is less than the preset threshold; A first scoring sub-subunit, which compares all the numbers of iterations and uses the minimum number of iterations as the highest scoring result.
4. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the image recognition method based on data augmentation strategy selection according to any one of claims 1 to 2 are implemented.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the image recognition method based on data augmentation strategy selection according to any one of claims 1 to 2 are implemented.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the image recognition method based on data augmentation strategy selection according to any one of claims 1 to 2 are implemented.
Citation Information
Patent Citations
Training method and training model for image enhancement network, and image enhancement method
CN109255769A
Bi-level optimization method for image deblurring
WO2020103171A1