A data augmentation strategy selection method, device and system

By calculating the target object detection performance of the detection model corresponding to the data augmentation strategy in an unsupervised learning scenario, and selecting the optimal strategy, the problem of insufficient generalization ability of the data augmentation strategy in an unsupervised learning scenario is solved, and the effective application of the target object detection model in an unsupervised scenario is realized.

CN114021634BActive Publication Date: 2025-11-04GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111275966.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-11-04
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

In unsupervised learning scenarios, due to the lack of validation datasets with labeling information, existing data augmentation techniques cannot effectively improve the generalization ability of target detection models, resulting in insufficient applicability of data augmentation strategies in unsupervised learning scenarios.

Method used

By acquiring detection models corresponding to multiple data augmentation strategies, these models are used to detect targets in sample images. The confidence and number of detection boxes are calculated, and the optimal data augmentation strategy is selected based on the target detection performance value, thus realizing target detection in unsupervised learning scenarios.

Benefits of technology

This improves the generalization ability of data augmentation strategies in unsupervised learning scenarios, enabling data augmentation techniques to be applied to the object detection process in unsupervised learning scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021634B_ABST
    Figure CN114021634B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a data augmentation strategy selection method, device and system. The scheme is as follows: obtaining a plurality of first detection models corresponding to a plurality of data augmentation strategies respectively; for each data augmentation strategy, using the first detection model corresponding to the data augmentation strategy to detect a target object in a third sample image in a third data set to obtain the confidence of each detection frame; calculating the target object detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frames corresponding to the data augmentation strategy; selecting a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy. Through the technical scheme provided by the embodiment of the application, the generalization ability of the determined target data augmentation strategy is improved, so that the data augmentation technology is suitable for the target object detection process in an unsupervised learning scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection, and in particular to a data augmentation strategy selection method, device and system. BACKGROUND

[0002] At present, in a target detection process in a supervised learning scenario, a target data augmentation strategy can be used to perform data augmentation processing on a pre-set training data set. This will effectively increase the number of training data in the training data set, improve the diversity of the training data set, and thus improve the generalization ability of a detection model trained according to the training data set after data augmentation processing.

[0003] In the above target detection process, the target data augmentation strategy needs to be determined based on a verification data set including identification information.

[0004] In the determination of the target data augmentation strategy, the verification data set including identification information is needed. However, in an unsupervised learning scenario, due to the high economic cost of obtaining identification information, a sufficient number of verification data sets including identification information cannot be obtained, resulting in that the above data set including identification information does not have high generalization ability, which seriously affects the generalization ability of the determined target data augmentation strategy, and thus the data augmentation technology in the related art is not applicable to the target detection process in the unsupervised learning scenario. SUMMARY

[0005] The embodiments of the present application aim to provide a data augmentation strategy selection method, device and system to improve the generalization ability of the determined target data augmentation strategy, so that the data augmentation technology is applicable to the target detection process in the unsupervised learning scenario. The specific technical solutions are as follows:

[0006] The embodiments of the present application provide a data augmentation strategy selection method, which comprises:

[0007] obtaining a plurality of data augmentation strategies each corresponding to a first detection model; the first detection model corresponding to each data augmentation strategy is obtained by training a first data set corresponding to the data augmentation strategy based on the data augmentation strategy; and the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy;

[0008] for each data augmentation strategy, using the first detection model corresponding to the data augmentation strategy to detect a target in a third sample image in a third data set, to obtain the confidence of each detection frame;

[0009] calculating a target detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frames corresponding to the data augmentation strategy.

[0010] select the target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0011] Optionally, the step of obtaining the first detection model corresponding to each data augmentation strategy comprises:

[0012] obtaining the second data set;

[0013] obtaining a plurality of data augmentation strategies;

[0014] for each data augmentation strategy, performing data augmentation processing on the second data set by using the data augmentation strategy to obtain a first data set corresponding to each data augmentation strategy;

[0015] for each data augmentation strategy, training a preset detection model by using the first data set corresponding to the data augmentation strategy to obtain a first detection model corresponding to each data augmentation strategy.

[0016] Optionally, the step of obtaining a plurality of data augmentation strategies comprises:

[0017] based on a plurality of data augmentation algorithms included in the candidate library, selecting a plurality of groups of seed strategies, each group of seed strategies including a first number of data augmentation algorithms;

[0018] for each group of seed strategies, based on a second number, and the maximum and minimum values of the random probability interval / amplitude value transformation interval corresponding to each data augmentation algorithm in the group of seed strategies, determining the random probability value / amplitude value corresponding to each data augmentation algorithm in the group of seed strategies to obtain a plurality of data augmentation strategies.

[0019] Optionally, for the first data set corresponding to any data augmentation strategy, the following steps are used to train the first detection model corresponding to the data augmentation strategy:

[0020] obtaining the first data set; the first data set includes a plurality of first sample images, and first identification information of a target object in each first sample image;

[0021] using a preset detection model to detect the target object in each first sample image to obtain a first detection result;

[0022] based on the first detection result and the first identification information, calculating a first loss value of the preset detection model;

[0023] if the first difference value is greater than the first preset threshold value, adjusting parameters of the preset detection model, and returning to perform the step of detecting the target object in each first sample image by using the preset detection model to obtain a first detection result; the first difference value is a difference value between the first loss value obtained in the current round of training and the first loss value obtained in the last round of training;

[0024] if the first difference value is not greater than the first preset threshold value, determining the current preset detection model as a trained first detection model.

[0025] Optionally, the first identification information includes a region identifier indicating a region where the target object is located, and a category identifier indicating a category of the target object.

[0026] The step of calculating the first loss value of the preset detection model based on the first detection result and the first identification information includes:

[0027] calculating a first error value between a region where a detection frame in the first detection result is located and a region where the target object indicated by the region identifier in the first data set is located;

[0028] calculating a second error value between a category of the target object corresponding to the detection frame in the first detection result and a category of the target object indicated by the category identifier in the first data set;

[0029] calculating a sum value of the first error value and the second error value as the first loss value of the preset detection model.

[0030] Optionally, the step of calculating the target object detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frame of the data augmentation strategy includes:

[0031] calculating a sum value of the number of detection frames corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy; and / or

[0032] calculating a sum value of the confidence of the detection frame corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

[0033] Optionally, the step of selecting a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy includes:

[0034] determining the data augmentation strategy with the highest target object detection performance value in the plurality of data augmentation strategies as the target data augmentation strategy.

[0035] Optionally, the method further includes:

[0036] obtain a fourth data set and a fifth data set; the fourth data set comprises a plurality of fourth sample images and third identification information corresponding to a target object in each fourth sample image, and the fifth data set comprises a plurality of fifth sample images;

[0037] perform data augmentation processing on the fourth data set by using the target data augmentation strategy to obtain a sixth data set;

[0038] detect the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result;

[0039] detect the target object in each fifth sample image in the fifth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a third detection result;

[0040] calculate a second loss value of the first detection model corresponding to the target data augmentation strategy based on the second detection result and the third detection result;

[0041] if the second difference value is greater than a second preset threshold, fine-tune the parameters of the first detection model corresponding to the target data augmentation strategy, and return to perform the step of detecting the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result; the second difference value is the difference between the second loss value obtained in the current round of training and the second loss value obtained in the last round of training;

[0042] if the second difference value is not greater than the second preset threshold, the current first target object detection model is determined as a trained second detection model.

[0043] Optionally, the method further comprises:

[0044] obtain a to-be-detected image;

[0045] detect the target object in the to-be-detected image by using the second detection model to obtain a fourth detection result.

[0046] Embodiments of the present application also provide a data augmentation strategy selection device, the device comprises:

[0047] a first obtaining module configured to obtain a plurality of first detection models corresponding to a plurality of data augmentation strategies respectively; each first detection model corresponding to a data augmentation strategy is trained based on a first data set corresponding to the data augmentation strategy, and the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy;

[0048] The first detection module is configured to, for each data augmentation strategy, detect the target object in the third sample image in the third data set by using the first detection model corresponding to the data augmentation strategy to obtain the confidence of each detection frame;

[0049] The first calculation module is configured to calculate the target object detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frames corresponding to the data augmentation strategy;

[0050] The selection module is configured to select the target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0051] Embodiments of the present application also provide a data augmentation strategy selection system, which comprises an image acquisition device and a model training device;

[0052] The image acquisition device is configured to obtain a plurality of second sample images and identification information corresponding to a target object in each second sample image to obtain a second data set, and obtain a plurality of third sample images as a third data set.

[0053] The model training device is configured to perform data augmentation processing on the second data set based on a plurality of data augmentation strategies to obtain a first data set corresponding to each data augmentation strategy, and train a first detection model corresponding to each data augmentation strategy based on the first data set corresponding to the data augmentation strategy.

[0054] The model training device is further configured to, for each data augmentation strategy, detect the target object in the third sample image in the third data set by using the first detection model corresponding to the data augmentation strategy to obtain the confidence of each detection frame, calculate the target object detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frames corresponding to the data augmentation strategy, and select the target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0055] Optionally, the system further comprises a client;

[0056] The image acquisition device is further configured to obtain a to-be-detected image after obtaining the target data augmentation strategy.

[0057] The client is configured to detect the target object in the to-be-detected image by using the first detection model corresponding to the target data augmentation strategy to obtain a fifth detection result.

[0058] Embodiments of the present application also provide an electronic device, which comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus;

[0059] a memory for storing a computer program;

[0060] a processor for implementing the steps of the data augmentation strategy selection method described above when executing the program stored in the memory.

[0061] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the data augmentation strategy selection method described above.

[0062] The embodiments of the present application also provide a computer program product containing instructions, which, when executed on a computer, cause the computer to perform the data augmentation strategy selection method described above.

[0063] The embodiments of the present application have the following beneficial effects:

[0064] The technical solution provided by the embodiments of the present application can determine the target object detection performance value corresponding to each data augmentation strategy according to the detection result of the target object in the sample image, that is, the confidence of the detected bounding box in the sample image, after obtaining the first detection model corresponding to each data augmentation strategy, and determine the target data augmentation strategy based on the target object detection performance value. Compared with the related art, the target object detection performance of the first detection model corresponding to each data augmentation strategy is verified by using the sample image without including the identification information as the verification data set, which makes it no longer necessary to include the identification information in the verification data set, improves the generalization ability of the verification data in the verification data set, and thus improves the generalization ability of the determined target data augmentation strategy, so that the data augmentation technology is applicable to the target object detection process in the unsupervised learning scenario.

[0065] Of course, implementing any product or method of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other embodiments can also be obtained by those skilled in the art based on these drawings.

[0067] Figure 1 A first flowchart of the data augmentation strategy selection method provided by the embodiments of the present application;

[0068] Figure 2 A schematic diagram of the detection result of the sample image provided by the embodiments of the present application;

[0069] Figure 3 A second flowchart of a data augmentation strategy selection method provided by an embodiment of the present application;

[0070] Figure 4 A first flowchart of a first detection model training method provided by an embodiment of the present application;

[0071] Figure 5 A first flowchart of a Faster-RCNN detection model provided by an embodiment of the present application;

[0072] Figure 6 A first flowchart of a first loss value calculation provided by an embodiment of the present application;

[0073] Figure 7 A first flowchart of a target object detection method provided by an embodiment of the present application;

[0074] Figure 8 A second flowchart of a second model training method provided by an embodiment of the present application;

[0075] Figure 9 A second flowchart of a target object detection method provided by an embodiment of the present application;

[0076] Figure 10 A first structural diagram of a data augmentation strategy selection device provided by an embodiment of the present application;

[0077] Figure 11-a A first structural diagram of a data augmentation strategy selection system provided by an embodiment of the present application;

[0078] Figure 11-b A second structural diagram of a data augmentation strategy selection system provided by an embodiment of the present application;

[0079] Figure 12 A first structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0080] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art based on the present application belong to the scope of protection of the present application.

[0081] In the related art, after a plurality of data augmentation strategies are used to perform data augmentation processing on a training data set, and a detection model corresponding to each data augmentation strategy is trained, the detection model corresponding to each data augmentation strategy is used to detect a target object in a sample image in a verification data set including identification information, to obtain a detection result. Then, according to the detection result and the identification information in the training data set, a performance index corresponding to each data augmentation strategy is determined, and a data augmentation strategy with the best performance index is determined as a target data augmentation strategy. Therefore, the target data augmentation strategy needs to be determined based on the verification data set including the identification information, which makes the data augmentation technology in the related art not applicable to a target object detection process in an unsupervised learning scenario.

[0082] To solve the technical problems in the related art, an embodiment of the present application provides a data augmentation strategy selection method. The method can be applied to any electronic device. As shown in Figure 1 Figure 1 The first flowchart of the data augmentation strategy selection method provided by the embodiment of the present application is shown. The method includes the following steps.

[0083] In step S101, a first detection model corresponding to each of a plurality of data augmentation strategies is obtained. The first detection model corresponding to each data augmentation strategy is trained based on a first data set corresponding to the data augmentation strategy. The first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy.

[0084] In step S102, for each data augmentation strategy, a target object in a third sample image in a third data set is detected by using the first detection model corresponding to the data augmentation strategy, to obtain a confidence of each detection box.

[0085] In step S103, a target object detection performance value corresponding to the data augmentation strategy is calculated according to the number of detection boxes and / or the confidence of the detection boxes corresponding to the data augmentation strategy.

[0086] In step S104, a target data augmentation strategy is selected based on the target object detection performance value corresponding to each data augmentation strategy.

[0087] Through Figure 1 ​According to the method, after obtaining the first detection model corresponding to each data augmentation strategy, the target object detection performance value corresponding to each data augmentation strategy is determined according to the confidence of the detection result of the target object in the sample image, that is, the confidence of the detected bounding box in the sample image, which is detected by the first detection model corresponding to each data augmentation strategy, so as to determine the target data augmentation strategy based on the target object detection performance value. Compared with the related art, the target object detection performance of the first detection model corresponding to each data augmentation strategy is verified by using the sample image without the identification information as the verification data set, which makes it unnecessary to include the identification information in the verification data set, improves the generalization ability of the verification data in the verification data set, and thus improves the generalization ability of the determined target data augmentation strategy, so that the data augmentation technology is applicable to the target object detection process in the unsupervised learning scene.

[0088] The embodiments of the present application will be described below by specific examples. For ease of description, the electronic device is described as the execution subject below, and does not have any limiting effect.

[0089] For the above step S101, that is, obtaining the first detection model corresponding to each data augmentation strategy; the first detection model corresponding to each data augmentation strategy is trained based on the first data set corresponding to the data augmentation strategy, and the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on the second data set based on the data augmentation strategy.

[0090] In this step, the electronic device can pre-obtain multiple second data sets as training data sets, and perform data augmentation processing on the second data sets by using different data augmentation strategies to obtain the first data set corresponding to each data augmentation strategy. The electronic device can train the preset detection model by using the first data set corresponding to each data augmentation strategy to obtain the trained target object detection model (denoted as the first detection model) corresponding to each data augmentation strategy. The electronic device can obtain the first detection model corresponding to each data augmentation strategy. The training process of the first detection model can be referred to the description below, and will not be described in detail here.

[0091] The above-mentioned second data set includes multiple sample images (denoted as second sample images) and identification information (denoted as second identification information) corresponding to the target object in each sample image.

[0092] After the above-mentioned data augmentation processing on the second data set by using the data augmentation strategy to obtain the first data set corresponding to the data augmentation strategy, the first data set includes multiple sample images (denoted as first sample images) and identification information (denoted as first identification information) corresponding to the target object in each sample image.

[0093] The first sample image in the first data set is obtained by performing data augmentation on a second sample image in a second data set, and the first identification information corresponding to each first sample image in the first data set is determined according to the second identification information corresponding to the second sample image before data augmentation. For details of the data augmentation manner, refer to the description below, which will not be described here.

[0094] The target object in the sample image includes, but is not limited to, a person, a vehicle, a license plate, and the like in the sample image. Here, the target object in the sample image is not specifically limited.

[0095] The number of first sample images included in the first data set is greater than the number of second sample images in the second data set. Here, the number of first sample images included in the first data set and the number of second sample images in the second data set are not specifically limited.

[0096] For the step S102, for each data augmentation strategy, the target object in the third sample image in the third data set is detected by using the first detection model corresponding to the data augmentation strategy, and the confidence of each detection frame is obtained.

[0097] In this step, after obtaining the first detection model corresponding to each of the plurality of data augmentation strategies, the electronic device can detect the target object included in each sample image (denoted as a third sample image) in the third data set by using the first detection model corresponding to each data augmentation strategy, and obtain the detection result (denoted as a sixth detection result) corresponding to each data augmentation strategy.

[0098] For ease of understanding, an example of Figure 2 is used for illustration. Figure 2 is a schematic diagram of the detection result of the sample image provided by the embodiment of the present application.

[0099] In Figure 2 , the sample image 201 includes two target objects, i.e., a target object 202 and a target object 203. After detecting the target objects in the sample image 201 by using the first detection model corresponding to a certain data augmentation strategy, the electronic device can obtain the detection result as shown in Figure 2 . The detection result includes the confidence corresponding to each recognized target object. For example, the detection frame 204 is the detection frame corresponding to the detected target object 203, the detection frame 205 is the detection frame corresponding to the detected target object 202, and the confidence corresponding to each detection frame is marked on the detection frame 204 and the detection frame 205.

[0100] The detection of the target object in the sample image can be represented as detection of attribute information of the target object in the sample image, such as detection of action information of a person in the sample image, or detection of license plate information of a vehicle in the sample image, and the like. In this embodiment, the detection of the target object in the sample image is not limited specifically.

[0101] In an optional embodiment, the confidence corresponding to the detection frame in the detection result can be used to indicate a probability value of the target object in the detection frame being a preset object. For example, if the preset object is a vehicle, and the confidence corresponding to the detection frame 204 is 0.9, it can be determined that the probability of the target object 203 in the detection frame 204 being a vehicle is 90%. Figure 2

[0102] In another optional embodiment, the confidence corresponding to the detection frame in the detection result can be used to indicate a probability value of attribute information corresponding to the target object in the detection frame matching preset attribute information. For example, the preset attribute information is action information of a person running, and when the confidence of the detection frame in the detection result corresponding to a sample image is 0.9, it can be determined that the probability of the action of the person in the detection frame being running is 90%.

[0103] According to different application scenarios of the detection model, the detection of the target object is different, and the detection result obtained by the detection is also different. In this embodiment, the detection process and the detection result are not limited specifically. For ease of understanding, only the detection of the target object as an example of detection of a position of the target object in an image and category detection of the target object is described below, and the description does not have any limiting effect.

[0104] In the embodiment of the present application, the third data set can include a plurality of sample images. The third data set does not include the identification information corresponding to the target object in each third sample data.

[0105] According to the number of detection frames and / or the confidence of the detection frame corresponding to the data augmentation strategy, the target object detection performance value corresponding to the data augmentation strategy is calculated.

[0106] ​In this step, after obtaining the sixth detection result, the electronic device can calculate, for each data augmentation strategy, a target object detection performance value corresponding to the data augmentation strategy according to the number of bounding boxes in the sixth detection result corresponding to the data augmentation strategy; or calculate, for each data augmentation strategy, a target object detection performance value corresponding to the data augmentation strategy according to the confidence of the bounding boxes in the sixth detection result corresponding to the data augmentation strategy; or calculate, for each data augmentation strategy, a target object detection performance value corresponding to the data augmentation strategy according to the number of bounding boxes in the sixth detection result corresponding to the data augmentation strategy and the confidence of the bounding boxes.

[0107] In an optional embodiment, the step S103 calculates the target object detection performance value corresponding to the data augmentation strategy according to the number of bounding boxes and / or the confidence of the bounding boxes corresponding to the data augmentation strategy, and can be specifically represented as:

[0108] The electronic device calculates the sum of the number of bounding boxes corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

[0109] The number of bounding boxes can indicate the number of target objects recognized in the sample image.

[0110] In the embodiments of the present application, since the sample images detected by the first detection model corresponding to each data augmentation strategy are completely the same, that is, all the sample images in the third data set, for each sample image, when the number of bounding boxes detected by the first detection model corresponding to a certain data augmentation strategy is larger, that is, the first detection model corresponding to the data augmentation strategy can detect more target objects, it can be determined that the accuracy of the first detection model corresponding to the data augmentation strategy in detecting target objects is higher, so that the first detection model corresponding to the data augmentation strategy can more accurately detect target objects in images in an unsupervised scene, so that the data augmentation technology can be applied to the target object detection process in an unsupervised learning scene.

[0111] In another optional embodiment, the step S103 calculates the target object detection performance value corresponding to the data augmentation strategy according to the number of bounding boxes and / or the confidence of the bounding boxes corresponding to the data augmentation strategy, and can be specifically represented as:

[0112] The electronic device calculates the sum of the confidence of the bounding boxes corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

[0113] In an optional embodiment, for each data augmentation strategy, the electronic device can calculate the target cumulative score (TACS) corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy using the following formula:

[0114]

[0115] wherein TACS is the target object detection performance value corresponding to any data augmentation strategy, box is the detected bounding box, N is the total number of detected bounding boxes, p box is the confidence corresponding to the bounding box.

[0116] In the embodiments of the present application, since the sample images detected by the first detection model corresponding to each data augmentation strategy are completely the same, i.e., all the sample images in the third data set, for each sample image, the more the number of bounding boxes detected by the first detection model corresponding to a certain data augmentation strategy, and the greater the confidence corresponding to the detected bounding box, it can be determined that the accuracy of the first detection model corresponding to the data augmentation strategy in detecting the target object is higher, thereby determining that the first detection model corresponding to the data augmentation strategy can more accurately detect the target object in the image in the unsupervised scene, so that the data augmentation technology can be applied to the target object detection process in the unsupervised learning scene.

[0117] For the above step S104, i.e., selecting a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0118] In an optional embodiment, the electronic device determines the data augmentation strategy with the highest target object detection performance value among the plurality of data augmentation strategies as the target data augmentation strategy.

[0119] In the embodiments of the present application, the greater the target object detection performance value, the higher the accuracy of the first detection model trained based on the data augmentation strategy corresponding to the target object detection performance value in detecting the target object in the image in the unsupervised scene, and the more suitable it is for detecting the target object in the image in the unsupervised learning scene.

[0120] In an optional embodiment, according to the data augmentation strategy selection method provided in the embodiments of the present application, the embodiments of the present application further provide a data augmentation strategy selection method. As shown in Figure 3 , and Figure 3 is a second flowchart of the data augmentation strategy selection method provided in the embodiments of the present application. Specifically, the above step S101 is refined into the following steps, i.e., step S1011 to step S1014.

[0121] Step S1011, a second data set is obtained.

[0122] The second data set includes a plurality of second sample images, and second identification information corresponding to a target object in each second sample image.

[0123] The second identification information includes region identification indicating a region of the target object in the second sample image, and category identification indicating a category of the target object in the second sample image.

[0124] The region identification can be in the form of a detection box. The category identification can be in the form of a number, a letter, etc. For example, the category identification of a certain detection box is 1, which can indicate that the target object in the detection box is a vehicle. For another example, the category identification of a certain detection box is 2, which can indicate that the target object in the detection box is a person. Here, the representation of the identification information is not specifically limited.

[0125] Step S1012, a plurality of data augmentation strategies are obtained.

[0126] In this step, the candidate library can store a plurality of data augmentation algorithms, and a random probability interval (also referred to as a scale interval) corresponding to each data augmentation algorithm. The electronic device can obtain a plurality of groups of seed strategies according to each data augmentation algorithm stored in the candidate library, to obtain a plurality of data augmentation strategies. Each group of seed strategies includes a first number of data augmentation algorithms.

[0127] For ease of understanding, the above-mentioned candidate library is exemplified, as shown in Table 1.

[0128] Table 1

[0129] Data augmentation algorithm Random type Min value Max value Random gray scale Probability 0.3 0.5 Color jitter Magnitude 0.3 0.5 Random HSV Probability 0.3 0.5 Random noise Probability 0.003 0.005 Random scale Magnitude 0.3 0.5

[0130] The candidate library shown in Table 1 includes five data augmentation algorithms, i.e. random grayscale, color jitter, random HSV, random noise, and random scale, a random type (i.e. amplitude transformation or probability transformation) corresponding to each data augmentation algorithm, and a random probability interval / amplitude transformation interval corresponding to each data augmentation algorithm. Among them, HSV is the hue (Hue), saturation (Saturation), and lightness (Value) of the image.

[0131] For ease of understanding, the random grayscale in Table 1 is taken as an example for illustration. In the candidate library shown in Table 1, when the data augmentation algorithm of random grayscale is used, the electronic device can adjust the gray value in the sample image according to a probability of 0.3 to 0.5.

[0132] Each of the above data augmentation strategies can be obtained by combination and sequencing of the first number of data augmentation algorithms. For example, when the first number is 3, the electronic device can combine the three data augmentation algorithms of random grayscale, color jittering, and random HSV in Table 1 into one data augmentation strategy, which is represented as: random grayscale-color jittering-random HSV.

[0133] In the embodiments of the present application, when the first number is small, the difference between the sample images before and after the data augmentation processing is small, and the generalization ability is not good; when the first number is large, the difference between the sample images before and after the data augmentation processing is large, which may cause errors. The first number can be set according to the experience value of the user, and in this case, the first number is not specifically limited.

[0134] In an optional embodiment, the step S1012 of obtaining a plurality of data augmentation strategies can be refined into the following steps, i.e., step one-step two.

[0135] Step one: based on the plurality of data augmentation algorithms included in the candidate library, a plurality of seed strategies are selected, each seed strategy including a first number of data augmentation algorithms.

[0136] Step two: for each seed strategy, based on the second number, and the maximum and minimum values of the random probability interval / amplitude transformation interval corresponding to each data augmentation algorithm in the seed strategy, the random probability value / amplitude value corresponding to each data augmentation algorithm in the seed strategy is determined, to obtain a plurality of data augmentation strategies.

[0137] In an optional embodiment, the electronic device can use the following formula to calculate the random probability value / amplitude value corresponding to each data augmentation algorithm:

[0138]

[0139] wherein val is the random probability value / amplitude value corresponding to any data augmentation algorithm, MAX is the maximum value in the random probability interval / amplitude transformation interval corresponding to the data augmentation algorithm, MIN is the minimum value in the random probability interval / amplitude transformation interval corresponding to the data augmentation algorithm, M is the second number, and m is any integer between 0 and M.

[0140] In the embodiments of the present application, each of the above data augmentation algorithms has a corresponding random probability interval / amplitude transformation interval, that is, when the sample image is processed by the data augmentation algorithm, the random probability value / amplitude can be selected as any value in the corresponding random probability interval / amplitude transformation interval. Therefore, in order to effectively reduce the value range of the random probability value / amplitude corresponding to each data augmentation algorithm, a limited number (i.e., second number+1) of random probability values / amplitudes can be calculated by the second value, so as to effectively control the number of selected data augmentation strategies while ensuring the comprehensiveness of the selected data augmentation strategies.

[0141] For ease of understanding, the number of the seed strategy is P and the second number is M. According to the calculation formula of the random probability value / amplitude, m has M+1 choices. Therefore, for each sub-strategy, the random probability value / amplitude corresponding to each data augmentation algorithm in the seed strategy has M+1 choices. That is, each sub-strategy has corresponding M+1 data augmentation strategies. The electronic device will obtain P*(M+1) data augmentation strategies.

[0142] The above step S1011 can be performed before or after step S1012, or the above step S1011 can be performed simultaneously with step S1012. Here, the execution order of the above step S1011 and step S1012 is not limited.

[0143] Step S1013, for each data augmentation strategy, the second data set is processed by the data augmentation strategy to obtain the first data set corresponding to each data augmentation strategy.

[0144] In this step, for each second sample image in the second data set, the electronic device can respectively process each data augmentation strategy obtained by the above step S1012, the sample image to obtain the data augmentation processed second sample image corresponding to the second sample image, and the electronic device can also determine the identification information (i.e., the above first identification information) corresponding to the target object in the data augmentation processed second sample image according to the second identification information corresponding to the second sample image. The electronic device can take each second sample image corresponding to the data augmentation processed second sample image and the identification information corresponding to the target object in the data augmentation processed second sample image as the first data set.

[0145] For ease of understanding, taking the data augmentation strategy only including the above random grayscale as an example.

[0146] It is assumed that the random probability value is 0.4. The electronic device can randomly adjust the gray values of 40% of the pixels in each second sample image in the second data set according to a probability of 0.4 to obtain a second sample image after data augmentation. Since the data augmentation process only changes the gray values of the pixels and does not change the target object in the second sample image, the identification information corresponding to the second sample image after data augmentation can remain unchanged.

[0147] In the embodiments of the present application, according to the different data augmentation algorithms included in the data augmentation strategy, the identification information corresponding to the second sample image after data augmentation may change. For example, the data augmentation strategy includes a data augmentation algorithm for target object enlargement, which will cause the position of the target object in the second sample image after data augmentation to change. At this time, the region identifier indicating the region of the target object in the identification information of the second sample image after data augmentation will change. Here, the changes of the first identification information in the first data set obtained by the above data augmentation process are not specifically described.

[0148] In step S1014, for each data augmentation strategy, the pre-set detection model is trained using the first data set corresponding to the data augmentation strategy to obtain a first detection model corresponding to each data augmentation strategy.

[0149] In this step, for each data augmentation strategy, the electronic device can train the pre-set detection model using the first data set corresponding to the data augmentation strategy to obtain a trained target object detection model corresponding to each data augmentation strategy, that is, to obtain a first detection model corresponding to each data augmentation strategy.

[0150] The above pre-set detection model includes but is not limited to a Faster-Regions Convolutional Neural Networks (Faster-RCNN) detection model, a RetinaNet detection model, and a You Only Look Once (YOLO) detection model. For example, YOLOV1, YOLOV2, YOLOV3, YOLOV4, or YOLOV5 detection model, wherein V1, V2, V3, V4, and V5 identify different versions of the YOLO detection model.

[0151] The above Faster-RCNN detection model, RetinaNet detection model, and YOLO detection model are target detection models in deep learning. According to the different application scenarios of the detection model, the above detection models are also different. Here, the above pre-set detection model is not specifically limited.

[0152] The training process of the preset detection model can be seen from the description below, and will not be described in detail here.

[0153] Based on the same inventive concept, according to the data augmentation strategy provided in the above embodiments of the application, for any first data set corresponding to a data augmentation strategy, the embodiment of the application further provides a first detection model training method. As shown in Figure 4 Figure 4 A flowchart of the first detection model training method provided by the embodiment of the application. The method comprises the following steps.

[0154] Step S401, obtaining a first data set; the first data set comprises a plurality of first sample images, and first identification information of a target object in each first sample image.

[0155] In this step, after the first data set is obtained by using the above-mentioned plurality of data augmentation strategies to perform data augmentation processing on the second data set, the electronic device can obtain the first data set.

[0156] Step S402, using a preset detection model to detect the target object in each first sample image to obtain a first detection result.

[0157] In this step, the electronic device can respectively input each first sample image in the above-mentioned first data set to the preset detection model, and detect the target object in the first sample image by using the preset detection model, determine the region where each target object is located in the first sample image, and the category corresponding to each target object, to obtain the detection result of each first sample image (denoted as the first detection result). The electronic device obtains the first detection result of each first sample image.

[0158] For the sake of understanding, the preset detection model is taken as the above-mentioned Faster-RCNN detection model as an example for description, and does not have any limiting effect.

[0159] The Faster-RCNN detection model comprises a feature extractor, a region proposal network (RPN) and a region of interest (ROI) classifier. As shown in Figure 5 Figure 5 A schematic diagram of the Faster-RCNN detection model provided by the embodiment of the application.

[0160] After the electronic device inputs the above-mentioned sample image to the Faster-RCNN detection model, the feature extractor in the Faster-RCNN detection model extracts the features of the input sample image, and sends the extracted features to the RPN and the ROI classifier. ​​

[0161] The RPN determines the region where the target object is located in the sample image according to the received features, and marks the region where the target object is located in the sample image with a detection frame.

[0162] The ROI classifier classifies each target object in the sample image according to the received features, that is, determines the category of each target object in the sample image, and adds the category to the detection frame corresponding to the target object.

[0163] The Faster-RCNN detection model outputs a detection result, which includes the detection frame corresponding to the detected target object in the sample image, and the category of the target object in each detection frame.

[0164] In the embodiment of the present application, the first detection result includes the detection frame corresponding to the region where the target object is located in the sample image, and the category of the target object in each detection frame.

[0165] In step S403, the first loss value of the preset detection model is calculated based on the first detection result and the first identification information.

[0166] In this step, the electronic device can calculate the loss value of the preset detection model according to the first detection result and the identification information corresponding to each first sample image, as the first loss value.

[0167] The first identification information includes a region identifier indicating the region where the target object is located in the first sample image, and a category identifier indicating the category of the target object in the first sample image.

[0168] In an optional embodiment, as shown in Figure 6 The present application also provides a first loss value calculation method. Figure 6 A flowchart of the first loss value calculation method provided by the embodiment of the present application. The method includes the following steps.

[0169] In step S4031, the first error value between the region where the detection frame is located in the first detection result and the region where the target object is located indicated by the region identifier in the first identification information is calculated.

[0170] In this step, for each first sample image, the electronic device can calculate the first error value, that is, the error corresponding to the RPN in the Faster-RCNN detection model, according to the region where each detection frame is located in the first detection result corresponding to the first sample image, and the region where the target object is located indicated by the region identifier in the first identification information corresponding to the first sample image.

[0171] The first error value can be represented as a sum of a region position offset between a region where each bounding box in the first detection result corresponding to each first sample image is located and a region indicated by a region identifier in the first identification information corresponding to the first sample image, and the like. In this regard, the calculation of the first error value is not specifically described.

[0172] In step S4032, a second error value between a category of the target object corresponding to the bounding box in the first detection result and a category of the target object indicated by the category identifier in the first identification information is calculated.

[0173] In this step, for each first sample image, the electronic device can calculate a second error value according to a category of the target object in the bounding box in the first detection result corresponding to the first sample image and a category of the target object indicated by the category identifier in the first identification information corresponding to the first sample image, which is an error corresponding to the ROI classifier in the Faster-RCNN detection model.

[0174] The second error value can be calculated by a loss function such as error sum of squares, and the calculation of the second error value is not specifically described.

[0175] In step S4033, a sum of the first error value and the second error value is calculated as a first loss value of the preset detection model.

[0176] In an optional embodiment, the electronic device can calculate the first loss value of the preset detection model by using the following formula.

[0177] L det = L rpn + L roi

[0178] wherein, L det is the first loss value of the preset detection model, L rpn is the first error value, and L roi is the second error value.

[0179] Through steps S4031-S4033, the electronic device can accurately calculate the loss value of the preset detection model according to the first detection result and the first identification information of each first sample image in the first data set.

[0180] In this embodiment, after calculating the first loss value, the electronic device can calculate the difference between the first loss value calculated in the current training round and the first loss value calculated in the previous training round, thus obtaining a first difference. The electronic device can compare the first difference with a first preset threshold. When the first difference is greater than the first preset threshold, the electronic device can execute step S404. When the first difference is not greater than the first preset threshold, the electronic device can execute step S405.

[0181] In step S404, if the first difference is greater than the first preset threshold, the parameters of the preset detection model are adjusted, and the process returns to step S402.

[0182] In this step, when the first difference is greater than the first preset threshold, the electronic device can determine that the preset detection model has not converged. At this time, the electronic device can adjust the parameters of the preset detection model and return to execute the above step S402, that is, return to execute the above step of using the preset detection model to detect the target object in each first sample image and obtain the first detection result.

[0183] The parameters of the aforementioned preset detection model include, but are not limited to, the weights and biases in the preset detection model.

[0184] In this embodiment, the electronic device can adjust the parameters of the preset detection model based on the first loss value calculated in this round, using methods such as gradient descent or inverse adjustment. Here, the specific method for adjusting the model parameters is not limited.

[0185] Step S405: If the first difference is not greater than the first preset threshold, then the current preset detection model is determined as the trained first detection model.

[0186] In this step, when the first difference is less than or equal to the first preset threshold, the electronic device can determine that the preset detection model has converged. At this time, the electronic device can determine the preset detection model at the current moment as the trained first detection model.

[0187] exist Figure 4 The method described herein uses only the training process of the first detection model corresponding to one data augmentation strategy as an example. The training of the first detection models corresponding to multiple data augmentation strategies can all be described using the same method. Figure 4 The training methods shown will not be elaborated upon here.

[0188] pass Figure 4 The method shown allows electronic devices to train a first detection model corresponding to each data augmentation strategy using the first dataset corresponding to that data augmentation strategy, ensuring the accuracy of the trained first detection model and the differences between each first detection model.

[0189] According to the method shown in the application embodiment, the application embodiment further provides a target object detection method. Figure 1 As shown in the application embodiment, the application embodiment further provides a target object detection method. Figure 7 As shown in the application embodiment, the application embodiment further provides a target object detection method. Figure 7 The first flowchart of the target object detection method provided by the application embodiment is shown. The method comprises the following steps.

[0190] In step S701, the image to be detected is obtained.

[0191] In this step, after the electronic device selects the target data augmentation strategy from the plurality of data augmentation strategies, the image to be detected can be obtained.

[0192] The image to be detected can be an image including a target object, or an image not including a target object. The target object in the image to be detected includes but is not limited to a person, an animal, a vehicle, etc.

[0193] The image to be detected can be an image in a supervised learning scene, or an image in an unsupervised learning scene.

[0194] In step S702, the first detection model corresponding to the target data augmentation strategy is used to detect the target object in the image to be detected, and a fifth detection result is obtained.

[0195] In this step, after the electronic device selects the target data augmentation strategy from the plurality of data augmentation strategies, it can be determined that the target data augmentation strategy is more suitable for target object detection in an unsupervised learning scene than other data augmentation strategies. That is, the accuracy of target object detection in an unsupervised learning scene based on the first detection model corresponding to the target data augmentation strategy is higher than the accuracy of target object detection in an unsupervised learning scene based on the first detection model corresponding to other data augmentation strategies. After obtaining the image to be detected, the electronic device can use the first detection model corresponding to the target data augmentation strategy to detect the target object in the image to be detected, and obtain a detection result (denoted as a fifth detection result).

[0196] According to the different first detection models, the detection process of the target object in the image to be detected is also different.

[0197] For example, the target object detection of the image to be detected can be the detection of the target object class in the image to be detected. At this time, the electronic device can use the first detection model corresponding to the target data augmentation strategy to detect the position of the target object in the image to be detected and the class of the target object, and obtain the fifth detection result.

[0198] For another example, the target object detection of the to-be-detected image can be detection of a behavior category of a target object in the to-be-detected image. At this time, the electronic device can use the first detection model corresponding to the target data augmentation strategy to detect the position of the target object in the to-be-detected image and the behavior category corresponding to the target object, and obtain a fifth detection result.

[0199] In the embodiments of the present application, the target object detection of the to-be-detected image can also be foreground / background detection, license plate detection, and the like. Here, the target object detection of the to-be-detected image is not specifically limited.

[0200] The electronic device for performing target object detection on the to-be-detected image (referred to as a first electronic device) and the electronic device for selecting the target data augmentation strategy (referred to as a second electronic device) can be the same electronic device or different electronic devices.

[0201] When the first electronic device and the second electronic device are different electronic devices, the first electronic device can call the first detection model corresponding to the target data augmentation strategy trained by the second electronic device, or the first electronic device integrates the first detection model corresponding to the target data augmentation strategy, so as to use the first detection model corresponding to the target data augmentation strategy to perform target object detection on the to-be-detected image. Here, the first electronic device and the second electronic device are not specifically limited. For ease of understanding, only the case where the first electronic device and the second electronic device are the same electronic device is taken as an example for description in the embodiments of the present application, and does not have any limiting effect.

[0202] Through the steps S701-S702, the electronic device directly determines the first detection model corresponding to the selected target data augmentation strategy as the detection model for performing target object detection on the to-be-detected image, which makes the detection model applicable to detection of images in an unsupervised learning scenario, and effectively improves the accuracy of the detection result.

[0203] Based on the same inventive concept, according to Figure 1 the method shown in FIG. 8, after the target data augmentation strategy is selected, the embodiments of the present application further provide a second model training method. As shown in Figure 8 , Figure 8 a flowchart of the second model training method provided by the embodiments of the present application. The method includes the following steps.

[0204] Step S801, acquiring a fourth data set and a fifth data set; wherein the fourth data set includes a plurality of fourth sample images and third identification information corresponding to a target object in each fourth sample image, and the fifth data set includes a plurality of fifth sample images.

[0205] In this step, the electronic device can obtain a fourth data set, which includes a plurality of sample images (denoted as fourth sample images) and identification information corresponding to the target object in each fourth sample image (denoted as third identification information).

[0206] The fourth identification information includes a region identifier indicating the region of the target object in the fourth sample image and a category identifier indicating the category of the target object in the fourth sample image.

[0207] The fifth data set does not include identification information corresponding to the target object in the fifth sample image.

[0208] The fourth data set can be the same as the second data set or different from the second data set. The fifth data set can be the same as the third data set or different from the third data set.

[0209] In step S802, the fourth data set is processed by a target data augmentation strategy to obtain a sixth data set.

[0210] The sixth data set includes a plurality of sample images (denoted as sixth sample images) and identification information corresponding to the target object in each sample image (denoted as fourth identification information).

[0211] The process of processing the fourth data set by data augmentation can refer to the process of data augmentation of the second data, which will not be described in detail here.

[0212] In step S803, the target object in each sixth sample image in the sixth data set is detected by a first detection model corresponding to the target data augmentation strategy to obtain a second detection result.

[0213] In this step, for each sixth sample image in the sixth data set, the electronic device can use the first detection model corresponding to the target data augmentation strategy to detect the target object in the sixth sample image to obtain a second detection result.

[0214] The second detection result includes a region identifier indicating the region of the target object in the sixth sample image and a category identifier indicating the category of the target object in the sixth sample image.

[0215] In step S804, the target object in each fifth sample image in the fifth data set is detected by a first detection model corresponding to the target data augmentation strategy to obtain a third detection result.

[0216] In this step, for each fifth sample image in the fifth data set, the electronic device can use the first detection model corresponding to the target data augmentation strategy to detect the target object in the fifth sample image to obtain a third detection result.

[0217] The third detection result includes the confidence of each detection box detected in the fifth sample image.

[0218] The second detection result and the third detection result can be obtained in the manner of the first detection result, which will not be described in detail.

[0219] The step S803 can be performed before or after the step S804. The step S803 can also be performed simultaneously with the step S804. Herein, the order of the steps S803 and S804 is not limited.

[0220] In step S805, based on the second detection result and the third detection result, a second loss value of the first detection model corresponding to the target data augmentation strategy is calculated.

[0221] In this step, after determining the second detection result and the third detection result, the electronic device can calculate the loss value of the first detection model corresponding to the target data augmentation strategy based on the second detection result and the third detection result as a second loss value.

[0222] In an optional embodiment, the electronic device can use the following formula to calculate the second loss value of the first detection model corresponding to the target data augmentation strategy:

[0223] L = L det + aL da

[0224] Wherein, L is the second loss value, L det is a loss value calculated according to the second detection result and the fourth identification information of the target object in each fifth sample data in the fifth data set, L da is a loss value calculated according to the second detection result and the third detection result, and a is a preset balance parameter between the loss value L det and the loss value L da .

[0225] The calculation method of the loss value L det will not be described in detail.

[0226] The loss value L da is calculated according to the second detection result and the third detection result.

[0227] In an optional embodiment, the L da is a discrimination loss. The electronic device discriminates the sixth sample image feature in the second detection result detection process and the fifth sample image feature in the third detection result detection process using the discriminator. For example, L da may be represented as a feature difference between the second detection result and the third detection result. Here, the calculation method of the L da is not specifically described.

[0228] In the embodiment of the present application, after the electronic device calculates the second loss value, the electronic device can calculate a difference between the second loss value obtained in the current round of training and the second loss value obtained in the last round of training to obtain a second difference value. The electronic device can compare the second difference value with a second preset threshold. When the second difference value is greater than the second preset threshold, the electronic device can perform step S806. When the second difference value is not greater than the second preset threshold, the electronic device can perform step S807.

[0229] In step S806, if the second difference value is greater than the second preset threshold, the parameters of the first detection model corresponding to the target data augmentation strategy are fine-tuned, and the step S803 is returned to be executed.

[0230] In this step, when the second difference value is greater than the second preset threshold, the electronic device can determine that the first detection model corresponding to the target data augmentation strategy has not converged. At this time, the electronic device can fine-tune the parameters of the first detection model corresponding to the target data augmentation strategy, and return to execute the step S803 of detecting the target object in each sixth sample image in the sixth data set using the first detection model corresponding to the target data augmentation strategy to obtain the second detection result.

[0231] The fine-tuning of the parameters of the first detection model corresponding to the target data augmentation strategy can refer to the adjustment method of the parameters of the preset detection model. Compared with the adjustment of the parameters of the preset detection model, the adjustment of the parameters of the first detection model corresponding to the target data augmentation strategy reduces the step size of parameter adjustment, and realizes the fine-tuning of the parameters of the first detection model corresponding to the target data augmentation strategy. Here, the fine-tuning process of the parameters of the first detection model corresponding to the target data augmentation strategy is not specifically described.

[0232] In step S807, if the second difference value is not greater than the second preset threshold, the current first target object detection model is determined as the trained second detection model.

[0233] In this step, when the second difference is less than or equal to the second preset threshold, the electronic device can determine that the first detection model corresponding to the target data augmentation strategy converges. At this time, the electronic device can determine the current first target object detection model as the trained second detection model.

[0234] By the method shown in Figure 8 The electronic device fine-tunes the parameters of the first detection model corresponding to the target data augmentation strategy in combination with the fourth data set and the fifth data set, further improves the accuracy of the trained second detection model, and improves the target object detection accuracy of the second detection model in the unsupervised learning scene, so that the data augmentation technology can be better applied to the supervised learning scene and the unsupervised learning scene.

[0235] Based on the same inventive concept, according to the above Figure 8 The second detection model is trained, and the embodiment of the present application further provides a target object detection method, as shown in Figure 9 , Figure 9 A second flowchart of a target object detection method provided by the embodiment of the present application is provided. The method comprises the following steps.

[0236] Step S901, obtaining a to-be-detected image.

[0237] In the embodiment of the present application, the to-be-detected image can be an image in a supervised learning scene or an image in an unsupervised learning scene.

[0238] Step S902, using the second detection model to detect the target object in the to-be-detected image to obtain a fourth detection result.

[0239] In the embodiment of the present application, the fourth detection result includes a region identifier indicating the region of the target object in the to-be-detected image, and a category identifier indicating the category of the target object in the to-be-detected image.

[0240] The detection process of the to-be-detected image using the second detection model can refer to the detection process of the to-be-detected image using the first detection model corresponding to the target data augmentation strategy, which will not be described in detail here.

[0241] By the method shown in Figure 9 The second detection model is trained Figure 8 When the target object in the to-be-detected image is identified, whether the to-be-detected image is an image in a supervised learning scene or an image in an unsupervised learning scene, the electronic device can accurately detect the target object in the to-be-detected image, and improve the accuracy of the detection result.

[0242] Based on the same inventive concept, according to the data augmentation strategy selection method provided in the embodiments of the present application, the embodiments of the present application further provide a data augmentation strategy selection device. As shown in Figure 10 Figure 10 FIG. 1 is a structural schematic diagram of a data augmentation strategy selection device provided in the embodiments of the present application. The device comprises the following modules.

[0243] The first acquisition module 1001 is configured to acquire a first detection model corresponding to each of a plurality of data augmentation strategies. The first detection model corresponding to each data augmentation strategy is trained based on a first data set corresponding to the data augmentation strategy. The first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy.

[0244] The first detection module 1002 is configured to, for each data augmentation strategy, perform detection on a target object in a third sample image in a third data set by using the first detection model corresponding to the data augmentation strategy to obtain a confidence of each detection frame.

[0245] The first calculation module 1003 is configured to calculate a target object detection performance value corresponding to the data augmentation strategy according to a number of detection frames and / or a confidence of the detection frames corresponding to the data augmentation strategy.

[0246] The selection module 1004 is configured to select a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0247] Optionally, the first acquisition module 1001 comprises:

[0248] The first acquisition sub-module is configured to acquire the second data set.

[0249] The second acquisition sub-module is configured to acquire a plurality of data augmentation strategies.

[0250] The processing sub-module is configured to, for each data augmentation strategy, perform data augmentation processing on the second data set by using the data augmentation strategy to obtain a first data set corresponding to each data augmentation strategy.

[0251] The training sub-module is configured to, for each data augmentation strategy, train a preset detection model by using the first data set corresponding to the data augmentation strategy to obtain a first detection model corresponding to each data augmentation strategy.

[0252] ​Optionally, the second obtaining sub-module is specifically configured to: select a plurality of groups of seed strategies based on a plurality of data augmentation algorithms included in the candidate library, each group of seed strategies including a first number of data augmentation algorithms; and determine, for each group of seed strategies, a random probability value / amplitude value corresponding to each data augmentation algorithm in the group of seed strategies based on the second number, and the maximum and minimum values of the random probability interval / amplitude transformation interval corresponding to each data augmentation algorithm in the group of seed strategies, to obtain a plurality of data augmentation strategies.

[0253] Optionally, the data augmentation strategy selection device can further include:

[0254] The second obtaining module is configured to obtain a first data set; the first data set includes a plurality of first sample images, and each first sample image includes first identification information of a target object;

[0255] The second detection module is configured to detect the target object in each first sample image by using a preset detection model to obtain a first detection result.

[0256] The second calculation module is configured to calculate a first loss value of the preset detection model based on the first detection result and the first identification information.

[0257] The adjustment module is configured to, if the first difference value is greater than a first preset threshold, adjust the parameters of the preset detection model, and return to call the second detection module to perform the step of detecting the target object in each first sample image by using the preset detection model to obtain the first detection result; the first difference value is a difference between the first loss value obtained in the current round of training and the first loss value obtained in the last round of training.

[0258] The first determination module is configured to, if the first difference value is not greater than the first preset threshold, determine the current preset detection model as a trained first detection model.

[0259] Optionally, the first identification information includes a region identifier indicating a region where the target object is located, and a category identifier indicating a category of the target object.

[0260] The second calculation module can be specifically configured to: calculate a first error value between a region where a detection frame in the first detection result is located and a region where the target object is located as indicated by the region identifier in the first data set; calculate a second error value between a category of the target object corresponding to the detection frame in the first detection result and a category of the target object as indicated by the category identifier in the first data set; and calculate a sum of the first error value and the second error value as the first loss value of the preset detection model.

[0261] Optionally, the first calculation module 1003 can be specifically configured to calculate a sum value of the number of bounding boxes corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy, and / or calculate a sum value of the confidence of the bounding boxes corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

[0262] Optionally, the selection module 1004 can be specifically configured to determine the data augmentation strategy with the highest target object detection performance value from the plurality of data augmentation strategies as the target data augmentation strategy.

[0263] Optionally, the data augmentation strategy selection apparatus can further include:

[0264] The third acquisition module is configured to acquire a fourth data set and a fifth data set, wherein the fourth data set includes a plurality of fourth sample images and third identification information corresponding to a target object in each fourth sample image, and the fifth data set includes a plurality of fifth sample images.

[0265] The processing module is configured to perform data augmentation processing on the fourth data set by using the target data augmentation strategy to obtain a sixth data set.

[0266] The third detection module is configured to detect the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result.

[0267] The fourth detection module is configured to detect the target object in each fifth sample image in the fifth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a third detection result.

[0268] The third calculation module is configured to calculate a second loss value of the first detection model corresponding to the target data augmentation strategy based on the second detection result and the third detection result.

[0269] The fine-tuning module is configured to fine-tune the parameters of the first detection model corresponding to the target data augmentation strategy if the second difference is greater than a second preset threshold, and return to call the third detection module to perform the step of detecting the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain the second detection result; the second difference is a difference between the second loss value obtained in the current round of training and the second loss value obtained in the last round of training.

[0270] The second determination module is configured to determine the current first target object detection model as the trained second detection model if the second difference is not greater than the second preset threshold.

[0271] Optionally, the data augmentation strategy selection apparatus can further include:

[0272] a fourth obtaining module, configured to obtain a to-be-detected image;

[0273] a fifth detecting module, configured to perform target object detection on the to-be-detected image by using the second detection model to obtain a fourth detection result.

[0274] Through the apparatus provided in the embodiments of the present application, after obtaining the first detection model corresponding to each data augmentation strategy, the target object detection performance value corresponding to each data augmentation strategy is determined according to the detection result of the target object in the sample image, that is, the confidence of the detected bounding box in the sample image, by using the first detection model corresponding to each data augmentation strategy, and then the target data augmentation strategy is determined based on the target object detection performance value. Compared with the related art, the target object detection performance of the first detection model corresponding to each data augmentation strategy is verified by using the sample image without including the identification information as the verification data set, which makes it no longer necessary to include the identification information in the verification data set, improves the generalization ability of the verification data in the verification data set, and thus improves the generalization ability of the determined target data augmentation strategy, so that the data augmentation technology is applicable to the target object detection process in the unsupervised learning scenario.

[0275] Based on the same inventive concept, according to the data augmentation strategy selection method provided in the embodiments of the present application, the embodiments of the present application further provide a data augmentation strategy selection system, as shown in Figure 11-a FIG. 1 is a first structural schematic diagram of the data augmentation strategy selection system provided in the embodiments of the present application. The data augmentation strategy selection system includes an image acquisition device 1101 and a model training device 1102. Figure 11-a The image acquisition device 1101 is configured to obtain a plurality of second sample images and the identification information corresponding to the target object in each second sample image to obtain a second data set, and obtain a plurality of third sample images as a third data set.

[0276] The model training device 1102 is configured to perform data augmentation processing on the second data set based on a plurality of data augmentation strategies to obtain a first data set corresponding to each data augmentation strategy, and train a first detection model corresponding to each data augmentation strategy based on the first data set corresponding to each data augmentation strategy.

[0277] The model training device 1102 is configured to perform data augmentation processing on the second data set based on a plurality of data augmentation strategies to obtain a first data set corresponding to each data augmentation strategy, and train a first detection model corresponding to each data augmentation strategy based on the first data set corresponding to each data augmentation strategy.

[0278] The model training device 1102 is further configured to, for each data augmentation strategy, detect the target object in the third sample image in the third data set by using the first detection model corresponding to the data augmentation strategy, to obtain the confidence of each detection box; calculate the target object detection performance value corresponding to the data augmentation strategy according to the number of detection boxes corresponding to the data augmentation strategy and / or the confidence of the detection boxes; and select the target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy.

[0279] Optionally, as shown in Figure 11-b The data augmentation strategy selection system can further include a client 1103.

[0280] The image acquisition device 1101 can further be configured to, after obtaining the target data augmentation strategy, acquire a to-be-detected image.

[0281] The client 1103 is configured to detect the target object in the to-be-detected image by using the first detection model corresponding to the target data augmentation strategy, to obtain a fifth detection result.

[0282] Optionally, the model training device 1102 can be specifically configured to acquire a second data set; acquire a plurality of data augmentation strategies; for each data augmentation strategy, perform data augmentation processing on the second data set by using the data augmentation strategy, to obtain a first data set corresponding to each data augmentation strategy; and for each data augmentation strategy, train a preset detection model by using the first data set corresponding to the data augmentation strategy, to obtain a first detection model corresponding to each data augmentation strategy.

[0283] Optionally, the model training device 1102 can be specifically configured to select a plurality of groups of seed strategies based on a plurality of data augmentation algorithms included in a candidate library, each group of seed strategies including a first number of data augmentation algorithms; for each group of seed strategies, determine a random probability value / amplitude value corresponding to each data augmentation algorithm in the group of seed strategies based on a second number, and a maximum value and a minimum value of a random probability interval / amplitude value transformation interval corresponding to each data augmentation algorithm in the group of seed strategies, to obtain a plurality of data augmentation strategies.

[0284] Optionally, for the first data set corresponding to any data augmentation strategy, the model training device 1102 can be specifically configured to acquire the first data set; the first data set includes a plurality of first sample images and first identification information of a target object in each first sample image.

[0285] Detect the target object in each first sample image by using the preset detection model, to obtain a first detection result.

[0286] Calculate a first loss value of the preset detection model based on the first detection result and the first identification information.

[0287] if the first difference value is greater than the first preset threshold value, adjusting parameters of the preset detection model, and returning to perform the step of detecting the target object in each first sample image by using the preset detection model to obtain a first detection result; the first difference value is a difference value between the first loss value obtained in the current round of training and the first loss value obtained in the last round of training;

[0288] if the first difference value is not greater than the first preset threshold value, determining the current preset detection model as the trained first detection model.

[0289] Optionally, the first identification information includes a region identifier indicating a region where the target object is located, and a category identifier indicating a category of the target object.

[0290] The model training device 1102 can be specifically configured to calculate a first error value between a region where a detection frame in the first detection result is located and a region where the target object indicated by the region identifier in the first data set is located; calculate a second error value between a category of the target object corresponding to the detection frame in the first detection result and a category of the target object indicated by the category identifier in the first data set; and calculate a sum of the first error value and the second error value as the first loss value of the preset detection model.

[0291] Optionally, the model training device 1102 can be specifically configured to calculate a sum of the number of detection frames corresponding to the data augmentation strategy as a target object detection performance value corresponding to the data augmentation strategy; and / or calculate a sum of the confidence of the detection frame corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

[0292] Optionally, the model training device 1102 can be specifically configured to determine a data augmentation strategy with the highest target object detection performance value from the plurality of data augmentation strategies as the target data augmentation strategy.

[0293] Optionally, the model training device 1102 can be further configured to obtain a fourth data set and a fifth data set; the fourth data set includes a plurality of fourth sample images and third identification information corresponding to a target object in each fourth sample image, and the fifth data set includes a plurality of fifth sample images.

[0294] performing data augmentation processing on the fourth data set by using the target data augmentation strategy to obtain a sixth data set;

[0295] detecting the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result;

[0296] The target data augmentation strategy corresponding to the first detection model is used to detect the target object in each fifth sample image in the fifth data set, to obtain a third detection result.

[0297] Based on the second detection result and the third detection result, a second loss value of the first detection model corresponding to the target data augmentation strategy is calculated.

[0298] If the second difference is greater than the second preset threshold, the parameters of the first detection model corresponding to the target data augmentation strategy are fine-tuned, and the step of using the first detection model corresponding to the target data augmentation strategy to detect the target object in each sixth sample image in the sixth data set to obtain the second detection result is returned. The second difference is the difference between the second loss value obtained in the current round of training and the second loss value obtained in the last round of training.

[0299] If the second difference is not greater than the second preset threshold, the current first target object detection model is determined as the trained second detection model.

[0300] Optionally, the image acquisition device 1101 can also be used to obtain the to-be-detected image.

[0301] The client 1103 can also be used to use the second detection model to detect the target object in the to-be-detected image to obtain a fourth detection result.

[0302] In the embodiments of the present application, the image acquisition device is used to obtain the training data set, the verification data set, and the to-be-detected image. The image acquisition device can be a camera or a device integrated with a camera module.

[0303] The model training device can be used to train the preset detection model and the first detection model corresponding to the target data augmentation strategy. The model training module can also be used to select the target data augmentation strategy.

[0304] The client can be used to detect the target object in the to-be-detected image.

[0305] The image acquisition device, the model training device, and the client can be different devices, or can be integrated in the same device, that is, the image acquisition device and the model training device can be integrated in the client (i.e., the electronic device).

[0306] Based on the same inventive concept, according to the data augmentation strategy selection method provided in the embodiments of the present application, the embodiments of the present application also provide an electronic device, such as Figure 12As shown, the electronic device includes a processor 1201, a communication interface 1202, a memory 1203 and a communication bus 1204, wherein the processor 1201, the communication interface 1202 and the memory 1203 communicate with each other through the communication bus 1204,

[0307] The memory 1203 is used to store a computer program.

[0308] The processor 1201 is used to execute the program stored in the memory 1203 to implement the following steps:

[0309] Obtain a plurality of data augmentation strategies each corresponding to a first detection model; the first detection model corresponding to each data augmentation strategy is obtained based on a first data set corresponding to the data augmentation strategy; the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy;

[0310] For each data augmentation strategy, use the first detection model corresponding to the data augmentation strategy to detect the target object in the third sample image in the third data set to obtain the confidence of each detection box;

[0311] According to the number of detection boxes and / or the confidence of the detection boxes corresponding to the data augmentation strategy, calculate the target object detection performance value corresponding to the data augmentation strategy;

[0312] Based on the target object detection performance value corresponding to each data augmentation strategy, select a target data augmentation strategy.

[0313] Through the electronic device provided by the embodiment of the present application, after obtaining a plurality of data augmentation strategies each corresponding to a first detection model, the target object detection performance value corresponding to each data augmentation strategy can be determined according to the detection result of the target object in the sample image, that is, the confidence of the detection box detected in the sample image, by using the first detection model corresponding to each data augmentation strategy. Then, the target data augmentation strategy is determined based on the target object detection performance value. Compared with the related art, the target object detection performance of the first detection model corresponding to each data augmentation strategy is verified by using the sample image without including the identification information as the verification data set, which makes it no longer necessary to include the identification information in the verification data set, improves the generalization ability of the verification data in the verification data set, and thus improves the generalization ability of the target data augmentation strategy determined, so that the data augmentation technology is applicable to the target object detection process in the unsupervised learning scenario.

[0314] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0315] The communication interface is used for communication between the above electronic device and other devices.

[0316] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.

[0317] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0318] Based on the same kind of invention concept, according to the data augmentation strategy selection method provided by the above embodiments of the application, the embodiments of the application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of any of the above data augmentation strategy selection methods.

[0319] Based on the same kind of invention concept, according to the data augmentation strategy selection method provided by the above embodiments of the application, the embodiments of the application further provide a computer program product containing instructions, which, when running on a computer, causes the computer to execute any of the data augmentation strategy selection methods in the above embodiments.

[0320] In the embodiments described above, all or some of the steps can be implemented by software, hardware or firmware, or any combination thereof. When implemented by software, all or some of the steps can be implemented in the form of one or more computer programs. The computer program can be stored in any computer readable medium, and when loaded into a computer system, causes the computer system to perform one or more of the steps of the computer program. The computer readable medium can be a machine readable storage medium such as a floppy drive, a ZIP disk, a hard disk, electronic storage devices, a magnetic tape, a CD-ROM, a DVD, a memory stick, or a computer database. The computer readable medium can also be a machine readable signal medium, such as a data signal embodied in a carrier wave, a digital signal converted into one or more types of RF or analog signals, or other type of modulated data signals.

[0321] It should be noted that, in the description, relative terms such as first and second are used merely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Also, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.

[0322] Each of the embodiments in the present specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the embodiments of an apparatus, a system, an electronic device, a computer readable storage medium, and a computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the description of the method embodiments.

[0323] The above merely provides the preferred embodiment of the present application, and not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for data augmentation policy selection, the method comprising: The method comprises: obtaining a plurality of data augmentation strategies each corresponding to a first detection model; the first detection model corresponding to each data augmentation strategy is obtained based on a first data set corresponding to the data augmentation strategy; the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy; for each data augmentation strategy, using the first detection model corresponding to the data augmentation strategy to detect the target object in the third sample image in the third data set to obtain the confidence of each detection box, wherein the third data set does not include the identification information corresponding to the target object in each third sample image; according to the number of detection boxes and / or the confidence of the detection boxes corresponding to the data augmentation strategy, calculating the target object detection performance value corresponding to the data augmentation strategy; based on the target object detection performance value corresponding to each data augmentation strategy, selecting a target data augmentation strategy; the obtaining method of the plurality of data augmentation strategies comprises: selecting a plurality of seed strategies based on a plurality of data augmentation algorithms included in a candidate library, each seed strategy including a first number of data augmentation algorithms; for each seed strategy, based on a second number, and the maximum and minimum values of the random probability interval / amplitude value transformation interval corresponding to each data augmentation algorithm in the seed strategy, determining the random probability value / amplitude value corresponding to each data augmentation algorithm in the seed strategy, to obtain a plurality of data augmentation strategies.

2. The method of claim 1, wherein, The step of obtaining a plurality of data augmentation strategies each corresponding to a first detection model comprises: obtaining the second data set; obtaining a plurality of data augmentation strategies; for each data augmentation strategy, using the data augmentation strategy to perform data augmentation processing on the second data set to obtain a first data set corresponding to each data augmentation strategy; for each data augmentation strategy, using the first data set corresponding to the data augmentation strategy to train a preset detection model to obtain a first detection model corresponding to each data augmentation strategy.

3. The method of claim 1, wherein, For any data augmentation strategy corresponding to the first data set, the following steps are used to train the first detection model corresponding to the data augmentation strategy: obtaining the first data set; the first data set includes a plurality of first sample images and first identification information of the target object in each first sample image; using a preset detection model to detect the target object in each first sample image to obtain a first detection result; based on the first detection result and the first identification information, calculating a first loss value of the preset detection model; if the first difference is greater than a first preset threshold, adjusting the parameters of the preset detection model, and returning to execute the step of using the preset detection model to detect the target object in each first sample image to obtain a first detection result; the first difference is the difference between the first loss value obtained in the current training and the first loss value obtained in the last training; if the first difference is not greater than the first preset threshold, the current preset detection model is determined as the trained first detection model.

4. The method of claim 3, wherein, The first identification information includes a region identifier indicating the region where the target object is located, and a category identifier indicating the category of the target object; The step of calculating the first loss value of the preset detection model based on the first detection result and the first identification information comprises: calculating a first error value between a region where a detection frame in the first detection result is located and a region where a target object indicated by region identification in the first data set is located; calculating a second error value between a category of the target object corresponding to the detection frame in the first detection result and a category of the target object indicated by category identification in the first data set; calculating a sum value of the first error value and the second error value as the first loss value of the preset detection model.

5. The method of claim 1, wherein, The step of calculating the target object detection performance value corresponding to the data augmentation strategy according to the number of detection frames and / or the confidence of the detection frames of the data augmentation strategy comprises: calculating a sum value of the number of detection frames corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy; and / or calculating a sum value of the confidence of the detection frames corresponding to the data augmentation strategy as the target object detection performance value corresponding to the data augmentation strategy.

6. The method of claim 1, wherein, The step of selecting a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy comprises: determining a data augmentation strategy with the highest target object detection performance value in the plurality of data augmentation strategies as the target data augmentation strategy.

7. The method of claim 1, wherein, The method further comprises: obtaining a fourth data set and a fifth data set; wherein the fourth data set comprises a plurality of fourth sample images, and third identification information corresponding to a target object in each fourth sample image, and the fifth data set comprises a plurality of fifth sample images; performing data augmentation processing on the fourth data set by using the target data augmentation strategy to obtain a sixth data set; detecting the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result; detecting the target object in each fifth sample image in the fifth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a third detection result; calculating a second loss value of the first detection model corresponding to the target data augmentation strategy based on the second detection result and the third detection result; if the second difference value is greater than a second preset threshold, fine-tuning the parameters of the first detection model corresponding to the target data augmentation strategy, and returning to perform the step of detecting the target object in each sixth sample image in the sixth data set by using the first detection model corresponding to the target data augmentation strategy to obtain a second detection result; the second difference value is a difference value between the second loss value obtained in the current round of training and the second loss value obtained in the last round of training; if the second difference value is not greater than the second preset threshold, determining the current first target object detection model as a trained second detection model.

8. The method of claim 7, wherein, The method further comprises: obtaining a to-be-detected image; detecting the target object in the to-be-detected image by using the second detection model to obtain a fourth detection result.

9. A data augmentation policy selection apparatus characterized by comprising: The device comprises: The first obtaining module is configured to obtain a first detection model corresponding to each data augmentation strategy; the first detection model corresponding to each data augmentation strategy is trained based on a first data set corresponding to the data augmentation strategy; the first data set corresponding to each data augmentation strategy is obtained by performing data augmentation processing on a second data set based on the data augmentation strategy; The first detection module is configured to, for each data augmentation strategy, perform detection on a target object in a third sample image in a third data set by using the first detection model corresponding to the data augmentation strategy to obtain a confidence of each detection frame, wherein the third data set does not include identification information corresponding to the target object in each third sample image; The first calculation module is configured to calculate a target object detection performance value corresponding to the data augmentation strategy according to a number of detection frames corresponding to the data augmentation strategy and / or a confidence of the detection frames; The selection module is configured to select a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy. The first obtaining module is specifically configured to select multiple groups of seed strategies based on multiple data augmentation algorithms included in a candidate library, each group of seed strategies including a first number of data augmentation algorithms; for each group of seed strategies, determine a random probability value / amplitude value corresponding to each data augmentation algorithm in the group of seed strategies based on a second number, a maximum value and a minimum value of a random probability interval / amplitude transformation interval corresponding to each data augmentation algorithm in the group of seed strategies, and obtain multiple data augmentation strategies.

10. A data augmentation policy selection system, comprising: The system includes an image acquisition device and a model training device; The image acquisition device is configured to obtain multiple second sample images and identification information corresponding to a target object in each second sample image to obtain a second data set, and obtain multiple third sample images as a third data set, wherein the third data set does not include identification information corresponding to a target object in each third sample image; The model training device is configured to perform data augmentation processing on the second data set based on multiple data augmentation strategies to obtain a first data set corresponding to each data augmentation strategy; and train a first detection model corresponding to each data augmentation strategy based on the first data set corresponding to the data augmentation strategy; The model training device is further configured to, for each data augmentation strategy, perform detection on a target object in a third sample image in the third data set by using the first detection model corresponding to the data augmentation strategy to obtain a confidence of each detection frame; calculate a target object detection performance value corresponding to the data augmentation strategy according to a number of detection frames corresponding to the data augmentation strategy and / or a confidence of the detection frames; and select a target data augmentation strategy based on the target object detection performance value corresponding to each data augmentation strategy. The model training device is specifically configured to select multiple groups of seed strategies based on multiple data augmentation algorithms included in the candidate library, each group of seed strategies including a first number of data augmentation algorithms; for each group of seed strategies, determine a random probability value / amplitude value corresponding to each data augmentation algorithm in the group of seed strategies based on a second number, and a maximum value and a minimum value of a random probability interval / amplitude transformation interval corresponding to each data augmentation algorithm in the group of seed strategies, to obtain multiple data augmentation strategies.

11. The system of claim 10, wherein, The system further includes a client; The image acquisition device is further configured to obtain a to-be-detected image after obtaining the target data augmentation strategy. The client is configured to use a first detection model corresponding to the target data augmentation strategy to perform target object detection on the to-be-detected image to obtain a fifth detection result.

Citation Information

Patent Citations

  • Data augmentation method and device, service processing method and device, computer equipment and storage medium

    CN111783902A

  • Detection model processing method and device

    CN112633496A