Target detection capability training method, device, storage medium and electronic device
By inheriting the parameters of the initial detection model and using fewer truth samples, the image object detection ability training method solves the problem of high acquisition cost of truth samples in the prior art, and achieves a significant improvement in object detection ability and self-learning of the model.
Patent Information
- Application Number
- CN202210164319.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-02-22
AI Technical Summary
The existing strongly supervised learning object detection algorithm relies on a large number of manual annotated truth-value sample data, resulting in high time and economic costs and difficult to guarantee quality.
Provide a training method for image object detection ability, which gradually enhances object detection ability by inheriting the parameters of the initial detection model and using fewer truth samples and simple image-level tags to achieve 'online learning'.
This method can significantly improve the target detection capability while using fewer truth samples, reduce the cost of truth samples acquisition, and realize self-learning and continuous enhancement of the model.
Smart Images

Figure CN114596484B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronics, and in particular to an image target detection capability training method, an image target detection device, a computer-readable storage medium, and an electronic device. Background Art
[0002] Existing strongly supervised learning target detection algorithms rely on a large amount of manually annotated true value sample data (GT data) to train target detection / segmentation models. The time and economic cost of producing GT data is very high and the quality is difficult to guarantee. Summary of the invention
[0003] In response to the above problems, an embodiment of the present application provides an image target detection capability training method, which only requires fewer true value samples and input samples containing only simple image-level labels to obtain stronger target detection capabilities. As the input samples continue to increase on their own, the number is increasing, the scenes are becoming more and more abundant, the model's self-labeling ability is becoming stronger and stronger, and the model's target detection ability can also continue to self-enhance, realizing "online learning".
[0004] The first aspect of the present application provides an image object detection capability training method, which includes:
[0005] Training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability;
[0006] Analyzing the second image with the second identifier using the first detection model to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0007] The second detection model determines, based on the first information, positive examples and negative examples of targets of each category among the multiple targets; and
[0008] The second detection model acquires the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model acquires a second detection capability.
[0009] The image target detection capability training method of the embodiment of the present application can inherit the model parameters of the first detection model, so that the second detection model can obtain the first detection capability of the first detection model. At the same time, the second detection model can also determine the positive examples and negative examples of targets of each category among the multiple targets based on the first information obtained by the first detection model, and train the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains the second detection capability, thereby improving the detection capability of the second detection model, so that only a small number of first images with the first identification (i.e., a small number of true value samples) are needed to obtain a stronger target detection capability.
[0010] Further, after the positive examples and the negative examples of the targets of each category in the targets are trained so that the second detection model acquires a second detection capability, the method further includes:
[0011] The second detection model is used to analyze the third image, the fourth image and the test image to obtain a first detection result, a second detection result and a third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, the third image is an unlabeled first image, the fourth image is an unlabeled second image, the fifth image is the first image with the third label, and the sixth image is the second image with the fourth label; and
[0012] When the accuracy of the third detection result meets the preset conditions or the training times of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for cyclic training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0013] Since the second detection model inherits the detection capability of the first detection model, the first detection capability of the first detection model is improved, and the second detection capability of the second detection model is also improved accordingly. As the detection capability of the second detection model is continuously improved, the accuracy of the first detection result and the second detection result obtained each time will also be continuously improved. The first detection result and the second detection result with improved accuracy are used as new training sets to train the first detection model, which can improve the first detection capability of the first detection model, and then improve the second detection capability of the second detection model. In this way, only a small number of first images and second images are needed. By looping through the first image, the second image, the first detection result and the second detection result, the detection capability and accuracy of the second detection model can be continuously improved. A large number of true value samples are not required, which greatly reduces the cost of obtaining true value samples (i.e., the cost of obtaining the first image with the first identifier). At the same time, the second detection model can be made to have the ability of self-learning, and the second detection capability of the second detection model can be continuously improved. The second image data set does not change in the complete process of this application, but when the algorithm is trained again, the second image data set will be a new data set. Since new second image data sets can always be continuously input during the use of the algorithm, at a certain stage the second image data set can automatically annotate image-level label information through the detection results, so the algorithm model can continuously provide positive feedback and self-enhancement to acquire the ability of self-learning.
[0014] Furthermore, after the first detection model and the second detection model complete the specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n=1, the method further includes:
[0015] Using the first detection model to detect the test image to obtain a fourth detection result;
[0016] The third detection result meets the preset conditions, including:
[0017] When the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold.
[0018] By limiting the loop conditions, the detection capability and accuracy of the second detection model can be improved according to actual needs. When using fewer true value samples, the detection capability of the second detection model can be continuously improved, the cost of true value samples can be effectively reduced, and continuous self-learning can be achieved.
[0019] Furthermore, after the first detection model and the second detection model complete the specified rounds of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n≥2, the method further includes:
[0020] Training the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model;
[0021] Using the second detection model to detect the test image to obtain an nth third detection result;
[0022] The third detection result meets the preset conditions, including:
[0023] When the difference between the accuracy of the third detection result of the nth time and the accuracy of the third detection result of the (n-1)th time is greater than a second preset threshold.
[0024] By limiting the loop conditions, the detection capability and accuracy of the second detection model can be improved according to actual needs. When using fewer true value samples, the detection capability of the second detection model can be continuously improved, the cost of true value samples can be effectively reduced, and continuous self-learning can be achieved.
[0025] Furthermore, before training the first image with the first identifier, the method further includes:
[0026] The data enhancement algorithm in the target detection algorithm is used to perform illumination distortion, geometric distortion and image occlusion on the first image, so as to improve the generalization of the scene in the first image.
[0027] Furthermore, the first information also includes the orientation information of each target in the second image and the confidence of each target, and the orientation information of each target includes the coordinates and vector of the target. In this way, not only the position information of each target can be obtained, but also the size of each target can be obtained, so that the target can be better identified to better achieve the obstacle avoidance function.
[0028] The second aspect of the present application provides an image object detection capability training device, which includes:
[0029] A first detection unit, configured to train a first image having a first identifier so that a first detection model acquires a first detection capability and forms a model parameter for the first detection capability; and
[0030] a second detection unit, configured to analyze the second image with the second identifier by using the first detection model to obtain first information, wherein the first information at least includes a plurality of targets and categories of the plurality of targets;
[0031] The second detection unit is further used for the second detection model to determine positive examples and negative examples of targets of each category among the multiple targets according to the first information;
[0032] The second detection unit is also used for the second detection model to obtain the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains second detection capability.
[0033] The image target detection capability training device of the embodiment of the present application can inherit the model parameters of the first detection model, so that the second detection model can obtain the first detection capability of the first detection model. At the same time, the second detection model can also determine the positive examples and negative examples of targets of each category among the multiple targets based on the first information obtained by the first detection model, and train the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains the second detection capability, thereby improving the detection capability of the second detection model, so that only a small number of first images with the first identification (i.e., a small number of true value samples) are needed to obtain a stronger target detection capability.
[0034] Further, the second detection unit is further used to analyze the third image, the fourth image and the test image using the second detection model to obtain a first detection result, a second detection result and a third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, the third image is an unlabeled first image, the fourth image is an unlabeled second image, the fifth image is the first image with a third label, and the sixth image is the second image with a fourth label;
[0035] The image target detection capability training device also includes an analysis and judgment unit, which is used to send the first detection result and the second detection result to the first detection model for cyclic training when the accuracy of the third detection result meets a preset condition or the training times of the second detection model is less than a first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0036] Since the second detection unit inherits the detection capability of the first detection model, the first detection capability of the first detection model is improved, and the second detection capability of the second detection unit is also improved accordingly. As the detection capability of the second detection unit is continuously improved, the accuracy of the first detection result and the second detection result obtained each time will also be continuously improved. The first detection result and the second detection result with improved accuracy are used as new training sets to train the first detection model, which can improve the first detection capability of the first detection model and then improve the second detection capability of the second detection unit. In this way, only a small number of first images and a small number of second images are needed. By looping the first image, the second image, the first detection result and the second detection result, the detection capability and accuracy of the second detection unit can be continuously improved. A large number of true value samples are not required, which greatly reduces the cost of obtaining true value samples (i.e., the cost of obtaining the first image with the first identifier). At the same time, the second detection model can be made to have the ability of self-learning, and the second detection capability of the second detection unit can be continuously improved.
[0037] Furthermore, after the first detection model and the second detection model complete the specified round of training, a test is performed on the first image without any mark, the second image and the test image, and the number of tests is n, where n is a positive integer. When n=1,
[0038] The first detection unit is further used to detect the test image using the first detection model to obtain a fourth detection result;
[0039] The analysis and judgment unit is also used to send the first detection result and the second detection result to the first detection model for further training when the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold; or the training times of the second detection model is less than the first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0040] By limiting the loop conditions, the detection capability and accuracy of the second detection unit can be improved according to actual needs. When using fewer true value samples, the detection capability of the second detection unit can be continuously improved, the cost of true value samples can be effectively reduced, and continuous self-learning can be achieved.
[0041] Furthermore, after the first detection model and the second detection model complete the specified round of training, the first image without any mark, the second image and the test image are tested once, and the number of tests is n, where n is a positive integer, and when n≥2,
[0042] The first detection unit is further used to train the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model;
[0043] Using the second detection model to detect the test image to obtain an nth third detection result;
[0044] The analysis and judgment unit is further configured to: when the difference between the accuracy of the third detection result of the nth time and the accuracy of the third detection result of the (n-1)th time is greater than a second preset threshold.
[0045] By limiting the loop conditions, the detection capability and accuracy of the second detection unit can be improved according to actual needs. When using fewer true value samples, the detection capability of the second detection unit can be continuously improved, the cost of true value samples can be effectively reduced, and continuous self-learning can be achieved.
[0046] Furthermore, the first detection unit is further used to: use a data enhancement algorithm in a target detection algorithm to perform illumination distortion, geometric distortion and image occlusion on the first image, so as to improve the generalization of scenes in the first image.
[0047] Furthermore, the first information also includes the orientation information of each target in the second image and the confidence of each target, and the orientation information of each target includes the coordinates and vector of the target. In this way, not only the position information of each target can be obtained, but also the size of each target can be obtained, so that the target can be better identified to better achieve the obstacle avoidance function.
[0048] The third aspect of the present application provides a computer-readable storage medium, which stores a computer-executable program code, and the computer-executable program code is used to enable a computer to execute the method described in the embodiment of the present application.
[0049] The third aspect of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores a program code executable by the processor, and when the program code is called and executed by the processor, the method described in the embodiment of the present application is executed.
[0050] The image target detection capability training method of the embodiment of the present application can inherit the model parameters of the first detection model, so that the second detection model can obtain the first detection capability of the first detection model. At the same time, the second detection model can also determine the positive examples and negative examples of targets of each category among the multiple targets based on the first information obtained by the first detection model, and train the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains the second detection capability, thereby improving the detection capability of the second detection model, so that only a small number of first images with the first identification (i.e., a small number of true value samples) are needed to obtain a stronger target detection capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0052] Figure 1 It is a flowchart of an image target detection capability training method according to an embodiment of the present application.
[0053] Figure 2 It is a flowchart of an image target detection capability training method according to another embodiment of the present application.
[0054] Figure 3 It is a flowchart of an image target detection capability training method according to another embodiment of the present application.
[0055] Figure 4 It is a flowchart of an image target detection capability training method according to another embodiment of the present application.
[0056] Figure 5 It is a flowchart of an image target detection capability training method according to another embodiment of the present application.
[0057] Figure 6 It is a structural block diagram of an image target detection capability training device according to an embodiment of the present application.
[0058] Figure 7 It is a structural block diagram of an image target detection capability training device according to another embodiment of the present application.
[0059] Figure 8 It is a circuit block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0061] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0062] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0063] It should be noted that, for the convenience of explanation, in the embodiments of the present application, the same figure marks represent the same components, and for the sake of brevity, detailed descriptions of the same components are omitted in different embodiments.
[0064] The current object detection analysis models applied to scene images, such as the strongly supervised object detection model, need to rely on a large amount of manually annotated true value sample data (GT data) to train the object detection / segmentation model. The annotation of true value sample data takes a long time, is costly, and the quality is difficult to guarantee. In addition, the strong supervised learning algorithm has insufficient generalization ability for unlearned scenes and is difficult to be universally applicable. Therefore, the traditional strongly supervised learning algorithm can only perform front-end optimization. Once the algorithm is transplanted and deployed, the strongly supervised object detection model cannot improve itself and cannot achieve "online learning".
[0065] The detection model obtained by the image target detection capability training method provided by the present application can be applied to the detection and segmentation of targets in scene images, such as cars, people, obstacles, etc., and the identification of the orientation information and categories of the targets. When applied to autonomous driving vehicles, it can achieve a good obstacle avoidance function and reduce the risk of autonomous driving. In addition, the image target detection capability training method of the present application can also be applied to the fields of face recognition, medical detection, etc.
[0066] See also Figure 1 The present application embodiment provides an image target detection capability training method, which includes:
[0067] S101, training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability;
[0068] In some embodiments, before training the first image with the first identification, the method further includes: acquiring the first image, and performing a first identification on the first image.
[0069] Optionally, a scene of the original data is selected according to the application scenario, and a first image of the scene is obtained. The scene of the original data can be selected according to the target categories that may exist in the scenario to which the method of the present application is applied, such as selecting a scene that includes all target categories, or a scene that includes main target categories, such as people, cars, road signs, trees, and other targets that need to be accurately identified during autonomous driving. In some embodiments, the first image can be obtained by an image acquisition device such as a camera, a camcorder, or a mobile phone. In other embodiments, the first pattern can also be obtained from a commonly used image database, a network database, and the like.
[0070] Furthermore, the first image is manually annotated so that the first image has a first identification. In other words, the image sample (i.e., the first image) is manually annotated to obtain true value sample data (i.e., the first image with the first identification). The first identification includes all targets in the first image, category information of all targets, orientation information of all targets, etc. Specifically, the orientation information of the target can be annotated in the form of a target candidate box. In other words, each target in the first image is selected using a candidate box, and the orientation of the candidate box is the orientation information of the target.
[0071] In some embodiments, before training the first image with the first identifier, the method also includes: using a data enhancement algorithm in an object detection algorithm (OD algorithm) to perform effective data enhancement operations such as illumination distortion, geometric distortion, and image occlusion on the first image to improve the generalization of the scene in the first image.
[0072] Optionally, training the first image with the first identifier so that the first detection model acquires the first detection capability and forms a model parameter of the first detection capability may be:
[0073] The manually annotated first image (ie, having a first identifier) is sent to the first detection model for training and learning by the first detection model, thereby obtaining a first detection capability for the target and forming model parameters having the first detection capability.
[0074] Specifically, a strongly supervised target detection model SSLOD (such as a target detection model based on the Mask_Rcnn algorithm) is selected, and the first image with the first identifier (or the original true value sample) is sent to the first detection model (SSLOD model) for training; at the same time, the attention mechanism technology (such as SENet, i.e., a neural network with an attention mechanism) is combined in the training process to add high weights to the key channels or regions of the feature map (i.e., the first image), and low weights to irrelevant or non-key channels or regions, so that the training process of the first detection model is more targeted, which is of obvious benefit to the detection of small and medium-sized targets (i.e., targets with small volume or size); at the same time, the SSP-Net (feature pyramid pooling network) feature pyramid structure can be used to perform multi-scale feature fusion, so that targets of different sizes can obtain fixed-size outputs after entering the first detection model without performing normalization scaling on the input target, which further improves the OD training results; at the same time, since the first detection model has true value sample data (the first image with the first identifier), the NMS (non-maximum suppression) technology can be used according to the true value sample data to filter the overlapping frame phenomenon in target detection. Therefore, by training and learning the first image with the first identification, the first detection model can obtain better target detection capability (i.e., the first detection capability), and can detect multiple categories of targets to be detected defined in the image (for example, K categories, where K is a positive integer), and output the target's orientation information, category information and confidence.
[0075] Optionally, the first detection model may be, but is not limited to, a strongly supervised target detection model (Semi-Supervised Learning Object Detection, SSLOD model). The number of first images is multiple, specifically, it may be, but is not limited to, 30,000, 50,000, 100,000, 200,000, 300,000, 500,000, etc. The more the number of first images, the better, but the more the number, the higher the cost. The specific number can be determined according to the usage scenario and the first detection capability that the initial model of the first detection model needs to achieve.
[0076] The existing strong supervision target detection model needs to meet the preset target detection capability and detection accuracy to be put into practical application. Before it is put into application, a lot of training and learning is required, so a large number of manually labeled true value samples need to be provided for its training and learning. In this embodiment, as long as the first detection model can obtain a certain first detection capability through the first image, the number of first images in this embodiment can be much less than the number of true value samples required for the formation of the strong supervision target detection model. For example, it can be 1%, 3%, 5%, 8%, 10%, etc. of the number of samples required by the existing model.
[0077] S102, using the first detection model to analyze the second image with the second identifier to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0078] In some embodiments, the first detection model is used to analyze a second image having a second identification, and the method further includes: acquiring the second image and performing a second identification on the second image.
[0079] Specifically, based on the actual possible typical application scenarios, a batch of new raw data without any manual annotation (i.e., the second image) is randomly sampled, and the second image is simply annotated at the image level; in other words, only the categories of the objects included in the second image are identified, and the orientation information of each category is not annotated. For example, if the second image includes several "cars" and several "persons", the label of the image is ["car", "person"], and there is no need to care about how many "cars" and "persons" there are and in which area of the image they are. Optionally, the second identification includes the categories of all objects in the second image; in other words, the second identification includes simple image-level labels.
[0080] In some embodiments, the first image is different from the second image, in other words, the first image and the second image are images of different scenes. In other embodiments, the second image is a part of the first image that is different, in other words, the second image may be the first image and may be selected from the first image.
[0081] Optionally, the first detection model is used to analyze the second image with the second identification to obtain the first information. This can be done by: using the first detection model that has obtained the first detection capability, analyzing and detecting the second image with the second identification to obtain the first information in each second image.
[0082] Optionally, the first information also includes the position information of each target in the second image and the confidence of each target, and the position information of each target includes the coordinates and vector of the target.
[0083] Specifically, the orientation information of the target can be annotated in the form of a target candidate box. In other words, each target in the second image is selected using a candidate box. Optionally, the candidate box can be a rectangular box, and the coordinates or vectors of the four corners of the rectangular box constitute the orientation information of the target. By obtaining the coordinates or vectors of the four points of the target candidate box, the specific position of the target can be accurately obtained. In addition, the size of the target can be calculated by the coordinates or vectors of the four corners of the candidate box, so that the target can be better identified to better achieve the obstacle avoidance function.
[0084] S103, determining, by a second detection model, positive examples and negative examples of targets of each category among the multiple targets according to the first information; and
[0085] Optionally, the second detection model may be, but is not limited to, a multi-instance learning module (MIL module), and the MIL module includes a multi-instance classifier (MILClassifier).
[0086] Optionally, the MIL module of the second detection model is used to obtain the first information obtained by the first detection model according to the second image, and based on the first information, the positive examples and negative examples of each category of the multiple targets are determined. For example, for "person", "person" is its positive bag (positive Box), and each pixel in the positive bag is a positive example of the category "person", and the category other than "person" is the negative bag (negative Box) of "person", and each pixel in the negative bag is a negative example of "person". In this way, positive examples and negative examples of targets of each category can be obtained. When the second image includes targets of K categories, all targets of each image can be converted into binary classification data sets, thereby obtaining K binary classification data sets.
[0087] S104, the second detection model obtains the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category of the multiple targets so that the second detection model obtains a second detection capability.
[0088] Optionally, transfer style transfer learning is used to transfer the model parameters of the first detection model to the MIL module of the second detection model, so that the second detection model acquires the model parameters of the first detection model to inherit the first detection capability of the first detection model. At the same time, the multi-instance classifier in the second detection model is used to train the positive examples and negative examples of each category of the multiple targets, so that the second detection model acquires the ability to distinguish the positive examples and negative examples of each category of targets, thereby acquiring the ability of target detection, that is, the second detection capability.
[0089] In order to obtain better training results, the MIL_classifier model of the second detection model still uses the attention mechanism and multi-scale feature fusion operation in the SSLOD module (the first detection model), and adopts an overlapping frame suppression algorithm similar to non-maximum suppression (for example, in the definition of the loss function loss_function, an additional penalty term is added to the loss_function of the target with a large IOU (intersection-over-union ratio) with the detection frame with the highest feature score but a small confidence C, so as to eliminate the influence of such overlapping frames). Finally, the MIL_classifier model obtains the ability to perform binary classification of K-class pixels in the image, and can identify whether each pixel in the image belongs to a certain class of targets in the K classes. In other words, it obtains the ability to classify positive examples and negative examples in the K-class targets.
[0090] The image target detection capability training method of the embodiment of the present application can inherit the model parameters of the first detection model, so that the second detection model can obtain the first detection capability of the first detection model. At the same time, the second detection model can also determine the positive examples and negative examples of targets of each category among the multiple targets based on the first information obtained by the first detection model, and train the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains the second detection capability, thereby improving the detection capability of the second detection model, so that only a small number of first images with the first identification (i.e., a small number of true value samples) are needed to obtain a stronger target detection capability.
[0091] See also Figure 2 The present application also provides an image target detection capability training method, which includes:
[0092] S201, training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability;
[0093] S202, using the first detection model to analyze the second image with the second identifier to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0094] S203, determining, by a second detection model, positive examples and negative examples of targets of each category among the multiple targets according to the first information;
[0095] S204: the second detection model acquires the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of the targets of each category of the multiple targets so that the second detection model acquires a second detection capability;
[0096] For a detailed description of steps S201 to S203, please refer to the description of the corresponding parts of the above embodiment, which will not be repeated here.
[0097] S205, using the second detection model to analyze the third image, the fourth image and the test image, to obtain a first detection result, a second detection result and a third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, wherein the third image is an unlabeled first image, the fourth image is an unlabeled second image, the fifth image is the first image with the third label, and the sixth image is the second image with the fourth label; and
[0098] Optionally, the first image without any manual annotation, the second image without any manual annotation, and the test image (testdata) without any manual annotation are sent to the MIL module of the second detection model for detection and analysis to obtain the first detection result of the first image, the second detection result of the second image, and the third detection result of the test image; wherein the first detection result includes the first image without identification and the first image with the third identification, and the second detection result includes the second image without identification and the second image with the fourth identification. Specifically, after the MIL module of the second detection model is used to perform pixel binary classification of K categories (i.e., positive example and negative example classification) on the first image, the second image, and the test image, the instantiation segmentation of each category of the first image, the second image, and the test image can be achieved through certain logical post-processing, that is, a certain target is significantly distinguished from the background and other targets, and finally the target segmented by the instance is surrounded by a minimum rectangle, and the enclosing box is the OD detection result of the target. Optionally, the second detection model also includes a pseudo label module (Pseude_GT module), which uses the pseudo label module of the second detection model to replace the OD detection results of the first image, the second image and the test image through template replacement, so that they become a data format that can be used for training and testing (such as .json or .xml format, etc.), that is, the first detection result, the second detection result and the third detection result are obtained respectively.
[0099] Optionally, the third identification includes the orientation information, category and confidence of each target in the first image detected by the second detection model. The fourth identification includes the orientation information, category and confidence of each target in the second image detected by the second detection model. The third detection result includes the orientation information, category and confidence of all targets in the test image.
[0100] It should be noted that the test images are only used for testing during the entire learning and training process, and are not used for model training.
[0101] S206, when the accuracy of the third detection result meets the preset conditions or the training times of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for cyclic training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0102] Optionally, the first preset threshold can be but not limited to 2 times, 3 times, 4 times, 5 times, 6 times, etc., and can be specifically determined based on the accuracy and economic cost of the second detection capability required to be obtained by the second detection model. This application does not make any specific limitations.
[0103] Optionally, the pseudo-label module of the second detection model is used to judge the accuracy of the third detection result of the test image. When the accuracy of the third detection result reaches a preset condition and is less than a preset value, such as but not limited to less than 90%, less than 95%, etc., or when the number of training times of the second detection model is less than the first preset threshold, it is considered that the detection accuracy of the model can be further improved. At this time, the first detection result and the second detection result are used as training sets and sent to the first detection model to continue training and learning to improve the first detection capability of the first detection model. In other words, the accuracy of the first detection model for target detection is improved, thereby improving the second detection capability of the second detection model, that is, improving the accuracy of the second detection model for target detection. Since the second detection model inherits the detection capability of the first detection model, the first detection capability of the first detection model is improved, and the second detection capability of the second detection model is also improved accordingly. As the detection capability of the second detection model continues to improve, the accuracy of the first detection result and the second detection result obtained each time will also continue to improve. The first detection result and the second detection result with improved accuracy are used as new training sets to train the first detection model, which can improve the first detection capability of the first detection model, and then improve the second detection capability of the second detection model. In this way, only a small number of first images and a small number of second images are needed. By looping through steps S201 to S206 for the first image, the second image, the first detection result, and the second detection result, the detection capability and accuracy of the second detection model can be continuously improved. A large number of true value samples are not required, which greatly reduces the cost of obtaining true value samples (i.e., the cost of obtaining the first image with the first identification). At the same time, the second detection model can be enabled to have the ability of self-learning, and the second detection capability of the second detection model can be continuously improved.
[0104] In addition, when the user continuously collects new scene images and needs to upgrade the version of the second detection model, the new scene image can be used as the second image, and steps S202 to S206 can be executed again in a loop. Since new second image data sets can always be continuously input during the use of the algorithm, at a certain stage the second image data set can automatically annotate image-level label information through the detection results. Therefore, the second detection model can continuously provide positive feedback and self-enhancement, acquire the ability of self-learning, and realize self-iterative version updates.
[0105] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0106] See also Figure 3 The present application also provides an image target detection capability training method, which includes:
[0107] S301, training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability;
[0108] S302: Analyze the second image with the second identifier using the first detection model to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0109] S303, determining, by a second detection model, positive examples and negative examples of targets of each category among the multiple targets according to the first information;
[0110] S304: the second detection model acquires the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of the targets of each category in the multiple targets so that the second detection model acquires a second detection capability;
[0111] S305, using the second detection model to analyze the unlabeled first image, the unlabeled second image, and the unlabeled test image, to obtain a first detection result, a second detection result, and a third detection result, respectively, wherein the first detection result includes the unlabeled first image and the first image with the third label, and the second detection result includes the unlabeled second image and the second image with the fourth label;
[0112] For a detailed description of steps S301 to S305, please refer to the description of the corresponding parts of the above embodiment, which will not be repeated here.
[0113] S306, using the first detection model to detect the test image to obtain a fourth detection result; and
[0114] Optionally, the test image is sent to the first detection model to perform detection analysis on it to obtain a fourth detection result of the test image, wherein the fourth detection result includes the category, orientation information and confidence of all targets included in the test image.
[0115] After the first detection model and the second detection model complete the specified rounds of training, the third image, the fourth image and the test image will be tested once. The number of tests is n, where n is a positive integer. When n=1, step S307 is executed.
[0116] S307, when the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than the second preset threshold or the training times of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0117] Optionally, after the first detection model and the second detection model complete a round of training and obtain the first detection result, the second detection result and the third detection result of the first unlabeled first image, the unlabeled second image and the unlabeled test image respectively, the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is compared with the second preset threshold. When the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than the second preset threshold or the training number of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0118] Optionally, the accuracy of the third detection result can be obtained by the confidence of each target obtained by detecting the test image with the first detection model, and the accuracy of the fourth detection result can be obtained by the confidence of each target obtained by detecting the test image with the second detection model.
[0119] Optionally, the second preset threshold may be, but is not limited to, 3%, 5%, 8%, 10%, 15%, etc.
[0120] In a specific embodiment, the first preset threshold is 3 times, and the second preset threshold is 5%. When the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than 5% or n is less than 3, the first detection result and the second detection result are sent to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0121] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0122] Please also see Figure 4 In some embodiments, after step S307, the method further includes: S308, training the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0123] Specifically, the first detection result and the second detection result are sent to the first detection model, and the first detection result and the second detection result are trained to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0124] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0125] See also Figure 5 The present application also provides an image target detection capability training method, which includes:
[0126] S401, training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability;
[0127] S402, using the first detection model to analyze the second image with the second identifier to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0128] S403, determining, by a second detection model, positive examples and negative examples of targets of each category among the multiple targets according to the first information;
[0129] S404: the second detection model acquires the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of the targets of each category of the multiple targets so that the second detection model acquires a second detection capability;
[0130] S405, using the second detection model to analyze the third image, the fourth image and the test image, to obtain a first detection result, a second detection result and a third detection result, respectively, wherein the first detection result includes the unlabeled first image and the first image with the third label, and the second detection result includes the unlabeled second image and the second image with the fourth label;
[0131] After the first detection model and the second detection model complete the specified rounds of training, the third image, the fourth image and the test image will be tested once. The number of tests is n, where n is a positive integer. When n=1, steps S406, S407 and S409 are executed; when n≥2, steps S408 and S409 are executed.
[0132] S406, using the first detection model to detect the test image to obtain a fourth detection result;
[0133] Optionally, the test image is sent to the first detection model to perform detection analysis on it to obtain a fourth detection result of the test image, wherein the fourth detection result includes the category, orientation information and confidence of all targets included in the test image.
[0134] S407, when the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold or the number of training times of the second detection model is less than a first preset threshold, sending the first detection result and the second detection result to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model; and
[0135] For detailed descriptions of steps S401 to S407, please refer to the descriptions of the corresponding parts of the above embodiments, which will not be repeated here. After step S407, step S409 is executed.
[0136] S408, using the second detection model to detect the test image to obtain the nth third detection result; when the difference between the accuracy of the nth third detection result and the accuracy of the n-1th third detection result is greater than a second preset threshold or the number of training times of the second detection model is less than a first preset threshold, the first detection result and the second detection result are sent to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model;
[0137] Optionally, after each round of training, the second detection model is used to detect the test image to obtain the nth third detection result. When the difference between the accuracy of the nth third detection result and the accuracy of the n-1th third detection result is greater than a second preset threshold, such as greater than 5% or 10%, or the number of training times of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for further training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model. After step S408, step S409 is executed.
[0138] S409: training the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0139] For a detailed description of step S409, please refer to the description of the corresponding part of the above embodiment, which will not be repeated here.
[0140] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0141] See also Figure 6 The present application also provides an image object detection capability training device 500, which includes:
[0142] A first detection unit 510 is used to train a first image having a first identifier so that a first detection model obtains a first detection capability and forms a model parameter of the first detection capability; and
[0143] A second detection unit 530 is used to analyze the second image with the second identifier using the first detection model to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets;
[0144] The second detection unit 530 is further configured to use the second detection model to determine positive examples and negative examples of targets of each category among the multiple targets according to the first information;
[0145] The second detection unit 530 is also used for the second detection model to obtain the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category of the multiple targets so that the second detection model obtains second detection capability.
[0146] Optionally, the first information also includes the position information of each target in the second image and the confidence of each target, and the position information of each target includes the coordinates and vector of the target.
[0147] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0148] The image target detection capability training device 500 of the embodiment of the present application can inherit the model parameters of the first detection model, so that the second detection model can obtain the first detection capability of the first detection model. At the same time, the second detection model can also determine the positive examples and negative examples of targets of each category among the multiple targets based on the first information obtained by the first detection model, and train the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains the second detection capability, thereby improving the detection capability of the second detection model, so that only a small number of first images with the first identification (i.e., a small number of true value samples) are needed to obtain a stronger target detection capability.
[0149] In some embodiments, the first detection unit 510 is further used to: use a data enhancement algorithm in a target detection algorithm to perform illumination distortion, geometric distortion and image occlusion on the first image.
[0150] See also Figure 7 In some embodiments, the second detection unit 530 is also used for analyzing the third image, the fourth image and the test image using the second detection model to obtain the first detection result, the second detection result and the third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, the third image is an unlabeled first image, the fourth image is an unlabeled second image, the fifth image is a first image with a third label, and the sixth image is a second image with a fourth label; the image target detection capability training device 500 also includes an analysis and judgment unit 550, and the analysis and judgment unit 550 is used to send the first detection result and the second detection result to the first detection model for cyclic training when the accuracy of the third detection result meets the preset conditions or the training times of the second detection model are less than the first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0151] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0152] In some embodiments, after the first detection model and the second detection model complete a specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n=1,
[0153] The first detection unit 510 is further configured to detect the test image using the first detection model to obtain a fourth detection result;
[0154] The analysis and judgment unit 550 is also used to send the first detection result and the second detection result to the first detection model for further training when the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold; or the training times of the second detection model is less than the first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
[0155] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0156] In some embodiments, after the first detection model and the second detection model complete a specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer, and when n≥2,
[0157] The first detection unit 510 is further used to train the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model;
[0158] Using the second detection model to detect the test image to obtain an nth third detection result;
[0159] The analysis and judgment unit 550 is further configured to determine when the difference between the accuracy of the third detection result for the nth time and the accuracy of the third detection result for the (n-1)th time is greater than a second preset threshold.
[0160] For the features of this embodiment that are the same as those of the above-mentioned embodiment, please refer to the description of the corresponding parts of the above-mentioned embodiment and no further details will be given here.
[0161] An embodiment of the present application also provides a computer-readable storage medium, which stores an executable program code, and the computer-executable program code is used to enable a computer to execute the image target detection capability training method of the embodiment of the present application.
[0162] See also Figure 8The embodiment of the present application also provides an electronic device 600, which includes a processor 610 and a memory 630. The memory 630 stores program codes that can be executed by the processor 610. When the program code is called and executed by the processor 610, the image target detection capability training method of the embodiment of the present application is executed.
[0163] The memory 630, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the image target detection capability training method in the embodiment of the present invention. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 630, that is, the image target detection capability training method in the above method embodiment is implemented.
[0164] It may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition. Any connection can be appropriately a computer-readable medium. For example, if the software is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technologies such as infrared, radio and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, wireless and microwave are included in the fixation of the medium. As used in the present invention, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and blue-ray disc, wherein disk usually reproduces data magnetically, while disc uses laser to reproduce data optically. The above combination should also be included in the scope of protection of computer-readable media.
[0165] The electronic device 600 of the present invention includes but is not limited to computers, laptops, tablet computers, mobile phones, cameras, smart bracelets, smart watches, smart glasses and other electronic devices with display screens.
[0166] Mentioning "embodiment" and "implementation method" in this application means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of the phrases in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments. In addition, it should also be understood that the features, structures or characteristics described in the various embodiments of the present application can be arbitrarily combined to form another embodiment that does not deviate from the spirit and scope of the technical solution of the present application, provided that there is no contradiction between them.
[0167] Finally, it should be noted that the above implementation modes are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the above preferred implementation modes, a person of ordinary skill in the art should understand that the technical solution of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present application.
Claims
1. A method for training image target detection capability, characterized in that: include: Training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability; Analyzing the second image with the second identifier using the first detection model to obtain first information, where the first information at least includes a plurality of targets and categories of the plurality of targets; The second detection model determines, according to the first information, positive examples and negative examples of targets of each category among the multiple targets; as well as The second detection model acquires the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model acquires a second detection capability.
2. The image target detection capability training method according to claim 1, characterized in that: After training the positive examples and the negative examples of the targets of each category in the multiple targets so that the second detection model acquires a second detection capability, the method further includes: The second detection model is used to analyze the third image, the fourth image and the test image to obtain a first detection result, a second detection result and a third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, the third image is an unlabeled first image, the fourth image is an unlabeled second image, the fifth image is the first image with the third label, and the sixth image is the second image with the fourth label; and When the accuracy of the third detection result meets the preset conditions or the training times of the second detection model is less than the first preset threshold, the first detection result and the second detection result are sent to the first detection model for cyclic training to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
3. The image target detection capability training method according to claim 2, characterized in that: After the first detection model and the second detection model complete the specified rounds of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n=1, the method further includes: Using the first detection model to detect the test image to obtain a fourth detection result; The third detection result meets the preset conditions, including: When the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold.
4. The image target detection capability training method according to claim 2, characterized in that: After the first detection model and the second detection model complete the specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n≥2, the method further includes: Training the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model; Using the second detection model to detect the test image to obtain an nth third detection result; The third detection result meets the preset conditions, including: When the difference between the accuracy of the nth third detection result and the accuracy of the (n-1)th third detection result is greater than a second preset threshold.
5. The image target detection capability training method according to any one of claims 1 to 4, characterized in that: Before training the first image with the first identifier, the method further includes: A data enhancement algorithm in the target detection algorithm is used to perform illumination distortion, geometric distortion and image occlusion on the first image.
6. The image target detection capability training method according to any one of claims 1 to 4, characterized in that: The first information also includes the position information of each target in the second image and the confidence of each target, and the position information of each target includes the coordinates and vector of the target.
7. An image target detection capability training device, characterized in that: include: A first detection unit, used for training a first image with a first identifier so that a first detection model acquires a first detection capability and forms a model parameter of the first detection capability; as well as a second detection unit, configured to analyze the second image with the second identifier by using the first detection model to obtain first information, wherein the first information at least includes a plurality of targets and categories of the plurality of targets; The second detection unit is further configured to determine, by a second detection model, positive examples and negative examples of targets of each category among the multiple targets according to the first information; The second detection unit is also used for the second detection model to obtain the model parameters of the first detection model so that the second detection model inherits the first detection capability of the first detection model, and trains the positive examples and the negative examples of targets of each category among the multiple targets so that the second detection model obtains second detection capability.
8. The image target detection capability training device according to claim 7, characterized in that: The second detection unit is further used to analyze the third image, the fourth image and the test image using the second detection model to obtain a first detection result, a second detection result and a third detection result respectively, wherein the first detection result includes the third image and the fifth image, the second detection result includes the fourth image and the sixth image, the third image is an unmarked first image, the fourth image is an unmarked second image, the fifth image is the first image with a third mark, and the sixth image is the second image with a fourth mark; The image target detection capability training device also includes an analysis and judgment unit, which is used to send the first detection result and the second detection result to the first detection model for cyclic training when the accuracy of the third detection result meets a preset condition or the training times of the second detection model is less than a first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
9. The image target detection capability training device according to claim 8, characterized in that: After the first detection model and the second detection model complete the specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer. When n=1, The first detection unit is further used to detect the test image using the first detection model to obtain a fourth detection result; The analysis and judgment unit is also used to send the first detection result and the second detection result to the first detection model for further training when the difference between the accuracy of the third detection result and the accuracy of the fourth detection result is greater than a second preset threshold; or the training times of the second detection model is less than the first preset threshold, so as to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model.
10. The image target detection capability training device according to claim 8, characterized in that: After the first detection model and the second detection model complete the specified round of training, the third image, the fourth image and the test image are tested once, and the number of tests is n, where n is a positive integer, and when n≥2, The first detection unit is further used to train the first detection result and the second detection result to improve the first detection capability of the first detection model, thereby improving the second detection capability of the second detection model; Using the second detection model to detect the test image to obtain an nth third detection result; The analysis and judgment unit is further configured to: when the difference between the accuracy of the nth third detection result and the accuracy of the (n-1)th third detection result is greater than a second preset threshold.
11. The image target detection capability training device according to any one of claims 7 to 10, characterized in that: The first detection unit is also used to: use a data enhancement algorithm in a target detection algorithm to perform illumination distortion, geometric distortion and image occlusion on the first image.
12. The image target detection capability training device according to any one of claims 7 to 10, characterized in that: The first information also includes the position information of each target in the second image and the confidence of each target, and the position information of each target includes the coordinates and vector of the target.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer-executable program code, and the computer-executable program code is used to enable a computer to execute the method according to any one of claims 1 to 6.
14. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory stores a program code executable by the processor, and when the program code is called and executed by the processor, the method according to any one of claims 1 to 6 is executed.
Citation Information
Patent Citations
Information classification method and device and information classification model training method and device
CN113178189A
Detection model training method and device, target detection method and device and electronic system
CN113239982A