Target detection model training method and device, target detection method and device, equipment and medium
Through the generation network in the generative adversarial network and iterative optimization of the discriminant network, the problem of unbalanced in the inference speed and accuracy of the object detection model in embedded devices is solved, and efficient object detection effect is achieved.
Patent Information
- Application Number
- CN202510198965.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing object detection model is difficult to balance the inference speed and accuracy in embedded devices. The second-stage network calculation is large, the One-stage network detection accuracy is not high enough, and the Transformer-based method is huge and difficult to deploy.
The generation network and discriminative network in the generative adversarial network are adopted to improve detection accuracy through iterative optimization of generation loss values and discriminative loss values, and the inference speed is controlled through lightweight algorithms and flexible embedding.
It achieves balancing inference speed and detection accuracy in embedded devices, improving the accuracy and efficiency of object detection.
Smart Images

Figure CN120219793A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method, apparatus, device, and medium for training a target detection model and performing target detection. Background Art
[0002] In recent years, with the continuous development of convolutional neural networks (CNNs) and Transformer models, object detection, as one of the core tasks in the field of computer vision, has become one of the most widely used algorithms. Currently, mainstream object detection algorithms are generally divided into two-stage networks, one-stage networks, and Transformer-based object detection algorithms.
[0003] Among them, relatively classic methods in two-stage networks include region-based convolutional neural networks (RCNNs) and fast region-based convolutional neural networks (Fast RCNNs) series families. This series of methods mainly use a backbone network to generate numerous region of interest (ROI) candidate boxes, and then classify and score these ROI candidate boxes to obtain the correct target boxes. Although this method has high object detection accuracy, its complex candidate box selection mechanism results in a large time consumption of the model in embedded devices and is difficult to be applied in practice.
[0004] Relatively classic methods in one-stage networks include the You Only Look Once (YOLO) series, single shot multibox detector (SSD), and centernet, etc. The target boxes and classification results of the model output can be directly obtained only by passing through the detection network once. Compared with two-stage algorithms, the number of model parameters and the amount of computation are greatly reduced. However, the detection accuracy is not high enough.
[0005] In addition, in the field of natural language processing (NLP), the Transformer-based object detection algorithms commonly used at present, such as the end-to-end object detection algorithm based on Transformer (DETR), input the input features and the features extracted by CNN into the Transform Encoder-Decoder, and utilize the global information extraction ability of Transformer to significantly improve the object detection effect of the model. Although the Transformer-based object detection method is convenient for extracting rich global information, its unique multi-head attention will generate a huge amount of computation, so it is difficult to deploy it to embedded devices.
[0006] In summary, how to obtain an object detection model that can balance the inference speed and accuracy is an urgent problem to be solved at present. Summary of the Invention
[0007] The embodiments of the present application provide an object detection model training method, an object detection method, a device, a device and a medium, which are used to solve the problem of how to obtain an object detection model that can balance the inference speed and accuracy in the prior art.
[0008] In the first aspect, the embodiments of the present application provide an object detection model training method, and the method includes:
[0009] Input the training images in the training set into the generation network in the generative adversarial network, and obtain the predicted position information and predicted category of the objects in the training images output by the generation network; determine the first generation loss value according to the predicted position information and predicted category of the objects, the true position information and true category of the objects in the training images;
[0010] Input the training images, the predicted position information and predicted category of the objects, and the true category of the objects in the training images into the discriminant network trained in the generative adversarial network, and determine the discriminant result of whether the predicted category of the objects is the true category based on the trained discriminant network;
[0011] Determine the first discriminant loss value according to the discriminant result of whether the predicted category of the objects is the true category; train the generation network based on the first generation loss value and the first discriminant loss value to obtain the trained generation network, and determine the trained generation network as the trained object detection model.
[0012] In the second aspect, the embodiments of the present application provide an object detection method based on the model trained by the above object detection model training method, and the method includes:
[0013] Input the image to be detected into the trained object detection model. Based on the trained object detection model, determine the detection result of whether there is an object in the image to be detected. If there is, determine the target position information and target category of the object in the image to be detected.
[0014] In a third aspect, an embodiment of the present application further provides an object detection model training device, and the device includes:
[0015] A detection module, configured to input the training images in the training set into the generation network in the generative adversarial network, and obtain the predicted position information and predicted category of the object in the training images output by the generation network; determine the first generation loss value according to the predicted position information and predicted category of the object, the true position information and true category of the object in the training images;
[0016] A discrimination module, configured to input the training images, the predicted position information and predicted category of the object, and the true category of the object in the training images into the trained discrimination network in the generative adversarial network, and based on the trained discrimination network, determine the discrimination result of whether the predicted category of the object is the true category;
[0017] A training module, configured to determine the first discrimination loss value according to the discrimination result of whether the predicted category of the object is the true category; based on the first generation loss value and the first discrimination loss value, train the generation network to obtain the trained generation network, and determine the trained generation network as the trained object detection model.
[0018] In a fourth aspect, an embodiment of the present application further provides an object detection device based on a model trained by the object detection model training method, and the device includes:
[0019] An input module, configured to input the image to be detected into the trained object detection model;
[0020] A processing module, configured to determine the detection result of whether there is an object in the image to be detected based on the trained object detection model. If there is, determine the target position information and target category of the object in the image to be detected.
[0021] In a fifth aspect, an embodiment of the present application further provides an electronic device, and the electronic device at least includes a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the object detection model training method or the object detection method described in any one of the above are implemented.
[0022] Sixth aspect, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the target detection model training method or the target detection method described in any one of the above.
[0023] In an embodiment of the present application, training images in a training set are input into a generation network in a generative adversarial network to obtain predicted position information and predicted categories of targets in the training images output by the generation network; according to the predicted position information and predicted categories of the targets, the true position information and true categories of the targets in the training images, a first generation loss value is determined; the training images, the predicted position information and predicted categories of the targets, and the true categories of the targets in the training images are input into a discriminative network that has been trained in the generative adversarial network, and based on the discriminative network that has been trained, a discriminative result of whether the predicted category of the target is the true category is determined; according to the discriminative result of whether the predicted category of the target is the true category, a first discriminative loss value is determined; based on the first generation loss value and the first discriminative loss value, the generation network is trained to obtain a trained generation network, and the trained generation network is determined as a trained target detection model. Since the characteristics of the discriminative network in the generative adversarial network are used to help identify whether the output result of the generation network, that is, the target detection model, is correct, and in the continuous optimization and iteration of the two, the output of the generation network gradually approaches the true result, thereby improving the accuracy of target detection; moreover, the generation network and the discriminative network in the embodiment of the present application are arbitrarily replaceable, providing flexible embeddability for the generation network, that is, the target detection model, so that the balance between the inference speed and the detection accuracy can be well controlled by this method. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0025] Figure 1 It is a schematic diagram of a target detection model training process provided by an embodiment of the present application;
[0026] Figure 2 It is a schematic diagram of a generative adversarial network structure provided by an embodiment of the present application;
[0027] Figure 3 It is a schematic diagram of a target detection process provided by an embodiment of the present application;
[0028] Figure 4 It is a schematic diagram of a target detection model training and target detection process provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of the structure of an object detection model training device provided by an embodiment of the present application;
[0030] Figure 6 Schematic diagram of the structure of an object detection device provided by an embodiment of the present application;
[0031] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0032] To make the objectives and implementation manners of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.
[0033] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.
[0034] The terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0035] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components that are not clearly listed or are inherent to these products or devices.
[0036] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the element.
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0038] For ease of explanation, the above description has been presented in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
[0039] An embodiment of the present application provides a method, apparatus, device, and medium for training a target detection model and for target detection. In this method, training images in a training set are input into a generation network in a generative adversarial network to obtain predicted position information and predicted categories of targets in the training images output by the generation network; according to the predicted position information and predicted categories of the targets, the true position information and true categories of the targets in the training images, a first generation loss value is determined; the training images, the predicted position information and predicted categories of the targets, and the true categories of the targets in the training images are input into a discriminative network that has been trained in the generative adversarial network, and based on the discriminative network that has been trained, a discriminative result is determined as to whether the predicted category of the target is the true category; according to the discriminative result as to whether the predicted category of the target is the true category, a first discriminative loss value is determined; based on the first generation loss value and the first discriminative loss value, the generation network is trained to obtain a trained generation network, and the trained generation network is determined as the trained target detection model. Since the characteristics of the discriminative network in the generative adversarial network are utilized to help identify whether the output result of the generation network, i.e., the target detection model, is correct, and in the continuous optimization and iteration of the two, the output of the generation network gradually approaches the true result, thereby improving the accuracy of target detection; moreover, the generation network and the discriminative network in the embodiment of the present application are arbitrarily replaceable, providing flexible embeddability for the generation network, i.e., the target detection model, so that the balance between the inference speed and the detection accuracy can be well controlled through this method.
[0040] Embodiment 1:
[0041] Figure 1 FIG. is a schematic diagram of a process for training a target detection model provided by an embodiment of the present application, and this process includes:
[0042] S101: Input training images in a training set into a generation network in a generative adversarial network to obtain predicted position information and predicted categories of targets in the training images output by the generation network; according to the predicted position information and predicted categories of the targets, the true position information and true categories of the targets in the training images, a first generation loss value is determined.
[0043] The target detection model training method provided by the embodiments of this application is applied to an electronic device, which can be a personal computer (PC), a server, etc.
[0044] The electronic device stores multiple original images, and there are targets in each original image. For the original images in the training set, operations such as Mosaic, Mixup, random cropping, horizontal and vertical flipping, and random scaling can be adopted to obtain multiple training images after data augmentation, and the training images are all saved to the training set, so that the model can be trained based on the training images in the training set.
[0045] For each training image in the training set, the training image is input into the generator network in the generative adversarial network. The generative adversarial network includes a generator network and a discriminator network. In the embodiments of this application, the generator network and the discriminator network are arbitrarily replaceable, providing flexible embeddability for the generator network, that is, the target detection model. Then, in order to better help the generator network balance the accuracy and speed, excellent and lightweight target detection algorithms can be arbitrarily designed and selected in the generator network. At the same time, in order to ensure the detection accuracy of the generator network, the most cutting-edge (State-Of-The-Art, SOTA) discriminator network can be selected to help the generator network improve the detection accuracy, and the model parameters of the discriminator network do not need to be considered.
[0046] Among them, the generator network can be any target detection model. Therefore, in the embodiments of this application, the generator network can perform target detection on the input image, determine the category information of the target in the image and the location information of the target, and mark the location information of the target with a detection box.
[0047] Specifically, considering the real-time performance and inference speed of target detection, in the embodiments of this application, the target detection model (You Only Look Once-Tiny, YOLOX-tiny) is used as the generator network. Then, as Figure 2 shown in the generator network part of the schematic diagram of the generative adversarial network structure, the generator network includes three modules: the first backbone module, the neck module, and the head.
[0048] Therefore, after the training image is input into the generation network, the first Backbone module in the generation network adopts an excellent lightweight feature extraction module, namely the Cross-Stage Partial Darknet (CSP-Darknet). This CSP-Darknet is used to help the generation network extract richer image features, obtain the training feature map corresponding to the training image, and its unique cross-stage connection reduces computational redundancy and improves the inference speed of the model. The Neck module uses the Feature Pyramid Networks (FPN) to construct bottom-up and top-down feature fusion paths, fuses the training feature maps of different scales, generates a fused training feature map with rich multi-scale information, and adopting this multi-scale fusion method can greatly improve the detection performance of small targets. The Head module uses a Decoupled Head to separate the regression and classification of the fused training feature map, that is, separates the prediction task of the position information of the target in the fused training feature map from the prediction task of the category, and obtains the predicted position information of the target included in the training image (i.e., the bbox shown in Figure 2 ), and the predicted category (i.e., the cls shown in Figure 2 ). This method not only reduces the mutual interference between tasks, but also enables the generation network to focus more on their respective tasks.
[0049] Among them, the training image may include multiple targets. For each target, the predicted position information and predicted category of the target will be determined. The number of targets in the training image is not specifically limited here.
[0050] Since for each training image in the electronic device, the true position information and true category of the target in the training image are also saved. Therefore, in order to determine the performance of the generation network, in the embodiments of the present application, after obtaining the predicted position information and predicted category of the target included in the training image output by the generation network, according to the predicted position information and true position information of the target, a position loss function is used to determine the position loss value (i.e., the first generation sub-loss value), and a cross-entropy loss function is used to determine the object loss value (i.e., the third generation sub-loss value); according to the predicted category and true category of the target, a cross-entropy loss function is used to determine the category loss value (i.e., the second generation sub-loss value); according to the position loss value, object loss value and category loss value, the first generation loss value corresponding to the generation network is determined.
[0051] S102: Input the training image, the predicted position information and predicted category of the target, and the true category of the target in the training image into the discriminant network trained in the generative adversarial network, and based on the trained discriminant network, determine the discriminant result of whether the predicted category of the target is the true category.
[0052] In an embodiment of the present application, after the generation network in the generative adversarial network obtains the output result, it will input the output result into the discriminative network, so that the discriminative network helps to identify whether the output result of the generation network is correct.
[0053] Specifically, after obtaining the predicted position information and predicted category of the target in the training image output by the generation network in the generative adversarial network, the training image, the predicted position information and predicted category of the target, and the true category of the target are input into the discriminative network that has been trained in the generative adversarial network. Based on the discriminative network that has been trained, the sub-image corresponding to the predicted position information of the target in the training image is determined, and based on this sub-image, it is judged whether the predicted category of the target is the true category, and the discriminative result of whether the predicted category of the target is the true category is output.
[0054] In addition, in a possible implementation, as Figure 2 shown, the true category and true position information of the target can also be taken as a whole and regarded as a label and jointly input into the discriminative network. That is to say, the training image, the predicted position information and predicted category of the target, and the true position information and true category of the target are input into the discriminative network that has been trained. Based on the discriminative network that has been trained, it is first determined whether the predicted category is the true category. If not, it is determined that the output result of the generation network is false; if so, it is then determined whether the predicted position information of the target is within the preset error range from the true position information. If so, it is determined that the output result of the generation network is true; if not, it is determined that the output result of the generation network is false. Among them, the dimension of the result output by the discriminative network is C×1 (as Figure 2 shown), where C is the number of results output by the discriminative network and is the same as the number of predicted categories of the target output by the generation network.
[0055] S103: Determine the first discriminative loss value according to the discriminative result of whether the predicted category of the target is the true category; based on the first generative loss value and the first discriminative loss value, train the generation network to obtain the trained generation network, and determine the trained generation network as the trained target detection model.
[0056] In an embodiment of the present application, after obtaining the discriminative result output by the trained discriminative network, according to the discriminative result of whether the predicted category of the target is the true category, the first discriminative loss value is calculated using the cross-entropy loss function.
[0057] After obtaining the first generation loss value corresponding to the generation network and the first discriminant loss value corresponding to the trained discriminant network, since the discriminant network has been trained, the parameters of the discriminant network are fixed. Based on the first generation loss value and the first discriminant loss value, the generation network is iteratively trained to make the first generation loss value corresponding to the generation network smaller and smaller, thereby increasing the difficulty for the discriminant network to distinguish the difference between the output result of the generation network and the real result, and making the output result of the generation network closer and closer to the real result, that is, improving the detection accuracy of the generation network. After meeting the training requirements, the trained generation network is obtained. Among them, the training requirements include but are not limited to the number of training times reaching a preset number, the first generation loss value corresponding to the generation network being less than a preset threshold, etc.
[0058] Since the generation network itself has the characteristic of relatively fast inference speed, and the detection accuracy of the generation network is improved through the above training method, in the embodiment of the present application, the trained generation network is determined as the trained target detection model.
[0059] In the embodiment of the present application, since the discriminant network characteristics in the generative adversarial network are used to help identify whether the output result of the generation network, that is, the target detection model, is correct, and in the continuous optimization and iteration of the two, the output of the generation network gradually approaches the real result, thereby improving the accuracy of target detection; moreover, the generation network and the discriminant network in the embodiment of the present application have arbitrary replaceability, providing flexible embeddability for the generation network, that is, the target detection model, so that the balance between the inference speed and the detection accuracy can be well controlled through this method.
[0060] Embodiment 2:
[0061] In order to further control the balance between the inference speed and the detection accuracy, on the basis of the above embodiment, in the embodiment of the present application, before inputting the training image, the predicted position information and predicted category of the target, and the real category of the target in the training image into the discriminant network trained in the generative adversarial network, the method further includes:
[0062] Input the sample image into the generation network, and based on the generation network, obtain the sample predicted position information and sample predicted category of the sample target in the sample image output by the generation network; determine the second generation loss value according to the sample predicted position information and sample predicted category of the sample target, the sample real position information and sample real category of the sample target in the sample image;
[0063] Input the sample image, the sample predicted position information and sample predicted category of the sample target, and the sample real category of the sample target in the sample image into the discriminant network to be trained, and based on the discriminant network to be trained, determine the sample discrimination result of whether the sample predicted category of the sample target is the sample real category;
[0064] Determine a second discrimination loss value according to the sample discrimination result of whether the sample prediction category of the sample target is the sample true category.
[0065] Based on the second generation loss value and the second discrimination loss value, train the discriminative network to be trained until the trained discriminative network is obtained.
[0066] In the embodiment of the present application, before training the generation network in the generative adversarial network, the parameters of the generation network can be fixed first, and at the same time, the discriminative network in the generative adversarial network is trained. By continuously iterating the parameters of the discriminative network, the loss value of the discriminative network becomes smaller and smaller, thereby improving the accuracy of the discriminative network in distinguishing real results and false results.
[0067] Specifically, in the embodiment of the present application, multiple sample images are also stored in the electronic device, and each sample image contains a sample target. When training the discriminative network in the generative adversarial network, the parameters of the generation network are fixed first. For each sample image, input the sample image into the generation network, and based on the generation network, determine the sample prediction position information and sample prediction category of the sample target in the sample image.
[0068] The electronic device stores the sample true position information and sample true category of the sample target in the sample image. After obtaining the sample prediction position information and sample prediction category of the sample target predicted by the generation network, according to the sample prediction position information and sample true position information of the sample target, use a position loss function to determine the sample position loss value, and use a cross-entropy loss function to determine the sample object loss value; according to the sample prediction category and sample true category of the sample target, use a cross-entropy loss function to determine the sample category loss value; according to the sample position loss value, sample object loss value and sample category loss value, determine the second generation loss value corresponding to the generation network.
[0069] After obtaining the sample prediction position information and sample prediction category of the sample target in the sample image output by the generation network, input the sample image, the sample prediction position information and sample prediction category of the sample target output by the generation network, and the sample true category of the sample target into the discriminative network to be trained. Based on the discriminative network to be trained, determine the sample sub-image corresponding to the sample prediction position information of the sample target in the sample image, and based on the sample sub-image, determine whether the sample prediction category of the sample target is the sample true category, and output the sample discrimination result of whether the sample prediction category of the sample target is the sample true category.
[0070] After obtaining the sample discrimination result output by the discriminative network to be trained, according to whether the sample prediction category of the sample target is the sample discrimination result of the sample true category, the cross-entropy loss function is used to calculate the second discrimination loss value corresponding to the discriminative network to be trained.
[0071] After obtaining the second generation loss value corresponding to the generation network and the second discrimination loss value corresponding to the discriminative network to be trained, fix the parameters of the generation network, and based on the second generation loss value and the second discrimination loss value, train the discriminative network to be trained. And, since multiple sample images are stored in the electronic device, the discriminative network to be trained can be iteratively trained based on the multiple sample images, so that the second discrimination loss value corresponding to the discriminative network becomes smaller and smaller, thereby improving the ability of the discriminative network to distinguish between true results and false results. After meeting the training requirements, the trained discriminative network is obtained. Among them, the training requirements include but are not limited to the number of training times reaching a preset number, and the second discrimination loss value corresponding to the discriminative network being less than a preset threshold, etc.
[0072] The following uses a specific embodiment to illustrate the training process of the above object detection model:
[0073] First, it can be understood that the loss function of the generative adversarial network satisfies the following formula:
[0074] L total =min G max D V(D,G)=min G max D V(1-L D ,L G )
[0075] Among them, L total is the total loss value of the generative adversarial network; G represents the generation network; D represents the discriminative network; V represents the defined model value function.
[0076] Specifically, after expanding the formula, it is:
[0077]
[0078] Among them, p data (x) represents the true position information and true category of the target; p z (z) represents the predicted position information and predicted category of the target output by the generation network; E represents expectation; D(x) = 1 - L D ; G(z) = L G ; in the formula It is an object loss function constructed based on real location information and real categories. That is, it is hoped that the discriminative network can correctly distinguish real results from false results. Therefore, when training the discriminative network, if the discriminative ability of the discriminative network becomes stronger and the value of D(x) gets closer and closer to 1, then the expected value of will become smaller and smaller, while the expected value of will become larger and larger. Therefore, when training the discriminative network, the final loss of the discriminative network is
[0079] In summary, the generative network hopes to reduce the value of V so that the discriminative network cannot determine the output result of the generative network as a false result, while the discriminative network hopes to increase the value of V to improve its accuracy in distinguishing true and false results. Therefore, in the embodiments of the present application, first, the parameters of the generative network are fixed, and the discriminative network is trained. And during the training, according to the obtained generative loss value corresponding to the generative network and the discriminative loss value corresponding to the discriminative network, the loss value calculated based on the above loss function formula of the generative adversarial network reaches the maximum value; then, the parameters of the discriminative network are fixed, and the generative network is trained. And during the training, according to the obtained generative loss value corresponding to the generative network and the discriminative loss value corresponding to the discriminative network, the loss value calculated based on the above loss function formula of the generative adversarial network reaches the minimum value.
[0080] In addition, in a possible implementation, the discriminative network and the generative network can be trained through multiple rounds of iterative training, that is, the steps of first fixing the parameters of the generative network and training the discriminative network, and then fixing the parameters of the discriminative network and training the generative network are repeated multiple times until the trained generative network and discriminative network are obtained.
[0081] In the embodiments of the present application, the discriminative network is first trained before training the generative network, and the joint iterative training of the generative network and the discriminative network can help the generative network, that is, the target detection model, further learn better parameters and improve the detection accuracy, so as to further control the balance between the inference speed and the detection accuracy while ensuring the detection speed.
[0082] Embodiment 3:
[0083] To further improve the detection accuracy, based on the above embodiments, in the embodiments of the present application, determining whether the predicted category of the target is the discriminative result of the real category based on the trained discriminative network includes:
[0084] Based on the discriminant network, determine the first probability value that the category of the target in the sub-image corresponding to the predicted position information of the target in the training image is the predicted category, the second probability value that it is a non-predicted category, the third probability value that it is the true category, and the fourth probability value that it is a non-true category;
[0085] According to the discriminant result of whether the predicted category of the target is the true category, determine the first discriminant loss value, including:
[0086] Determine the first discriminant sub-loss value according to the first probability value that the category of the target in the sub-image is the predicted category and the second probability value that it is a non-predicted category;
[0087] Determine the second discriminant sub-loss value according to the third probability value that the category of the target in the sub-image is the true category and the fourth probability value that it is a non-true category;
[0088] Determine the first discriminant loss value according to the first discriminant sub-loss value and the second discriminant sub-loss value.
[0089] In order to enable the discriminant network to capture more complex and rich semantic features, thereby distinguishing the differences between the true result and the predicted result, and prompting the generation network to iteratively optimize step by step based on this difference, and finally generating a predicted category and predicted position information that the discriminant network cannot distinguish from the true category and true position information, so in the embodiments of the present application, a classification model based on the Residual Network (ResNet101) that can learn complex and advanced semantic features as the backbone can be used as the discriminant network.
[0090] Specifically, as Figure 2 shown, the discriminant network mainly consists of four modules: the second Backbone module, the Average Pooling module, the Fully Connected (FC) module, and the Softmax module.
[0091] Then, after inputting the training image, the predicted position information and predicted category of the target, and the true category of the target in the training image into the discriminant network, the second Backbone module in the discriminant network determines the sub-image corresponding to the predicted position information in the training image according to the predicted position information of the target, and extracts the semantic features in the sub-image to obtain the backbone feature (BF) of the sub-image, that is, the semantic feature.
[0092] Input the backbone features of the sub-image into the Average Pooling module for downsampling to obtain the first feature of the sub-image; input the first feature of the sub-image into the FC module for classification to obtain the category corresponding to the target in the sub-image; input the category corresponding to the target in the sub-image into the Softmax module, and based on this Softmax module, determine the first probability value that the category corresponding to the target in the sub-image is the predicted category, the second probability value that it is not the predicted category, the third probability value that the category corresponding to the target in the sub-image is the true category, and the fourth probability value that it is not the true category.
[0093] Moreover, according to the first probability value that the category corresponding to the target in the sub-image is the predicted category and the second probability value that it is not the predicted category, use the cross-entropy loss function to calculate and obtain the first discriminant sub-loss value.
[0094] According to the third probability value that the category corresponding to the target in the sub-image is the true category and the fourth probability value that it is not the true category, use the cross-entropy loss to calculate and obtain the second discriminant sub-loss value.
[0095] After obtaining the first discriminant sub-loss value and the second discriminant sub-loss value, determine the sum value of the first discriminant sub-loss value and the second discriminant sub-loss value, and determine this sum value as the discriminant category loss value, and determine this discriminant category loss value as the first discriminant loss value.
[0096] Among them, the first discriminant loss value satisfies the following formula:
[0097] L D =L cls
[0098] Among them, L D represents the first discriminant loss value, and L cls represents the discriminant category loss value.
[0099] In the embodiments of the present application, the determination method of the discriminant network is designed, and the category and position information of the target in the training image are combined as the basis for determining authenticity. This not only provides rich information for the discriminant network to more accurately determine the authenticity of the target, but also helps the generation network to more accurately predict the target category and position information, thereby further improving the detection accuracy.
[0100] Embodiment 4:
[0101] On the basis of the above embodiments, in the embodiments of the present application, according to the predicted position information and predicted category of the target, the true position information and true category of the target in the training image, determine the first generation loss value, including:
[0102] Determine the first generation sub-loss value according to the predicted category and the true category of the target;
[0103] Determine the second generation sub-loss value according to the predicted position information and the true position information of the target;
[0104] Determine that the sub-image corresponding to the predicted position information of the target has the predicted object label of the target; according to the predicted position information of the target, determine whether the sub-image corresponding to the predicted position information has the true object label of the target; according to the predicted object label and the true object label, determine the third generation sub-loss value;
[0105] Determine the first generation loss value according to the first generation sub-loss value, the second generation sub-loss value, and the third generation sub-loss value.
[0106] In the embodiment of the present application, before determining the first generation loss value corresponding to the generation network, a label assignment strategy can be used to determine positive and negative samples first.
[0107] Specifically, first perform a preliminary screening: obtain multiple anchor boxes (i.e., multiple predicted position information) generated by the generation network for the training image and the predicted category corresponding to each anchor box, and determine the center point of each anchor box. For each anchor box, judge whether the center point of the anchor box is within the true rectangle of any target (i.e., the true position information). If so, determine the anchor box as the first candidate positive sample. If not, judge whether the anchor box falls within any square box with the center point of the true rectangle as the benchmark and n as the side length. If so, determine the anchor box as the first candidate positive sample. If not, determine the anchor box as a negative sample. Here, n can be the longer side of the true rectangle, or the shorter side, or a multiple greater than 1 of the longer side, or a multiple greater than 1 of the shorter side, and no specific limitation is made here.
[0108] Then perform a refined screening (Sim Optimal Transport Assignment, SimOTA): for each true rectangle, calculate the intersection over union (IOU) between the anchor box corresponding to the first candidate positive sample and each true rectangle, add the top m IOU values with the largest IOU values to determine the corresponding sum value, perform operations such as rounding or truncating on the sum value to obtain an integer k, and determine that there are k anchor boxes with the largest degree of overlap with the true rectangle. Here, m is a positive integer. For example, m is 10.
[0109] For each ground truth rectangle, according to the ground truth category corresponding to the ground truth rectangle and the predicted category corresponding to each first candidate positive sample, the cross-entropy loss function is used to determine the category prediction loss value corresponding to the predicted category of each candidate positive sample and the ground truth category; for each first candidate positive sample, it is determined whether the center point of the ground truth rectangle is within the range with the center point corresponding to the first candidate positive sample as the benchmark and a as the radius. If so, the first candidate positive sample is determined as the second candidate positive sample. If not, the first candidate positive sample is determined as the negative sample of the target corresponding to the ground truth rectangle.
[0110] For each second candidate positive sample, according to the anchor-box (i.e., predicted position information) corresponding to the second candidate positive sample and the ground truth rectangle (i.e., ground truth position information), the position loss function is used to determine the position prediction loss value corresponding to the second candidate positive sample and the ground truth rectangle; according to the category prediction loss value and the position prediction loss value corresponding to the second candidate positive sample and the ground truth rectangle, the cost function value corresponding to the second candidate positive sample and the ground truth rectangle is determined.
[0111] After obtaining the cost function values corresponding to each second candidate positive sample and the ground truth rectangle, the second candidate positive samples are sorted in the order of decreasing cost function values, and the first k second candidate positive samples are determined as the positive samples of the target corresponding to the ground truth rectangle, and the other second candidate positive samples are determined as the negative samples of the target corresponding to the ground truth rectangle.
[0112] After obtaining the positive and negative samples of each target, according to the predicted category corresponding to the positive sample of the target output by the generation network and the ground truth category of the target, the cross-entropy loss function is used to determine the first generation sub-loss value. Among them, the process of using the cross-entropy loss function to determine the loss value according to the predicted category and the ground truth category corresponding to the positive sample of the target is the prior art and will not be elaborated here.
[0113] According to the predicted position information corresponding to the positive and negative samples of the target output by the generation network and the ground truth position information of the target, the position loss function is used to determine the second generation sub-loss value. Among them, the process of using the position loss function to determine the loss value according to the predicted position information corresponding to the positive and negative samples of the target and the ground truth position information is the prior art and will not be elaborated here.
[0114] Since the generation network determines the area where the target is located as the predicted position information when detecting the target, it means that the generation network predicts that there is a target in the sub-image corresponding to the predicted position information. Therefore, in the embodiments of the present application, the predicted object label determined based on the generation network is determined as that there is a target in the sub-image corresponding to the predicted position information of the target. However, in actual situations, there may be no target or other objects in the sub-image corresponding to the predicted position information output by the generation network. Therefore, it is also necessary to obtain the true object label indicating whether there is a target in the sub-image corresponding to the predicted position information of the target output by the generation network. After obtaining the predicted object label and the true object label, according to the predicted object label and the true object label, the cross-entropy loss function is used to determine the third generated sub-loss value. Among them, the process of determining the loss value by using the position loss function according to the predicted position information and the true position information corresponding to each positive and negative sample of the target is the prior art and will not be elaborated here.
[0115] After obtaining the first generated sub-loss value, the second generated sub-loss value, and the third generated sub-loss value, according to the first generated sub-loss value, the second generated sub-loss value, and the third generated sub-loss value, use L G = L cls + σ × L reg + ρ × L obj to determine the first generated loss value; where L G is the first generated loss value corresponding to the generation network, L cls is the first generated sub-loss value, L reg is the second generated sub-loss value, L obj is the third generated sub-loss value, σ and ρ are parameters, and the value ranges of σ and ρ are 0 - 1.
[0116] In the embodiments of the present application, the label assignment method SimOTA adopted can help the generation network finely screen out valuable positive samples. At the same time, its method of dynamically assigning and taking the top k positive samples greatly reduces the training time.
[0117] Embodiment 5:
[0118] Based on the above embodiments, in the embodiments of the present application, a target detection method using a model trained by the above target detection model training method is further provided. As Figure 3 shown in the schematic diagram of the target detection process, this process includes:
[0119] S301: Input the image to be detected into the trained target detection model.
[0120] The target detection method provided in the embodiments of the present application is applied to the above electronic device. After the electronic device obtains the image to be detected, it inputs the image to be detected into the trained target detection model.
[0121] Among them, the trained object detection model is a model obtained by training based on the object detection model training method provided in the above embodiments.
[0122] S302: Based on the trained object detection model, determine whether there is a detection result of an object in the image to be detected. If so, determine the target position information and target category of the object in the image to be detected.
[0123] In the embodiments of the present application, since the object detection model includes a first Backbone module, a Neck module, and a Head module, after the image to be detected is input into the trained object detection model, the first Backbone uses CSP-Darknet to extract features from the image to be detected, obtains the feature map corresponding to the image to be detected, and inputs the feature maps of each scale into the Neck module; the Neck module uses FPN to fuse the feature maps of different scales, generates a fused feature map with rich multi-scale information, and inputs the fused feature map into the Head module; the Head module uses Decoupled Head to respectively perform the task of predicting position information and the task of predicting category on the fused feature map, determine whether there is an object in the image to be detected, and if there is an object, output the target position information and target category of the object in the image to be detected.
[0124] In the embodiments of the present application, since the object detection model obtained by training based on the object detection model training method provided in the above embodiments can take into account both the inference speed and the detection accuracy, using this object detection model to process the image to be detected can well control the balance between the inference speed and the detection accuracy.
[0125] The following uses a specific embodiment to illustrate the above embodiments. Refer to Figure 4 The object detection model training and object detection process shown in the figure includes the following steps:
[0126] First, the training process of the object detection model will be described:
[0127] Step 1, obtain an input image, use the input image as a training image, and preprocess the training image. Among them, the preprocessing includes but is not limited to operations such as Mosaic, Mixup, random cropping, horizontal and vertical flipping, and random scaling.
[0128] Step 2, input the preprocessed training image into the generation network (i.e., the object detection model), and based on the generation network, determine the predicted position information and predicted category of the object in the training image; determine the generation loss value according to the predicted position information and predicted category of the object, the true position information and true category of the object.
[0129] Step 3: Input the training images, the predicted position information and predicted categories of the targets in the training images output by the generation network, the true position information and true categories of the targets into the discriminant network (i.e., the classification model). Based on the discriminant network, determine whether the output result of the generation network is the true result; according to whether the output result of the generation network is the true result, determine the discriminant loss value.
[0130] Step 4: First, fix the parameters of the generation network. Based on the generation loss value corresponding to the generation network and the discriminant loss value corresponding to the discriminant network, train the discriminant network; then fix the parameters of the discriminant network. Based on the generation loss value corresponding to the generation network and the discriminant loss value corresponding to the discriminant network, train the generation network. Perform multiple rounds of iterative training on the discriminant network and the generation network as described above to obtain the trained generation network and discriminant network.
[0131] Secondly, determine the trained generation network as the trained object detection model, and only this object detection model needs to be used during the inference process, without using the discriminant network. Specifically, the inference process of the object detection model will be described next:
[0132] Step 1: Obtain the input image, and use this input image as the image to be detected. Input the image to be detected into the trained object detection model; and before inputting the image to be detected into the object detection model, the image to be detected can also be preprocessed. Among them, this input image is different from the input image used during the training process of the object detection model; the preprocessing performed on the image to be detected is different from the preprocessing performed on the training images. By way of example, the preprocessing of the image to be detected includes but is not limited to denoising, supplementing missing information, etc., and no specific limitation is made here.
[0133] Step 2: The first Backbone in the trained object detection model uses CSP-Darknet to extract features from the image to be detected, obtain the feature map corresponding to the image to be detected, and input the feature maps of each scale into the Neck module; the Neck module uses FPN to fuse the feature maps of different scales to generate a fused feature map with rich multi-scale information, and input this fused feature map into the Head module; the Head module uses Decoupled Head to perform the tasks of predicting position information and predicting categories on the fused feature map respectively to determine the output result. Among them, the output result is whether there is a target in the image to be detected. If there is a target, the output result also includes the target position information and target category of the target in the image to be detected.
[0134] Example 6:
[0135] Based on the same inventive concept and on the basis of the above embodiments, the present application provides a target detection model training device. Figure 5 It is a schematic structural diagram of a target detection model training device provided by an embodiment of the present application. As Figure 5 shown, the device includes:
[0136] A detection module 501, configured to input training images in a training set into a generation network in a generative adversarial network, and obtain predicted position information and predicted categories of targets in the training images output by the generation network; determine a first generation loss value according to the predicted position information and predicted categories of the targets, the true position information and true categories of the targets in the training images;
[0137] A discrimination module 502, configured to input the training images, the predicted position information and predicted categories of the targets, and the true categories of the targets in the training images into a trained discrimination network in the generative adversarial network, and based on the trained discrimination network, determine a discrimination result as to whether the predicted categories of the targets are true categories;
[0138] A training module 503, configured to determine a first discrimination loss value according to the discrimination result as to whether the predicted categories of the targets are true categories; based on the first generation loss value and the first discrimination loss value, train the generation network to obtain a trained generation network, and determine the trained generation network as a trained target detection model.
[0139] In a possible implementation manner, before inputting the training images, the predicted position information and predicted categories of the targets, and the true categories of the targets in the training images into a trained discrimination network in the generative adversarial network, the detection module 501 is further configured to input sample images into the generation network, and based on the generation network, obtain sample predicted position information and sample predicted categories of sample targets in the sample images; determine a second generation loss value according to the sample predicted position information and sample predicted categories of the sample targets, the sample true position information and sample true categories of the sample targets in the sample images;
[0140] The discrimination module 502 is further configured to input the sample images, the sample predicted position information and sample predicted categories of the sample targets, and the sample true categories of the sample targets in the sample images into an untrained discrimination network, and based on the untrained discrimination network, determine a sample discrimination result as to whether the sample predicted categories of the sample targets are sample true categories;
[0141] The training module 503 is further configured to determine a second discrimination loss value according to the sample discrimination result as to whether the sample predicted categories of the sample targets are sample true categories; based on the second generation loss value and the second discrimination loss value, train the untrained discrimination network until a trained discrimination network is obtained.
[0142] In a possible implementation manner, the discrimination module 502 is specifically configured to, based on a discrimination network, determine a first probability value that the category of the target in the sub-image corresponding to the predicted position information of the target in the training image is the predicted category, a second probability value that it is a non-predicted category, a third probability value that it is the true category, and a fourth probability value that it is a non-true category;
[0143] The training module 503 is specifically configured to determine a first discrimination sub-loss value according to the first probability value that the category of the target in the sub-image is the predicted category and the second probability value that it is a non-predicted category; determine a second discrimination sub-loss value according to the third probability value that the category of the target in the sub-image is the true category and the fourth probability value that it is a non-true category; and determine a first discrimination loss value according to the first discrimination sub-loss value and the second discrimination sub-loss value.
[0144] In a possible implementation manner, the detection module 501 is specifically configured to determine a first generation sub-loss value according to the predicted category and the true category of the target; determine a second generation sub-loss value according to the predicted position information and the true position information of the target; determine a predicted object label indicating the existence of a target in the sub-image corresponding to the predicted position information of the target; determine whether there is a true object label of the target in the sub-image corresponding to the predicted position information according to the predicted position information of the target; determine a third generation sub-loss value according to the predicted object label and the true object label; and determine a first generation loss value according to the first generation sub-loss value, the second generation sub-loss value, and the third generation sub-loss value.
[0145] Based on the same technical concept, on the basis of the above embodiments, the present application further provides an object detection device. Figure 6 As shown in Figure 6 the following is a schematic structural diagram of an object detection device provided by an embodiment of the present application. The device includes:
[0146] An input module 601, configured to input an image to be detected into a trained object detection model;
[0147] A processing module 602, configured to, based on the trained object detection model, determine a detection result indicating whether there is a target in the image to be detected. If there is, determine the target position information and the target category of the target in the image to be detected.
[0148] Embodiment 7:
[0149] Based on the same technical concept, the present application further provides an electronic device. Figure 7 As shown in Figure 7As shown in the figure, it includes: a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704;
[0150] The memory 703 stores a computer program. When the program is executed by the processor 701, the processor 701 is caused to execute the following steps:
[0151] Input the training images in the training set into the generator network in the generative adversarial network, and obtain the predicted position information and predicted category of the target in the training images output by the generator network; according to the predicted position information and predicted category of the target, the true position information and true category of the target in the training images, determine the first generation loss value;
[0152] Input the training images, the predicted position information and predicted category of the target, and the true category of the target in the training images into the discriminator network that has been trained in the generative adversarial network. Based on the discriminator network that has been trained, determine the discriminant result of whether the predicted category of the target is the true category;
[0153] According to the discriminant result of whether the predicted category of the target is the true category, determine the first discriminant loss value; based on the first generation loss value and the first discriminant loss value, train the generator network to obtain a trained generator network, and determine the trained generator network as the trained object detection model.
[0154] In a possible implementation manner, the processor 701 is further configured to, before inputting the training images, the predicted position information and predicted category of the target, and the true category of the target in the training images into the discriminator network that has been trained in the generative adversarial network, input the sample images into the generator network, and based on the generator network, obtain the sample predicted position information and sample predicted category of the sample target in the sample images; according to the sample predicted position information and sample predicted category of the sample target, the sample true position information and sample true category of the sample target in the sample images, determine the second generation loss value; input the sample images, the sample predicted position information and sample predicted category of the sample target, and the sample true category of the sample target in the sample images into the discriminator network to be trained, and based on the discriminator network to be trained, determine the sample discriminant result of whether the sample predicted category of the sample target is the sample true category; according to the sample discriminant result of whether the sample predicted category of the sample target is the sample true category, determine the second discriminant loss value; based on the second generation loss value and the second discriminant loss value, train the discriminator network to be trained until a trained discriminator network is obtained.
[0155] In a possible implementation, the processor 701 is specifically configured to, based on a discrimination network, determine a first probability value that the category of the target in the sub-image corresponding to the predicted position information of the target in the training image is the predicted category, a second probability value that it is a non-predicted category, a third probability value that it is the true category, and a fourth probability value that it is a non-true category; determine a first discrimination sub-loss value according to the first probability value that the category of the target in the sub-image is the predicted category and the second probability value that it is a non-predicted category; determine a second discrimination sub-loss value according to the third probability value that the category of the target in the sub-image is the true category and the fourth probability value that it is a non-true category; and determine a first discrimination loss value according to the first discrimination sub-loss value and the second discrimination sub-loss value.
[0156] In a possible implementation, the processor 701 is specifically configured to determine a first generation sub-loss value according to the predicted category and the true category of the target; determine a second generation sub-loss value according to the predicted position information and the true position information of the target; determine that there is a predicted object label of the target in the sub-image corresponding to the predicted position information of the target; determine whether there is a true object label of the target in the sub-image corresponding to the predicted position information according to the predicted position information of the target; determine a third generation sub-loss value according to the predicted object label and the true object label; and determine a first generation loss value according to the first generation sub-loss value, the second generation sub-loss value, and the third generation sub-loss value.
[0157] In addition, when the program is executed by the processor 701, the processor 701 can also perform the following steps:
[0158] Input the image to be detected into the trained object detection model, and based on the trained object detection model, determine whether there is a detection result of the target in the image to be detected. If so, determine the target position information and the target category of the target in the image to be detected.
[0159] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0160] The communication interface 702 is used for communication between the above electronic device and other devices.
[0161] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0162] The aforementioned processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0163] Embodiment 8:
[0164] Based on the same technical concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by an electronic device. When the program runs on the electronic device, it enables the electronic device to implement any of the above embodiments when executed.
[0165] The aforementioned computer-readable storage medium may be any available medium or data storage device accessible by the processor in the electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSD), etc.
[0166] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0167] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0170] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A target detection model training method, characterized in that: The method comprises: Inputting a training image in a training set into a generative network in a generative adversarial network, obtaining predicted position information and predicted category of a target in the training image output by the generative network; determining a first generation loss value according to the predicted position information and predicted category of the target, and the actual position information and actual category of the target in the training image; Inputting the training image, the predicted position information and predicted category of the target, and the true category of the target in the training image into the trained discriminant network in the generative adversarial network, and determining whether the predicted category of the target is a discriminant result of the true category based on the trained discriminant network; A first discrimination loss value is determined according to a discrimination result of whether the predicted category of the target is the true category; based on the first generation loss value and the first discrimination loss value, the generation network is trained to obtain a trained generation network, and the trained generation network is determined as a trained target detection model.
2. The method according to claim 1, characterized in that Before inputting the training image, the predicted position information and predicted category of the target, and the real category of the target in the training image into the trained discriminant network in the generative adversarial network, the method further includes: Input a sample image into a generation network, and based on the generation network, obtain sample prediction position information and sample prediction category of a sample target in the sample image output by the generation network; determine a second generation loss value according to the sample prediction position information and sample prediction category of the sample target, and the sample true position information and sample true category of the sample target in the sample image; Inputting the sample image, the sample predicted position information and the sample predicted category of the sample target, and the sample true category of the sample target in the sample image into the discriminant network to be trained, and determining whether the sample predicted category of the sample target is a sample discrimination result of the sample true category based on the discriminant network to be trained; Determining a second discrimination loss value according to a sample discrimination result of whether the sample prediction category of the sample target is the sample true category; Based on the second generation loss value and the second discrimination loss value, the discriminant network to be trained is trained until a trained discriminant network is obtained.
3. The method according to claim 1, characterized in that: The determining, based on the trained discriminant network, whether the predicted category of the target is a discriminant result of the real category includes: Based on the discriminant network, determining that the category corresponding to the target in the sub-image corresponding to the predicted position information of the target in the training image is a first probability value of the predicted category, a second probability value of the non-predicted category, a third probability value of the true category, and a fourth probability value of the non-true category; The determining of the first discrimination loss value according to the discrimination result of whether the predicted category of the target is the true category includes: Determining a first discriminant sub-loss value according to a first probability value that the category corresponding to the target in the sub-image is a predicted category and a second probability value that the category is a non-predicted category; Determine a second discriminant sub-loss value according to a third probability value that the category corresponding to the target in the sub-image is a real category and a fourth probability value that the category is a non-real category; A first discriminant loss value is determined according to the first discriminant sub-loss value and the second discriminant sub-loss value.
4. The method according to claim 1, characterized in that The determining of a first generation loss value according to the predicted position information and the predicted category of the target and the real position information and the real category of the target in the training image comprises: Determine a first generation sub-loss value according to the predicted category and the true category of the target; Determining a second generation sub-loss value according to the predicted position information and the actual position information of the target; Determine whether the sub-image corresponding to the predicted position information of the target has a predicted object label of the target; determine whether the sub-image corresponding to the predicted position information has a real object label of the target based on the predicted position information of the target; determine a third generated sub-loss value based on the predicted object label and the real object label; The first generation loss value is determined according to the first generation sub-loss value, the second generation sub-loss value, and the third generation sub-loss value.
5. A method for performing target detection using a model trained by the target detection model training method according to any one of claims 1 to 4, characterized in that: The method comprises: The image to be detected is input into the trained target detection model, and based on the trained target detection model, the detection result of whether there is a target in the image to be detected is determined. If so, the target position information and target category of the target in the image to be detected are determined.
6. A target detection model training device, characterized in that: The device comprises: A detection module, configured to input a training image in a training set into a generative network in a generative adversarial network, obtain predicted position information and predicted category of a target in the training image output by the generative network; and determine a first generation loss value according to the predicted position information and predicted category of the target, and the actual position information and actual category of the target in the training image; A discriminant module, used for inputting the training image, the predicted position information and predicted category of the target, and the real category of the target in the training image into the trained discriminant network in the generative adversarial network, and determining whether the predicted category of the target is a discriminant result of the real category based on the trained discriminant network; A training module is used to determine a first discrimination loss value according to a discrimination result of whether the predicted category of the target is the true category; based on the first generation loss value and the first discrimination loss value, the generation network is trained to obtain a trained generation network, and the trained generation network is determined as a trained target detection model.
7. The device according to claim 6, characterized in that Before inputting the training image, the predicted position information and predicted category of the target, and the real category of the target in the training image into the trained discriminant network in the generative adversarial network, the detection module is further used to input the sample image into the generative network, and based on the generative network, obtain the sample predicted position information and sample predicted category of the sample target in the sample image output by the generative network; Determine a second generation loss value according to the sample predicted position information and the sample predicted category of the sample target, the sample true position information and the sample true category of the sample target in the sample image; The discrimination module is further used to input the sample image, the sample predicted position information and the sample predicted category of the sample target, and the sample true category of the sample target in the sample image into the discriminant network to be trained, and determine whether the sample predicted category of the sample target is a sample discrimination result of the sample true category based on the discriminant network to be trained; The training module is further used to determine a second discrimination loss value according to a sample discrimination result of whether the sample prediction category of the sample target is the sample true category; Based on the second generation loss value and the second discrimination loss value, the discriminant network to be trained is trained until a trained discriminant network is obtained.
8. A device for performing target detection based on a model trained by the target detection model training method according to any one of claims 1 to 4, characterized in that: The device comprises: An input module is used to input the image to be detected into the trained object detection model; The processing module is used to determine whether there is a detection result of a target in the image to be detected based on the trained target detection model, and if so, determine the target position information and target category of the target in the image to be detected.
9. An electronic device, characterized in that: The electronic device includes at least a processor and a memory, and the processor is used to implement the target detection model training method as described in any one of claims 1 to 4 or the steps of the target detection method as described in claim 5 when executing the computer program stored in the memory.
10. A computer storage medium, characterized in that: It stores a computer program that can be executed by an electronic device. When the program runs on the electronic device, the electronic device executes the target detection model training method described in any one of claims 1 to 4 or the steps of the target detection method described in claim 5.