Object Detection Method and Related Devices
By using a fusion neural network model to label multiple images twice in object detection, the problem of low accuracy and efficiency in detecting a large number of images in the prior art is solved, and automated labeling and efficient detection are realized.
Patent Information
- Application Number
- CN202011188078.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-29
AI Technical Summary
The prior art is inaccurate and efficient when detecting multiple set targets present in a large number of pictures, especially an annotation process that requires manual participation.
The fusion neural network model is adopted, including the convolutional neural network model and the generative adversarial network model, and multiple images to be detected are marked twice, and the set targets in the picture are automatically identified and marked.
Without manual participation, it significantly improves the accuracy and efficiency of detecting multiple set targets present in a large number of images and improves the degree of automation of image annotations.
Smart Images

Figure CN112464985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network technologies, and particularly to an object detection method and related devices. Background Art
[0002] Currently, the process of object detection is as follows: First, the pictures to be detected are preliminarily labeled through machine learning; then, the preliminarily labeled images are relabeled manually to obtain an object detection result including at least one set target.
[0003] However, when there are a large number of pictures, the manual method will result in low accuracy and efficiency in detecting multiple set targets existing in a large number of pictures. Summary of the Invention
[0004] Embodiments of this application provide an object detection method and related devices for improving the accuracy and efficiency in detecting multiple set targets existing in a large number of pictures.
[0005] In a first aspect, embodiments of this application provide an object detection method, and the method includes:
[0006] Obtain multiple pictures to be detected;
[0007] Input the multiple pictures to be detected into a fusion neural network model for processing, and output at least one object detection result each including one or more set targets, where at least one object detection result corresponds to at least one picture to be detected one by one; wherein, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model;
[0008] Perform a display operation on at least one object detection result.
[0009] In some possible embodiments, inputting the multiple pictures to be detected into a fusion neural network model for processing and outputting at least one object detection result each including one or more set targets includes:
[0010] Input the multiple pictures to be detected into a convolutional neural network model for primary processing, and output multiple first labeled images, where the multiple first labeled images correspond to the multiple pictures to be detected one by one;
[0011] Input the multiple first labeled images into a generative adversarial network model for secondary processing, and output at least one object detection result each including one or more set targets.
[0012] In some possible embodiments, inputting the multiple first labeled images into a generative adversarial network model for secondary processing and outputting at least one object detection result each including one or more set targets includes:
[0013] Input multiple first labeled images into a generative adversarial network model;
[0014] Label the multiple first labeled images to obtain multiple second labeled images. The area of each second labeled image in the multiple second labeled images, which includes one or more set targets, is smaller than the area of its corresponding first labeled image. The multiple second labeled images correspond one-to-one with the multiple first labeled images;
[0015] Use one or more set target recognition algorithms to perform target recognition on the multiple second labeled images, obtaining multiple first set target quantities and multiple first set target sets. The multiple first set target quantities and the multiple first set target sets respectively correspond one-to-one with the multiple second labeled images;
[0016] Select at least one target second labeled image from the multiple second labeled images. The first set target quantity corresponding to each target second labeled image in the at least one target second labeled image is non-zero. The at least one target second labeled image and the at least one first set target set are at least one target detection result. The at least one target second labeled image and the at least one first set target set respectively correspond one-to-one with the at least one target detection result.
[0017] In a second aspect, an embodiment of the present application provides a target detection device, which includes:
[0018] An acquisition unit for acquiring multiple pictures to be detected;
[0019] A processing unit for inputting the multiple pictures to be detected into a fusion neural network model for processing and outputting at least one target detection result that includes one or more set targets. The at least one target detection result corresponds one-to-one with at least one picture to be detected; wherein the fusion neural network model includes a convolutional neural network model and a generative adversarial network model;
[0020] A display unit for performing a display operation on the at least one target detection result.
[0021] In a third aspect, an embodiment of the present application provides a target detection device. The above target detection device includes a processor, a memory, a communication interface, and a target detection program stored in the memory. The above target detection program is configured to be executed by the processor to implement the steps in the method of the first aspect of the embodiments of the present application.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The above computer-readable storage medium is used to store a target detection program. The above target detection program is executed by a processor to implement the steps in the method of the first aspect of the embodiments of the present application.
[0023] Fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a target detection program. The target detection program is operable to cause a computer to execute the steps in the method of the first aspect of the embodiments of the present application. This computer program product can be a software installation package.
[0024] It can be seen that compared with initially annotating a large number of pictures through machine learning and re-annotating a large number of images after initial annotation manually, in the embodiments of the present application, multiple pictures obtained are annotated twice through a convolutional neural network model and a generative adversarial network model to obtain multiple set targets existing in the multiple pictures and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of pictures.
[0025] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.
[0027] Figure 1 is a schematic structural diagram of a target detection system provided by an embodiment of the present application;
[0028] Figure 2A is a schematic flowchart of a target detection method provided by an embodiment of the present application;
[0029] Figure 2B is a schematic structural diagram of a fusion neural network model provided by an embodiment of the present application;
[0030] Figure 3 is a schematic flowchart of another target detection method provided by an embodiment of the present application;
[0031] Figure 4 is a block diagram of the functional units of a target detection device provided by an embodiment of the present application;
[0032] Figure 5 is a schematic structural diagram of a target detection system provided by an embodiment of the present application.
[0033] Specific implementation method
[0034] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0035] The following will be described in detail respectively.
[0036] The terms "first", "second", "third", "fourth", etc. in the specification, claims and drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.
[0037] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0038] The following will introduce the embodiments of this application in detail.
[0039] Please refer to Figure 1 , Figure 1 which is a schematic architecture diagram of an object detection system 100 provided by an embodiment of this application. The object detection system 100 includes a processor 101 and a display screen 102. The processor 101 is connected to the display screen 102, wherein:
[0040] The processor 101 is configured to obtain multiple pictures to be detected;
[0041] The processor 101 is further configured to input the multiple pictures to be detected into a fusion neural network model for processing, and output at least one object detection result including one or more set objects. The at least one object detection result corresponds to at least one picture to be detected one by one; wherein, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model;
[0042] The display screen 102 is configured to perform a display operation on at least one object detection result.
[0043] It can be seen that, compared with initially annotating a large number of images through machine learning and re-annotating a large number of images after initial annotation manually, in the embodiments of the present application, multiple images obtained are annotated twice through a convolutional neural network model and a generative adversarial network model to obtain multiple set targets existing in the multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0044] Please refer to Figure 2A , Figure 2A which is a schematic flowchart of a target detection method provided by an embodiment of the present application. The target detection method can be applied to an intelligent monitoring scenario or an autonomous driving scenario. The target detection method includes steps S201 - S203, specifically as follows:
[0045] S201. The target detection device acquires multiple images to be detected.
[0046] Among them, the target detection device can be an intelligent terminal device or a server, which is not limited herein.
[0047] In some possible embodiments, before the target detection device acquires multiple images to be detected, the method further includes:
[0048] The target detection device acquires a training data set, and the training data set includes multiple pairs of training data;
[0049] The target detection device trains an untrained generative adversarial network model using the training data set to obtain a trained generative adversarial network model, and the trained generative adversarial network model is a generative adversarial network model.
[0050] Among them, the generative adversarial network model is obtained by training an untrained generative adversarial network model using multiple pairs of training data, and the generative adversarial network model has higher accuracy in annotating images.
[0051] It can be seen that, in this example, multiple images obtained are annotated twice through a convolutional neural network model and a generative adversarial network model to obtain multiple set targets existing in the multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images; in addition, since the generative adversarial network model is obtained by training an untrained generative adversarial network using multiple pairs of training data, the accuracy of detecting multiple set targets existing in a large number of images is further improved.
[0052] S202. The target detection device inputs multiple images to be detected into the fusion neural network model for processing, and outputs at least one target detection result including one or more set targets. The at least one target detection result corresponds one-to-one to at least one image to be detected.
[0053] As Figure 2B shown, Figure 2B FIG. is a schematic structural diagram of a fusion neural network model provided by the present application. The fusion neural network model includes a convolutional neural network model and a generative adversarial network model.
[0054] Among them, the set target can be a face or a license plate, which is not limited herein.
[0055] In some possible embodiments, the target detection device inputs multiple images to be detected into the fusion neural network model for processing, and outputs at least one target detection result including one or more set targets, including:
[0056] The target detection device inputs multiple images to be detected into the convolutional neural network model for initial processing, and outputs multiple first labeled images. The multiple first labeled images correspond one-to-one to the multiple images to be detected;
[0057] The target detection device inputs the multiple first labeled images into the generative adversarial network model for further processing, and outputs at least one target detection result including one or more set targets.
[0058] It can be seen that in this example, the convolutional neural network model is first used to perform initial labeling on multiple images to be detected to obtain multiple first labeled images, and then the generative adversarial network model is used to perform further labeling on the multiple first labeled images to obtain multiple set targets existing in the multiple images and perform a display operation. In this way, the process of image labeling does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0059] In some possible embodiments, the target detection device inputs the multiple first labeled images into the generative adversarial network model for further processing, and outputs at least one target detection result including one or more set targets, including:
[0060] The target detection device inputs the multiple first labeled images into the generative adversarial network model;
[0061] The target detection device labels the multiple first labeled images to obtain multiple second labeled images. The area of each second labeled image including one or more set targets in the multiple second labeled images is smaller than the area of its corresponding first labeled image. The multiple second labeled images correspond one-to-one to the multiple first labeled images;
[0062] The target detection device uses one or more set target recognition algorithms to perform target recognition on multiple second annotated images, obtaining multiple first set target quantities and multiple first set target sets, where the multiple first set target quantities and the multiple first set target sets respectively correspond one-to-one to the multiple second annotated images;
[0063] The target detection device selects at least one target second annotated image from the multiple second annotated images. For each target second annotated image among the at least one target second annotated images, the corresponding first set target quantity is non-zero. The at least one target second annotated image and the at least one first set target set are at least one target detection result, and the at least one target second annotated image and the at least one first set target set respectively correspond one-to-one to the at least one target detection result.
[0064] Among them, the area of the second annotated image is smaller than the area of its corresponding first annotated image, indicating that the accuracy of annotating the second annotated image is higher than the accuracy of annotating the first annotated image corresponding to the second annotated image.
[0065] Among them, the set target recognition algorithm can be a face recognition algorithm or a license plate recognition algorithm, which is not limited here.
[0066] It can be seen that in this example, multiple first annotated images are first obtained by initially annotating multiple images to be detected through a convolutional neural network model, and then multiple second annotated images are obtained by re-annotating the multiple first annotated images through a generative adversarial network model. One or more set target recognition algorithms are used to perform target recognition on the multiple second annotated images to obtain multiple first set target quantities and multiple first set target sets, and at least one target second annotated image is selected from the multiple second annotated images to obtain multiple set targets existing in multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0067] In some possible embodiments, the target detection device inputs the multiple first annotated images into a generative adversarial network model for further processing, and the output includes at least one target detection result of one or more set targets, including:
[0068] The target detection device inputs the multiple first annotated images into the generative adversarial network model;
[0069] The target detection device uses one or more set target recognition algorithms to perform target recognition on the multiple first annotated images, obtaining multiple third set target quantities and multiple third set target sets, where the multiple third set target quantities and the multiple third set target sets respectively correspond one-to-one to the multiple first annotated images;
[0070] The target detection device selects at least one target first labeled image from multiple first labeled images, and the third set target quantity corresponding to each target first labeled image in the at least one target first labeled image is non-zero;
[0071] The target detection device labels the at least one target first labeled image to obtain at least one fourth labeled image. The area of each fourth labeled image in the at least one fourth labeled image is smaller than the area of its corresponding target first labeled image. The at least one fourth labeled image corresponds one-to-one with the at least one target first labeled image. The at least one fourth labeled image and the at least one third set target set are at least one target detection result, and the at least one fourth labeled image and the at least one third set target set respectively correspond one-to-one with the at least one target detection result.
[0072] Among them, the set target recognition algorithm can be a face recognition algorithm or a license plate recognition algorithm, which is not limited here.
[0073] Among them, the area of the fourth labeled image being smaller than the area of its corresponding target first labeled image means that the accuracy of labeling the fourth labeled image is higher than the accuracy of labeling the target first labeled image corresponding to the fourth labeled image.
[0074] It can be seen that in this example, multiple images to be detected are first initially labeled by a convolutional neural network model to obtain multiple first labeled images, and then a generative adversarial network model uses one or more set target recognition algorithms to perform target recognition on the multiple first labeled images to obtain multiple third set target quantities and multiple third set target sets. At least one target first labeled image is selected from the multiple first labeled images, and at least one target first labeled image is labeled to obtain at least one fourth labeled image, so as to obtain multiple set targets existing in multiple images and perform a display operation. In this way, the process of image labeling does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0075] In some possible embodiments, after the target detection device inputs multiple images to be detected into a convolutional neural network model for initial processing and outputs multiple first labeled images, the method further includes:
[0076] The target detection device uses one or more set target recognition algorithms to perform target recognition on the multiple first labeled images to obtain multiple second set target quantities and multiple second set target sets. The multiple second set target quantities and the multiple second set target sets respectively correspond one-to-one with the multiple first labeled images;
[0077] The target detection device selects at least one target first-annotated image from multiple first-annotated images, and the number of second-set targets corresponding to each target first-annotated image in the at least one target first-annotated image is non-zero;
[0078] The target detection device inputs the multiple first-annotated images into a generative adversarial network model for further processing, and outputs at least one target detection result including one or more set targets, including:
[0079] The target detection device inputs the at least one target first-annotated image into the generative adversarial network model;
[0080] The target detection device annotates the at least one target first-annotated image to obtain at least one third-annotated image. The area of each third-annotated image in the at least one third-annotated image is smaller than the area of its corresponding target first-annotated image. The at least one third-annotated image corresponds one-to-one with the at least one target first-annotated image. The at least one third-annotated image and the at least one second-set target set are the at least one target detection result, and the at least one third-annotated image and the at least one second-set target set correspond one-to-one with the at least one target detection result respectively.
[0081] Among them, the set target recognition algorithm can be a face recognition algorithm or a license plate recognition algorithm, which is not limited here.
[0082] Among them, the area of the third-annotated image being smaller than the area of its corresponding target first-annotated image means that the accuracy of annotating the third-annotated image is higher than the accuracy of annotating the target first-annotated image corresponding to the third-annotated image.
[0083] It can be seen that in this example, multiple images to be detected are first initially annotated by a convolutional neural network model to obtain multiple first-annotated images, then one or more set target recognition algorithms are used to recognize the targets in the multiple first-annotated images to obtain multiple second-set target quantities and multiple second-set target sets, and at least one target first-annotated image is selected from the multiple first-annotated images. Then, the at least one target first-annotated image is annotated by a generative adversarial network model to obtain at least one third-annotated image, so as to obtain multiple set targets existing in multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0084] S203. The target detection device performs a display operation on at least one target detection result.
[0085] It can be seen that, compared with initially annotating a large number of images through machine learning and re-annotating a large number of images after initial annotation manually, in the embodiments of the present application, multiple images obtained are twice annotated through a convolutional neural network model and a generative adversarial network model to obtain multiple set targets existing in the multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0086] In some possible embodiments, after the target detection device inputs multiple images to be detected into a fusion neural network model for processing and outputs at least one target detection result each including one or more set targets, the method further includes:
[0087] The target detection device performs identity recognition on the second set target set corresponding to the third annotated image A to obtain a first identity set. The first identity set has the same number as and corresponds one-to-one with the second set target set corresponding to the third annotated image A, and the third annotated image A is any one of at least one third annotated image;
[0088] The target detection device performs a display operation on at least one target detection result, including:
[0089] The target detection device performs a display operation on at least one third annotated image, at least one second set target set, and at least one first identity set.
[0090] Among them, the identity can be a person's name, a student, a teacher, or a license plate number, a vehicle brand, a vehicle type (such as a bicycle, an electric bicycle, a car), which is not limited herein.
[0091] Specifically, the second set target set corresponding to the third annotated image A includes a second set target 1, a second set target 2, and a second set target 3. The target detection device performs identity recognition on the second set target set corresponding to the third annotated image A to obtain a first identity set, including:
[0092] The target detection device determines the first identity 1 corresponding to the second set target 1 according to the pre-stored mapping relationship between the set target and the identity;
[0093] The target detection device determines the first identity 2 corresponding to the second set target 2 according to the mapping relationship between the set target and the identity;
[0094] The target detection device determines the first identity 3 corresponding to the second set target 3 according to the mapping relationship between the set target and the identity. The first identity 1, the first identity 2, and the first identity 3 are the first identity set.
[0095] Among them, the set targets and the identities in the mapping relationship between the set target and the identity correspond one-to-one.
[0096] Among them, the order in which the target detection device performs identity recognition on the second set target 1, the second set target 2, and the second set target 3 is not in any particular order. The target detection device can also use a parallel method to perform identity recognition on the second set target 1, the second set target 2, and the second set target 3.
[0097] It can be seen that in this example, multiple images to be detected are first initially labeled by a convolutional neural network model to obtain multiple first labeled images. Then, one or more set target recognition algorithms are used to perform target recognition on the multiple first labeled images to obtain multiple second set target quantities, multiple second set target sets, and at least one target first labeled image is selected from the multiple first labeled images. Subsequently, the generative adversarial network model is used to label at least one target first labeled image to obtain at least one third labeled image, and the identity of the second set target set corresponding to each third labeled image is recognized to obtain at least one first identity set, so as to obtain multiple set targets existing in multiple images and the identity of each set target among the multiple set targets and perform a display operation. In this way, the process of image labeling does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images; at the same time, the identity of each set target among the multiple set targets can be obtained conveniently.
[0098] In some possible embodiments, after the target detection device inputs multiple images to be detected into the fusion neural network model for processing and outputs at least one target detection result each including one or more set targets, the method further includes:
[0099] The target detection device inputs at least one third labeled image into the classification neural network model;
[0100] The target detection device performs identity recognition on the second set target set corresponding to the third labeled image B to obtain a second identity set. The second identity set has the same quantity as and is in one-to-one correspondence with the second set target set corresponding to the third labeled image A. The third labeled image B is any one of the at least one third labeled image;
[0101] The target detection device classifies at least one third labeled image and at least one second set target set in one-to-one correspondence with the at least one third labeled image according to at least one second identity set to obtain N third labeled image groups and N second set target set groups. Each second set target set corresponding to each third labeled image in the N third labeled image groups corresponds to the same identity. The N second set target set groups are in one-to-one correspondence with the N third labeled image groups, and N is an integer greater than or equal to 1;
[0102] The target detection device performs a display operation on at least one target detection result, including:
[0103] The target detection device performs a display operation on N groups of third-annotated images, N groups of second-set target sets, and N second identities, and the N groups of third-annotated images and the N groups of second-set target sets correspond to the N second identities one by one.
[0104] Among them, the fusion neural network model includes a convolutional neural network model, a generative adversarial network model, and a classification neural network model.
[0105] Among them, for the target detection device to perform identity recognition on the second-set target set corresponding to the third-annotated image B to obtain the second identity set, the implementation manner can refer to the implementation manner in which the target detection device performs identity recognition on the second-set target set corresponding to the third-annotated image A to obtain the first identity set, which will not be described here again.
[0106] Among them, the identity can be a person's name, a student, a teacher, or a license plate number, a vehicle brand, a vehicle type (such as a bicycle, an electric bicycle, a car), which is not limited here.
[0107] It can be seen that in this example, first, the convolutional neural network model is used to perform primary annotation on multiple images to be detected to obtain multiple first-annotated images, then one or more set target recognition algorithms are used to perform target recognition on the multiple first-annotated images to obtain multiple second-set target quantities and multiple second-set target sets, and at least one target first-annotated image is selected from the multiple first-annotated images. Then, the generative adversarial network model is used to annotate at least one target first-annotated image to obtain at least one third-annotated image. The classification neural network model is used to perform identity recognition on the second-set target set corresponding to each third-annotated image to obtain at least one second identity set. Classification is performed according to at least one second identity set to obtain N groups of third-annotated images and N groups of second-set target sets, and a display operation is performed on the N groups of third-annotated images, the N groups of second-set target sets, and the N second identities. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images; at the same time, the identity of each set target in the multiple set targets can be obtained conveniently.
[0108] In some possible embodiments, after the target detection device inputs multiple images to be detected into the fusion neural network model for processing and outputs at least one target detection result including one or more set targets, the method further includes:
[0109] The target detection device uses a preset edge segmentation algorithm to perform edge segmentation on at least one second-set target set to obtain at least one fourth-set target set, and the at least one fourth-set target set corresponds to the at least one second-set target set one by one;
[0110] The target detection device performs a display operation on at least one target detection result, including:
[0111] The target detection device performs a display operation on at least one third labeled image and at least one fourth set of preset targets.
[0112] The preset edge segmentation algorithm can be an image-based edge segmentation algorithm, which is not limited here.
[0113] Among them, the fourth set of preset targets is obtained by performing edge segmentation on the second set of preset targets corresponding to the fourth set of preset targets using the preset edge segmentation algorithm, and the edge details of the fourth set of preset targets are richer.
[0114] It can be seen that in this example, multiple images to be detected are initially labeled through a convolutional neural network model to obtain multiple first labeled images, and then one or more preset target recognition algorithms are used to perform target recognition on the multiple first labeled images to obtain multiple second preset target quantities and multiple second sets of preset targets, and at least one target first labeled image is selected from the multiple first labeled images. Then, the at least one target first labeled image is labeled through a generative adversarial network model, and the preset edge segmentation algorithm is used to perform edge segmentation on at least one second set of preset targets to obtain at least one fourth set of preset targets, so as to obtain multiple preset targets existing in multiple images and perform a display operation. In this way, the process of image labeling does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple preset targets existing in a large number of images; in addition, since the fourth set of preset targets is obtained by performing edge segmentation on the second set of preset targets corresponding to the fourth set of preset targets using the preset edge segmentation algorithm, the edge details of the fourth set of preset targets are richer.
[0115] Consistent with the above Figure 2A shown embodiment, please refer to Figure 3 , Figure 3 is a flowchart of another target detection method provided by an embodiment of the present application. The target detection method includes steps S301-S309, specifically as follows:
[0116] S301. The target detection device obtains a training data set, and the training data set includes multiple pairs of training data.
[0117] S302. The target detection device uses the training data set to train an untrained generative adversarial network model to obtain a trained generative adversarial network model, and the trained generative adversarial network model is a generative adversarial network model.
[0118] S303. The target detection device obtains multiple images to be detected.
[0119] S304. The target detection device inputs multiple images to be detected into a convolutional neural network model for initial processing, and outputs multiple first labeled images, where the multiple first labeled images correspond one-to-one to the multiple images to be detected.
[0120] S305. The target detection device inputs the multiple first labeled images into a generative adversarial network model.
[0121] S306. The target detection device labels the multiple first labeled images to obtain multiple second labeled images. For each second labeled image among the multiple second labeled images that includes one or more set targets, the area of the second labeled image is smaller than the area of its corresponding first labeled image. The multiple second labeled images correspond one-to-one to the multiple first labeled images.
[0122] S307. The target detection device uses one or more set target recognition algorithms to perform target recognition on the multiple second labeled images, and obtains multiple first set target quantities and multiple first set target sets, where the multiple first set target quantities and the multiple first set target sets correspond one-to-one to the multiple second labeled images respectively.
[0123] S308. The target detection device selects at least one target second labeled image from the multiple second labeled images. For each target second labeled image in the at least one target second labeled image, the corresponding first set target quantity is non-zero. The at least one target second labeled image and the at least one first set target set are at least one target detection result, and the at least one target second labeled image and the at least one first set target set correspond one-to-one to the at least one target detection result respectively.
[0124] S309. The target detection device performs a display operation on the at least one target detection result.
[0125] It should be noted that Figure 3 For the specific implementation processes of the steps in the method shown, reference can be made to the specific implementation process of the above method, which will not be described here again.
[0126] The above embodiments mainly introduce the solutions of the embodiments of the present application from the perspective of the execution process on the method side. It can be understood that in order for the target detection device to implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0127] The embodiments of the present application can divide the functional units of the target detection device according to the method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0128] The following is the device embodiment of the present application. The device embodiment of the present application is used to execute the method implemented by the method embodiment of the present application. Please refer to Figure 4 , Figure 4 which is a block diagram of the functional unit composition of a target detection device 400 provided by the embodiments of the present application. The target detection device 400 includes:
[0129] An acquisition unit 401, configured to acquire multiple pictures to be detected;
[0130] A processing unit 402, configured to input the multiple pictures to be detected into a fusion neural network model for processing, and output at least one target detection result including one or more set targets, and at least one target detection result corresponds to at least one picture to be detected one by one; wherein, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model;
[0131] A display unit 403, configured to perform a display operation on at least one target detection result.
[0132] It can be seen that, compared with initially annotating a large number of images through machine learning and re-annotating a large number of images after initial annotation manually, in the embodiments of the present application, multiple acquired images are twice annotated through a convolutional neural network model and a generative adversarial network model to obtain multiple set targets existing in the multiple images and perform a display operation. In this way, the process of image annotation does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple set targets existing in a large number of images.
[0133] In some possible embodiments, in terms of inputting multiple images to be detected into a fusion neural network model for processing and outputting at least one target detection result each including one or more set targets, the above-mentioned processing unit 402 is specifically configured to:
[0134] Input multiple images to be detected into a convolutional neural network model for initial processing, and output multiple first annotated images, where the multiple first annotated images correspond to the multiple images to be detected one by one;
[0135] Input the multiple first annotated images into a generative adversarial network model for re-processing, and output at least one target detection result each including one or more set targets.
[0136] In some possible embodiments, in terms of inputting the multiple first annotated images into a generative adversarial network model for re-processing and outputting at least one target detection result each including one or more set targets, the above-mentioned processing unit 402 is specifically configured to:
[0137] Input the multiple first annotated images into the generative adversarial network model;
[0138] Annotate the multiple first annotated images to obtain multiple second annotated images, where the area of each second annotated image including one or more set targets in the multiple second annotated images is smaller than the area of its corresponding first annotated image, and the multiple second annotated images correspond to the multiple first annotated images one by one;
[0139] Use one or more set target recognition algorithms to perform target recognition on the multiple second annotated images to obtain multiple first set target quantities and multiple first set target sets, where the multiple first set target quantities and the multiple first set target sets correspond to the multiple second annotated images one by one;
[0140] Select at least one target second annotated image from the multiple second annotated images, where the first set target quantity corresponding to each target second annotated image in the at least one target second annotated image is non-zero, and the at least one target second annotated image and the at least one first set target set are at least one target detection result, and the at least one target second annotated image and the at least one first set target set correspond to the at least one target detection result one by one.
[0141] In some possible embodiments, the above-mentioned target detection device 400 further includes:
[0142] A target recognition unit 404, configured to perform target recognition on multiple first labeled images by using one or more set target recognition algorithms, to obtain multiple second set target quantities and multiple second set target sets, where the multiple second set target quantities and the multiple second set target sets respectively correspond one-to-one to the multiple first labeled images;
[0143] A selection unit 405, configured to select at least one target first labeled image from the multiple first labeled images, where the second set target quantity corresponding to each target first labeled image in the at least one target first labeled image is non-zero;
[0144] In terms of inputting the multiple first labeled images into the generative adversarial network model for further processing and outputting at least one target detection result including one or more set targets, the above-mentioned processing unit 402 is specifically configured to:
[0145] Input at least one target first labeled image into the generative adversarial network model;
[0146] Label at least one target first labeled image to obtain at least one third labeled image, where the area of each third labeled image in the at least one third labeled image is smaller than the area of its corresponding target first labeled image, the at least one third labeled image corresponds one-to-one to the at least one target first labeled image, the at least one third labeled image and the at least one second set target set are at least one target detection result, and the at least one third labeled image and the at least one second set target set respectively correspond one-to-one to the at least one target detection result.
[0147] In some possible embodiments, the above-mentioned target detection device 400 further includes:
[0148] A training unit 406, configured to obtain a training data set, where the training data set includes multiple pairs of training data; use the training data set to train an untrained generative adversarial network model to obtain a trained generative adversarial network model, and the trained generative adversarial network model is the generative adversarial network model.
[0149] In some possible embodiments, the above-mentioned target detection device 400 further includes:
[0150] An identity recognition unit 407, configured to perform identity recognition on the second set target set corresponding to the third labeled image A to obtain a first identity set, where the first identity set has the same quantity as and corresponds one-to-one to the second set target set corresponding to the third labeled image A, and the third labeled image A is any one of the at least one third labeled image;
[0151] In terms of performing a display operation on at least one target detection result, the above display unit 403 is specifically configured to:
[0152] Perform a display operation on at least one third labeled image, at least one second set of preset targets, and at least one first identity set.
[0153] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a target detection system provided by an embodiment of the present application. The target detection system includes a processor, a memory, a communication interface, and a target detection program stored in the memory. The above target detection program is configured to be executed by the above processor. The above target detection program includes instructions for performing the following steps:
[0154] Obtain multiple pictures to be detected;
[0155] Input the multiple pictures to be detected into a fusion neural network model for processing, and output at least one target detection result each including one or more preset targets. The at least one target detection result corresponds one-to-one with at least one picture to be detected; wherein, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model;
[0156] Perform a display operation on at least one target detection result.
[0157] It can be seen that compared with initially labeling a large number of pictures through machine learning and then relabeling a large number of the initially labeled pictures manually, in the embodiment of the present application, the convolutional neural network model and the generative adversarial network model are used to perform two labelings on the obtained multiple pictures, obtain multiple preset targets existing in the multiple pictures, and perform a display operation. In this way, the process of image labeling does not require manual participation, greatly improving the accuracy and efficiency of detecting multiple preset targets existing in a large number of pictures.
[0158] In some possible embodiments, in terms of inputting the multiple pictures to be detected into a fusion neural network model for processing and outputting at least one target detection result each including one or more preset targets, the above target detection program includes instructions specifically for performing the following steps:
[0159] Input the multiple pictures to be detected into a convolutional neural network model for initial processing, and output multiple first labeled images. The multiple first labeled images correspond one-to-one with the multiple pictures to be detected;
[0160] Input the multiple first labeled images into a generative adversarial network model for further processing, and output at least one target detection result each including one or more preset targets.
[0161] In some possible embodiments, in the aspect of inputting multiple first annotated images into a generative adversarial network model for further processing and outputting at least one object detection result that includes one or more set targets, the above object detection program includes instructions specifically for performing the following steps:
[0162] Input multiple first annotated images into the generative adversarial network model;
[0163] Annotate the multiple first annotated images to obtain multiple second annotated images. The area of each second annotated image in the multiple second annotated images that includes one or more set targets is smaller than the area of its corresponding first annotated image. The multiple second annotated images correspond one-to-one with the multiple first annotated images;
[0164] Use one or more set target recognition algorithms to perform object recognition on the multiple second annotated images to obtain multiple first set target quantities and multiple first set target sets. The multiple first set target quantities and the multiple first set target sets correspond one-to-one with the multiple second annotated images respectively;
[0165] Select at least one target second annotated image from the multiple second annotated images. The first set target quantity corresponding to each target second annotated image in the at least one target second annotated image is non-zero. The at least one target second annotated image and the at least one first set target set are at least one object detection result. The at least one target second annotated image and the at least one first set target set correspond one-to-one with the at least one object detection result respectively.
[0166] In some possible embodiments, the above object detection program further includes instructions for performing the following steps:
[0167] Use one or more set target recognition algorithms to perform object recognition on the multiple first annotated images to obtain multiple second set target quantities and multiple second set target sets. The multiple second set target quantities and the multiple second set target sets correspond one-to-one with the multiple first annotated images respectively;
[0168] Select at least one target first annotated image from the multiple first annotated images. The second set target quantity corresponding to each target first annotated image in the at least one target first annotated image is non-zero;
[0169] In the aspect of inputting multiple first annotated images into a generative adversarial network model for further processing and outputting at least one object detection result that includes one or more set targets, the above object detection program includes instructions specifically for performing the following steps:
[0170] Input at least one target first annotated image into the generative adversarial network model;
[0171] Annotate at least one target first annotated image to obtain at least one third annotated image. The area of each third annotated image in the at least one third annotated image is smaller than the area of its corresponding target first annotated image. The at least one third annotated image corresponds one-to-one with the at least one target first annotated image. The at least one third annotated image and the at least one second set of preset targets are at least one target detection result, and the at least one third annotated image and the at least one second set of preset targets correspond one-to-one with the at least one target detection result respectively.
[0172] In some possible embodiments, the above-mentioned target detection program further includes instructions for performing the following steps:
[0173] Obtain a training data set, where the training data set includes multiple pairs of training data;
[0174] Train an untrained generative adversarial network model using the training data set to obtain a trained generative adversarial network model. The trained generative adversarial network model is a generative adversarial network model.
[0175] In some possible embodiments, the above-mentioned target detection program further includes instructions for performing the following steps:
[0176] Perform identity recognition on the second set of preset targets corresponding to the third annotated image A to obtain a first identity set. The first identity set has the same number as and corresponds one-to-one with the second set of preset targets corresponding to the third annotated image A. The third annotated image A is any one of the at least one third annotated image;
[0177] In terms of performing a display operation on at least one target detection result, the above-mentioned target detection program includes instructions specifically for performing the following steps:
[0178] Perform a display operation on at least one third annotated image, at least one second set of preset targets, and at least one first identity set.
[0179] An embodiment of the present application provides a computer storage medium. The above-mentioned computer storage medium is used to store a target detection program. The above-mentioned target detection program is executed by a processor to implement the steps in any of the methods described in the above method embodiments.
[0180] An embodiment of the present application further provides a computer program product. The above-mentioned computer program product includes a non-transitory computer-readable storage medium storing a target detection program. The above-mentioned target detection program is operable to cause a computer to execute the steps in any of the methods described in the above method embodiments. This computer program product can be a software installation package.
[0181] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0182] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0183] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical or other forms.
[0184] The units described as separate components above may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0185] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0186] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of this application. The aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), external hard drives, magnetic disks, or optical discs that can store program codes.
[0187] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs, etc.
[0188] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A target detection method, characterized in that, The method includes: Obtain multiple pictures to be detected; Input the multiple pictures to be detected into a fusion neural network model for processing, and output at least one object detection result including one or more set targets, where the at least one object detection result corresponds to at least one picture to be detected one by one; among them, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model; Perform a display operation on the at least one object detection result; the inputting the multiple pictures to be detected into the fusion neural network model for processing and outputting at least one object detection result including one or more set targets includes: Input the multiple pictures to be detected into the convolutional neural network model for initial processing, and output multiple first labeled images, where the multiple first labeled images correspond to the multiple pictures to be detected one by one; Input the multiple first labeled images into the generative adversarial network model for further processing, and output at least one object detection result including one or more set targets; The inputting the multiple first labeled images into the generative adversarial network model for further processing and outputting at least one object detection result including one or more set targets includes: Input the multiple first labeled images into the generative adversarial network model; Label the multiple first labeled images to obtain multiple second labeled images, where the area of each second labeled image including one or more set targets in the multiple second labeled images is smaller than the area of its corresponding first labeled image, and the multiple second labeled images correspond to the multiple first labeled images one by one; Use one or more set target recognition algorithms to perform target recognition on the multiple second labeled images to obtain multiple first set target quantities and multiple first set target sets, where the multiple first set target quantities and the multiple first set target sets correspond to the multiple second labeled images one by one; Select at least one target second labeled image from the multiple second labeled images, where the first set target quantity corresponding to each target second labeled image in the at least one target second labeled image is non-zero, the at least one target second labeled image is the object detection result, and the at least one first set target set is the object detection result, so as to obtain at least one object detection result.
2. The method according to claim 1, characterized in that After the inputting the multiple pictures to be detected into the convolutional neural network model for initial processing and outputting multiple first labeled images, the method further includes: Use one or more set target recognition algorithms to perform target recognition on the multiple first labeled images to obtain multiple second set target quantities and multiple second set target sets, where the multiple second set target quantities and the multiple second set target sets correspond to the multiple first labeled images one by one; Select at least one target first labeled image from the multiple first labeled images, where the second set target quantity corresponding to each target first labeled image in the at least one target first labeled image is non-zero; Inputting the multiple first labeled images into the generative adversarial network model for reprocessing, and outputting at least one object detection result including one or more set targets, includes: Inputting the at least one object first labeled image into the generative adversarial network model; Labeling the at least one object first labeled image to obtain at least one third labeled image, where the area of each third labeled image in the at least one third labeled image is smaller than the area of its corresponding object first labeled image, the at least one third labeled image corresponds one-to-one with the at least one object first labeled image, the at least one third labeled image is the object detection result, and the at least one second set target set is the object detection result.
3. The method according to claim 1 or 2, characterized in that, Before obtaining the multiple images to be detected, the method further includes: Obtaining a training data set, where the training data set includes multiple pairs of training data; Training an untrained generative adversarial network model using the training data set to obtain a trained generative adversarial network model, and the trained generative adversarial network model is the generative adversarial network model.
4. The method according to claim 2, wherein After inputting the multiple images to be detected into the fusion neural network model for processing and outputting at least one object detection result including one or more set targets, the method further includes: Performing identity recognition on the second set target set corresponding to the third labeled image A to obtain a first identity set, where the first identity set has the same number as and corresponds one-to-one with the second set target set corresponding to the third labeled image A, and the third labeled image A is any one of the at least one third labeled image; The performing a display operation on the at least one object detection result includes: Performing a display operation on the at least one third labeled image, the at least one second set target set, and the at least one first identity set.
5. A target detection device, characterized in that, The device includes: An obtaining unit, configured to obtain multiple images to be detected; A processing unit, configured to input the multiple images to be detected into a fusion neural network model for processing and output at least one object detection result including one or more set targets, where the at least one object detection result corresponds one-to-one with at least one image to be detected; wherein, the fusion neural network model includes a convolutional neural network model and a generative adversarial network model; A display unit, configured to perform a display operation on the at least one object detection result; In terms of inputting the multiple images to be detected into the fusion neural network model for processing and outputting at least one object detection result including one or more set targets, the processing unit is specifically configured to: Input the multiple images to be detected into the convolutional neural network model for initial processing and output multiple first labeled images, where the multiple first labeled images correspond one-to-one with the multiple images to be detected; Input the multiple first labeled images into the generative adversarial network model for reprocessing and output at least one object detection result including one or more set targets; When inputting multiple first annotated images into a generative adversarial network model for further processing and outputting at least one object detection result that includes one or more set targets, the processing unit is specifically configured to: Input multiple first annotated images into the generative adversarial network model; Annotate the multiple first annotated images to obtain multiple second annotated images, where the area of each second annotated image in the multiple second annotated images that includes one or more set targets is smaller than the area of its corresponding first annotated image, and the multiple second annotated images correspond to the multiple first annotated images one by one; Use one or more set target recognition algorithms to perform object recognition on the multiple second annotated images to obtain multiple first set target quantities and multiple first set target sets, where the multiple first set target quantities and the multiple first set target sets correspond to the multiple second annotated images one by one; Select at least one target second annotated image from the multiple second annotated images, where the first set target quantity corresponding to each target second annotated image in the at least one target second annotated image is non-zero, the at least one target second annotated image is the object detection result, and the at least one first set target set is the object detection result, so as to obtain at least one object detection result.
6. A target detection device, characterized in that, The object detection device includes a processor, a memory, a communication interface, and an object detection program stored in the memory, and the object detection program is configured to be executed by the processor to implement the steps in the method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store the object detection program, and the object detection program is executed by the processor to implement the steps in the method according to any one of claims 1-4.
Citation Information
Patent Citations
Weed recognition method and device and terminal equipment
CN110135341A