Target detection method, device, equipment and readable storage medium
By clustering and fusing multiple initial detection results, the problem of limited object detection accuracy of a single model is solved, and a higher object detection accuracy is achieved.
Patent Information
- Application Number
- CN202310259992.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-17
AI Technical Summary
In the prior art, the object detection accuracy of a single model is limited, and it is difficult to effectively improve the detection accuracy of the object in the image.
By acquiring multiple initial detection results of the same image, determining the reference results, and using each object in the reference results as the clustering center, clustering multiple initial detection results to obtain the clustering results corresponding to each object, and then fusing each initial detection result based on the weight value to obtain new position information and new category information.
This method can comprehensively consider multiple detection results, reduce the probability that individual targets are not detected, and improve the detection accuracy of targets in the image.
Smart Images

Figure CN116258883B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a target detection method, device, equipment and readable storage medium. Background Art
[0002] At present, models such as convolutional neural networks can be used to detect various targets in images. In order to improve the detection accuracy, the stochastic gradient descent method is generally used to make the network model converge to the optimal solution. However, the detection ability of a single model depends on the structure of the model itself and the training method, resulting in limited detection accuracy of targets in the image.
[0003] Therefore, how to improve the detection accuracy of objects in images is a problem that needs to be solved by those skilled in the art. Summary of the invention
[0004] In view of this, the purpose of this application is to provide a target detection method, device, equipment and readable storage medium to improve the detection accuracy of targets in images. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a target detection method, comprising:
[0006] Acquire multiple initial detection results of the same image; each initial detection result includes: location information and category information of at least one target in the image;
[0007] The initial detection result with the most detected targets among all the initial detection results is determined as the benchmark result;
[0008] Taking each target in the benchmark result as a cluster center, taking the location information and category information of each target included in each initial detection result as a clustering basis, clustering the multiple initial detection results to obtain a clustering result corresponding to each target in the benchmark result;
[0009] The initial detection results included in each clustering result are fused according to the corresponding weight values to obtain new location information and new category information of the target corresponding to each clustering result.
[0010] Optionally, obtaining multiple initial detection results of the same image includes:
[0011] Performing image transformation on the image to obtain a plurality of transformation images;
[0012] Using the same model to detect multiple transformation images, to obtain the multiple initial detection results;
[0013] or
[0014] Using different models to detect the image to obtain the multiple initial detection results;
[0015] or
[0016] Performing image transformation on the image to obtain a plurality of transformation images;
[0017] Different transformation graphs are detected using different models to obtain the multiple initial detection results.
[0018] Optionally, taking each target in the benchmark result as a cluster center, taking the location information and category information of each target included in each initial detection result as a clustering basis, clustering the multiple initial detection results to obtain a clustering result corresponding to each target in the benchmark result, including:
[0019] Taking the position information of each target in the benchmark result as the benchmark position information, and selecting the position information closest to the current benchmark position information from each initial detection result except the benchmark result, to obtain a first set;
[0020] Determine the category information L with the largest number of repetitions among the category information corresponding to each position information included in the first set, and form each position information corresponding to the category information L into a second set;
[0021] Selecting the position information corresponding to the category information with the highest credibility in the second set, and calculating the intersection and union ratio of the selected position information and other position information in the second set;
[0022] The position information whose intersection-and-union ratio is not 0 is formed into a third set, and the position information whose intersection-and-union ratio is 0 is formed into a fourth set;
[0023] The position information whose intersection-union ratio is not 0 in the fourth set is used to form a fifth set;
[0024] A union of the third set and the fifth set is determined, and initial detection results to which each piece of position information in the union belongs are combined into a clustering result of the current target.
[0025] Optionally, the initial detection results included in each clustering result are fused according to corresponding weight values to obtain new location information and new category information of the target corresponding to each clustering result, including:
[0026] For each clustering result, the position information in each initial detection result included in the current clustering result is used as a fusion object, and a weight value of each fusion object is determined;
[0027] The fusion objects are fused according to the weight value of each fusion object to obtain the position fusion result corresponding to the current clustering result;
[0028] The position fusion result and the category information in each initial detection result included in the current clustering result are used as the new position information and new category information of the target corresponding to the current clustering result.
[0029] Optionally, determining a weight value of each fusion object includes:
[0030] Calculate the initial weight coefficient of each fusion object using a preset algorithm;
[0031] Determine the evaluation weight coefficient of the model corresponding to each fusion object;
[0032] The weight value of the corresponding fusion object is determined based on the initial weight coefficient of each fusion object and the evaluation weight coefficient.
[0033] Optionally, determining the evaluation weight coefficient of the model corresponding to each fusion object includes:
[0034] The same test data set is processed using the model corresponding to each fusion object to obtain multiple test results;
[0035] The evaluation weight coefficient of the corresponding model is calculated based on the evaluation score of each test result.
[0036] Optionally, it also includes:
[0037] If the multiple initial detection results are obtained by detecting multiple transformation images of the image by the same model, the new position information and new category information of each target in the reference result are determined as the credible detection result of the model for the image;
[0038] Determining credible detection results of different models for the image;
[0039] All credible detection results are determined as the multiple initial detection results, and the step of determining the initial detection result with the most detected targets among all the initial detection results as the reference result is performed.
[0040] In a second aspect, the present application provides a target detection device, comprising:
[0041] An acquisition module, used to acquire multiple initial detection results of the same image; each initial detection result includes: location information and category information of at least one target in the image;
[0042] A selection module is used to determine an initial detection result with the most detected targets among all initial detection results as a benchmark result;
[0043] A clustering module, used to take each target in the benchmark result as a cluster center, take the location information and category information of each target included in each initial detection result as a clustering basis, cluster the multiple initial detection results, and obtain a clustering result corresponding to each target in the benchmark result;
[0044] The fusion module is used to fuse the initial detection results included in each clustering result according to the corresponding weight value to obtain the new position information and new category information of the target corresponding to each clustering result.
[0045] In a third aspect, the present application provides an electronic device, including:
[0046] Memory for storing computer programs;
[0047] A processor is used to execute the computer program to implement the target detection method disclosed above.
[0048] In a fourth aspect, the present application provides a readable storage medium for storing a computer program, wherein the computer program implements the aforementioned disclosed target detection method when executed by a processor.
[0049] It can be seen from the above scheme that the present application provides a target detection method, including: obtaining multiple initial detection results of the same image; each initial detection result includes: the location information and category information of at least one target in the image; determining the initial detection result with the most targets detected in all the initial detection results as the benchmark result; taking each target in the benchmark result as the clustering center, and taking the location information and category information of each target included in each initial detection result as the clustering basis, clustering the multiple initial detection results to obtain the clustering result corresponding to each target in the benchmark result; fusing the initial detection results included in each clustering result according to the corresponding weight value to obtain the new location information and new category information of the target corresponding to each clustering result.
[0050] It can be seen that when the present application fuses multiple initial detection results of the same image, the initial detection result with the most detected targets among all the initial detection results is determined as the benchmark result, and then each target in the benchmark result is used as the clustering center, and the location information and category information of each target included in each initial detection result is used as the clustering basis to cluster the multiple initial detection results to obtain the clustering results corresponding to each target in the benchmark result. That is: taking each target in the image as the clustering center, clustering the multiple initial detection results of the image, and then determining the comprehensive result of each target in the image based on each clustering result, thus obtaining the new location information and new category information of each target. Among them, the multiple initial detection results of the same image can be: the detection results of different models for the image, or: the detection results of different transformation maps of the same model for the image, or: the detection results of different transformation maps of the image by different models. That is to say, no matter how the multiple initial detection results of the same image are obtained, the scheme provided by the present application can be used to determine the benchmark results among all the initial detection results, and each target in the benchmark results can be used as the cluster center to cluster all the initial detection results, and then the new position information and new category information of each target in the benchmark results can be determined based on the clustering results, thereby completing the fusion of multiple initial detection results. This scheme can comprehensively consider multiple detection results to obtain the optimal result, reduce the probability of individual targets in the image not being detected, and the position and category of each target output in the end are more accurate, so the detection accuracy of the target in the image can be improved.
[0051] Correspondingly, a target detection device, equipment and readable storage medium provided by the present application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0053] Figure 1 A flow chart of a target detection method disclosed in this application;
[0054] Figure 2 This is a flow chart of another target detection method disclosed in this application;
[0055] Figure 3 A schematic diagram of comprehensive calculation of model weights and algorithm output weights disclosed in this application;
[0056] Figure 4A schematic diagram of a single model fusion disclosed in this application;
[0057] Figure 5 A schematic diagram of a target detection device disclosed in this application;
[0058] Figure 6 This is a schematic diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0060] At present, the detection capability of a single model depends on the structure of the model itself and the training method, resulting in limited detection accuracy of targets in images. To this end, the present application provides a target detection solution that can fuse multiple initial detection results of the same image, comprehensively consider multiple detection results to obtain the optimal result, reduce the probability of individual targets in the image not being detected, and finally output the position and category of each target more accurately, thereby improving the detection accuracy of targets in the image.
[0061] See also Figure 1 As shown, the embodiment of the present application discloses a target detection method, including:
[0062] S101, obtaining multiple initial detection results of the same image; each initial detection result includes: location information and category information of at least one target in the image.
[0063] In this embodiment, multiple initial detection results of the same image can be: detection results of different models for the image, or: detection results of different transformation maps of the same model for the image, or: detection results of different transformation maps of the image by different models. Therefore, in a specific implementation, obtaining multiple initial detection results of the same image includes: performing image transformation on the image to obtain multiple transformation maps; using the same model to detect multiple transformation maps to obtain multiple initial detection results; or using different models to detect the image to obtain multiple initial detection results; or performing image transformation on the image to obtain multiple transformation maps; using different models to detect different transformation maps respectively to obtain multiple initial detection results.
[0064] S102: Determine the initial detection result with the most detected targets among all the initial detection results as a benchmark result.
[0065] S103, taking each target in the benchmark result as a cluster center, taking the location information and category information of each target included in each initial detection result as a clustering basis, clustering multiple initial detection results to obtain a clustering result corresponding to each target in the benchmark result.
[0066] In a specific implementation, each target in the benchmark result is used as a clustering center, and the location information and category information of each target included in each initial detection result is used as a clustering basis, and multiple initial detection results are clustered to obtain a clustering result corresponding to each target in the benchmark result, including: taking the location information of each target in the benchmark result as the benchmark location information, selecting the location information closest to the current benchmark location information from each initial detection result except the benchmark result to obtain a first set; determining the category information L with the largest number of repetitions among the category information corresponding to each location information included in the first set, and forming each location information corresponding to the category information L into a second set; selecting the location information corresponding to the category information with the highest credibility in the second set, and calculating the intersection-and-union ratio of the selected location information and other location information in the second set; forming a third set with location information whose intersection-and-union ratio is not 0, and forming a fourth set with location information whose intersection-and-union ratio is 0; forming a fifth set with location information whose intersection-and-union ratio is not 0 in the fourth set; determining the union of the third set and the fifth set, and combining the initial detection results to which each location information in the union belongs into the clustering result of the current target.
[0067] In a specific implementation, each initial detection result included in each clustering result is fused according to a corresponding weight value to obtain new position information and new category information of a target corresponding to each clustering result, including: for each clustering result, taking the position information in each initial detection result included in the current clustering result as a fusion object, and determining the weight value of each fusion object; fusing each fusion object according to the weight value of each fusion object to obtain a position fusion result corresponding to the current clustering result; taking the position fusion result and the category information in each initial detection result included in the current clustering result as the new position information and new category information of the target corresponding to the current clustering result.
[0068] In a specific implementation, determining the weight value of each fusion object includes: calculating the initial weight coefficient of each fusion object using a preset algorithm; determining the evaluation weight coefficient of the model corresponding to each fusion object; and determining the weight value of the corresponding fusion object based on the initial weight coefficient and the evaluation weight coefficient of each fusion object. The preset algorithm is, for example, AABBFI (Axis-Aligned Bounding Box Fuzzy Integral).
[0069] In a specific implementation, determining the evaluation weight coefficient of the model corresponding to each fusion object includes: using the model corresponding to each fusion object to process the same test data set to obtain multiple test results; and calculating the evaluation weight coefficient of the corresponding model based on the evaluation score of each test result.
[0070] S104: The initial detection results included in each clustering result are fused according to corresponding weight values to obtain new position information and new category information of the target corresponding to each clustering result.
[0071] In a specific implementation, if multiple initial detection results are obtained by detecting multiple transformation graphs of an image by the same model, the new position information and new category information of each target in the benchmark result are determined as the credible detection result of the model for the image; the credible detection results of different models for the image are determined; all credible detection results are determined as multiple initial detection results, and the initial detection result with the most targets detected in all initial detection results is determined as the benchmark result; each target in the benchmark result is used as a clustering center, and the position information and category information of each target included in each initial detection result is used as a clustering basis to cluster the multiple initial detection results to obtain the clustering result corresponding to each target in the benchmark result; each initial detection result included in each clustering result is fused according to the corresponding weight value to obtain the new position information and new category information of the target corresponding to each clustering result. That is: each model in the multiple models is made to detect multiple transformation graphs of the same image to obtain a detection result, and then the multiple detection results output by each model are fused according to S101-S104 provided in this embodiment, so as to obtain the fusion result corresponding to each model. Afterwards, the fusion results corresponding to each model are used as the initial detection results in step S101, and S101-S104 are executed again to fuse the fusion results corresponding to all models. This can combine the detection capabilities of different models for the same image and the detection capabilities of the same model for different transformation maps of the same image, so that the final output target detection result can integrate more factors, thereby further improving the detection accuracy.
[0072] It can be seen that, for multiple initial detection results of the same image, this embodiment determines the benchmark result among all the initial detection results, takes each target in the benchmark result as the cluster center, and clusters all the initial detection results, and then determines the new position information and new category information of each target in the benchmark result based on the clustering result, thereby completing the fusion of multiple initial detection results. This solution can comprehensively consider multiple detection results to obtain the optimal result, reduce the probability of individual targets in the image not being detected, and the position and category of each target finally output are more accurate, so the detection accuracy of the target in the image can be improved.
[0073] The multi-model fusion solution provided by this application is further introduced below.
[0074] See also Figure 2 The implementation steps of the multi-model fusion solution include: selecting multiple target detection models of different types based on the criterion of the maximum differentiation model structure; selecting the corresponding data expansion algorithm for any selected target detection model; training each target detection model separately according to the expanded training set; after the training is completed, using the same validation set to calculate the weights corresponding to different models; and finally performing multi-model fusion.
[0075] In order to improve the diversity of the fused results, multiple models with large differences are selected to detect the same image. Specifically, models can be selected from anchor-base models, anchor-free models, and collection models. Both anchor-base models and anchor-free models classify targets based on boxes, while collection models recognize targets based on image blocks without preset boxes. Models based on anchor-base include YOLOv1-v7, RetinaNet, Faster R-CNN, etc., models based on anchor-free include CenterNet, FCOS, etc.; collection models include DETR, Deformable-DETR, etc. To ensure model diversity, models can be selected from anchor-base models, anchor-free models, and collection models. The selection criteria are: using the same data set and the same evaluation criteria to select the best evaluated model.
[0076] In practical applications, although it is impossible to obtain massive data, data expansion can be performed on limited data sets to enhance sample diversity and improve the generalization performance of the algorithm. Generally, data expansion algorithms are divided into two categories: those based on geometric transformation and those based on color transformation. Those based on geometric transformation include random cropping, random expansion, random horizontal flip, random resize, etc., and those based on color transformation include color jitter and Fancy PCA. Based on this, the appropriate data expansion algorithm can be selected for the above selected model to expand the training samples and improve the robustness of the model.
[0077] Based on the training set expanded by the above method, the corresponding models are trained respectively. After each model is trained, the weight of each model is calculated using the same validation set. Figure 3, A1, A2, A3 represent the weights of the three models, where the calculation process of A1, A2, A3 is: let each model process the same validation set, and use the same evaluation criteria to determine the score of each model, and determine the weight of each model based on the corresponding score. If there are N models, the corresponding N validation result scores are: m1, m2, ..., mN, then the weight Wi of any model i = mi / (m1+m2+...+mN), mi is the score of the validation result of model i.
[0078] When training a model, you can use methods such as stochastic gradient descent to make the model converge to a local optimal solution, and then adjust the learning rate to make the model converge to a local optimal solution again. In this way, you can use the same training set and the same initial model to train models of the same structure with different detection capabilities.
[0079] Among them, the learning rate η is generally a number related to the number of iterations t, that is: Among them, t is the current iteration number, T is the total training iteration number, η 0 is the initial learning rate, and M is the number of times the learning rate is adjusted. This formula uses the properties of the cosine function to cyclically update the learning rate in a cycle. As t increases, the learning rate η is increased from the initial learning rate η 0 Gradually decreases to 0, and it is considered that an optimal solution is obtained. Then adjust the learning rate to amplify the learning rate again so as to jump out of the previous local optimal solution and start the next cycle of training. After this cycle ends, the network model converges to a new local optimal solution. After M such cycles, multiple different local optimal solutions will be obtained, and models of the same structure with different detection capabilities will be obtained. The different models described in this application refer to: models with different detection capabilities or different structures.
[0080] See also Figure 4 , the score of each model is determined according to the following process. For any model, an image is selected in the validation set, and the image is transformed to obtain multiple transformation maps. The model is used to process the multiple transformation maps to obtain multiple detection results, and the multiple detection results are fused to obtain the fusion result of the model for a single image. After that, the current model is used to process all images in the validation set, and then the fusion results of the model for all images in the validation set are combined to obtain the score of the model.
[0081] When detecting each target in a certain image, the current image is transformed to obtain multiple transformation maps, and then a model is selected from the above-mentioned trained models, and these transformation maps are respectively input into the currently selected model to obtain multiple results output by the current model, and then single model fusion is performed according to S101-S104 provided in the above-mentioned embodiment. After obtaining the single model fusion result of each model, the fusion result corresponding to each model is used as the initial detection result in step S101 of the above-mentioned embodiment, and S101-S104 are executed again to fuse the fusion results corresponding to all models, and finally obtain a multi-model fusion result.
[0082] Specifically, the fusion step includes: selecting the one with the most detected targets from the N detection results currently being fused, denoted as R max , and R max Each target in is used as the K cluster centers of k-means or other clustering algorithms. max For each target in, take the position Box and category label of the target X currently traversed, and perform the following operations: Use k-means or other clustering algorithms to select a target closest to the Box in other detection results. N-1 targets can be selected, and the selected N-1 targets and the position Box and category label of target X are recorded as A. Then use the voting method to select the one with the most occurrences of the category label in A, recorded as L. Obtain all detection boxes and corresponding labels of category L in A, and record them as B, B∈A. The next step is to sort the labels in B from large to small scores, and select the detection box X corresponding to the label with the highest score. Calculate the intersection and union ratio of detection box X and other detection boxes in B, obtain the detection box corresponding to the intersection and union ratio when it is not 0, recorded as C; obtain the detection box corresponding to the intersection and union ratio when it is 0, recorded as D, calculate the intersection and union ratio of different detection boxes in D, and take the detection box corresponding to the intersection and union ratio when it is not 0, recorded as E, D∈A, Finally, for F=C∪E, a weighted calculation of the detection box is performed.
[0083] See also Figure 3 , perform weighted calculation of the detection frame for F = C∪E, including: using the AABBF algorithm to calculate the horizontal and vertical coordinate weights of each detection frame in F, and determining the model weight corresponding to each detection frame in F, calculating the final weight based on the above two weights, and calculating the new detection frame based on the final weight, and at the same time determining the category L and the average of the label scores of all categories L in F, and obtaining the fusion result of the target X: the new detection frame + category L + the average of the label scores of category L. Based on this, R max The fusion result of each target in is the fusion result of N detection results. If the above N detection results are output by the same model, the model weight corresponding to each detection box in F is equal.
[0084] It can be seen that this embodiment comprehensively considers the model weight and the weight output by the AABBFI algorithm when performing weighted calculation on the detection frame, and this embodiment selects the detection frame benchmark during clustering, thereby performing targeted processing on the outliers, so that the quality of the final output detection frame is higher and the final accuracy can be improved. At the same time, this embodiment performs multiple transformations on the input image, which can make full use of the differences between images that are similar but not nearly identical to improve accuracy. The fusion method of the AABBFI algorithm + model weight is successively adopted to make full use of the differences after image enhancement and the diversity between different models to achieve the purpose of improving target detection accuracy.
[0085] An object detection device provided in an embodiment of the present application is introduced below. The object detection device described below and the object detection method described above can be referenced to each other.
[0086] See also Figure 5 As shown, the embodiment of the present application discloses a target detection device, including:
[0087] The acquisition module 501 is used to acquire multiple initial detection results of the same image; each initial detection result includes: location information and category information of at least one target in the image;
[0088] A selection module 502 is used to determine the initial detection result with the most detected targets among all the initial detection results as a reference result;
[0089] The clustering module 503 is used to take each target in the benchmark result as a cluster center, take the location information and category information of each target included in each initial detection result as a clustering basis, cluster the multiple initial detection results, and obtain a clustering result corresponding to each target in the benchmark result;
[0090] The fusion module 504 is used to fuse the initial detection results included in each clustering result according to the corresponding weight value to obtain the new position information and new category information of the target corresponding to each clustering result.
[0091] In a specific implementation, the acquisition module is specifically used to:
[0092] Performing image transformation on the image to obtain multiple transformation images;
[0093] Using the same model to detect multiple transformation images, multiple initial detection results are obtained;
[0094] or
[0095] Use different models to detect images and obtain multiple initial detection results;
[0096] or
[0097] Performing image transformation on the image to obtain multiple transformation images;
[0098] Different models are used to detect different transformation images and obtain multiple initial detection results.
[0099] In a specific implementation, the clustering module is specifically used to:
[0100] Taking the position information of each target in the benchmark result as the benchmark position information, selecting the position information closest to the current benchmark position information from each initial detection result except the benchmark result, to obtain a first set;
[0101] Determine the category information L with the largest number of repetitions among the category information corresponding to each position information included in the first set, and form each position information corresponding to the category information L into a second set;
[0102] Selecting the position information corresponding to the category information with the highest credibility in the second set, and calculating the intersection and union ratio of the selected position information and other position information in the second set;
[0103] The position information whose intersection-and-union ratio is not 0 is formed into a third set, and the position information whose intersection-and-union ratio is 0 is formed into a fourth set;
[0104] The position information whose intersection and union ratio is not 0 in the fourth set is used to form a fifth set;
[0105] A union of the third set and the fifth set is determined, and the initial detection results to which each piece of position information in the union belongs are combined into a clustering result of the current target.
[0106] In a specific implementation, the clustering module is specifically used to:
[0107] For each clustering result, the position information in each initial detection result included in the current clustering result is used as a fusion object, and a weight value of each fusion object is determined;
[0108] The fusion objects are fused according to the weight value of each fusion object to obtain the position fusion result corresponding to the current clustering result;
[0109] The position fusion result and the category information in each initial detection result included in the current clustering result are used as the new position information and new category information of the target corresponding to the current clustering result.
[0110] In a specific implementation, the clustering module is specifically used to:
[0111] Calculate the initial weight coefficient of each fusion object using a preset algorithm;
[0112] Determine the evaluation weight coefficient of the model corresponding to each fusion object;
[0113] The weight value of the corresponding fusion object is determined based on the initial weight coefficient of each fusion object and the evaluation weight coefficient.
[0114] In a specific implementation, the clustering module is specifically used to:
[0115] The same test data set is processed using the model corresponding to each fusion object to obtain multiple test results;
[0116] The evaluation weight coefficient of the corresponding model is calculated based on the evaluation score of each test result.
[0117] In a specific embodiment, it also includes:
[0118] The re-fusion module is used to determine the new position information and new category information of each target in the benchmark result as the credible detection result of the model for the image if multiple initial detection results are obtained by multiple transformation maps of the same model detection image; determine the credible detection results of different models for the image; determine all credible detection results as multiple initial detection results, and execute the steps in the acquisition module, selection module, clustering module, and fusion module.
[0119] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0120] It can be seen that this embodiment provides a target detection device that can fuse multiple initial detection results of the same image, and can comprehensively consider multiple detection results to obtain the optimal result, thereby reducing the probability of individual targets in the image not being detected. The position and category of each target finally output are also more accurate, thereby improving the detection accuracy of targets in the image.
[0121] An electronic device provided in an embodiment of the present application is introduced below. The electronic device described below and the target detection method and device described above can be referenced to each other.
[0122] See also Figure 6 As shown, the embodiment of the present application discloses an electronic device, including:
[0123] Memory 601, used for storing computer programs;
[0124] The processor 602 is used to execute the computer program to implement the method disclosed in any of the above embodiments.
[0125] Furthermore, an embodiment of the present application also provides a server as the above-mentioned electronic device. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the target detection method disclosed in any of the above-mentioned embodiments.
[0126] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0127] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.
[0128] The operating system is used to manage and control the hardware devices and computer programs on the server to realize the operation and processing of the data in the memory by the processor, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the target detection method disclosed in any of the above embodiments, the computer program can further include computer programs that can be used to complete other specific tasks. In addition to data such as virtual machines, data can also include data such as the developer information of the virtual machine.
[0129] Furthermore, the embodiment of the present application also provides a terminal as the above electronic device. The terminal may specifically include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer.
[0130] Generally, the terminal in this embodiment includes: a processor and a memory.
[0131] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0132] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is at least used to store the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the target detection method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of the application.
[0133] In some embodiments, the terminal may also include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0134] A readable storage medium provided in an embodiment of the present application is introduced below. The readable storage medium described below and the target detection method, device and apparatus described above can be referenced to each other.
[0135] A readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the target detection method disclosed in the above-mentioned embodiment. The readable storage medium is a computer-readable storage medium, which, as a carrier for storing resources, may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system, a computer program and data, etc., and the storage method may be temporary storage or permanent storage.
[0136] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0137] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of readable storage medium known in the art.
[0138] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A target detection method, It is characterized in that include: Get multiple initial detection results for the same image; Each initial detection result includes: location information and category information of at least one target in the image; The initial detection result with the most detected targets among all the initial detection results is determined as the benchmark result; Taking each target in the benchmark result as a cluster center, taking the location information and category information of each target included in each initial detection result as a clustering basis, clustering the multiple initial detection results to obtain a clustering result corresponding to each target in the benchmark result; The initial detection results included in each clustering result are fused according to the corresponding weight values to obtain the new position information and new category information of the target corresponding to each clustering result; The method of taking each target in the benchmark result as a cluster center, taking the location information and category information of each target included in each initial detection result as a clustering basis, clustering the multiple initial detection results, and obtaining a clustering result corresponding to each target in the benchmark result includes: Taking the position information of each target in the benchmark result as the benchmark position information, and selecting the position information closest to the current benchmark position information from each initial detection result except the benchmark result, to obtain a first set; Determine the category information L with the largest number of repetitions among the category information corresponding to each position information included in the first set, and form each position information corresponding to the category information L into a second set; Selecting the position information corresponding to the category information with the highest credibility in the second set, and calculating the intersection and union ratio of the selected position information and other position information in the second set; The position information whose intersection-and-union ratio is not 0 is formed into a third set, and the position information whose intersection-and-union ratio is 0 is formed into a fourth set; The position information whose intersection-union ratio is not 0 in the fourth set is used to form a fifth set; A union of the third set and the fifth set is determined, and initial detection results to which each piece of position information in the union belongs are combined into a clustering result of the current target.
2. The method according to claim 1, It is characterized in that The obtaining of multiple initial detection results of the same image includes: Performing image transformation on the image to obtain a plurality of transformation images; Using the same model to detect multiple transformation images, to obtain the multiple initial detection results; or Using different models to detect the image to obtain the multiple initial detection results; or Performing image transformation on the image to obtain a plurality of transformation images; Different transformation images are detected using different models to obtain the multiple initial detection results.
3. The method according to claim 1, It is characterized in that The initial detection results included in each clustering result are fused according to the corresponding weight values to obtain new location information and new category information of the target corresponding to each clustering result, including: For each clustering result, the position information in each initial detection result included in the current clustering result is used as a fusion object, and a weight value of each fusion object is determined; The fusion objects are fused according to the weight value of each fusion object to obtain the position fusion result corresponding to the current clustering result; The position fusion result and the category information in each initial detection result included in the current clustering result are used as the new position information and new category information of the target corresponding to the current clustering result.
4. The method according to claim 1, It is characterized in that The step of determining the weight value of each fusion object comprises: Calculate the initial weight coefficient of each fusion object using a preset algorithm; Determine the evaluation weight coefficient of the model corresponding to each fusion object; The weight value of the corresponding fusion object is determined based on the initial weight coefficient of each fusion object and the evaluation weight coefficient.
5. The method according to claim 4, It is characterized in that The step of determining the evaluation weight coefficient of the model corresponding to each fusion object includes: The same test data set is processed using the model corresponding to each fusion object to obtain multiple test results; The evaluation weight coefficient of the corresponding model is calculated based on the evaluation score of each test result.
6. The method according to any one of claims 1 to 4, It is characterized in that Also includes: If the multiple initial detection results are obtained by detecting multiple transformation images of the image by the same model, the new position information and new category information of each target in the reference result are determined as the credible detection result of the model for the image; Determining credible detection results of different models for the image; All credible detection results are determined as the multiple initial detection results, and the step of determining the initial detection result with the most detected targets among all the initial detection results as the reference result is performed.
7. A target detection device, It is characterized in that include: An acquisition module, used for acquiring multiple initial detection results of the same image; Each initial detection result includes: location information and category information of at least one target in the image; A selection module is used to determine an initial detection result with the most detected targets among all initial detection results as a benchmark result; A clustering module, used to take each target in the benchmark result as a cluster center, take the location information and category information of each target included in each initial detection result as a clustering basis, cluster the multiple initial detection results, and obtain a clustering result corresponding to each target in the benchmark result; A fusion module is used to fuse the initial detection results included in each clustering result according to the corresponding weight value to obtain the new position information and new category information of the target corresponding to each clustering result; Wherein, the clustering module is specifically used for: Taking the position information of each target in the benchmark result as the benchmark position information, and selecting the position information closest to the current benchmark position information from each initial detection result except the benchmark result, to obtain a first set; Determine the category information L with the largest number of repetitions among the category information corresponding to each position information included in the first set, and form each position information corresponding to the category information L into a second set; Selecting the position information corresponding to the category information with the highest credibility in the second set, and calculating the intersection and union ratio of the selected position information and other position information in the second set; The position information whose intersection-and-union ratio is not 0 is formed into a third set, and the position information whose intersection-and-union ratio is 0 is formed into a fourth set; The position information whose intersection-union ratio is not 0 in the fourth set is used to form a fifth set; A union of the third set and the fifth set is determined, and initial detection results to which each piece of position information in the union belongs are combined into a clustering result of the current target.
8. An electronic device, It is characterized in that include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 6.
9. A readable storage medium, It is characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Target detection method and device based on multi-model fusion, equipment and medium
CN113688957A
Target detection method and device, equipment and storage medium
CN115376054A