Target object prediction method, device, electronic device and storage medium

By combining the two-stage detection model and cross-border threshold judgment, the recall and stability problems of the target object prediction model in the prior art in the background noise processing are solved, and higher accuracy and recall are achieved.

CN114419390BActive Publication Date: 2025-05-13BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111604280.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-05-13
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

When the existing target object prediction model deals with negative background noise samples, it is difficult to maintain a high recall rate, and the parameter settings are unstable, resulting in a decrease in recognition ability and insufficient generalization ability.

Method used

Using a combination of two-stage detection model, first obtains the coarse position and category information of the target object through the first detection model, and then crops the original image to generate the target object area image, and inputs it to the second detection model to obtain the fine position and category information. Determine whether the target object is a target category object by crossing and comparing the threshold value, and calculate the confidence to select the predicted object.

Benefits of technology

It improves the accuracy and recall of target object prediction, reduces interference from background noise, and enhances the stability and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419390B_ABST
    Figure CN114419390B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and device for predicting a target object, wherein the method comprises: inputting an original image into a first detection model, outputting first position information and rough category information of the target object; cropping the original image to obtain multiple target object area images, inputting the multiple target object area images into a second detection model, outputting multiple second position information and multiple fine category information of each target object; selecting a target prediction object if a target category object exists; generating a position prediction result according to the first position information and the second position information, and using the fine category information as the category prediction result. The embodiment of the present invention adds context information of the target object, thereby improving the accuracy of the fine category information output from the second detection model. The target prediction object is selected for the target category object, avoiding interference from non-target prediction objects, and further improving the accuracy of target object prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a method, device, electronic device and computer-readable storage medium for predicting a target object. Background Art

[0002] As a basic technology in computer vision, target object prediction is generally divided into two subtasks. The first subtask is to determine the position of the target object by judging the foreground and background, and the second subtask is to classify the target object after the position is determined. Under the premise of meeting these two subtasks, the prediction model needs to have strong enough ability to judge the appearance difference and scale difference of the target object after the position is determined. The quality of a prediction model is usually measured by recall rate and accuracy rate. Among them, the recall rate reflects the ratio of the number of target objects of a certain category predicted by the prediction model that truly belong to the category to the number of ground truth data of the category, and the accuracy rate reflects the ratio of the number of target objects predicted by the prediction model that truly belong to the category to the number of target objects of all categories. An indicator similar to the accuracy rate is the false alarm rate, which reflects the ratio of the number of target objects predicted by the prediction model that do not belong to the category (falsely detected target objects) to the number of target objects of all categories.

[0003] The problem of removing negative samples such as background noise in the prediction process of the target object by the prediction model is usually solved by balancing the ratio of positive and negative samples. For example, online difficult sample mining designed in the network learning process and methods designed at the loss function level. In addition, the negative samples predicted by the Region Proposal Network (RPN) layer can be collected for offline learning of subsequent prediction models. These methods have obvious advantages in reducing false alarm rates and improving accuracy, but for scenarios where the prediction model needs to ensure high recall, the online difficult sample mining method only retains samples with higher loss functions and completely ignores simple samples. This essentially changes the input distribution during training (only includes difficult samples), which will cause the prediction model to lose the ability to distinguish easy-to-classify samples during learning and cannot fully guarantee the full recall of simple samples. For methods designed at the loss function level, the setting of parameters will play a decisive role in the learning of positive and negative samples, which is not conducive to stable result output. For the scheme of collecting negative samples for offline learning of prediction models, the judgment of inter-class error of prediction models will increase. Although most negative samples can be identified, because negative samples themselves have a high appearance similarity with positive samples, the inter-class error caused in the absence of contextual reference will reduce the recall rate. On the other hand, the mined negative samples are limited and can only meet the false detection target of the current training set, and the generalization ability is insufficient. Summary of the invention

[0004] In view of the above problems, embodiments of the present invention are proposed to provide a method, device, electronic device and computer-readable storage medium for predicting a target object that overcome the above problems or at least partially solve the above problems.

[0005] In order to solve the above problem, according to a first aspect of an embodiment of the present invention, a method for predicting a target object is disclosed, the method comprising: obtaining an original image to be processed, the original image comprising at least one target object; inputting the original image into a trained first detection model, and outputting first position information and rough category information of at least one target object; for each target object, cropping the original image according to the first position information to obtain a plurality of target object area images, inputting the plurality of target object area images into a trained second detection model, and outputting a plurality of second position information and a plurality of fine category information of each target object; for each target object, judging whether each target object belongs to a target category object of a preset category according to the rough category information, the first position information, a plurality of the second position information and a preset intersection-over-union threshold; if the target category object exists in at least one of the target objects, calculating the confidence of the target category object, and selecting a target prediction object from the target category objects according to the confidence; generating a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object, and taking the fine category information corresponding to the target prediction object as the category prediction result.

[0006] Optionally, judging whether each of the target objects belongs to a target category object of a preset category based on the coarse category information, the first position information, multiple pieces of the second position information and a preset intersection-and-union threshold includes: when the coarse category information belongs to the preset category, judging whether the target object is located in the central area of ​​multiple target object area images based on the first position information, multiple pieces of the second position information and a preset intersection-and-union threshold; and determining the target object located in the central area of ​​at least one of the target object area images as the target category object.

[0007] Optionally, judging whether the target object is located in the central area of ​​a plurality of target object area images based on the first position information, a plurality of the second position information and a preset intersection-and-union ratio threshold comprises: calculating a plurality of intersection-and-union ratio parameters of the target object based on the first position information and a plurality of the second position information; if there is at least one intersection-and-union ratio parameter greater than or equal to the preset intersection-and-union ratio threshold, confirming that the target object is located in the central area of ​​at least one of the target object area images.

[0008] Optionally, calculating the confidence of the target category object includes: counting the number of the target category objects located in the central area, and calculating the confidence according to the number and the number of target object area images containing the target category objects.

[0009] Optionally, selecting a target prediction object from the target category objects according to the confidence level includes: selecting the target prediction object whose confidence level satisfies a preset condition from the target category objects.

[0010] Optionally, generating a position prediction result of the target prediction object based on the first position information and the second position information of the target prediction object includes: calculating an average value of the first position information and the second position information, and using the average value as the position prediction result.

[0011] Optionally, the first detection model includes a Two-stage network model, and the second detection model includes a One-stage network model.

[0012] According to a second aspect of an embodiment of the present invention, a prediction device for a target object is also disclosed, the device comprising: an image acquisition module, used to acquire an original image to be processed, the original image containing at least one target object; a first detection module, used to input the original image into a trained first detection model, and output first position information and rough category information of at least one target object; a second detection module, used to crop the original image according to the first position information for each target object to obtain a plurality of target object area images, input the plurality of target object area images into the trained second detection model, and output a plurality of second position information and a plurality of fine category information for each target object; An object judgment module is used to judge whether each target object belongs to a target category object of a preset category according to the coarse category information, the first position information, multiple second position information and a preset intersection-and-union ratio threshold for each target object; an object selection module is used to calculate the confidence of the target category object if the target category object exists in at least one of the target objects, and select a target prediction object from the target category objects according to the confidence; a result determination module is used to generate a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object, and use the fine category information corresponding to the target prediction object as the category prediction result.

[0013] Optionally, the object judgment module includes: a center judgment module, which is used to judge whether the target object is located in the center area of ​​multiple target object area images based on the first position information, multiple second position information and a preset intersection-and-union ratio threshold when the rough category information belongs to the preset category; and an object determination module, which is used to determine the target object located in the center area of ​​at least one of the target object area images as the target category object.

[0014] Optionally, the center judgment module includes: a parameter calculation module, used to calculate multiple intersection-and-union parameters of the target object based on the first position information and multiple second position information; a center determination module, used to confirm that the target object is located in the central area of ​​at least one of the target object area images if there is at least one intersection-and-union parameter greater than or equal to the preset intersection-and-union threshold.

[0015] Optionally, the object selection module is used to count the number of the target category objects located in the central area, and calculate the confidence level according to the number and the number of target object area images containing the target category objects.

[0016] Optionally, the object selection module is used to select the target prediction object whose confidence meets a preset condition from the target category objects.

[0017] Optionally, the result determination module is used to calculate an average value of the first position information and the second position information, and use the average value as the position prediction result.

[0018] Optionally, the first detection model includes a Two-stage network model, and the second detection model includes a One-stage network model.

[0019] According to a third aspect of an embodiment of the present invention, an electronic device is also disclosed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the target object prediction method described in the first aspect when executing the computer program.

[0020] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is further disclosed, on which a computer program is stored. When the program is executed by a processor, the method for predicting a target object described in the first aspect is implemented.

[0021] Compared with the prior art, the technical solution provided by the embodiment of the present invention has the following advantages:

[0022] A prediction scheme for a target object provided by an embodiment of the present invention obtains an original image containing at least one target object, first inputs the original image into a trained first detection model, outputs the first position information and rough category information of at least one target object, and then for each target object, crops the original image according to the first position information to obtain multiple target object area images, and then inputs the multiple target object area images into a trained second detection model, and outputs multiple second position information and multiple fine category information of each target object. Next, for each target object, according to the rough category information, the first position information, the multiple second position information and the preset intersection-and-union ratio threshold, it is determined whether each target object belongs to a target category object of a preset category. If there is a target category object, the confidence of the target category object is calculated, and a target prediction object is selected from the target category object according to the confidence. Finally, a position prediction result of the target prediction object is generated according to the first position information and the second position information of the target prediction object, and the fine category information corresponding to the target prediction object is used as the category prediction result.

[0023] After the first position information of the target object is output from the first detection model, the embodiment of the present invention crops the original image according to the first position information to obtain a target object area image, and inputs the target object area image into the second detection model, so that the image input into the second detection model adds contextual information of the target object, thereby improving the accuracy of the fine category information output from the second detection model.

[0024] The embodiment of the present invention determines whether each target object is a target category object, and subsequently selects a target prediction object for the target category object, thereby avoiding interference from non-target prediction objects, such as background noise, and further improving the accuracy of target object prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of a method for predicting a target object according to an embodiment of the present invention;

[0026] Figure 2a , 2b 2c and 2d are schematic diagrams of training sample data obtained by clipping in an embodiment of the present invention;

[0027] Figure 3a , 3b 3c and 3d are schematic diagrams of training sample data obtained by replacing the background in an embodiment of the present invention;

[0028] Figure 4a is a schematic diagram of a warning sign in a traffic sign according to an embodiment of the present invention;

[0029] Figure 4bis a schematic diagram of a prohibition sign in a traffic sign according to an embodiment of the present invention;

[0030] Figure 4c is a schematic diagram of an indication sign in a traffic sign according to an embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of a target object prediction solution based on local coordinate alignment and global coordinate alignment according to an embodiment of the present invention;

[0032] Figure 6 is a structural block diagram of a target object prediction device according to an embodiment of the present invention;

[0033] Figure 7 It is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Reference Figure 1 , shows a flowchart of a method for predicting a target object according to an embodiment of the present invention. The method for predicting a target object may specifically include the following steps:

[0036] Step 101: Obtain an original image to be processed.

[0037] In the embodiment of the present invention, the original image may contain at least one target object. In practical applications, the target object may be a building, an animal, a human body, a vehicle, a ship, a traffic sign, etc. The embodiment of the present invention does not specifically limit the type, quantity, color, size, etc. of the target object.

[0038] Step 102: input the original image into a trained first detection model, and output first position information and rough category information of at least one target object.

[0039] In an embodiment of the present invention, the first detection model can output first position information and rough category information of each target object in the original image, and the first detection model has a high recall characteristic.

[0040] In practical applications, the training sample data of the first detection model may be conventional training sample data in the field of target object prediction. For example, the conventional training sample data is an image containing a traffic sign, and the conventional training sample data may also contain coordinate information and rough category information of the traffic sign in the image. The traffic sign may be understood as a target object in the conventional training sample data.

[0041] Step 103, for each target object, the original image is cropped according to the first position information to obtain multiple target object area images, the multiple target object area images are input into the trained second detection model, and multiple second position information and multiple fine category information of each target object are output.

[0042] In an embodiment of the present invention, the second detection model can output the second position information and fine category information of the target object in the target object area image, and the second detection model has a high accuracy characteristic. Moreover, before the target object area image is input into the second detection model, the target object area image needs to be obtained. Specifically, the original image can be cropped according to the first position information to obtain multiple target object area images. For example, the first position information represents a rectangular area, and the original image is cropped into multiple target object area images according to different proportions with the rectangular area as the center area. It should be noted that the target object can be located in the center area of ​​the target object area image.

[0043] In practical applications, the training sample data of the second detection model can be obtained by converting the training sample data of the first detection model. For example, the training sample data of the first detection model is an image containing a traffic sign. For each image in the training sample data of the first detection model, the area where the traffic sign is located is used as the center area, and each image in the training sample data of the first detection model is cropped according to different proportions to obtain a part of the training sample data of the second detection model. Figure 2a , 2b and 2c, Figure 2a , 2b 2c and 2d are schematic diagrams showing the training sample data obtained by cropping. Due to the limited number of traffic signs, if each image in the training sample data of the first detection model is cropped according to different proportions, the training sample data of the second detection model obtained will not be diverse. Therefore, for each traffic sign in the training sample data of the first detection model, another part of the training sample data of the second detection model can be obtained by replacing different backgrounds. Figure 3a , 3b and 3c, Figure 3a , 3b 3 and 3c are schematic diagrams showing training sample data obtained by replacing the background.

[0044] Moreover, the training sample data of the second detection model may also include coordinate information and detailed category information of the traffic sign in the image.

[0045] Step 104 : for each target object, judging whether each target object belongs to a target category object of a preset category according to the rough category information, the first position information, the plurality of second position information and a preset intersection-over-union threshold.

[0046] In an embodiment of the present invention, not every target object is an object to be predicted, and the first position information output by the first detection model and the second position information output by the second detection model are not absolutely accurate. Therefore, it is necessary to filter out target category objects from the target objects in preparation for subsequent predictions.

[0047] Step 105: If there is a target category object in at least one target object, the confidence of the target category object is calculated, and a target prediction object is selected from the target category objects according to the confidence.

[0048] In an embodiment of the present invention, if there is a target category object in the target object, a target prediction object can be further selected from the target category object according to the confidence level, wherein the confidence level can represent the probability that the target category object is the target prediction object that ultimately needs to be predicted.

[0049] Step 106: Generate a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object, and use the fine category information corresponding to the target prediction object as the category prediction result.

[0050] In an embodiment of the present invention, after determining the target prediction object, a position prediction result can be generated based on the first position information and the second position information of the target prediction object, and the fine category information corresponding to the target prediction object can be used as the category prediction result. Finally, the position prediction result and the category prediction result can be used as the prediction result of the target prediction object.

[0051] A prediction scheme for a target object provided by an embodiment of the present invention obtains an original image containing at least one target object, first inputs the original image into a trained first detection model, outputs the first position information and rough category information of at least one target object, and then for each target object, crops the original image according to the first position information to obtain multiple target object area images, and then inputs the multiple target object area images into a trained second detection model, and outputs multiple second position information and multiple fine category information of each target object. Next, for each target object, according to the rough category information, the first position information, the multiple second position information and the preset intersection-and-union ratio threshold, it is determined whether each target object belongs to a target category object of a preset category. If there is a target category object, the confidence of the target category object is calculated, and a target prediction object is selected from the target category object according to the confidence. Finally, a position prediction result of the target prediction object is generated according to the first position information and the second position information of the target prediction object, and the fine category information corresponding to the target prediction object is used as the category prediction result.

[0052] After the first position information of the target object is output from the first detection model, the embodiment of the present invention crops the original image according to the first position information to obtain a target object area image, and inputs the target object area image into the second detection model, so that the image input into the second detection model adds contextual information of the target object, thereby improving the accuracy of the fine category information output from the second detection model.

[0053] The embodiment of the present invention determines whether each target object is a target category object, and subsequently selects a target prediction object for the target category object, thereby avoiding interference from non-target prediction objects, such as background noise, and further improving the accuracy of target object prediction.

[0054] In a preferred embodiment of the present invention, one implementation method of determining whether each target object belongs to a target category object of a preset category is to determine whether the rough category information belongs to a predicted category based on the rough category information, the first position information, the plurality of second position information and the preset intersection-and-union ratio threshold. When the rough category information does not belong to the preset category, it is determined that the target object is not a target category object; when the rough category information belongs to the preset category, it is further determined whether the target object is located in the central area of ​​the plurality of target object area images based on the first position information, the plurality of second position information and the preset intersection-and-union ratio threshold, and then the target object located in the central area of ​​at least one target object area image is determined as a target category object.

[0055] In practical applications, if the preset category is a traffic sign category. First, it can be determined whether the rough category information of the target object a belongs to the traffic sign category. If the rough category information does not belong to the traffic sign category, the target object a may not be subsequently processed. If the rough category information belongs to the traffic sign category, it is further determined whether the target object a is located in the central area of ​​multiple target object area images. If the target object a is not located in the central area of ​​the target object area image, since the target object images input to the second detection model all have target objects in the central area, it means that the first position information and rough category information output by the first detection model do not correspond to the target object belonging to the traffic sign category, but to background noise. In this case, it is necessary to determine whether the first position information and rough category information output by the first detection model correspond to the target object belonging to the traffic sign category based on the coordinate alignment of the target object area image. If the first position information and rough category information of the target object b output by the first detection model do not correspond to the target object belonging to the traffic sign category, the target object b is removed.

[0056] In a preferred embodiment of the present invention, one implementation method of determining whether a target object is located in the central area of ​​multiple target object area images according to the first position information, multiple second position information and a preset intersection over union threshold is to calculate multiple intersection over union (IoU) parameters of the target object according to the first position information and multiple second position information; if there is at least one IoU parameter greater than or equal to the preset IoU threshold, it is confirmed that the target object is located in the central area of ​​at least one target object area image. If the IoU parameter of a target object is less than the preset IoU threshold, the target object is considered to be background noise and is directly removed.

[0057] In a preferred embodiment of the present invention, one implementation method for calculating the confidence of the target category object is to count the number of target category objects located in the central area, and calculate the confidence based on the number and the number of target object area images containing the target category object. For example, if a target category object c is located in the central area of ​​10 target object area images, that is, the number of target category objects c located in the central area is 10, and the number of target object area images containing the target category object c is 15, then the number of target category objects c located in the central area, 10, can be divided by the number of target object area images containing the target category object c, 15, to obtain the confidence of the target category object c, 0.667.

[0058] In a preferred embodiment of the present invention, one implementation method of selecting a target prediction object from a target category object according to a confidence level is to select a target prediction object whose confidence level satisfies a preset condition from the target category object. In practical applications, the preset condition may be one or more of the target category objects with the highest confidence level, that is, a target category object with the highest confidence level is selected from the target category objects as the target prediction object, or several target category objects with the highest confidence levels are selected from the target category objects as the target prediction object. The highest confidence level in the embodiment of the present invention may be 1.

[0059] In a preferred embodiment of the present invention, an implementation method of generating a position prediction result of a target prediction object according to the first position information and the second position information of the target prediction object is to calculate the average value of the first position information and the second position information, and use the average value as the position prediction result. For example, the first position information of the target prediction object d is (xd1, yd1), and the second position information of the target prediction object d is (xd2, yd2), then the average value of the horizontal coordinate xd1 in the first position information of the target prediction object d and the horizontal coordinate xd2 in the second position information of the target prediction object d is calculated as xd, and the average value of the vertical coordinate yd1 in the first position information of the target prediction object d and the vertical coordinate yd2 in the second position information of the target prediction object d is calculated as yd. The average value (xd, yd) is used as the position prediction result.

[0060] In a preferred embodiment of the present invention, the first detection model may include a Two-stage network model, and the second detection model may include a One-stage network model. Among them, the Two-stage network model and the One-stage network model are mainstream target detection algorithm models. The One-stage network model directly classifies and regresses the anchor, and has the characteristics of fast speed. The Two-stage network model first generates the anchor, and then classifies and regresses the anchor, and has the characteristics of high precision.

[0061] Based on the above description of an embodiment of a target object prediction method, a target object prediction scheme based on local coordinate alignment and global coordinate alignment is introduced below. The target object prediction scheme is mainly used to predict the position prediction result and category prediction result of the traffic sign from the image. Figure 4a , 4b and 4c, Figure 4a A schematic diagram showing a warning sign in a traffic sign. Figure 4b A schematic diagram showing a prohibition sign in a traffic sign. Figure 4c A schematic diagram showing an indication sign in a traffic sign.

[0062] Reference Figure 5 , Figure 5 A schematic diagram of a target object prediction scheme based on local coordinate alignment and global coordinate alignment of an embodiment of the present invention is shown. The target object prediction scheme uses the high recall characteristics of the coarse category information of the target object regressed by the Two-stage network model (i.e., the warning sign, prohibition sign and indication sign are regressed, and the fine categories under each sign are not distinguished), combined with the high accuracy characteristics of the One-stage network model regressing the fine category information (distinguishing each small category under the warning sign, prohibition sign and indication sign), and combines the prediction results of the two network models as the output result in the manner of global coordinate alignment and local coordinate alignment. Not only are the simple samples and difficult samples re-discriminated twice, but there is no difficulty in parameter setting. The One-stage network model is used for the judgment of fine category information, and combined with contextual information, it can reduce the false detection of inter-class errors and background noise. The target object prediction scheme does not have limited negative example mining, and only judges the background noise based on the idea that the regression results of the two network models for the same background noise are inconsistent. Experiments have shown that the target object prediction scheme can effectively solve the interference of negative samples, while improving the accuracy rate without reducing the recall rate.

[0063] The original image to be processed is input into the Two-stage network model, and the first position information and rough category information of the target object in the original image are output from the Two-stage network model. The original image is cropped according to the first position information to obtain the target object area image (patch) with context information, and the patch is input into the One-stage network model.

[0064] The second position information and fine category information of the target object in the patch are output from the One-stage network model. If the global coordinate alignment method and the local coordinate alignment method in the embodiment of the present invention are not added to the second position information and the fine category information output by the One-stage network model, the accuracy and recall are directly evaluated, and the results are shown in the second row of Table 1. The first row in Table 1 shows the results without adding the One-stage network model and relying only on the output of the Two-stage network model. It can be seen from Table 1 that because context information is added to the One-stage network model, the inter-class error and background noise can be reduced, which improves the accuracy and recall to a certain extent.

[0065] Recall Accuracy 96.07% 92.83% 96.36% 94.3% 95.92% 94.68% 96.36% 94.71%

[0066] Table 1

[0067] The local coordinate alignment method is used to constrain the second position information output by the One-stage network model. Because the target object only exists in the central area of ​​the image during the training of the One-stage network model, if the predicted target object is not in the central area, it means that the target object is background noise in the output result of the Two-stage network model. In this case, it is necessary to determine whether the Two-stage network model predicts a positive sample or a negative sample based on the coordinate alignment of the input and output targets in the patch. A positive sample indicates that the target object is a traffic sign, and a negative sample indicates that the target object is not a traffic sign. If it is a positive sample, the fine category information of the sample in the One-stage network model is output, and if it is a negative sample, it is removed. Among them, the method of local coordinate alignment is to determine whether the IoU meets a given threshold. When the IoU is less than the threshold, the local coordinate alignment method determines that the target object is the background and removes it directly. The results obtained by the local coordinate alignment method are shown in the third row of Table 1. It can be seen that the local coordinate alignment method can remove background false detections, but the target object detected by the One-stage network model that is not in the central area of ​​the patch is not necessarily the background, and the direct removal method will affect the recall.

[0068] The second position information of the target object output by the one-stage network model needs to be combined with the first position information output by the two-stage network model for judgment. When the one-stage network model regresses a target object that is not in the central area, and the target object is regressed to the central area in another image input to the one-stage network model, the target object is not removed, but retained with a higher confidence, because the target object is likely to be a positive sample. The local coordinate alignment method considers that for the same negative sample, the two detection results of the one-stage network model and the two-stage network model will be different, so that the background noise can be judged. Global coordinate alignment considers that for the same positive sample, the two detection results of the one-stage network model and the two-stage network model are basically the same, and the sample with the highest confidence is retained by merging to improve the recall rate. As shown in the fourth row of Table 1, the results after merging the global coordinate alignment and the local coordinate alignment are improved in both recall and precision compared with the results in the first row.

[0069] Finally, the position prediction result of the target object is the average of the second position information and the first position information output by the one-stage network model and the two-stage network model respectively, and the category prediction result is the refined category result output by the one-stage network model.

[0070] The two-stage network model in the embodiment of the present invention is responsible for the recall of rough category information. For the output target object, the original image is cropped to obtain a patch, and the patch is input into the one-stage network model. The one-stage network model is responsible for the recognition of fine classification information. On the one hand, the recognition of the one-stage network model can add context information to obtain more accurate classification results. On the other hand, the output results of the one-stage network model are combined with the output results of the two-stage network model through global coordinate alignment and local coordinate alignment to effectively judge the background noise. The accuracy and recall rate can be significantly improved in the sample set required for map automation production.

[0071] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0072] Reference Figure 6 , shows a structural block diagram of a target object prediction device according to an embodiment of the present invention, and the target object prediction device may specifically include the following modules:

[0073] An image acquisition module 61 is used to acquire an original image to be processed, wherein the original image contains at least one target object;

[0074] A first detection module 62, configured to input the original image into a trained first detection model, and output first position information and rough category information of at least one target object;

[0075] A second detection module 63 is used to crop the original image according to the first position information for each target object to obtain multiple target object area images, input the multiple target object area images into a trained second detection model, and output multiple second position information and multiple fine category information for each target object;

[0076] An object determination module 64 is used to determine, for each of the target objects, whether each of the target objects belongs to a target category object of a preset category according to the rough category information, the first position information, a plurality of the second position information and a preset intersection-over-union ratio threshold;

[0077] An object selection module 65, configured to calculate the confidence of the target category object if the target category object exists in at least one of the target objects, and select a target prediction object from the target category objects according to the confidence;

[0078] The result determination module 66 is used to generate a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object, and use the fine category information corresponding to the target prediction object as the category prediction result.

[0079] In a preferred embodiment of the present invention, the object determination module 64 includes:

[0080] a center judgment module, configured to judge whether the target object is located in a center area of ​​a plurality of target object area images according to the first position information, a plurality of the second position information and a preset intersection-over-union ratio threshold value when the rough category information belongs to the preset category;

[0081] The object determination module is used to determine the target object located in the central area of ​​at least one of the target object area images as the target category object.

[0082] In a preferred embodiment of the present invention, the center determination module includes:

[0083] A parameter calculation module, used for calculating a plurality of intersection-over-union ratio parameters of the target object according to the first position information and a plurality of the second position information;

[0084] The center determination module is used to confirm that the target object is located in the center area of ​​at least one of the target object area images if there is at least one IoU parameter greater than or equal to the preset IoU threshold.

[0085] In a preferred embodiment of the present invention, the object selection module 65 is used to count the number of the target category objects located in the central area, and calculate the confidence level according to the number and the number of target object area images containing the target category objects.

[0086] In a preferred embodiment of the present invention, the object selection module 65 is used to select the target prediction object whose confidence meets a preset condition from the target category objects.

[0087] In a preferred embodiment of the present invention, the result determination module 66 is used to calculate an average value of the first position information and the second position information, and use the average value as the position prediction result.

[0088] In a preferred embodiment of the present invention, the first detection model includes a Two-stage network model, and the second detection model includes a One-stage network model.

[0089] The embodiment of the present invention further provides an electronic device, see Figure 7 , including: a processor 701, a memory 702, and a computer program 7021 stored in the memory 702 and executable on the processor 701, wherein when the processor 701 executes the program 7021, the target object prediction method of the aforementioned embodiment is implemented.

[0090] An embodiment of the present invention further provides a readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for predicting a target object of the foregoing embodiment is implemented.

[0091] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0092] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0093] Those skilled in the art will appreciate that the embodiments of the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0094] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0095] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0097] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0098] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0099] The above is a detailed introduction to a method and device for predicting a target object provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for predicting a target object, characterized in that: The method comprises: Acquire an original image to be processed, wherein the original image contains at least one target object; Inputting the original image into a trained first detection model, and outputting first position information and rough category information of at least one target object; For each of the target objects, the original image is cropped according to the first position information to obtain a plurality of target object area images, the plurality of target object area images are input into a trained second detection model, and a plurality of second position information and a plurality of fine category information of each of the target objects are output; For each of the target objects, judging whether each of the target objects belongs to a target category object of a preset category according to the rough category information, the first position information, a plurality of the second position information and a preset intersection-over-union ratio threshold; If the target category object exists in at least one of the target objects, then calculating the confidence of the target category object, and selecting a target prediction object from the target category objects according to the confidence; A position prediction result of the target prediction object is generated according to the first position information and the second position information of the target prediction object, and the fine category information corresponding to the target prediction object is used as a category prediction result.

2. The method according to claim 1, characterized in that The determining, based on the rough category information, the first position information, the plurality of second position information and a preset intersection-over-union threshold, whether each of the target objects belongs to a target category object of a preset category includes: When the rough category information belongs to the preset category, judging whether the target object is located in a central area of ​​a plurality of target object area images according to the first position information, a plurality of the second position information and a preset intersection-over-union threshold; The target object located in the central area of ​​at least one of the target object area images is determined as the target category object.

3. The method according to claim 2, characterized in that The step of judging whether the target object is located in a central area of ​​a plurality of target object area images according to the first position information, a plurality of the second position information and a preset intersection-over-union ratio threshold comprises: Calculate a plurality of intersection-over-joint parameters of the target object according to the first position information and a plurality of the second position information; If there is at least one IoU parameter that is greater than or equal to the preset IoU threshold, it is confirmed that the target object is located in a central area of ​​at least one of the target object area images.

4. The method according to claim 2 or 3, characterized in that: The calculating the confidence of the target category object comprises: The number of the target category objects located in the central area is counted, and the confidence is calculated according to the number and the number of the target object area images containing the target category objects.

5. The method according to claim 1, characterized in that The selecting a target prediction object from the target category objects according to the confidence level includes: The target prediction object whose confidence meets a preset condition is selected from the target category objects.

6. The method according to claim 1, characterized in that The generating a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object includes: Calculate an average value of the first position information and the second position information, and use the average value as the position prediction result.

7. The method according to claim 1, characterized in that The first detection model includes a Two-stage network model, and the second detection model includes a One-stage network model.

8. A prediction device for a target object, characterized in that: The device comprises: An image acquisition module, used to acquire an original image to be processed, wherein the original image contains at least one target object; A first detection module, used for inputting the original image into a trained first detection model, and outputting first position information and rough category information of at least one target object; A second detection module is used to crop the original image according to the first position information for each target object to obtain multiple target object area images, input the multiple target object area images into a trained second detection model, and output multiple second position information and multiple fine category information for each target object; An object determination module, configured to determine, for each of the target objects, whether each of the target objects belongs to a target category object of a preset category according to the rough category information, the first position information, a plurality of the second position information and a preset intersection-over-union ratio threshold; an object selection module, configured to calculate the confidence of the target category object if the target category object exists in at least one of the target objects, and select a target prediction object from the target category objects according to the confidence; A result determination module is used to generate a position prediction result of the target prediction object according to the first position information and the second position information of the target prediction object, and use the fine category information corresponding to the target prediction object as the category prediction result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for predicting a target object according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for predicting a target object described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Neural network training and image processing method and device

    CN110009090A

  • Target detection method, device and equipment and storage medium

    CN110517262A