Image-based Object Detection Method, Device, Equipment and Readable Storage Medium

By generating candidate boxes and their confidence in the image, and iterating the candidate boxes according to the target confidence to generate the object detection box, the problem of inaccurate positioning in the image is solved, and higher positioning accuracy is achieved.

CN111291717BActive Publication Date: 2025-06-10WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010128120.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-28
Publication Date
2025-06-10
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

In the prior art, the positioning of objects in the image may be inaccurate because it depends only on the position with the greatest possibility.

Method used

By transferring the image to a preset model for detection, a candidate box and its confidence are generated, and the candidate box is iterated according to the target confidence, a target detection box is generated, and the object in the image is finally positioned.

Benefits of technology

Through iterative fusion of candidate boxes, the positioning accuracy of objects in the image is improved, and positioning inaccuracy caused by relying only on the most likely location is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111291717B_ABST
    Figure CN111291717B_ABST
Patent Text Reader

Abstract

The present invention discloses an image-based object detection method, device, equipment and readable storage medium. The method includes the following steps: transmitting an image to a preset model for detection to generate candidate boxes corresponding to the objects in the image and the confidence levels of each of the candidate boxes; determining a target confidence level among the confidence levels and iterating on each of the candidate boxes according to the target confidence level to generate a target detection box; and positioning the objects in the image according to the target detection box. The present invention locates the objects in the image by iterative fusion of each candidate box, improving the accuracy of object positioning in the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial technology (Fintech), and particularly to an image-based object detection method, device, equipment, and readable storage medium. Background Art

[0002] With the continuous development of financial technology (Fintech), especially Internet technology finance, more and more technologies (such as artificial intelligence, big data, cloud storage, etc.) are applied in the financial field. However, the financial field also puts forward higher requirements for various technologies, such as accurately detecting the position of objects in images.

[0003] Currently, when determining the position of an object in an image, first determine the possible positions of the object and the likelihood of the object existing at each position, and then determine the position with the highest likelihood as the position where the object is located.

[0004] However, the position with the highest likelihood is not necessarily the true position where the object is located, and a position with a low likelihood does not mean that it completely deviates from the true position where the object is located. Therefore, the method of completely relying on the position with the highest likelihood to determine the position of the object in the image may lead to inaccurate positioning of the object in the image due to the inaccuracy of the position with the highest likelihood. Summary of the Invention

[0005] The main purpose of the present invention is to provide an image-based object detection method, device, equipment, and readable storage medium, aiming to solve the technical problem of inaccurate positioning of objects in images in the prior art.

[0006] To achieve the above object, the present invention provides an image-based object detection method, and the image-based object detection method includes the following steps:

[0007] Transmit the image to a preset model for detection, generate candidate boxes corresponding to the objects in the image, and the confidence levels of each candidate box;

[0008] Determine the target confidence level among the confidence levels, and iterate each candidate box according to the target confidence level to generate a target detection box;

[0009] Locate the object in the image according to the target detection box.

[0010] Optionally, the step of iterating each candidate box according to the target confidence level to generate a target detection box includes:

[0011] Take the candidate box corresponding to the target confidence level as the target candidate box, and take the candidate boxes other than the target candidate box among each candidate box as the candidate boxes to be iterated;

[0012] Iteratively compare the target candidate box with each of the candidate boxes to be iterated one by one according to the target confidence level and the confidence levels of the candidate boxes to be iterated, and generate a target detection box.

[0013] Optionally, the step of iteratively comparing the target candidate box with each of the candidate boxes to be iterated one by one according to the target confidence level and the confidence levels of the candidate boxes to be iterated, and generating a target detection box includes:

[0014] Update each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of the candidate boxes to be iterated;

[0015] Arrange the updated candidate boxes to be iterated in descending order of their confidence levels to generate a candidate box sequence;

[0016] Iteratively obtain each candidate box to be iterated in the candidate box sequence as the current candidate box, and iterate between the target candidate box and the current candidate box to update the target candidate box;

[0017] If all candidate boxes to be iterated in the candidate box sequence have been obtained and iterated, generate the target detection box.

[0018] Optionally, the step of updating each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of the candidate boxes to be iterated includes:

[0019] Compare the target confidence level with the confidence levels of the candidate boxes to be iterated respectively to generate confidence ratios;

[0020] Obtain the overlap degrees between the target candidate box and the candidate boxes to be iterated, determine the target confidence ratios greater than a preset ratio among the confidence ratios, and the target overlap degrees less than a preset value among the overlap degrees;

[0021] Determine the first candidate boxes to be iterated corresponding to the target confidence ratios, and the second candidate boxes to be iterated corresponding to the target overlap degrees;

[0022] Exclude all the first candidate boxes to be iterated and all the second candidate boxes to be iterated from the candidate boxes to be iterated to update each of the candidate boxes to be iterated.

[0023] Optionally, the step of iterating between the target candidate box and the current candidate box to update the target candidate box includes:

[0024] Read the target coordinate values of the target candidate box, and the current coordinate values and the current confidence level of the current candidate box;

[0025] Determine the rectangular overlap degree between the target candidate box and the current candidate box, as well as the first offset ratio in the first preset direction and the second offset ratio in the second preset direction between the target candidate box and the current candidate box according to the target coordinate value and the current coordinate value;

[0026] Determine the target data value of the target candidate box according to the target coordinate value, and determine the current data value of the current candidate box according to the current coordinate value;

[0027] Update the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target data value, and the current data value.

[0028] Optionally, the step of updating the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target width value, the target height value, the current width value, and the current height value includes:

[0029] Update the target coordinate value of the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, and the rectangular overlap degree;

[0030] Update the target width value and the target height value of the target candidate box according to the target confidence, the current confidence, the target width value, the target height value, the current width value, and the current height value.

[0031] Optionally, the step of determining the target confidence among the confidences includes:

[0032] Compare among the confidences to determine the confidence with the largest value among the confidences;

[0033] Determine the confidence with the largest value as the target confidence.

[0034] Furthermore, to achieve the above object, the present invention further provides an object detection device based on an image, and the object detection device based on an image includes:

[0035] A generation module, configured to transmit an image to a preset model for detection, generate candidate boxes corresponding to the objects in the image, and confidences of the candidate boxes;

[0036] An iteration module, configured to determine the target confidence among the confidences and iterate on the candidate boxes according to the target confidence to generate a target detection box;

[0037] A positioning module, configured to position the objects in the image according to the target detection box.

[0038] Furthermore, to achieve the above object, the present invention also provides an image-based object detection device, which includes a memory, a processor, and an image-based object detection program stored on the memory and executable on the processor. When the image-based object detection program is executed by the processor, it implements the steps of the image-based object detection method as described above.

[0039] Furthermore, to achieve the above object, the present invention also provides a readable storage medium, on which an image-based object detection program is stored. When the image-based object detection program is executed by a processor, it implements the steps of the image-based object detection method as described above.

[0040] In the image-based object detection method of the present invention, an image is first transmitted to a preset model for detection, generating a plurality of candidate boxes representing possible positions of objects in the image, and confidence levels indicating the probability that each candidate box may contain an object. Then, the target confidence level with the largest value is determined from each confidence level, and each candidate box is iteratively processed based on the target confidence level to generate target detection boxes for each candidate box. Furthermore, based on the target detection boxes, the objects in the image are located. By iteratively fusing each candidate box to locate the objects in the image, it avoids locating the objects only based on the most likely position, making the location of the objects in the image more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic structural diagram of the hardware operating environment of the device related to the embodiment of the image-based object detection device of the present invention;

[0042] Figure 2 It is a schematic flowchart of the first embodiment of the image-based object detection method of the present invention;

[0043] Figure 3 It is a schematic diagram of the functional modules of the preferred embodiment of the image-based object detection device of the present invention;

[0044] Figure 4 It is an iterative diagram of the target candidate box and the current candidate box in the image-based object detection method of the present invention.

[0045] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] The present invention provides an image-based object detection device, refer toFigure 1 , Figure 1 This is a schematic structural diagram of the hardware operating environment of the device according to the embodiment of the object detection device based on images of the present invention.

[0048] As shown in Figure 1 , the object detection device based on images may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.

[0049] Those skilled in the art can understand that Figure 1 the hardware structure of the object detection device based on images shown in

[0050] does not constitute a limitation on the object detection device based on images, and may include more or fewer components than shown in the figure, or combine certain components, or arrange different components. Figure 1 As shown in

[0051] In Figure 1 the hardware structure of the object detection device based on images shown, the network interface 1004 is mainly used to connect to the background server and perform data communication with the background server; the user interface 1003 is mainly used to connect to the client (user side) and perform data communication with the client; the processor 1001 can call the object detection program stored in the memory 1005 and perform the following operations:

[0052] Transmit the image to a preset model for detection, generate candidate boxes corresponding to the objects in the image, and the confidence levels of each of the candidate boxes;

[0053] Determine the target confidence level among the respective confidence levels, and iterate through each of the candidate bounding boxes according to the target confidence level to generate a target detection bounding box;

[0054] Locate the object in the image according to the target detection bounding box.

[0055] Further, the step of iterating through each of the candidate bounding boxes according to the target confidence level to generate a target detection bounding box includes:

[0056] Take the candidate bounding box corresponding to the target confidence level as the target candidate bounding box, and take the candidate bounding boxes other than the target candidate bounding box among each of the candidate bounding boxes as the candidate bounding boxes to be iterated;

[0057] Iterate the target candidate bounding box one by one with each of the candidate bounding boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate bounding boxes to be iterated to generate a target detection bounding box.

[0058] Further, the step of iterating the target candidate bounding box one by one with each of the candidate bounding boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate bounding boxes to be iterated to generate a target detection bounding box includes:

[0059] Update each of the candidate bounding boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate bounding boxes to be iterated;

[0060] Arrange the updated candidate bounding boxes to be iterated in descending order of their confidence levels to generate a sequence of candidate bounding boxes;

[0061] Successively obtain the candidate bounding boxes to be iterated in the sequence of candidate bounding boxes as the current candidate bounding box, and iterate the target candidate bounding box and the current candidate bounding box to update the target candidate bounding box;

[0062] If all the candidate bounding boxes to be iterated in the sequence of candidate bounding boxes have been obtained and iterated, generate the target detection bounding box.

[0063] Further, the step of updating each of the candidate bounding boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate bounding boxes to be iterated includes:

[0064] Compare the target confidence level with the confidence levels of each of the candidate bounding boxes to be iterated to generate a confidence ratio;

[0065] Obtain the overlap degree between the target candidate bounding box and each of the candidate bounding boxes to be iterated, and determine the target confidence ratios greater than a preset ratio among the respective confidence ratios, and the target overlap degrees less than a preset value among the respective overlap degrees;

[0066] Determine the first candidate box to be iterated corresponding to each of the target confidence ratios, and the second candidate box to be iterated corresponding to each of the target overlap degrees;

[0067] Exclude each of the first candidate boxes to be iterated and each of the second candidate boxes from each of the candidate boxes to be iterated, so as to update each of the candidate boxes to be iterated.

[0068] Further, the step of iterating the target candidate box and the current candidate box to update the target candidate box includes:

[0069] Read the target coordinate values of the target candidate box, and the current coordinate values and the current confidence of the current candidate box;

[0070] According to the target coordinate values and the current coordinate values, determine the rectangular overlap degree between the target candidate box and the current candidate box, and the first offset ratio of the target candidate box and the current candidate box in the first preset direction and the second offset ratio in the second preset direction;

[0071] Determine the target data value of the target candidate box according to the target coordinate values, and determine the current data value of the current candidate box according to the current coordinate values;

[0072] Update the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target data value, and the current data value.

[0073] Further, the step of updating the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target width value, the target height value, the current width value, and the current height value includes:

[0074] Update the target coordinate values of the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence, the current confidence, and the rectangular overlap degree;

[0075] Update the target width value and the target height value of the target candidate box according to the target confidence, the current confidence, the target width value, the target height value, the current width value, and the current height value.

[0076] Further, the step of determining the target confidence among the confidences includes:

[0077] Compare among the confidences to determine the confidence with the largest value among the confidences;

[0078] Determine the confidence with the largest value as the target confidence.

[0079] The specific implementation manner of the object detection device based on images in the present invention is basically the same as each embodiment of the following object detection method based on images, and will not be described in detail herein.

[0080] The present invention also provides an object detection method based on images.

[0081] Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the object detection method based on images in the present invention.

[0082] The embodiments of the present invention provide embodiments of the object detection method based on images. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here. Specifically, the object detection method based on images in this embodiment includes:

[0083] Step S10, transmitting the image to a preset model for detection, generating candidate boxes corresponding to the objects in the image, and the confidence levels of each of the candidate boxes.

[0084] The object detection method based on images in this embodiment is applied to a server and is suitable for detecting the objects in an image through the server to determine the positions of the objects in the image; such as the positions of trees, houses, people, etc. in the image. A preset model for detecting the positions of objects in an image is pre-trained in the server, and the image is transmitted to this preset model for detection, generating candidate boxes corresponding to the objects in the image, and the possible positions of the objects are reflected by the candidate boxes.

[0085] Furthermore, due to the differences in the objects between images, when the preset model is detecting, usually multiple candidate boxes will be generated for the same object in the image to represent the possible positions of the object. And in order to determine the likelihood of the object existing at each position, in addition to generating candidate boxes for the objects in the image, the preset model also generates confidence levels representing the likelihood of the object existing at each position, that is, the confidence levels of each candidate box, to reflect the probability of each candidate box completely containing the object through the confidence levels.

[0086] Step S20, determining the target confidence level among the confidence levels, and iterating on each of the candidate boxes according to the target confidence level to generate a target detection box;

[0087] It can be understood that the likelihood of each candidate box completely containing the object is different. In order to determine the candidate box that is most likely to completely contain the object, the target confidence level can be determined through the relationship of the confidence levels between the candidate boxes, and then the candidate box that is most likely to completely contain the object is represented by the target confidence level. Specifically, the steps of determining the target confidence level among the confidence levels include:

[0088] Step S21: Compare among the respective confidence levels to determine the confidence level with the largest value among the respective confidence levels.

[0089] Step S22: Determine the confidence level with the largest value as the target confidence level.

[0090] The server compares among the respective confidence levels, filters out the confidence level with the largest data among them, and determines the confidence level with the largest value as the target confidence level, which represents the position where the object is most likely to exist in the picture. It should be noted that the process of filtering out the target confidence level through comparison can be implemented by the NMS (Non-maximum-suppression) algorithm. The NMS algorithm completes clustering division through spatial distance combined with the intersection over union (IOU), obtains the candidate box with the highest score, and the candidate box with the highest score is the confidence level with the largest value, which is determined as the target confidence level.

[0091] Further, after determining the target confidence level, the respective candidate boxes can be iterated based on the target confidence level, and the respective candidate boxes are gradually fused into one candidate box as the target detection box representing the position of the object. Among them, the iteration of the respective candidate boxes is essentially to fuse two candidate boxes into one, and then use the candidate box obtained by this fusion to fuse with other candidate boxes until all the candidate boxes have been fused. Specifically, the steps of generating the target detection box by iterating the respective candidate boxes according to the target confidence level include:

[0092] Step S23: Use the candidate box corresponding to the target confidence level as the target candidate box, and use the candidate boxes other than the target candidate box among the respective candidate boxes as the candidate boxes to be iterated.

[0093] Step S24: According to the target confidence level and the confidence levels of the respective candidate boxes to be iterated, iterate the target candidate box one by one with the respective candidate boxes to be iterated to generate the target detection box.

[0094] Even further, the server assigns a target identifier to the candidate box corresponding to the target confidence level to define it as the target candidate box, and assigns an iteration identifier to the other candidate boxes other than this target candidate box among the respective candidate boxes to define the other candidate boxes as the candidate boxes to be iterated. Then, based on the target confidence level of the target candidate box and the confidence levels of the respective candidate boxes to be iterated, the target candidate box is iterated one by one with the respective candidate boxes to be iterated. After all the candidate boxes to be iterated have been iterated, the target detection box is obtained.

[0095] Step S30: Locate the object in the image according to the target detection box.

[0096] Further, the object in the target detection box is the position where the object in the image is located, and the positioning of the object in the image is realized based on the target detection box. It should be noted that there are many objects in the image. For each object, a corresponding candidate box is generated, and the candidate boxes generated by each object are iterated by their respective object confidence levels to generate a target detection box representing the position of each object, thereby realizing the positioning of each object in the image.

[0097] The object detection method based on an image according to the present invention first transmits the image to a preset model for detection, generates a plurality of candidate boxes representing the possible positions of the objects in the image, and the confidence levels of the probabilities that the objects may be included in each candidate box; then determines the target confidence level with the largest value from each confidence level, and iterates each candidate box according to the target confidence level to generate each candidate box into a target detection box; and then locates the objects in the image based on the target detection box. By iteratively fusing each candidate box to locate the objects in the image, it is avoided to locate the objects only based on the position with the highest possibility, making the positioning of the objects in the image more accurate.

[0098] Further, based on the first embodiment of the object detection method based on an image according to the present invention, a second embodiment of the object detection method based on an image according to the present invention is proposed.

[0099] The difference between the second embodiment of the object detection method based on an image and the first embodiment of the object detection method based on an image is that the step of iteratively generating the target detection box by iteratively comparing the target candidate box with each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate boxes to be iterated includes:

[0100] Step S241, update each candidate box to be iterated according to the target confidence level and the confidence levels of each candidate box to be iterated;

[0101] It can be understood that since there are many objects included in the image, during the iteration of the candidate box of a certain object, the candidate boxes of other objects may also be fused, resulting in inaccurate iteration. To avoid such a situation, in this embodiment, a rejection mechanism is set during the iteration of each candidate box to be iterated. Specifically, taking the target confidence level as a benchmark, according to the deviation degree between the confidence level of each candidate box to be iterated and the target confidence level, the candidate boxes to be removed in each iterated candidate box are determined, so as to update each candidate box to be iterated. Specifically, the step of updating each candidate box to be iterated according to the target confidence level and the confidence levels of each candidate box to be iterated includes:

[0102] Step a1, compare the target confidence level with the confidence levels of each candidate box to be iterated respectively to generate a confidence level ratio;

[0103] Step a2, obtain the overlap degrees between the target candidate box and each of the candidate boxes to be iterated, and determine the target confidence ratios greater than a preset ratio among the confidence ratios, and the target overlap degrees less than a preset value among the overlap degrees;

[0104] Step a3, determine the first candidate boxes to be iterated corresponding to the target confidence ratios, and the second candidate boxes to be iterated corresponding to the target overlap degrees;

[0105] Step a4, remove each of the first candidate boxes to be iterated and each of the second candidate boxes to be iterated from the candidate boxes to be iterated, so as to update the candidate boxes to be iterated.

[0106] Furthermore, in this embodiment, in addition to removing the candidate boxes to be iterated according to the deviation degree between the confidence of each candidate box to be iterated and the target confidence; it is also set to remove the candidate boxes to be iterated according to the overlap degree between each candidate box to be iterated and the target candidate box respectively. When the deviation degree is large or the overlap degree is low, they are all candidate boxes to be iterated that need to be removed; only the candidate boxes to be iterated with a small deviation degree and a high overlap degree are iterated.

[0107] Even further, a preset ratio representing the level of deviation degree and a preset value representing the level of overlap degree are preset in advance. Compare the confidence of each candidate box to be iterated with the target confidence respectively to generate the confidence ratios of each candidate box to be iterated, and obtain the overlap degrees between the position area where the target candidate box is located and the position areas where each candidate box to be iterated is located respectively. Then compare each confidence ratio with the preset ratio one by one to determine the target confidence ratios greater than the preset ratio among the confidence ratios. Such target confidence ratios represent that the deviation degree between the confidence of the candidate box to be iterated and the target confidence is large, and the candidate box to be iterated that generates such target confidence ratios needs to be removed. At the same time, compare each overlap degree with the preset value one by one to determine the target overlap degrees less than the preset value among the overlap degrees. Such target overlap degrees represent that the overlap degree between the candidate box to be iterated and the target candidate box is low, and the candidate box to be iterated that generates such target overlap degrees needs to be removed. Specifically, according to the confidence generating each target confidence ratio, determine the first candidate boxes to be iterated corresponding to the target confidence ratios, and use the candidate boxes to be iterated generating each target overlap degree as the second candidate boxes to be iterated. Then remove the first candidate boxes to be iterated and the second candidate boxes to be iterated from the candidate boxes to be iterated, and complete the update of the candidate boxes to be iterated to ensure that each candidate box to be iterated is for the same object.

[0108] Step S242, arrange the updated candidate boxes to be iterated in descending order according to the confidence of the updated candidate boxes to be iterated, and generate a candidate box sequence;

[0109] Step S243: Iteratively obtain each candidate box to be iterated in the candidate box sequence as the current candidate box, and iterate between the target candidate box and the current candidate box to update the target candidate box.

[0110] Step S244: If all candidate boxes to be iterated in the candidate box sequence have been obtained and iterated, generate the target detection box.

[0111] Furthermore, after each candidate box to be iterated is updated by elimination, the updated candidate boxes to be iterated are sorted in descending order according to their respective confidence levels to generate a candidate box sequence. Thereafter, according to the arrangement order of the candidate boxes to be iterated in the candidate box sequence, each candidate box to be iterated in the candidate box sequence is obtained one by one as the current candidate box, and iteration is performed between the current candidate box and the target candidate box to update the target candidate box. Thereafter, the next candidate box to be iterated in the candidate box sequence is obtained as the new current candidate box, and iteration is performed between the updated target candidate box and the new current candidate box to update the target candidate box again. In this way, after all candidate boxes to be iterated in the candidate box sequence have been obtained and iterated, the finally updated target candidate box is determined as the target detection box.

[0112] In this embodiment, by eliminating each candidate box to be iterated, the update of each candidate box to be iterated is realized, so that each candidate box targets the same object, ensuring the accuracy of object positioning. At the same time, by forming a candidate box sequence with the updated candidate boxes to be iterated, each candidate box to be iterated is read one by one as the current candidate box and the target candidate box for iteration to generate the target detection box, ensuring that all candidate boxes to be iterated are iterated, avoiding omission; and iterating in the order from large to small according to the confidence level is conducive to ensuring the accuracy of iteration.

[0113] Further, based on the first or second embodiment of the object detection method based on images of the present invention, a third embodiment of the object detection method based on images of the present invention is proposed.

[0114] The difference between the third embodiment of the object detection method based on images and the first or second embodiment of the object detection method based on images is that the step of iterating between the target candidate box and the current candidate box to update the target candidate box includes:

[0115] Step b1: Read the target coordinate value of the target candidate box, as well as the current coordinate value and current confidence level of the current candidate box.

[0116] Step b2: Determine the rectangular overlap degree between the target candidate box and the current candidate box, as well as the first offset ratio in the first preset direction and the second offset ratio in the second preset direction between the target candidate box and the current candidate box according to the target coordinate value and the current coordinate value;

[0117] Step b3: Determine the target data value of the target candidate box according to the target coordinate value, and determine the current data value of the current candidate box according to the current coordinate value;

[0118] Step b4: Update the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target data value, and the current data value.

[0119] In this embodiment, on the basis of the NMS algorithm, the Fine-tuned NMS algorithm (FTNMS) is adopted to realize the iteration between the target candidate box and the current candidate box. Specifically, please refer to Figure 4 , where the candidate box is represented by a rectangle, and each rectangle can be represented by four parameters, that is, the x coordinate value and y coordinate value of a certain corner point in the rectangle, as well as the width width and height height of the rectangle. The corner point is preferably the upper left corner point. If the candidate box is represented by bb (bounding box), then bb = (x, y, width, height). Specifically in Figure 4 , the rectangle ABCD is the target candidate box, bb ABCD = (x1, y1, width1, height1), A 1 B 1 C 1 D 1 is the current candidate box,

[0120] Furthermore, the target confidence of the target candidate box is represented by score ABCD , read the confidence of the current candidate box as the current confidence, and represent it with , and at the same time read the target coordinate value of the target candidate box, and this target coordinate value includes the coordinate values of four corner points, which are A (x1, y1), B(x1, y2), C(x2, y2) and D(x2, y1) respectively. Similarly, read the current coordinate value of the current candidate box, including the coordinate values of the four corner points of the current candidate box, which are A 1 (x1’, y1’), B 1 (x1’, y2’), C 1 (x2’, y2’) and D 1(x2’, y1’). Thereafter, the rectangular overlap degree between the target candidate box and the current candidate box is calculated based on the target coordinate value and the current coordinate value, and the rectangular overlap degree can be represented by the formula wherein, represents the area of the overlapping part between the target candidate box and the current candidate box, dw = x2 - x1’, dh = y2 - y1’; represents the sum of the areas between the target candidate box and the current candidate box,

[0121] Furthermore, based on the target coordinate value and the current coordinate value, the first offset ratio of the target candidate box and the current candidate box in the first preset direction and the second offset ratio in the second preset direction are calculated. Among them, the first preset direction and the second preset direction are preset coordinate axis directions. The first preset direction is preferably the x-axis direction, and the second preset direction is preferably the y-axis direction. The first offset ratio is represented by and W1 = x1’ - x1, W2 = x2’ - x1, and the second offset ratio is represented by and H1 = y1’ - y1, H2 = y2’ - y1.

[0122] Further, the target data values of the target candidate box are determined through the target coordinate value. The target data values include the target width value and the target height value. Among them, the target width value of the target candidate box can be determined by width1 = x2 - x1, and the target height value of the target candidate box can be determined by height1 = y2 - y1. At the same time, the current data values of the current candidate box are determined through the current coordinate value. The current data values include the current width value and the current height value. Among them, the current width value of the current candidate box is determined by width2 = x2’ - x1’, and the current height value of the current candidate box is determined by height2 = y2’ - y1’. Then, based on the rectangular overlap degree, the first offset ratio, the second offset ratio, the target width value, the target height value, the current width value, the current height value obtained from the above calculations, combined with the target coordinate value, the target confidence level and the current confidence level, the target candidate box is updated, that is, the x1, y1, width1, height1 in bb ABCD =(x1, y1, width1, height1) are updated. Specifically, according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence level, the current confidence level, the rectangular overlap degree, the target width value, the target height value, the current width value and the current height value, the steps of updating the target candidate box include:

[0123] Step b41, updating the target coordinate value of the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence level, the current confidence level, and the rectangular overlap degree;

[0124] Step b42: Update the target width value and target height value of the target candidate box according to the target confidence, current confidence, target width value, target height value, current width value, and current height value.

[0125] Furthermore, a first preset formula for updating the upper left corner coordinate point in the target candidate box and a second preset formula for updating the width value and height value of the target candidate box are preset. The first preset formula includes formula (1) and formula (2) to respectively update the x coordinate value and y coordinate value of the upper left corner coordinate point, and the update of the target coordinate value of the target candidate box is achieved through the first preset formula. The second preset formula includes formula (3) and formula (4) to respectively update the width value and height value of the target candidate box. Formula (1), formula (2), formula (3), and formula (4) are shown as follows:

[0126]

[0127]

[0128]

[0129]

[0130] Transmit the x coordinate value of the upper left corner coordinate point, w1, the first offset ratio, the current confidence, the target confidence, and the rectangle overlap degree in the target coordinate value into formula (1) for calculation, and the new x coordinate value of the upper left corner coordinate point in the target coordinate value can be obtained. Transmit the y coordinate value of the upper left corner coordinate point, H1, the second offset ratio, the current confidence, the target confidence, and the rectangle overlap degree in the target coordinate value into formula (2) for calculation, and the new y coordinate value of the upper left corner coordinate point in the target coordinate value can be obtained. Transmit the current confidence, the target confidence, the target width value, and the current width value into formula (3) for calculation, and the new target width value of the target candidate box can be obtained. Transmit the current confidence, the target confidence, the target height value, and the current height value into formula (4) for calculation, and the new target height value of the target candidate box can be obtained.

[0131] Further, use the newly obtained x coordinate value, y coordinate value, new target width value, and new target height value through the above calculations to update the upper left corner coordinate point, target width value, and target height value in the target coordinate value of the target candidate box, that is, form a new target candidate box. Subsequently, based on the new target candidate box, obtain the next candidate box to be iterated in the candidate box sequence as the new current candidate box, and continue to iterate in the above manner until all the candidate boxes to be iterated in the candidate sequence box are iterated in the above manner to generate a target detection box, realizing the positioning of the object in the image.

[0132] In this embodiment, the NMS algorithm is fine-tuned to implement the iteration between the target candidate box and the current candidate box, so that the positioning of the object in the image is related to multiple candidate boxes, improving the accuracy of positioning.

[0133] The present invention also provides an object detection device based on an image.

[0134] Referring to Figure 3 , Figure 3 is a schematic diagram of the functional modules of the first embodiment of the object detection device based on an image according to the present invention. The object detection device based on an image includes:

[0135] A generation module 10, configured to transmit an image to a preset model for detection, generate candidate boxes corresponding to the objects in the image, and the confidence levels of the candidate boxes;

[0136] An iteration module 20, configured to determine a target confidence level among the confidence levels, and iterate on the candidate boxes according to the target confidence level to generate a target detection box;

[0137] A positioning module 30, configured to position the objects in the image according to the target detection box.

[0138] Further, the iteration module 20 further includes:

[0139] A determination unit, configured to use the candidate box corresponding to the target confidence level as the target candidate box, and use the candidate boxes other than the target candidate box among the candidate boxes as the candidate boxes to be iterated;

[0140] An iteration unit, configured to iterate the target candidate box one by one with each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of the candidate boxes to be iterated, to generate a target detection box.

[0141] Further, the iteration unit is further configured to:

[0142] Update each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of the candidate boxes to be iterated;

[0143] Arrange the updated candidate boxes to be iterated in descending order of the confidence levels of the updated candidate boxes to be iterated, to generate a candidate box sequence;

[0144] Obtain each of the candidate boxes to be iterated in the candidate box sequence as the current candidate box one by one, and iterate on the target candidate box and the current candidate box to update the target candidate box;

[0145] If all the candidate boxes to be iterated in the candidate box sequence have been obtained and iterated, generate the target detection box.

[0146] Further, the iteration unit is further configured to:

[0147] Compare the target confidence level with the confidence levels of the candidate boxes to be iterated respectively to generate confidence ratios;

[0148] Obtain the overlap degrees between the target candidate box and the candidate boxes to be iterated, and determine the target confidence ratios greater than a preset ratio among the confidence ratios, and the target overlap degrees less than a preset value among the overlap degrees;

[0149] Determine the first candidate boxes to be iterated corresponding to the target confidence ratios, and the second candidate boxes to be iterated corresponding to the target overlap degrees;

[0150] Exclude all the first candidate boxes to be iterated and all the second candidate boxes to be iterated from the candidate boxes to be iterated, so as to update the candidate boxes to be iterated.

[0151] Further, the iteration unit is further configured to:

[0152] Read the target coordinate values of the target candidate box, and the current coordinate values and the current confidence level of the current candidate box;

[0153] According to the target coordinate values and the current coordinate values, determine the rectangular overlap degree between the target candidate box and the current candidate box, and the first offset ratio in the first preset direction and the second offset ratio in the second preset direction between the target candidate box and the current candidate box;

[0154] Determine the target data value of the target candidate box according to the target coordinate values, and determine the current data value of the current candidate box according to the current coordinate values;

[0155] Update the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence level, the current confidence level, the rectangular overlap degree, the target data value, and the current data value.

[0156] Further, the iteration unit is further configured to:

[0157] Update the target coordinate values of the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence level, the current confidence level, and the rectangular overlap degree;

[0158] Update the target width value and the target height value of the target candidate box according to the target confidence level, the current confidence level, the target width value, the target height value, the current width value, and the current height value.

[0159] Further, the iteration module 20 further includes:

[0160] A comparison unit, configured to compare among the respective confidence levels to determine the confidence level with the largest value among the respective confidence levels;

[0161] The determination unit is further configured to determine the confidence level with the largest value as the target confidence level.

[0162] The specific implementation manner of the image-based object detection device of the present invention is basically the same as that of the above-described embodiments of the image-based object detection method, and will not be described in detail herein.

[0163] In addition, an embodiment of the present invention further provides a readable storage medium.

[0164] An image-based object detection program is stored on the readable storage medium. When the image-based object detection program is executed by a processor, the steps of the above-described image-based object detection method are implemented.

[0165] The readable storage medium of the present invention may be a computer-readable storage medium, and its specific implementation manner is basically the same as that of the above-described embodiments of the image-based object detection method, and will not be described in detail herein.

[0166] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope of the present invention as protected by the claims. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, all fall within the protection scope of the present invention.

Claims

1. An image-based object detection method, characterized in that, the image-based object detection method comprises the following steps: Transmit the image to a preset model for detection, generate candidate boxes corresponding to the objects in the image, and the confidence levels of each of the candidate boxes; Determine the target confidence level among each of the confidence levels, and iterate each of the candidate boxes according to the target confidence level to generate a target detection box; The step of iterating each of the candidate boxes according to the target confidence level to generate a target detection box comprises: Take the candidate box corresponding to the target confidence level as the target candidate box, and take the candidate boxes other than the target candidate box among each of the candidate boxes as the candidate boxes to be iterated; According to the target confidence level and the confidence levels of each of the candidate boxes to be iterated, iterate the target candidate box one by one with each of the candidate boxes to be iterated to generate a target detection box; The step of iterating the target candidate box one by one with each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate boxes to be iterated to generate a target detection box comprises: Update each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate boxes to be iterated; Arrange the updated candidate boxes to be iterated in descending order according to the confidence levels of the updated candidate boxes to be iterated to generate a candidate box sequence; Obtain each of the candidate boxes to be iterated in the candidate box sequence one by one as the current candidate box, and iterate the target candidate box and the current candidate box to update the target candidate box, wherein the iteration of each candidate box is to fuse two candidate boxes into one, and fuse the fused candidate box with other candidate boxes until each candidate box has been fused; If each of the candidate boxes to be iterated in the candidate box sequence has been obtained and iterated, generate the target detection box; Locate the objects in the image according to the target detection box; The step of iterating the target candidate box and the current candidate box to update the target candidate box comprises: Read the target coordinate values of the target candidate box, and the current coordinate values and current confidence level of the current candidate box; According to the target coordinate values and the current coordinate values, determine the rectangular overlap degree between the target candidate box and the current candidate box, and the first offset ratio in the first preset direction and the second offset ratio in the second preset direction between the target candidate box and the current candidate box; Determine the target data value of the target candidate box according to the target coordinate values, and determine the current data value of the current candidate box according to the current coordinate values; Update the target candidate box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence level, the current confidence level, the rectangular overlap degree, the target data value, and the current data value.

2. The image-based object detection method according to claim 1, characterized in that, the step of updating each of the candidate boxes to be iterated according to the target confidence level and the confidence levels of each of the candidate boxes to be iterated comprises: Compare the target confidence level with the confidence levels of each of the candidate boxes to be iterated respectively to generate a confidence ratio; Obtain the overlap degrees between the target candidate box and each of the candidate boxes to be iterated, and determine the target confidence ratio among each of the confidence ratios that is greater than a preset ratio, and the target overlap degree among each of the overlap degrees that is less than a preset value; Determine the first candidate boxes to be iterated corresponding to each of the target confidence ratios, and the second candidate boxes to be iterated corresponding to each of the target overlap degrees; Exclude each of the first candidate boxes to be iterated and each of the second candidate boxes to be iterated from each of the candidate boxes to be iterated, so as to update each of the candidate boxes to be iterated.

3. The image-based object detection method according to claim 1, characterized in that the step of updating the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangle overlap degree, the target width value, the target height value, the current width value, and the current height value includes: Update the target coordinate value of the target candidate box according to the target coordinate value, the first offset ratio, the second offset ratio, the target confidence, the current confidence, and the rectangle overlap degree; Update the target width value and the target height value of the target candidate box according to the target confidence, the current confidence, the target width value, the target height value, the current width value, and the current height value.

4. The image-based object detection method according to any one of claims 1-3, characterized in that the step of determining the target confidence among each of the confidences includes: Compare among each of the confidences to determine the confidence with the largest value among each of the confidences; Determine the confidence with the largest value as the target confidence.

5. An image-based object detection device, characterized in that the image-based object detection device includes: A generation module, configured to transmit an image to a preset model for detection, generate candidate boxes corresponding to objects in the image, and the confidence of each of the candidate boxes; An iteration module, configured to determine the target confidence among each of the confidences, and iterate each of the candidate boxes according to the target confidence to generate a target detection box; The iteration module includes: A determination unit, configured to use the candidate box corresponding to the target confidence as the target candidate box, and use the candidate boxes other than the target candidate box among each of the candidate boxes as the candidate boxes to be iterated; An iteration unit, configured to iterate the target candidate box one by one with each of the candidate boxes to be iterated according to the target confidence and the confidence of each of the candidate boxes to be iterated, to generate a target detection box; The iteration unit is further configured to: Update each of the candidate boxes to be iterated according to the target confidence and the confidence of each of the candidate boxes to be iterated; Arrange the updated candidate boxes to be iterated in descending order of the confidence of the updated candidate boxes to be iterated to generate a candidate box sequence; Iteratively obtain each candidate bounding box to be iterated in the candidate bounding box sequence as the current candidate bounding box, and iterate between the target candidate bounding box and the current candidate bounding box to update the target candidate bounding box, where the iteration of each candidate bounding box is to fuse two candidate bounding boxes into one, and fuse the fused candidate bounding box with other candidate bounding boxes until all candidate bounding boxes have been fused; If all candidate bounding boxes to be iterated in the candidate bounding box sequence have been obtained and iterated, generate the target detection bounding box; A positioning module, configured to locate an object in the image according to the target detection bounding box; The iteration unit is further configured to: Read the target coordinate values of the target candidate bounding box, as well as the current coordinate values and current confidence of the current candidate bounding box; According to the target coordinate values and the current coordinate values, determine the rectangular overlap degree between the target candidate bounding box and the current candidate bounding box, as well as the first offset ratio of the target candidate bounding box and the current candidate bounding box in the first preset direction and the second offset ratio in the second preset direction; Determine the target data value of the target candidate bounding box according to the target coordinate values, and determine the current data value of the current candidate bounding box according to the current coordinate values; Update the target candidate bounding box according to the target coordinate values, the first offset ratio, the second offset ratio, the target confidence, the current confidence, the rectangular overlap degree, the target data value, and the current data value.

6. An image-based object detection device, characterized in that, the image-based object detection device includes a memory, a processor, and an image-based object detection program stored on the memory and executable on the processor, and when the image-based object detection program is executed by the processor, it implements the steps of the image-based object detection method according to any one of claims 1-4.

7. A readable storage medium, characterized in that, an image-based object detection program is stored on the readable storage medium, and when the image-based object detection program is executed by a processor, it implements the steps of the image-based object detection method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Object detection method, device and system

    CN108268869A

  • Small target detection method based on feature fusion and depth learning

    CN109344821A