A method and system for obtaining a stable target detection frame based on target detection

By acquiring and correcting the pixel coordinates and movement speed of the target detection box, the problems of low accuracy and unstable jumping of the detection box in dense vehicle environments are solved, achieving a higher detection accuracy.

CN115760945BActive Publication Date: 2026-01-02INTELLIGENT INTER CONNECTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211474470.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2026-01-02
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

Existing target detection boxes have low accuracy in densely trafficked conditions and suffer from instability due to jumping back and forth.

Method used

By obtaining the pixel coordinates of the four corner points of the detection box in the target detection result of each frame of the image, performing perspective transformation, calculating the true coordinates of the points in the detection box, and performing stabilization processing based on the matching relationship between image frames and the movement speed, the position of the detection box is corrected.

Benefits of technology

It improves the accuracy of target detection, reduces the jumping of the detection frame, and enhances the detection accuracy in densely trafficked environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760945B_ABST
    Figure CN115760945B_ABST
Patent Text Reader

Abstract

The application discloses a target detection-based stable target detection frame acquisition method and system, and relates to the field of intelligent traffic management.The method comprises the following steps: acquiring the pixel horizontal moving speed and the pixel vertical moving speed of the detection frame that is matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame that is not matched successfully in the current frame, and the pixel coordinates of the four corner points of the detection frame that is not matched successfully in the last frame, the pixel horizontal moving speed and the pixel vertical moving speed of the detection frame that is not matched successfully in the last frame, and the real coordinates of the center point of the detection frame, according to the matching relationship of the detection frames between different image frames and the real coordinates of the center point of each detection frame; and correcting the position of the detection frame according to the pixel horizontal moving speed and the pixel vertical moving speed of the target in different situations, so that the position of the detection frame that has a jitter can be corrected, and the accuracy of target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent traffic management, and particularly relates to a method and system for obtaining a stable target detection frame based on target detection. BACKGROUND

[0002] With the increasing number of vehicles in cities, the road conditions are becoming more and more complex, especially in various intersection areas, where vehicles, non-motor vehicles, pedestrians, and the like converge together. Therefore, a target detection algorithm is usually used to track and detect vehicle targets at multiple intersections. The accuracy of target detection is closely related to the precision of the target detection frame. However, the target detection frame obtained by the target detection algorithm is unstable due to various reasons, such as target frame jumping back and forth, and sometimes the detection frame can detect and sometimes it cannot.

[0003] Currently, when obtaining a target detection frame, the network-predicted bounding box is usually grouped according to different categories, and then the intersection over union of all bounding boxes of each category with other categories is calculated. According to the intersection over union and the confidence, the bounding boxes with high overlap and different categories are removed. However, this method does not address the problem of target detection frame jumping back and forth and unstable detection, and the object processed is the intersection over union of each category and other categories. When the vehicles are dense, the detection frame in the same area is too dense, and the precision of the detection frame obtained by this method is low. SUMMARY

[0004] To solve the above technical problems, the present application provides a method and system for obtaining a stable target detection frame based on target detection, which can solve the problem of severe detection frame jumping and difficulty in accurately obtaining the detection frame.

[0005] To achieve the above purpose, on the one hand, the present application provides a method for obtaining a stable target detection frame based on target detection, which comprises:

[0006] According to the detection frame in the target detection result of each frame of image, the pixel coordinates corresponding to the four corner points of the detection frame in the orthoscopic scene of each frame of image are obtained;

[0007] According to the pixel coordinates corresponding to the four corner points of the detection frame in the orthoscopic scene, the real coordinates corresponding to the center point of each detection frame and the matching relationship of the detection frame between different image frames are obtained;

[0008] According to the matching relationship of the detection boxes between different image frames and the real coordinates corresponding to the points in each detection box, the pixel horizontal moving speed and the pixel vertical moving speed of the detection boxes matched successfully in the current frame, the pixel coordinates of the four corner points of the detection boxes not matched successfully in the current frame, the pixel coordinates of the four corner points of the detection boxes not matched successfully in the previous frame, the pixel horizontal moving speed and the pixel vertical moving speed of the detection boxes not matched successfully in the previous frame, and the real coordinates of the points in the detection boxes are obtained.

[0009] According to the pixel coordinates corresponding to the four corner points of the detection boxes not matched successfully in the previous frame, the pixel horizontal moving speed and the pixel vertical moving speed, and the real coordinates of the points in the detection boxes, the detection boxes not matched successfully in the previous frame are stabilized.

[0010] According to the pixel coordinates of the four corner points corresponding to the detection boxes matched successfully in the current frame and the previous frame in the orthographic scene, the pixel horizontal moving speed and the pixel vertical moving speed of the detection boxes, the detection boxes matched successfully in the current frame are stabilized.

[0011] Further, the step of obtaining the real coordinates corresponding to the points in each detection box according to the pixel coordinates of the four corner points of the detection boxes in the orthographic scene comprises:

[0012] The pixel coordinates of the four corner points of the detection boxes in the orthographic scene are subjected to perspective transformation, and the real coordinates of the four corner points are obtained according to the preset calibration parameters MPPH and MPPW and the preset matrix fa.

[0013] The real coordinates corresponding to the points in each detection box are obtained according to the real coordinates of the four corner points.

[0014] Further, the step of stabilizing the detection boxes not matched successfully in the previous frame according to the pixel coordinates corresponding to the four corner points of the detection boxes not matched successfully in the previous frame, the pixel horizontal moving speed and the pixel vertical moving speed, and the real coordinates of the points in the detection boxes comprises:

[0015] If it is judged that the target detection is abnormal according to the pixel horizontal moving speed and the pixel vertical moving speed of the target corresponding to the detection boxes not matched successfully in the previous frame and the real coordinates of the points in the detection boxes, the position of the detection boxes not matched successfully in the previous frame is moved according to the pixel horizontal moving speed and the pixel vertical moving speed.

[0016] Further, the step of stabilizing the detection boxes matched successfully in the current frame according to the pixel coordinates of the four corner points corresponding to the detection boxes matched successfully in the current frame and the previous frame in the orthographic scene, the pixel horizontal moving speed and the pixel vertical moving speed of the detection boxes comprises:

[0017] According to the pixel coordinates of the four corners of the detection frame corresponding to the current frame and the previous frame, the pixel transverse moving speed and the pixel longitudinal moving speed of the detection frame, it is determined whether the detection frame jumps or not.

[0018] If the detection frame jumps, the detection frame is stabilized according to the pixel coordinates of the four corners of the detection frame corresponding to the current frame and the previous frame, the pixel transverse moving speed and the pixel longitudinal moving speed of the detection frame.

[0019] Further, the step of stabilizing the detection frame according to the pixel coordinates of the four corners of the detection frame corresponding to the current frame and the previous frame, the pixel transverse moving speed and the pixel longitudinal moving speed of the detection frame includes:

[0020] If the pixel longitudinal moving speed is greater than a preset moving threshold, the position of the detection frame is moved according to the pixel transverse moving speed and the pixel longitudinal moving speed.

[0021] On the other hand, the present application provides a system for obtaining a stable target detection frame based on target detection, which comprises: an obtaining unit for obtaining the pixel coordinates of the four corners of the detection frame in the orthoviewing scene corresponding to each frame of image according to the detection frame in the target detection result of each frame of image; obtaining the real coordinates of the center point of each detection frame and the matching relationship of the detection frame between different image frames according to the pixel coordinates of the four corners of the detection frame in the orthoviewing scene;

[0022] The obtaining unit is further used for obtaining the pixel transverse moving speed and the pixel longitudinal moving speed of the detection frame matched successfully in the current frame, the pixel coordinates of the four corners of the detection frame not matched successfully in the current frame, the pixel coordinates of the four corners of the detection frame not matched successfully in the previous frame, the pixel transverse moving speed, the pixel longitudinal moving speed and the real coordinates of the center point of the detection frame not matched successfully in the previous frame according to the matching relationship of the detection frame between different image frames and the real coordinates of the center point of each detection frame.

[0023] A processing unit is used for stabilizing the detection frame not matched successfully in the previous frame according to the pixel coordinates of the four corners of the detection frame not matched successfully in the previous frame, the pixel transverse moving speed, the pixel longitudinal moving speed and the real coordinates of the center point of the detection frame.

[0024] The processing unit is further used for stabilizing the detection frame matched successfully in the current frame according to the pixel coordinates of the four corners of the detection frame matched successfully in the current frame and the previous frame, the pixel transverse moving speed and the pixel longitudinal moving speed of the detection frame.

[0025] Further, the acquisition unit is specifically configured to perform perspective transformation on pixel coordinates corresponding to four corner points of the detection box in the orthoview scene, and obtain real coordinates of the four corner points according to the preset calibration parameters MPPH and MPPW and the preset matrix fa; and obtain real coordinates corresponding to each center point of the detection box according to the real coordinates of the four corner points.

[0026] Further, the processing unit is specifically configured to move the position of the detection box in the last frame that fails to match if it is determined that the target detection is abnormal according to the horizontal and vertical moving speeds of the target pixel and the real coordinates of the center point of the detection box.

[0027] Further, the processing unit is specifically configured to determine whether the detection box jumps according to the pixel coordinates in the orthoview scene corresponding to each pixel point in the detection box that matches in the current frame and the last frame; and perform stabilization processing on the detection box that matches in the current frame if the detection box jumps, according to the pixel coordinates in the orthoview scene of the four corner points corresponding to the detection box that matches in the current frame and the last frame, the horizontal and vertical moving speeds of the detection box.

[0028] Further, the processing unit is specifically configured to move the position of the detection box that matches in the current frame according to the horizontal and vertical moving speeds of the target pixel if the vertical moving speed is greater than a preset moving threshold.

[0029] The application provides a method and system for obtaining a stable target detection box based on target detection, which obtains the horizontal and vertical moving speeds of the detection box that matches in the current frame, the pixel coordinates of the four corner points of the detection box that fails to match in the current frame, the pixel coordinates of the four corner points of the detection box that fails to match in the last frame, the horizontal and vertical moving speeds of the detection box that fails to match in the last frame, and the real coordinates of the center point of the detection box, according to the matching relationship between the detection boxes in different image frames and the real coordinates corresponding to each center point of the detection box, and corrects the position of the detection box according to the horizontal and vertical moving speeds of the target pixel, thereby correcting the position of the detection box that jumps and improving the accuracy of target detection. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a flowchart of the method for obtaining a stable target detection box based on target detection provided by the application;

[0031] Figure 2 is a structural schematic diagram of the system for obtaining a stable target detection box based on target detection provided by the application. DETAILED DESCRIPTION

[0032] The technical solutions of the present application are described in further detail below with reference to the accompanying drawings and examples.

[0033] As shown in the drawings, Figure 1 The method for obtaining a stable target detection frame based on target detection provided by the embodiments of the present application comprises the following steps:

[0034] 101. Obtain the pixel coordinates of the four corner points of the detection frame in the orthographic scene according to the detection frame in the target detection result of each image frame.

[0035] Specifically, first, obtain the pixel coordinates (px_leftup, py_leftup) of the upper left corner point of the detection frame in the orthographic scene, the width W of the detection frame, and the height H of the detection frame according to the target detection result, so as to obtain the pixel coordinates (px_rightup = px_leftup + W, py_rightup = py_leftup) of the upper right corner point, the pixel coordinates (px_rightdown = px_leftup + W, py_rightdown = py_leftup + H) of the lower right corner point, and the pixel coordinates (px_leftdown = px_leftup, py_leftdown = py_leftup + H) of the lower left corner point.

[0036] 102. Obtain the real coordinates corresponding to the center point of each detection frame and the matching relationship of the detection frame between different image frames according to the pixel coordinates corresponding to the four corner points of the detection frame in the orthographic scene.

[0037] For the embodiments of the present application, the step of obtaining the real coordinates corresponding to the center point of each detection frame according to the pixel coordinates corresponding to the four corner points of the detection frame in the orthographic scene comprises: performing perspective transformation on the pixel coordinates corresponding to the four corner points of the detection frame in the orthographic scene, and obtaining the real coordinates of the four corner points according to the preset calibration parameters MPPH and MPPW and the preset matrix fa; and obtaining the real coordinates corresponding to the center point of each detection frame according to the real coordinates of the four corner points.

[0038] Specifically, for example, the pixel coordinates px_leftup of the upper left corner point is taken as x and py_leftup as y, which are substituted into the formula c = [yx1], cut = c * fa, Y = cut(1) / cut(3), X = cut(2) / cut(3), and the results of the perspective transformation X and Y are respectively assigned to the pixel coordinates (Px_leftup, Py_leftup) of the upper left corner point in the top view of the scene. Similarly, the pixel coordinates (Px_rightup, Py_rightup) of the upper right corner point, the pixel coordinates (Px_rightdown, Py_rightdown) of the lower right corner point, and the pixel coordinates (Px_leftdown, Py_leftdown) of the lower left corner point in the top view of the scene can be obtained.

[0039] Y = (-Y + H) * MPPH + edgeH

[0040] X = X * MPPTW - edgeW where H is the height of the perspective transformation image; edgeH is the longitudinal distance in the camera coordinate system, that is, the actual measured distance between the straight line at the bottom of the perspective transformation and the X axis, edgeW is the transverse distance in the camera coordinate system, that is, the actual measured distance between the straight line at the left of the perspective transformation and the Y axis; MPPH is the real distance represented by each pixel in the x direction in the camera coordinate system, and MPPW is the real distance represented by each pixel in the y direction in the camera coordinate system. The results of the coordinate transformation X and Y are respectively assigned to the real coordinates (Rx_leftup, Ry_leftup) of the upper left corner point. Similarly, the real coordinates (Rx_rightup, Ry_rightup) of the upper right corner point, the real coordinates (Rx_rightdown, Ry_rightdown) of the lower right corner point, and the real coordinates (Rx_leftdown, Ry_leftdown) of the lower left corner point can be obtained. Rx = ((Rx_leftup + Rx_rightup) / 2 + (Rx_leftdown + Rx_rightdown) / 2) / 2, Ry = ((Ry_leftup + Ry_leftdown) / 2 + (Ry_rightup + Ry_rightdown) / 2) / 2, and the obtained Rx and Ry are the real coordinates of the center point of the detection frame.

[0041] 103. According to the matching relationship of the detection frame between different image frames and the real coordinates corresponding to the center point of each detection frame, the pixel transverse and longitudinal moving speeds of the detection frame matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame not matched successfully in the current frame, and the pixel coordinates of the four corner points of the detection frame not matched successfully in the previous frame, the pixel transverse and longitudinal moving speeds of the detection frame not matched successfully in the previous frame, and the real coordinates of the center point of the detection frame are obtained.

[0042] Specifically, first, the detection frame region of the target of the current frame can be obtained, if the detection frame of the last frame is in the region, as long as one corner point is in the region, the detection frame in the region is considered to meet the condition, and the number of detection frames meeting the condition is recorded as N. Then, the overlapping regions of the N detection frames and the detection frame of the target of the current frame are calculated respectively, the target of the last frame detection frame with the maximum overlapping region and exceeding the overlapping region threshold is the same target as the target of the current frame detection frame, that is, the target matching relationship of multiple frames is completed; if the above conditions are not met, it is considered to be a new target.If the matching is successful, the horizontal pixel moving speed Pxv and the vertical pixel moving speed Pyv of the same target camera frame are calculated, taking the camera frame center point as the reference point: the pixel coordinates of the four corners of the camera frame in the last frame are as follows: the pixel coordinates of the upper left corner (Pxm_leftup, Pym_leftup), the pixel coordinates of the upper right corner (Pxm_rightup, Pym_rightup), the pixel coordinates of the lower right corner (Pxm_rightdown, Pym_rightdown), and the pixel coordinates of the lower left corner (Pxm_leftdown, Pym_leftdown), the camera frame center point coordinates (Pxm, Pym) are Pxm = ((Pxm_leftup + Pxm_rightup) / 2 + (Pxm_leftdown + Pxm_rightdown) / 2) / 2, Pym = ((Pym_leftup + Pym_leftdown) / 2 + (Pym_rightup + Pym_rightdown) / 2) / 2, the pixel coordinates of the four corners of the camera frame in the current frame are as follows: the pixel coordinates of the upper left corner (Pxn_leftup, Pyn_leftup), the pixel coordinates of the upper right corner (Pxn_rightup, Pyn_rightup), the pixel coordinates of the lower right corner (Pxn_rightdown, Pyn_rightdown), and the pixel coordinates of the lower left corner (Pxn_leftdown, Pyn_leftdown), the camera frame center point coordinates (Pxn, Pyn) are Pxn = ((Pxn_leftup + Pxn_rightup) / 2 + (Pxn_leftdown + Pxn_rightdown) / 2) / 2, Pyn = ((Pyn_leftup + Pyn_leftdown) / 2 + (Pyn_rightup + Pyn_rightdown) / 2) / 2, the horizontal pixel moving speed Pxvn of the target camera frame is (Pxn-Pxm) / T, wherein T is the time of each frame, the vertical pixel moving speed Pyvn is (Pyn-Pym) / T, wherein T is the time of each frame, and thus the horizontal moving speed Pxv_vector and the vertical moving speed Pyv_vector of the multiple target pixels matched successfully can be obtained, wherein the real-time update is performed according to each frame.

[0043] 104. According to the pixel coordinates corresponding to the four corners of the detection frame which is not matched successfully in the last frame, the horizontal moving speed, the vertical moving speed, and the real coordinates of the detection frame center point, the detection frame which is not matched successfully in the last frame is stably processed.

[0044] For the embodiment of the present application, step 104 can specifically include: if it is determined that the target detection is abnormal according to the horizontal moving speed, the vertical moving speed of the target pixel and the real coordinates of the center point of the detection frame corresponding to the detection frame of the previous frame which fails to match, then moving the position of the detection frame of the previous frame which fails to match according to the horizontal moving speed and the vertical moving speed of the target pixel.

[0045] Specifically, for example, the pixel coordinates Pxm of the center point of the detection frame are obtained from the four corner points of the detection frame of the previous frame which fails to match, Pxm=(Pxm_leftup+Pxm_rightup) / 2+(Pxm_leftdown+Pxm_rightdown) / 2) / 2, Pym=(Pym_leftup+Pym_leftdown) / 2+(Pym_rightup+Pym_rightdown) / 2) / 2, the empirical pixel speed threshold is gatepx, gatepy, if abs(Pxvm)<gatepx+abs(Pyvm)>gatepy and abs(Rxm)<50+0<Rym<200 are satisfied at the same time, it indicates that the target is moving straight in the scene at this time, but the target is not detected in the current frame, which means that the target is missed. At this time, the four corner points of the new camera frame are predicted as follows: taking the left upper corner point as an example, the other corner points are the same Pxn_leftup=Pxm_leftup+Pxvm*T, Pyn_leftup=Pym_leftup+Pyvm*T.

[0046] 105. Stabilizing the detection frame matched successfully in the current frame according to the pixel coordinates in the orthoscopic scene, the horizontal moving speed and the vertical moving speed of the detection frame of the four corner points corresponding to the detection frame matched successfully in the current frame and the detection frame of the previous frame.

[0047] For the embodiment of the present application, step 105 can specifically include: judging whether the detection frame is jumping according to the pixel coordinates in the orthoscopic scene of each pixel point in the detection frame matched successfully in the current frame and the detection frame of the previous frame; if yes, then stabilizing the detection frame matched successfully in the current frame according to the pixel coordinates in the orthoscopic scene, the horizontal moving speed and the vertical moving speed of the detection frame of the four corner points corresponding to the detection frame matched successfully in the current frame and the detection frame of the previous frame. The step of stabilizing the detection frame matched successfully in the current frame according to the pixel coordinates in the orthoscopic scene, the horizontal moving speed and the vertical moving speed of the detection frame of the four corner points corresponding to the detection frame matched successfully in the current frame and the detection frame of the previous frame includes: if the vertical moving speed is greater than a preset moving threshold, then moving the position of the detection frame matched successfully in the current frame according to the horizontal moving speed and the vertical moving speed of the target pixel.

[0048] Specifically, for example, the camera frame of the target in the current frame and the previous frame are both divided into 1000*1000 small grids, it is considered that the pixel points (i, j) located at the i-th column and the j-th row in the camera frame, the coordinates of which are (i, j) (i=500 takes the middle column to determine, j takes values from 1 to 1001), are the same spatial positions of the vehicle in the camera frame of the previous and next frames, and thus the longitudinal pixel difference of the coordinates can be calculated to determine whether the camera jumps forward and backward. Vector (i, j) = (Pyn_leftup+Hn / 1000*(i-1))-(Pym_leftup+Hm / 1000*(i-1)), thus 1001 longitudinal pixel differences can be obtained, if Pyvm*Vector (i, 1)*Vector (i, 2)*…*Vector (i, 1001)>0, it is indicated that the directions are the same and the longitudinal pixel movement direction is the same as that of the previous frame, if it is not satisfied, it is considered that the forward and backward jumps occur, and the detection frame is predicted: taking the upper left corner point as an example, the other corner points are the same, if abs (Pyvm)>gatepy, it is considered that the moving vehicle jumps forward and backward, and the prediction is performed: Pxn_leftup=Pxm_leftup+Pxvm*T, Pyn_leftup=Pym_leftup+Pyvm*T; if abs (Pyvm)>gatepy is not satisfied, it is considered that the stationary vehicle jumps forward and backward, and the detection frame position of the previous frame is kept unchanged: Pxn_leftup=Pxm_leftup, Pyn_leftup=Pym_leftup.

[0049] Further, the detection frame information of the current frame matched successfully after the stable processing, and / or the target detection frame information of the previous frame unmatched successfully after the stable processing, and / or the detection frame information of the current frame unmatched successfully can be used as the final target camera frame information of the current frame.

[0050] The method for obtaining a stable target detection frame provided by the embodiment of the application is based on target detection, and the pixel horizontal movement speed and the pixel longitudinal movement speed of the detection frame matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame unmatched successfully in the current frame, and the pixel coordinates of the four corner points of the detection frame unmatched successfully in the previous frame, the pixel horizontal movement speed and the pixel longitudinal movement speed of the detection frame unmatched successfully in the previous frame, and the real coordinates of the center point of the detection frame are obtained according to the matching relationship between the detection frames in different image frames and the real coordinates of the center point of each detection frame, and the position of the detection frame is corrected according to the pixel horizontal movement speed and the pixel longitudinal movement speed in different cases, so that the position of the detection frame with jumps can be corrected, and the accuracy of target detection is improved.

[0051] To realize the method provided by the embodiment of the present application, the embodiment of the present application provides an acquisition system for a stable target detection frame based on target detection, as shown in Figure 2 The system comprises an acquisition unit 21 and a processing unit 22.

[0052] The acquisition unit 21 is configured to acquire pixel coordinates of four corner points of a detection frame in an orthographic scene of each frame of image according to the detection frame in the target detection result of each frame of image, and acquire real coordinates of a center point of each detection frame and a matching relationship of the detection frames between different image frames according to the pixel coordinates of the four corner points of the detection frame in the orthographic scene.

[0053] The acquisition unit 21 is further configured to acquire a horizontal moving speed and a vertical moving speed of a detection frame that is successfully matched in a current frame, pixel coordinates of four corner points of a detection frame that is not successfully matched in the current frame, and pixel coordinates of the four corner points of the detection frame that is not successfully matched in a previous frame, the horizontal moving speed, the vertical moving speed, and the real coordinates of the center point of the detection frame that is not successfully matched in the previous frame according to the matching relationship of the detection frames between different image frames and the real coordinates of the center point of each detection frame.

[0054] The processing unit 22 is configured to perform a stabilization process on the detection frame that is not successfully matched in the previous frame according to the pixel coordinates of the four corner points of the detection frame that is not successfully matched in the previous frame, the horizontal moving speed, the vertical moving speed, and the real coordinates of the center point of the detection frame.

[0055] The processing unit 22 is further configured to perform a stabilization process on the detection frame that is successfully matched in the current frame according to the pixel coordinates of the four corner points of the detection frame that is successfully matched in the current frame and the previous frame, the horizontal moving speed, and the vertical moving speed of the detection frame.

[0056] Further, the acquisition unit 21 is specifically configured to perform a perspective transformation on the pixel coordinates of the four corner points of the detection frame in the orthographic scene, and obtain real coordinates of the four corner points according to preset calibration parameters MPPH and MPPW and a preset matrix fa; and obtain the real coordinates of the center point of each detection frame according to the real coordinates of the four corner points.

[0057] Further, the processing unit 22 is specifically configured to move the position of the detection frame that is not successfully matched in the previous frame according to the horizontal moving speed and the vertical moving speed of the target if it is determined that the target detection is abnormal according to the horizontal moving speed, the vertical moving speed, and the real coordinates of the center point of the detection frame corresponding to the target.

[0058] Further, the processing unit 22 is specifically further configured to determine whether there is detection frame jitter according to the pixel coordinates of each pixel point in the detection frame matched successfully between the current frame and the previous frame in the orthoview scene; if there is, then perform stabilization processing on the detection frame matched successfully in the current frame according to the pixel coordinates of the four corner points of the detection frame matched successfully between the current frame and the previous frame, the horizontal movement speed and the vertical movement speed of the detection frame.

[0059] Further, the processing unit 22 is specifically further configured to move the position of the detection frame matched successfully in the current frame according to the horizontal movement speed and the vertical movement speed of the target pixel if the vertical movement speed is greater than a preset movement threshold.

[0060] The embodiment of the application provides an acquisition system for a stable target detection frame based on target detection, which acquires the horizontal movement speed and the vertical movement speed of the detection frame matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame not matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame not matched successfully in the previous frame, the horizontal movement speed and the vertical movement speed of the detection frame not matched successfully in the previous frame, and the real coordinates of the center point of the detection frame, and corrects the position of the detection frame according to the horizontal movement speed and the vertical movement speed of the target pixel, so that the position of the detection frame with jitter can be corrected, and the accuracy of target detection is improved.

[0061] It should be understood that the specific order or hierarchy of steps in the processes disclosed should not be taken as a reflection of the exclusive logic flow for implementing the subject matter. Rather, the specific order or hierarchy of steps in the processes disclosed can be re-arranged or reordered in other implementations without departing from the scope of the disclosure. The accompanying method claims present elements of the various steps in a sample order, and are not meant to be limited to the specific order or hierarchy presented.

[0062] In the above detailed description, various features are grouped together in single embodiments for the purpose of streamlining the disclosure. Such disclosed methods should not be interpreted as reflecting an intention that the claimed subject matter requires more features than are explicitly stated in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of single disclosed embodiments. Thus, the following claims are hereby expressly incorporated into this detailed description, with each claim acting as a separate embodiment of the inventive subject matter.

[0063] The foregoing description of the exemplary embodiments of this application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description, but rather by the claims appended hereto.

[0064] The above description includes exemplary embodiments. Of course, not all possible combinations of components or method steps are described above, but one of ordinary skill in the art will recognize that further combinations are possible. Persons of ordinary skill in the art will also recognize that the various embodiments described above can be further modified than described, and thus all modifications and further combinations are believed to be encompassed in the scope of the claims.

[0065] Those of skill would further appreciate that the various illustrative logical blocks, modules, and steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present embodiments.

[0066] The various illustrative logical blocks, modules, and steps described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the general purpose processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other such configuration.

[0067] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.

[0068] In one or more exemplary designs, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. Storage media can be any available media that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code means in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or data

[0069] The above detailed description describes the purpose, technical solutions and advantages of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the present application.

Claims

1. An acquisition method of a stable target detection frame based on target detection, characterized in that, The method comprises: According to the detection frame in each frame image target detection result, the pixel coordinates corresponding to the four corner points of the detection frame in the orthoscopic scene are obtained; According to the pixel coordinates corresponding to the four corner points of the detection frame in the orthoscopic scene, the real coordinates corresponding to the center point of each detection frame and the matching relationship of the detection frames between different image frames are obtained; According to the matching relationship of the detection frames between different image frames and the real coordinates corresponding to the center point of each detection frame, the pixel horizontal moving speed and the vertical moving speed of the detection frame matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame not matched successfully in the current frame, the pixel coordinates of the four corner points of the detection frame not matched successfully in the last frame, the pixel horizontal moving speed and the vertical moving speed of the detection frame not matched successfully in the last frame, and the real coordinates of the center point of the detection frame are obtained; According to the pixel coordinates corresponding to the four corner points of the detection frame not matched successfully in the last frame, the pixel horizontal moving speed and the vertical moving speed of the target, and the real coordinates of the center point of the detection frame, the detection frame not matched successfully in the last frame is processed stably. According to the pixel coordinates of the four corner points corresponding to the detection frame matched successfully in the current frame and the last frame in the orthoscopic scene, the pixel horizontal moving speed and the vertical moving speed of the detection frame, the detection frame matched successfully in the current frame is processed stably.

2. The method of claim 1, wherein the method further comprises: The step of obtaining the real coordinates corresponding to the center point of each detection frame according to the pixel coordinates of the four corner points of the detection frame in the orthoscopic scene comprises: The perspective transformation is performed on the pixel coordinates of the four corner points of the detection frame in the orthoscopic scene, and the real coordinates of the four corner points are obtained according to the preset calibration parameters MPPH and MPPW and the preset matrix fa; The real coordinates corresponding to the center point of each detection frame are obtained according to the real coordinates of the four corner points.

3. The method of claim 1, wherein the method further comprises: The step of processing the detection frame not matched successfully in the last frame stably according to the pixel coordinates corresponding to the four corner points of the detection frame not matched successfully in the last frame, the pixel horizontal moving speed and the vertical moving speed of the target, and the real coordinates of the center point of the detection frame comprises: If it is judged that the target detection is abnormal according to the pixel horizontal moving speed and the vertical moving speed of the target corresponding to the detection frame not matched successfully in the last frame and the real coordinates of the center point of the detection frame, the position of the detection frame not matched successfully in the last frame is moved according to the pixel horizontal moving speed and the vertical moving speed of the target.

4. The method of claim 1, wherein the method further comprises: The step of processing the detection frame matched successfully in the current frame stably according to the pixel coordinates of the four corner points corresponding to the detection frame matched successfully in the current frame and the last frame in the orthoscopic scene, the pixel horizontal moving speed and the vertical moving speed of the detection frame comprises: Whether the detection frame jumps is judged according to the pixel coordinates of the four corner points corresponding to the detection frame matched successfully in the current frame and the last frame in the orthoscopic scene; If yes, the detection frame matched successfully in the current frame is processed stably according to the pixel coordinates of the four corner points corresponding to the detection frame matched successfully in the current frame and the last frame in the orthoscopic scene, the pixel horizontal moving speed and the vertical moving speed of the detection frame.

5. The method of claim 4, wherein the method further comprises: The four corner points of the detection frame matched successfully between the current frame and the previous frame correspond to pixel coordinates in the orthographic scene, a horizontal movement speed and a vertical movement speed of the detection frame, and the step of performing stabilization processing on the detection frame matched successfully in the current frame includes: If the vertical movement speed is greater than a preset movement threshold, then the position of the detection frame matched successfully in the current frame is moved according to the horizontal movement speed and the vertical movement speed of the target pixel.

6. An acquisition system for a stable target detection box based on target detection, characterized in that, The system includes: An acquisition unit is configured to acquire, according to the detection frame in the target detection result of each frame of image, pixel coordinates corresponding to four corner points of the detection frame in the orthographic scene; and acquire, according to the pixel coordinates corresponding to the four corner points of the detection frame in the orthographic scene, real coordinates corresponding to a center point of each detection frame and a matching relationship between the detection frames in different image frames. The acquisition unit is further configured to acquire, according to the matching relationship between the detection frames in different image frames and the real coordinates corresponding to the center point of each detection frame, a horizontal movement speed and a vertical movement speed of the detection frame matched successfully in the current frame, pixel coordinates of the four corner points of the detection frame not matched successfully in the current frame, and pixel coordinates of the four corner points of the detection frame not matched successfully in the previous frame, a horizontal movement speed and a vertical movement speed of the detection frame not matched successfully in the previous frame, and real coordinates of the center point of the detection frame. A processing unit is configured to perform stabilization processing on the detection frame not matched successfully in the previous frame according to the pixel coordinates corresponding to the four corner points of the detection frame not matched successfully in the previous frame, the horizontal movement speed and the vertical movement speed of the target pixel, and the real coordinates of the center point of the detection frame. The processing unit is further configured to perform stabilization processing on the detection frame matched successfully in the current frame according to the pixel coordinates of the four corner points of the detection frame matched successfully between the current frame and the previous frame in the orthographic scene, the horizontal movement speed and the vertical movement speed of the detection frame.

7. The acquisition system for stabilizing a target detection frame based on target detection according to claim 6, wherein The acquisition unit is specifically configured to perform perspective transformation on the pixel coordinates corresponding to the four corner points of the detection frame in the orthographic scene, and obtain real coordinates of the four corner points according to preset calibration parameters MPPH and MPPW and a preset matrix fa; and obtain real coordinates corresponding to a center point of each detection frame according to the real coordinates of the four corner points.

8. The acquisition system for stabilizing a target detection frame based on target detection according to claim 6, wherein The processing unit is specifically configured to move the position of the detection frame not matched successfully in the previous frame according to the horizontal movement speed and the vertical movement speed of the target pixel if it is determined that the target detection is abnormal according to the horizontal movement speed and the vertical movement speed of the target pixel corresponding to the detection frame not matched successfully in the previous frame and the real coordinates of the center point of the detection frame.

9. The system for obtaining a stable bounding box based on target detection according to claim 6, wherein, The processing unit is further configured to determine whether there is detection frame jitter according to the pixel coordinates of the front view scene corresponding to each pixel point in the detection frame successfully matched between the current frame and the previous frame; if there is, perform stabilization processing on the detection frame successfully matched in the current frame according to the pixel coordinates of the front view scene corresponding to the four corner points of the detection frame successfully matched between the current frame and the previous frame, the horizontal movement speed and the vertical movement speed of the detection frame.

10. The system for obtaining a stable detection frame based on target detection according to claim 9, characterized in that, The processing unit is further configured to move the position of the detection frame successfully matched in the current frame according to the horizontal movement speed and the vertical movement speed of the target pixel if the vertical movement speed is greater than a preset movement threshold.

Citation Information

Patent Citations

  • A target tracking method and computing device

    CN109544590A

  • An improved method for improving the stability of video object detection

    CN109902620A