Flow statistics method, electronic device and storage medium for target object
By setting the statistical cross-section in the video and setting the frame threshold, we can determine whether the position of the target object intersects the statistical cross-section, which solves the problem of counting errors in the prior art, and improves the accuracy and reliability of traffic statistics.
Patent Information
- Application Number
- CN202211003215.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-08-19
AI Technical Summary
Existing video detection methods are prone to errors in target object counting in traffic-intensive video data, resulting in inaccurate statistical results.
By setting the statistical cross-section in the video, we can judge whether the position of the target object intersects the statistical cross-section between continuous frame images, and set the frame threshold for logical judgment to prevent the target object from repeatedly jumping horizontally near the statistical cross-section or misjudgment caused by screen jitter, and improve counting accuracy.
It effectively prevents the target object from repeatedly jumping horizontally near the statistical cross-section or screen jitter, and improves the accuracy and reliability of traffic statistics.
Smart Images

Figure CN115471798B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic statistics, and particularly to a traffic statistics method for target objects, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of science and technology and the advent of the big data era, empowering transportation through computer vision technology and video data can assist traffic system managers in achieving efficient traffic dispatching, regulating traffic behavior, and thus reducing traffic accidents, which can effectively improve the urban traffic situation.
[0003] In a traffic scenario, a method of counting moving target objects based on a video stream is also a way to quantify traffic congestion or load conditions.
[0004] For video data with a large traffic volume, due to the high traffic density, in the detection process of such data by existing video detection, there are often errors in counting target objects, resulting in inaccurate statistical results. Summary of the Invention
[0005] The present invention provides a traffic statistics method for target objects, an electronic device, and a storage medium to solve the problem of inaccurate traffic statistics results.
[0006] To solve the above technical problems, the present invention provides a traffic statistics method for target objects, including: obtaining a video including a plurality of consecutive frames of images and setting a statistical cross-section of a target area in the video; performing target detection on the target area of each frame of image to obtain target objects located in the target area on each frame of image and the positions of the target objects; in response to the lines connecting the positions of the target objects on each image from the previous preset frame of image to the previous frame of image to the position of the target object in the current frame of image all intersecting the statistical cross-section, counting the target object in the current frame of image as a valid count; counting the target objects on the next frame of image until the target objects in the target area of each frame of image are counted to obtain the traffic statistics result of the video.
[0007] Among them, in response to the lines connecting the positions of the target objects on each image from the previous preset frame of image to the previous frame of image to the position of the target object in the current frame of image all intersecting the statistical cross-section, counting the target object in the current frame of image as a valid count, includes: predicting the predicted position of the target object in the current frame of image; matching the predicted position with the position of the target object in the current frame of image to determine the matching degree; when the matching degree exceeds the preset matching degree, in response to the lines connecting the positions of the target objects on each image from the previous preset frame of image to the previous frame of image to the position of the target object in the current frame of image all intersecting the statistical cross-section, counting the target object in the current frame of image as a valid count.
[0008] Among them, predicting the predicted position of the target object in the current frame image includes: predicting the predicted position of the target object in the current frame image based on the historical information of the target object; where the historical information includes the position of the target object in the historical frame images of the target object; after counting the target object in the current frame image as a valid count, it includes: updating the position of the target object in the current frame image to the historical information.
[0009] Among them, performing object detection on the target regions of each frame image to obtain the target objects and their positions located in the target regions of each frame image includes: dividing each frame image to obtain multiple image blocks of each frame image; respectively performing object detection on the multiple image blocks of each frame image to obtain the target objects on each image block; synthesizing the target objects on the image blocks and filtering out duplicate target objects to obtain the target objects of each frame image; determining the target objects located in the target regions of each frame image and the positions of the target objects.
[0010] Among them, dividing each frame image to obtain multiple image blocks of each frame image includes: respectively dividing each frame image to obtain multiple initial image blocks of the same size for each frame image; on each frame image, respectively expanding the sizes of the initial image blocks in a first direction and a second direction to obtain multiple image blocks of each frame image; where the first direction and the second direction are perpendicular to each other.
[0011] Among them, on each frame image, respectively expanding the sizes of the initial image blocks in a first direction and a second direction to obtain multiple image blocks of each frame image further includes: in response to the expanded initial image block exceeding the boundary of the corresponding frame image, determining the pixel values of the area where the expanded initial image block exceeds the boundary as 0.
[0012] Among them, obtaining a video including consecutive multiple frame images and setting a statistical cross-section of the target region in the video includes: obtaining a video including consecutive multiple frame images and at least one initial target region; in response to the initial target region being a convex polygon, determining the initial target region as the target region; determining the plane perpendicular to the moving direction of the target object in the target region as the statistical cross-section; counting the target objects in the next frame image until the target objects in the target regions of all frame images are counted to obtain the traffic statistical result of the video, including: counting the target objects in the next frame image until all the target objects in the target regions of all frame images are counted to obtain the traffic statistical results of each target region in the video.
[0013] Among them, before counting the target object in the current frame image as a valid count in response to the fact that the connection lines between the positions of the target objects on each image from the previous preset frame image to the previous frame image and the position of the target object in the current frame image all intersect the statistical cross-section, it further includes: in response to the connection line between the position of the target object in the current frame image and the position of the target object in the previous frame image of the current frame image intersecting the statistical cross-section, recording the position of the target object in the current frame image; recording the target objects on the next frame image until the positions of the target objects in each frame image are recorded.
[0014] To solve the above technical problems, the present invention also provides an electronic device, which includes: a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the traffic statistics method of the target object in any one of the above.
[0015] To solve the above technical problems, the present invention also provides a computer-readable storage medium, which stores program data that can be executed to implement the traffic statistics method of the target object as described in any one of the above.
[0016] The beneficial effects of the present invention are: different from the prior art, the present invention determines the statistical cross-section of the target area in the video; performs target detection on the target areas of each frame image to obtain the target objects located in the target area on each frame image and their positions; in response to the connection lines between the positions of the target objects on each image from the previous preset frame image to the previous frame image and the position of the target object in the current frame image all intersecting the statistical cross-section, counting the target object in the current frame image as a valid count; counting the target objects on the next frame image until the target objects in the target areas of each frame image are counted to obtain the traffic statistics result of the video, thereby setting a frame number threshold for logical judgment, and further preventing the target object from repeatedly jumping near the statistical cross-section, or being misjudged due to factors such as vehicle or picture jitter, and further improving the accuracy of the traffic statistics of the target object. Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of an embodiment of the traffic statistics method of the target object provided by the present invention;
[0018] Figure 2 It is a schematic flowchart of another embodiment of the traffic statistics method of the target object provided by the present invention;
[0019] Figure 3 It is Figure 2 A schematic diagram of an implementation manner of cropping each frame image in the embodiment;
[0020] Figure 4 It is Figure 2Schematic diagram of an embodiment of effective count statistics in an embodiment;
[0021] Figure 5 It is a schematic framework diagram of an embodiment of a traffic statistics device for the target object of the present invention;
[0022] Figure 6 It is a schematic structural diagram of an embodiment of an electronic device provided by the present invention;
[0023] Figure 7 It is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present invention. Specific embodiments
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of a traffic statistics method for the target object provided by the present invention.
[0026] Step S11: Obtain a video including a plurality of consecutive frames of images, and set a statistical cross-section of the target area in the video.
[0027] First, obtain a video including a plurality of consecutive frames of images. Among them, the target area to be counted for traffic can be photographed by a drone, a camera from an aerial perspective, or other imaging devices to obtain a continuous video. The target area in the video can be specified based on the actual situation, which is not limited here. The number of target areas in a video can be one or more, which is not limited here.
[0028] Set a statistical cross-section of the target area in the video. Among them, the statistical cross-section in this embodiment refers to the plane perpendicular to the flow direction of the target object in the target area. The specific position of the statistical cross-section in the direction perpendicular to the target area can be determined based on the actual situation, which is not limited here. Among them, in the traffic statistics of a specific video, once the statistical cross-section is determined, it is fixed and will not change. For example: when this embodiment is applied to traffic flow statistics, the target area can be a lane, and the statistical cross-section can be a plane perpendicular to the lane.
[0029] Step S12: Perform target detection on the target area of each frame of image to obtain the target object located in the target area on each frame of image and the position of the target object.
[0030] The target object of this embodiment is an object that needs to be counted for traffic, which may include moving objects such as vehicles, people, animals, etc., and is not specifically limited here.
[0031] Target detection is performed on the target areas of each image frame in the video to obtain the target objects located within the target areas on each frame image and their positions.
[0032] In a specific application scenario, the target areas on each frame image can be detected by a pre-trained target detection model to obtain the target objects located within the target areas on each frame image and their positions. In another specific application scenario, the target areas on each frame image can be detected by target detection technology to obtain the target objects located within the target areas on each frame image and their positions. Among them, the target detection technology can include YOLO-v4 target detection technology, Faster R-CNN target detection technology or other target detection technologies, etc., and is not limited here.
[0033] Step S13: In response to the lines connecting the positions of the target objects on each image between the pre-set previous frame image of the current frame image and the previous frame image of the current frame image all intersecting the statistical cross-section, count the target object in the current frame image as a valid count.
[0034] After obtaining the target objects located within the target areas on each frame image and their positions, in response to the lines connecting the positions of the target objects on each image between the pre-set previous frame image of the current frame image and the previous frame image of the current frame image all intersecting the statistical cross-section, count the target object in the current frame image as a valid count. Among them, the number of frames of the pre-set previous frame image can include the previous fifteen frames, the previous twenty frames, the previous twenty-two frames, etc., and the specific number of frames can be set based on the actual situation and is not limited here. Among them, the larger the number of frames of the pre-set previous frame image, the more accurate the valid count.
[0035] By determining that the lines connecting the positions of the target objects on each image between the pre-set previous frame image of the current frame image and the previous frame image of the current frame image all intersect the statistical cross-section, the target object is counted as a valid count, thereby strictifying the counting conditions, reducing the influence of factors such as the target object repeatedly jumping near the statistical cross-section, or vehicle or picture jitter on the valid count, and improving the reliability of traffic statistics.
[0036] In response to the lines connecting the positions of the target objects on each image between the pre-set previous frame image and the previous frame image not all intersecting the statistical cross-section, do not count the target object in the current frame image.
[0037] In a specific application scenario, when this embodiment is applied to traffic flow statistics and the previous preset frame images are the first twenty frame images, in response to the fact that the lines connecting the positions of the target objects in each of the 20 frame images between the first twenty frame images and the previous frame image to the positions of the corresponding target objects in the current frame image, that is, all 20 lines intersect the statistical section, the target objects in the current frame image are counted as valid counts.
[0038] Based on the above statistical method for setting the frame number threshold, determine whether all the target objects on the current frame image are valid counts, and complete the traffic flow statistics of the target objects on the current frame image.
[0039] In this step, the target object is counted as a valid count only when the lines connecting the positions of the target objects on each image between the previous preset frame image and the previous frame image to the positions of the corresponding target objects in the current frame image all intersect the statistical section, so as to set the frame number threshold for logical judgment and improve the accuracy of the traffic flow statistics of the target objects.
[0040] Step S14: Count the target objects on the next frame image until the target objects in the target areas of all frames of images are counted, and obtain the traffic flow statistics result of the video.
[0041] Continue to count the target objects on the next frame image according to the method in step S13, that is, use the next frame image as the current frame image for statistics, and so on, until the target objects in the target areas of all frames of images in the video are counted, and obtain the traffic flow statistics result of the entire video.
[0042] Among them, the traffic flow statistics result of the video in this embodiment can be displayed in the form of a table, a line chart or other forms based on each frame, for the traffic flow statistics numbers corresponding to each frame; it can also be displayed after post-processing based on the traffic flow statistics numbers corresponding to each frame. Among them, the post-processing can include processing such as averaging and selecting some representative data, which is not limited here.
[0043] In a specific application scenario, if the target object is a valid count in the current frame image, it means that the position of the target object in the current frame image is on one side of the statistical cross-section, and all positions in the previous preset frame images are on the other side of the statistical cross-section. When counting the target object in the next frame image, regardless of which side of the statistical cross-section the target object in the next frame image is on, there will be at least one connection line between the positions of the target objects on the images from the previous preset frame image of the next frame image to the previous frame of the next frame image and the position of the target object in the next frame image that is completely on the same side of the statistical cross-section and cannot intersect the statistical cross-section, thus not meeting the statistical rules of valid counting. The counting of subsequent frames is similar. Therefore, this embodiment can achieve the effect of only performing a valid count on a target object once and preventing repeated counting of the same target object. Furthermore, it can prevent the target object from repeatedly jumping near the statistical cross-section or being misjudged due to factors such as vehicle or picture jitter, improving the accuracy and reliability of the traffic statistics result.
[0044] Through the above steps, the traffic statistics method for the target object in this embodiment sets a statistical cross-section for the target area in the video; performs target detection on the target areas of each frame image to obtain the target objects located in the target area of each frame image and their positions; in response to the connection lines between the positions of the target objects on the images from the previous preset frame image of the current frame image to the previous frame of the current frame image and the position of the target object in the current frame image all intersecting the statistical cross-section, counts the target object in the current frame image as a valid count; counts the target objects in the next frame image until the target objects in the target areas of all frame images are counted to obtain the traffic statistics result of the video, thereby setting a frame number threshold for logical judgment, and further preventing the target object from repeatedly jumping near the statistical cross-section or being misjudged due to factors such as vehicle or picture jitter, and further improving the accuracy rate of the traffic statistics of the target object.
[0045] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the traffic statistics method for the target object provided by the present invention.
[0046] Step S21: Obtain a video including a series of consecutive frame images and set a statistical cross-section for the target area in the video.
[0047] First, obtain a video including a series of consecutive frame images and determine the statistical cross-section of the target area in the video. Among them, the statistical cross-section refers to the plane perpendicular to the flow direction of the target object in the target area. The specific position of the statistical cross-section in the direction perpendicular to the target area can be determined based on the actual situation and is not limited here. The target area in the video can be specified based on the actual situation and is not limited here.
[0048] In a specific application scenario, this embodiment can obtain a video including a series of consecutive frames of images and receive at least one initial target area specified by the user at the front end; in response to the initial target area being a convex polygon, determine the initial target area as the target area, and respectively determine the planes perpendicular to the moving directions of the target objects in each target area as the statistical cross-sections; in response to the initial target area not being a convex polygon, do not determine the initial target area as the target area. When there are multiple target areas, there are also the same number of corresponding statistical cross-sections.
[0049] Among them, a convex polygon is a polygon with an interior that is a convex set, which can include rectangles, circles, polygons, etc. When the target area is a convex polygon, when the target object moves within the target area, it will only enter and exit the convex polygon singly, which is convenient for determining the flow direction of the target object. When the target area is not a convex polygon, when the target object moves within the target area, it may enter and exit multiple times, affecting the determination of the flow direction of the target object.
[0050] By receiving the user's division of the initial target area at the front end, it is possible to provide the user with the function of customizing the flow area and range of interest, improving the pertinence of the flow statistics method for target objects.
[0051] In a specific application scenario, after obtaining a video including a series of consecutive frames of images, the coordinate axes can be set based on the video. For example: set the upper left corner of the image as the origin (0, 0) of the image coordinate system, the horizontal direction as the x-axis, and the vertical direction as the y-axis.
[0052] Step S22: Divide each frame of the image to obtain multiple image blocks of each frame of the image; respectively perform target detection on the multiple image blocks of each frame of the image to obtain the target objects on each image block; synthesize the target objects on the image blocks and filter out duplicate target objects to obtain the target objects of each frame of the image; determine the target objects located within the target area in each frame of the image.
[0053] After obtaining a video including a series of consecutive frames of images, each frame of the image in the video is divided respectively to obtain multiple initial image blocks of the same size for each frame of the image.
[0054] In a specific application scenario, for each input frame of image I with width W and height H, first perform a division operation to obtain n initial image blocks of a fixed size w×h, where the values of w and h can be determined according to the specific chip performance.
[0055] Denote the i-th initial image block in image I as I i , with width w and height h. Among them, i ∈ [1, n], then the calculation formula for the number n of initial image blocks in image I is:
[0056]
[0057] Among them, n is the number of initial image blocks in image I, W and H are the width and height of image I, w and h are the width and height of the initial image blocks, and all divisions are rounded down.
[0058] After dividing each frame of image into multiple initial image blocks, in order to avoid that the target object on the boundary may be split into two initial image blocks during the process of cropping into initial image blocks, which affects the accuracy of target detection. Therefore, on each frame of image, the sizes of the initial image blocks are respectively expanded in the first direction and the second direction to obtain multiple image blocks of each frame of image. Among them, the first direction and the second direction are perpendicular to each other. For example: the first direction can be the horizontal direction of the video, and the second direction can be the vertical direction of the video.
[0059] In response to the fact that the expanded initial image block exceeds the boundary of the corresponding frame of image, the pixel values in the area where the expanded initial image block exceeds the boundary are determined to be 0.
[0060] Please refer to Figure 3 , Figure 3 is Figure 2 a schematic diagram of an implementation manner of cropping each frame of image in the embodiment.
[0061] The width and height of image 30 in this embodiment are W and H respectively. First, image 30 is divided into multiple initial image blocks 31. Among them, the width and height of the initial image block 31 are w and h respectively.
[0062] On the basis of the initial image block 31, it is expanded by dw to the right in the horizontal direction and by dh downward in the vertical direction. When it exceeds the original image boundary, it is filled with 0 to obtain the image block 32.
[0063] Among them, dw and dh in this implementation manner can be respectively set to 10%-15% of the corresponding width and height of the initial image block 31. For example: 10%, 11%, 12%, 13%, 14% or 15%, etc., which are not limited here.
[0064] When the video is a high-resolution video, its viewing angle is relatively wide, the target object accounts for a relatively small proportion, the amount of image data for full-image detection is too large, and the requirements for device performance are relatively high. And when the image is scaled proportionally, the target is very likely to be lost, which affects the accuracy. Therefore, through the above image block method, this step can increase the proportion of small targets in the detected input image and improve the detection rate of target objects in high-resolution images. And the above method can also be compatible with video images of conventional resolutions, improving the accuracy and reliability of target detection for video images of conventional resolutions.
[0065] In a specific embodiment, after cropping each frame of the image into multiple image patches, preprocessing can be performed on each image patch based on the requirements of subsequent object detection, including operations such as padding and scaling, which are not limited herein.
[0066] After obtaining the respective image patches of the image, object detection is performed on the multiple image patches of each frame of the image to obtain the target objects on each image patch. In a specific application scenario, object detection can be performed on the target regions of each frame of the image through a pre-trained object detection model to obtain the target objects located within the target regions of each frame of the image and their positions. In another specific application scenario, object detection can be performed on the target regions of each frame of the image through object detection technology to obtain the target objects located within the target regions of each frame of the image and their positions. Among them, the object detection technology can include YOLO-v4 object detection technology, Faster R-CNN object detection technology, or other object detection technologies, etc., which are not limited herein.
[0067] In a specific application scenario, when performing object detection on each image patch, the maximum number M of targets that can be recognized on each image patch can be set first, where the specific value of M depends on the device performance and the target scenario. By setting the maximum number M, the situation where the device is overloaded during object detection, affecting the efficiency and accuracy of object detection, is reduced. The reliability of object detection is improved.
[0068] After object detection obtains all the target objects on each image patch, the target objects on the image patches are combined and duplicate target objects are filtered to obtain the target objects of each frame of the image.
[0069] In a specific embodiment, object detection is performed on the multiple image patches of each frame of the image to obtain the detection boxes of the target objects on each image patch. According to the cropping rule in step S21, based on the positions of the respective image patches in the original image, the coordinates of the detection boxes in the original image are calculated, and then the non-maximum suppression method is used to combine the detection boxes of the entire image to filter out the overlapping detection boxes, obtaining all the target objects on each frame of the image. The target objects located within the target regions of each frame of the image are determined based on the above method.
[0070] In a specific application scenario, when a certain frame of the image is divided into 4 image patches, where there are 3 detection boxes of target objects on the first image patch, 2 detection boxes of target objects on the second image, 0 detection boxes of target objects on the third image, and 4 detection boxes of target objects on the fourth image, and there are two overlapping and identical detection boxes of target objects on the first image patch and the second image, then after using the non-maximum suppression method to combine the detection boxes of the entire image and filtering out the overlapping detection boxes, the number of detection boxes of the entire image of this frame of the image is 7.
[0071] In a specific application scenario, during the target tracking phase, the target object is traversed and detected in the image patch. For the detection box of the target object numbered j, the information at the t-th moment is denoted as where is the center coordinate of the detection box, and are the width and height of the detection box respectively. The specific position of the detection box can be restored through the center point coordinates and the width and height of the detection box. Among them, j ∈ [1, M], that is, j is any one of the 1 - M target objects in the image patch.
[0072] Based on the target objects in each frame of the image, the target objects located within the target area in each frame of the image are determined. In a specific application scenario, the target objects located within the target area in each frame of the image can be determined from the target objects in the entire image through methods such as coordinate judgment and similarity comparison, which are not limited herein.
[0073] Step S23: Predict the predicted position of the target object in the current frame of the image; match the predicted position with the position of the target object in the current frame of the image to determine the matching degree; when the matching degree exceeds the preset matching degree, in response to the connection lines between the positions of the target objects on each image from the previous preset frame of the image to the previous frame of the image and the position of the target object in the current frame of the image all intersecting the statistical section, the target object in the current frame of the image is counted as a valid count.
[0074] Before counting, first predict the predicted position of the target objects located within the target area in each frame of the image in the current frame of the image. In a specific implementation manner, the predicted position of the target object in the current frame of the image can be predicted based on the historical information of the target object; among them, the historical information includes the position of the target object in the historical frame of the image, that is, the position of the target object in the current frame of the image is predicted based on its position in the historical frame of the image. In another specific implementation manner, the predicted position of the target object in the current frame of the image can also be predicted based on the flow direction of the target area and the position of the target object in the historical frame of the image. This is not limited herein.
[0075] Among them, the prediction method can adopt a trained regression model or common prediction algorithms, such as: Kalman filtering method, Naive Bayes algorithm, etc., which are not limited herein.
[0076] After obtaining the predicted position of the target object in the current frame of the image, the predicted position is matched with the position of the target object in the current frame of the image to determine the matching degree. When matching, methods such as the intersection - over - union matching method, NCC matching algorithm, etc. can be adopted, which are not limited herein.
[0077] In a specific embodiment, the state of the target object in the current frame image can be determined based on the matching degree between the predicted position of the target object in the current frame image and the position of the target object in the current frame image in reality. Among them, when the matching degree does not exceed the preset matching degree and there is no historical matching information for the target object, it indicates that the target object appears for the first time in the current frame image, and its state is the creation state. When the matching degree exceeds the preset matching degree, it indicates that there is historical matching information for the target object, which means that the target object does not appear for the first time in the current frame image, and its state is the to-be-updated state. When the target object in the current frame image in reality does not exist, its state is the lost state.
[0078] In a specific application scenario, the statistical cross-section is generally set at a relatively central position in the target area. When the target object is in the to-be-updated state, it is possible for it to be near the statistical cross-section, and then it can be determined whether it is a valid count.
[0079] In a specific application scenario, in response to the intersection of the line connecting the position of the target object in the target area in the current frame image t and the position of the target object in the previous frame image t - 1 with the statistical cross-section, record the position of the target object in the current frame image t; based on the above steps, continue to record the target object in the next frame image until the positions of the target object in each frame image are recorded. Specifically, the center point coordinates of the detection box of the target object can be recorded, which is convenient for subsequent counting. Among them, the positions of the target object in each frame image can also be recorded in the historical information.
[0080] When the target object is in the creation state or the lost state, no counting is performed on the target object. When the target object is in the to-be-updated state, in response to the intersection of the lines connecting the positions of the target object in each image from the pre-preset frame image to the previous frame image with the position of the target object in the current frame image, the target object in the current frame image is counted as a valid count.
[0081] In a specific application scenario, when counting a certain target object in the target area of the current frame image t, first determine that the lines connecting the positions of the target object in each image from the pre-preset frame image t - s to the previous frame image t - 1 with the position of the target object in the current frame image t intersect with the statistical cross-section, and count the target object in the current frame image as a valid count. That is, s is the number of pre-preset frames.
[0082] In a specific application scenario, that is, to judge the center point of the detection box of the target object in each image from the pre-preset frame image t - s to the previous frame image t - 1 and the center point of the detection box of the target object in the current frame image t Whether all s connection lines between them intersect with the statistical cross-section. If all intersect, the target object in the current frame image is counted as a valid count. If there is at least one connection line that does not intersect, the target object is not counted. s can be 15, 18, 19, 20, etc., and can be specifically set based on the actual situation and is not limited here.
[0083] Please refer to Figure 4 , Figure 4 Yes Figure 2 Schematic diagram of the valid count statistics of an implementation manner in the embodiment.
[0084] In this implementation manner, the target objects in the target area 40 are counted. Among them, 42 is the center point of the detection box of the target object on the current frame image t, 43 is the center point of the detection box of the target object on the previous frame image t - 1, 44 is the center point of the detection box of the target object on the previous second frame image t - 2, 45 is the center point of the detection box of the target object on the previous third frame image t - 3, 46 is the center point of the detection box of the target object on the previous fourth frame image t - 4, and 47 is the center point of the detection box of the target object on the previous fifth frame image t - 5. s in this implementation manner is 5.
[0085] Then when the connection lines between 47 and 42, between 46 and 42, between 45 and 42, between 44 and 42, and between 43 and 42 all intersect with the statistical cross-section 41, it is determined that the target object on the current frame image t is a valid count.
[0086] Among them, after counting the target object in the current frame image as a valid count, it includes: updating the position of the target object in the current frame image to the historical information. Specifically, the center point coordinates of the detection box of the target object can be updated to the historical information.
[0087] Step S24: Count the target objects on the next frame image until all the target objects in each target area of each frame image are counted, and obtain the traffic statistics results of each target area in the video.
[0088] Count the target objects on the next frame image based on the statistical method in step S23, that is, take the next frame image as the current frame image for statistics, and so on, until all the target objects in each target area of each frame image in the video are counted, and obtain the traffic statistics results of each target area in the video.
[0089] In a specific application scenario, if the target object is a valid count in the current frame image, it means that the position of the target object in the current frame image is on one side of the statistical cross-section, and all positions in the previous preset frame images are on the other side of the statistical cross-section. When counting the target object in the next frame image, no matter which side of the statistical cross-section the target object in the next frame image is on, there will be at least one connection line between the positions of the target object on each image from the previous preset frame image of the next frame image to the previous frame image of the next frame image and the position of the target object in the next frame image that is completely on the same side of the statistical cross-section and cannot intersect the statistical cross-section, thus failing to meet the statistical rules of valid counting. The counting of subsequent frames is similar. Therefore, this embodiment can achieve the effect of only performing a valid count on a target object once and preventing repeated counting of the target object, thereby preventing the target object from repeatedly jumping near the statistical cross-section or causing misjudgment due to factors such as vehicle or image jitter, improving the accuracy and reliability of the traffic statistics result.
[0090] In a specific application scenario, when there are multiple target areas in the video, the traffic statistics results in each target area are separately counted and reported, and the traffic of the entire video may not be comprehensively counted.
[0091] Through the above steps, the traffic statistics method for the target object in this embodiment sets a frame number threshold for logical judgment by counting the target object in the current frame image as a valid count in response to the connection lines between the positions of the target object on each image from the previous preset frame image to the previous frame image and the position of the target object in the current frame image all intersecting the statistical cross-section, thereby preventing the target object from repeatedly jumping near the statistical cross-section or causing misjudgment due to factors such as vehicle or image jitter, and then improving the accuracy of the traffic statistics of the target object. Before detecting the target object in this embodiment, the original image frames of the video are first divided into multiple image blocks, and then the image blocks are detected, so that the proportion of the target object in the image block is increased, thereby improving the accuracy and reliability of the target detection. Moreover, the method of target detection after cutting the image in this embodiment is applicable to high-resolution videos collected in drone aerial photography and mid-air perspective scenarios. It improves the problems of wide field of view, small target proportion, and easy loss in the high-altitude perspective, and also is compatible with conventional-resolution videos collected in general perspective scenarios, expanding the applicable range of the traffic statistics method for the target object.
[0092] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an embodiment of a traffic statistics device for the target object of the present invention. The traffic statistics device 50 for the target object includes an acquisition module 51, a detection module 52, a counting module 53, and a statistics module 54. The acquisition module 51 is used to acquire a video including a series of consecutive frames of images and set a statistical cross-section for the target area within the video. The detection module 52 is used to perform target detection on the target area of each frame of image to obtain the target object located within the target area on each frame of image and the position of the target object. The counting module 53 is used to, in response to the fact that the lines connecting the positions of the target objects on each image between the previous preset frame of image and the previous frame of the current frame of image to the position of the target object in the current frame of image all intersect the statistical cross-section, count the target object in the current frame of image as a valid count. The statistics module 54 is used to count the target objects on the next frame of image until the target objects within the target area of each frame of image are counted, and obtain the traffic statistics result of the video.
[0093] The counting module 53 is further used to predict the predicted position of the target object in the current frame of image; match the predicted position with the position of the target object in the current frame of image to determine the matching degree; when the matching degree exceeds the preset matching degree, in response to the fact that the lines connecting the positions of the target objects on each image between the previous preset frame of image and the previous frame of the current frame of image to the position of the target object in the current frame of image all intersect the statistical cross-section, count the target object in the current frame of image as a valid count.
[0094] The counting module 53 is further used to predict the predicted position of the target object in the current frame of image based on the historical information of the target object; wherein, the historical information includes the position of the target object in the historical frame of images; update the position of the target object in the current frame of image to the historical information.
[0095] The detection module 52 is further used to divide each frame of image to obtain multiple image blocks of each frame of image; perform target detection on the multiple image blocks of each frame of image respectively to obtain the target objects on each image block; synthesize the target objects on the image blocks and filter out the duplicate target objects to obtain the target objects of each frame of image; determine the target objects located within the target area of each frame of image and the position of the target object.
[0096] The detection module 52 is further used to divide each frame of image respectively to obtain multiple initial image blocks of the same size for each frame of image; on each frame of image, expand the size of each initial image block in the first direction and the second direction respectively to obtain multiple image blocks of each frame of image; wherein, the first direction and the second direction are perpendicular to each other.
[0097] The detection module 52 is further used to, in response to the fact that the expanded initial image block exceeds the boundary of the corresponding frame of image, determine the pixel value of the area where the expanded initial image block exceeds the boundary as 0.
[0098] The obtaining module 51 is further configured to obtain a video including a plurality of consecutive frames of images and at least one initial target area; in response to the initial target area being a convex polygon, determine the initial target area as the target area; determine the plane perpendicular to the moving direction of the target object within the target area as the statistical cross-section; count the target objects in the next frame of image until the target objects in the target areas of all frames of images are counted, so as to obtain the traffic statistics result of the video, including: counting the target objects in the next frame of image until all the target objects in the target areas of all frames of images are counted, so as to obtain the traffic statistics results of each target area in the video.
[0099] The statistics module 54 is further configured to record the position of the target object in the current frame of image in response to the connection line between the position of the target object in the current frame of image and the position of the target object in the previous frame of image of the current frame of image intersecting with the statistical cross-section; record the target objects in the next frame of image until the positions of the target objects in all frames of images are recorded.
[0100] The above solution can improve the accuracy of traffic statistics of target objects.
[0101] Based on the same inventive concept, the present invention also provides an electronic device, which can be executed to implement the traffic statistics method of the target object in any of the above embodiments. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the electronic device provided by the present invention. The electronic device includes a processor 61 and a memory 62.
[0102] The processor 61 is configured to execute the program instructions stored in the memory 62 to implement the steps of the traffic statistics method of the target object in any of the above. In a specific implementation scenario, the electronic device may include, but is not limited to: a microcomputer, a server. In addition, the electronic device may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0103] Specifically, the processor 61 is used to control itself and the memory 62 to implement the steps of any of the above embodiments. The processor 61 may also be referred to as a CPU (Central Processing Unit). The processor 61 may be an integrated circuit chip with the ability to process signals. The processor 61 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 61 may be implemented jointly by integrated circuit chips.
[0104] The above solution can improve the accuracy of traffic statistics of the target object.
[0105] Based on the same inventive concept, the present invention also proposes a computer-readable storage medium. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present invention. At least one program data 71 is stored in the computer-readable storage medium 70, and the program data 71 is used to implement any of the above methods. In one embodiment, the computer-readable storage medium 70 includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0106] In several embodiments provided by the present invention, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other may be through some interfaces. The indirect couplings or communication connections of devices or units may be in electrical, mechanical, or other forms.
[0107] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0109] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product, and this computer software product is stored in a storage medium.
[0110] The above is only the embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present invention.
[0111] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up information or asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A traffic statistics method for a target object, characterized in that Including: Obtain a video including a series of consecutive frames of images, and set a statistical cross-section of a target area within the video; wherein, the target area is a convex polygon, and the statistical cross-section is a plane perpendicular to the flow direction of the target object within the target area; Perform target detection on the target area of each frame of the images to obtain the target objects located within the target area on each frame of the images and the positions of the target objects; In response to the lines connecting the positions of the target objects on each image between the previous preset frame of the current frame image and the previous frame of the current frame image intersecting the statistical cross-section, count the target object in the current frame image as a valid count; Count the target objects on the next frame of the image until the target objects within the target area of each frame of the images are counted, obtaining the traffic statistics result of the video.
2. The traffic statistics method for the target object according to claim 1, wherein The step of, in response to the lines connecting the positions of the target objects on each image between the previous preset frame of the current frame image and the previous frame of the current frame image intersecting the statistical cross-section, counting the target object in the current frame image as a valid count, includes: Predict the predicted position of the target object in the current frame image; Match the predicted position with the position of the target object in the current frame image to determine the matching degree; When the matching degree exceeds a preset matching degree, in response to the lines connecting the positions of the target objects on each image between the previous preset frame and the previous frame intersecting the statistical cross-section, count the target object in the current frame image as a valid count.
3. The flow statistics method for the target object according to claim 2, wherein The step of predicting the predicted position of the target object in the current frame image includes: Predict the predicted position of the target object in the current frame image based on the historical information of the target object; wherein, the historical information includes the positions in the historical frame images of the target object; After counting the target object in the current frame image as a valid count, it includes: Update the position of the target object in the current frame image to the historical information.
4. The traffic statistics method for the target object according to claim 1, wherein The step of performing target detection on the target area of each frame of the images to obtain the target objects located within the target area on each frame of the images and the positions of the target objects includes: Divide each frame of the images to obtain a plurality of image blocks for each frame of the images; Perform target detection on the plurality of image blocks of each frame of the images respectively to obtain the target objects on each of the image blocks; Integrate the target objects on the image blocks and filter out duplicate target objects to obtain the target objects of each frame of the images; Determine the target objects located within the target area in each frame of the images and the positions of the target objects.
5. The flow statistics method for the target object according to claim 4, wherein The step of dividing each frame of the images to obtain a plurality of image blocks for each frame of the images includes: Divide each frame of the images respectively to obtain a plurality of initial image blocks of the same size for each frame of the images; On each frame of the images, expand the size of each initial image block in a first direction and a second direction respectively to obtain a plurality of image blocks for each frame of the images; Wherein, the first direction and the second direction are perpendicular to each other.
6. The flow statistics method for the target object according to claim 5, characterized in that, The expanding the size of each initial image block in a first direction and a second direction respectively on each frame of the images to obtain a plurality of image blocks for each frame of the images further includes: In response to the expanded initial image block exceeding the boundary of the corresponding frame of the image, determining the pixel values of the area where the expanded initial image block exceeds the boundary as 0.
7. The method for traffic statistics of the target object according to claim 1, wherein, The obtaining a video including a plurality of consecutive frames of images and setting a statistical cross-section of a target area in the video includes: Obtaining a video including a plurality of consecutive frames of images and at least one initial target area; In response to the initial target area being a convex polygon, determining the initial target area as the target area; Respectively setting the planes perpendicular to the moving directions of the target objects in each of the target areas as the statistical cross-sections; The counting the target objects in the next frame of the images until the target objects in the target areas of all the frames of the images are counted to obtain the traffic statistics result of the video includes: Counting the target objects in the next frame of the images until all the target objects in each of the target areas of all the frames of the images are counted to obtain the traffic statistics results of each target area in the video.
8. The method for traffic statistics of the target object according to claim 1, characterized in that, Before counting the target object in the current frame of the image as a valid count in response to the fact that the lines connecting the positions of the target objects in the images from the previous preset number of frames of the current frame of the image to the previous frame of the current frame of the image all intersect the statistical cross-section, further includes: In response to the line connecting the position of the target object in the current frame of the image and the position of the target object in the previous frame of the current frame of the image intersecting the statistical cross-section, recording the position of the target object in the current frame of the image; Recording the target objects in the next frame of the images until the positions of the target objects in all the frames of the images are recorded.
9. An electronic device, characterized in that, The electronic device includes: a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the traffic statistics method of the target object as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data, and the program data can be executed to implement the traffic statistics method of the target object as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Automatic labeling method based on edge end traffic audio and video synchronization samples
CN112100435A
Traffic flow statistical method based on vehicle detection and multi-target tracking
CN112750150A