Object Recognition Method, Device, Equipment and Storage Medium Based on Motion Detection
By setting confidence thresholds in the video stream data and combining motion detection and feature matching algorithms, the problem of inaccurate recognition of small target objects in traditional methods is solved, achieving higher recognition accuracy and more comprehensive target object detection.
Patent Information
- Application Number
- CN202510552220.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing object recognition algorithm based on single-frame visual features is difficult to accurately identify small target objects, such as birds, drones, etc. at long distances or high noise, resulting in high missed detection and false detection rates.
Using a motion detection method, the candidate objects are filtered by setting the first and second confidence thresholds, combined with the motion detection algorithm, the target tracking algorithm and the feature matching algorithm, the motion characteristics and category characteristics of the target object are obtained, and the weighted fusion mechanism is used to improve the recognition accuracy.
It significantly improves the accuracy of recognition of small target objects, reduces static background interference, and can more accurately determine the category of target objects, making up for the shortcomings of traditional target detection algorithms.
Smart Images

Figure CN120071031B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to an object recognition method, device, equipment and storage medium based on motion detection. Background Art
[0002] With the continuous development of ultra-high resolution camera array technology, it is possible to collect extremely wide and detailed image or video data at the same time. However, in the actual application scenarios of ultra-high resolution array cameras, although small target objects (such as vehicles) within 2 - 3 kilometers can be clearly seen through ultra-high resolution technology, there are still smaller objects (especially small birds, drones, etc.) at farther distances, which often only occupy very few pixel points, resulting in difficulty for traditional single-frame vision feature-based detection algorithms to accurately identify. Taking birds as an example, even if a deep neural network (such as the YOLO series algorithm) is used to detect them, due to the too small target or too high similarity with the background, there are still problems of both missed detection and false detection. Especially in the case of extremely far distances of birds or a large amount of noise in the picture, the accuracy and confidence of existing related recognition methods are difficult to meet the actual requirements.
[0003] Therefore, how to improve the accuracy of image recognition has become an urgent technical problem to be solved currently. Summary of the Invention
[0004] This application provides an object recognition method, device, equipment and storage medium based on motion detection, aiming to improve the accuracy of image recognition.
[0005] In a first aspect, this application provides an object recognition method based on motion detection, and the object recognition method based on motion detection includes the following steps:
[0006] Obtain video stream data to be detected;
[0007] Based on a target detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence of each object to be detected corresponding to its category;
[0008] When the first confidence corresponding to the object to be detected is greater than or equal to a preset first confidence threshold, determine the object to be detected as a target object, and obtain the target detection frame of the target object;
[0009] When the first confidence corresponding to the object to be detected is less than the first confidence threshold and greater than a second confidence threshold, determine the object to be detected as a candidate object;
[0010] Based on a motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result;
[0011] When the motion detection result conforms to the characteristics of a moving object, determine the candidate object as the target object and obtain the target detection frame of the target object;
[0012] Based on the target tracking algorithm, perform tracking operations on the target detection frames of the target objects screened by the target detection algorithm and the motion detection algorithm to obtain the first motion features corresponding to each target object;
[0013] Based on the feature matching algorithm, perform feature matching calculations on the first motion features and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object and the feature similarity between the target object and the preset objects corresponding to each candidate category;
[0014] Based on the preset weighted fusion mechanism, perform weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object;
[0015] Based on the category matching degree between the target object and each preset object, determine the category recognition result of the target object.
[0016] In a second aspect, the present application further provides an object recognition device based on motion detection, and the object recognition device based on motion detection includes:
[0017] A data acquisition module, configured to acquire video stream data to be detected;
[0018] A target detection module, configured to identify at least one object to be detected in the video stream data based on a target detection algorithm and obtain the first confidence level of each category corresponding to the object to be detected;
[0019] A first target object determination module, configured to determine the object to be detected as a target object when the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtain the target detection frame of the target object;
[0020] A candidate object determination module, configured to determine the object to be detected as a candidate object when the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold;
[0021] A motion detection module, configured to perform motion feature detection on the candidate object based on a motion detection algorithm to obtain a motion detection result;
[0022] A second target object determination module, configured to determine the candidate object as a target object when the motion detection result conforms to the characteristics of a moving object, and obtain the target detection frame of the target object;
[0023] A target tracking module, configured to perform tracking operations on the target detection boxes of the target objects screened by the target detection algorithm and the motion detection algorithm based on a target tracking algorithm, and obtain first motion features corresponding to each of the target objects;
[0024] A feature matching module, configured to perform feature matching calculations on the first motion features and second motion features of at least one preset object based on a feature matching algorithm, determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each of the candidate categories;
[0025] A feature fusion module, configured to perform weighted fusion on the first confidence level and the feature similarity based on a preset weighted fusion mechanism, and determine the category matching degree between the target object and each of the preset objects;
[0026] A category recognition module, configured to determine the category recognition result of the target object based on the category matching degree between the target object and each of the preset objects.
[0027] In a third aspect, the present application further provides a computer device, where the computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the object recognition method based on motion detection as described above are implemented.
[0028] In a fourth aspect, the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the object recognition method based on motion detection as described above are implemented.
[0029] The present application provides an object recognition method, apparatus, computer device and storage medium based on motion detection. The method of the present application accurately recognizes moving target objects in a video stream through a target detection algorithm and obtains their first confidence level and first motion features, initially screening out target objects with motion characteristics and reducing interference from factors such as static backgrounds. By setting a first confidence threshold and a second confidence threshold, small-volume candidate objects ignored by the target detection algorithm are screened out. Through a motion detection algorithm, further motion detection is performed on the candidate objects, and small target objects ignored by the target detection algorithm are screened and recognized, making up for the deficiencies of the target detection algorithm, thereby detecting target objects more comprehensively and accurately. Through a target tracking algorithm, the target objects are tracked, the first motion features of the target objects are extracted, and through a feature matching algorithm, the first motion features are matched and calculated with the second motion features of preset objects to determine the candidate categories and feature similarities, further limiting and screening the target object categories from the perspective of motion features, increasing the dimension and accuracy of recognition. Through a preset weighted fusion mechanism, the first confidence level and the feature similarity are weighted and fused, comprehensively considering the information of the target object in terms of both category confidence and motion feature similarity, determining the category matching degree, and then obtaining the category recognition result, avoiding errors that may be caused by solely relying on a certain feature for recognition, thereby significantly improving the accuracy of image recognition and being able to more accurately determine the category of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 It is a schematic flowchart of the first embodiment of an object recognition method based on motion detection provided by the present application;
[0032] Figure 2 It is a schematic flowchart of the detection of small target moving objects provided by the present application;
[0033] Figure 3 It is a schematic structural diagram of the first embodiment of an object recognition apparatus based on motion detection provided by the present application;
[0034] Figure 4 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.
[0035] The realization, functional characteristics and advantages of the objectives of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0037] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0038] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. Without conflict, the embodiments described below and the features in the embodiments can be combined with each other.
[0039] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of the first embodiment of an object recognition method based on motion detection provided by the present application.
[0040] As Figure 1 shown, the object recognition method based on motion detection includes steps S101 to S110.
[0041] S101. Obtain video stream data to be detected;
[0042] In one embodiment, the video stream data can be a real-time video stream collected by a video capture device in real time, or a historical video stream collected by the video capture device.
[0043] Exemplarily, the video capture device can capture the dynamic scene of the monitoring area in real time. For example, in a traffic monitoring scenario, array cameras deployed at positions such as road intersections can continuously transmit real-time moving images of vehicles, pedestrians, etc. to the background detection system in the form of a video stream; or, in an airport monitoring scenario, the array cameras can continuously collect aerial images and transmit them to the background detection system in the form of a video stream to facilitate real-time monitoring of possible obstacles such as birds, drones, and balloons in the airspace above the airport. The real-time video stream can be used for real-time analysis of events in the current scene of the monitoring area.
[0044] Exemplarily, the video stream data can also be historical video stream data that has been collected and stored by a video acquisition device, such as video stream data periodically uploaded by the video acquisition device. In some scenarios, when the amount of video stream data is too large or in non-real-time analysis requirements, the data upload period of the video acquisition device can be set. For example, the video stream data is transmitted once every hour, that is, the video stream data collected by the video acquisition device in the past hour is obtained for analyzing the video stream data in this time period. At this time, the video stream data in this time period can be regarded as historical video stream data.
[0045] Exemplarily, the video acquisition device can be an ultra-high-resolution array camera.
[0046] S102. Based on the object detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence level of each object to be detected corresponding to a category.
[0047] In one embodiment, a preliminary screening can be performed by running a common object detection method (such as algorithms like the YOLO series, EfficientDet, or Nano Det).
[0048] In one embodiment, detection algorithms such as YOLO can be used to process and analyze the video stream data, and initially output the classes, boxes, and confidence levels of the target objects.
[0049] Generally, the confidence level reflects the credibility of the model that there is indeed such a target object in the detection box. Its value is between 0 and 1, and the closer the value is to 1, the more credible the model's judgment of the target object in the detection box is. For example, if the confidence level of the detection box is 0.95, it means that the model believes that the probability of there being a target object in this box is very high.
[0050] In one embodiment, the object detection algorithm can process each frame of the video stream data frame by frame, identify the coordinates, contours, classes, and corresponding confidence levels of the target objects in the image. The object detection algorithm can output multiple possible classes corresponding to the target object. For example, for an object flying in the air, the possible classes it may correspond to can include birds, drones, etc.
[0051] Exemplarily, after the object detection algorithm processes each frame of the video stream data, a detection box corresponding to each target object is obtained. For example, for a video image containing multiple pedestrians, the YOLOv7 model can generate a rectangular box (detection box) for each pedestrian to represent the position of the pedestrian in the image.
[0052] Exemplarily, a confidence threshold can be preset according to the requirements for the accuracy of the detection results in the actual application scenario, such as 0.7 or 0.8. For example, in the security monitoring scenario, in order to reduce false alarms, the confidence threshold can be set to 0.8 to ensure that only the detection results with a high level of confidence in the model are retained.
[0053] Specifically, the first confidence of each target detection box is compared with the preset confidence threshold. When the first confidence is greater than or equal to the confidence threshold, the object within the target detection box is considered a target object; otherwise, it is considered that the object within the detection box is not a target object or the detection is not accurate enough, and it is filtered out. For example, when the confidence threshold is set to 0.7, if the confidence of a certain detection box is 0.6, then this detection box will not be retained as a target object.
[0054] S103. When the first confidence corresponding to the object to be detected is greater than or equal to the preset first confidence threshold, determine that the object to be detected is a target object, and obtain the target detection box of the target object;
[0055] S104. When the first confidence corresponding to the object to be detected is less than the first confidence threshold and greater than the second confidence threshold, determine that the object to be detected is a candidate object;
[0056] In one embodiment, when the first confidence is less than the second confidence threshold, determine that the object in the target detection box is a non-target object.
[0057] In one embodiment, two thresholds can be set: the first confidence threshold Th1 and the second confidence threshold Th2. Among them, Th2 < Th1.
[0058] If the first confidence of a certain detection box is higher than Th1, it can be considered that the detection result is relatively reliable, and directly mark the object in this detection box as "target object".
[0059] If the confidence of the detection box is lower than Th2, it is temporarily abandoned, unless subsequent motion detection gives new clues to the same area.
[0060] In one embodiment, for the detection results determined to be non-target objects, it can be considered that they have nothing to do with the target category and can be excluded from subsequent processing and analysis, thereby reducing the waste of computing resources. However, it should be noted that sometimes non-target objects may have a certain similarity in appearance to target objects. For example, in a detection task where the target category is a car, some trucks or motorcycles may be misdetected as cars, but due to their low confidence, they are excluded from the target objects. To avoid this situation, the image can be filtered and enhanced in the preprocessing stage, such as using image segmentation technology to divide the image into different regions, or performing operations such as histogram equalization on the image to improve the detectability of target objects and the distinguishability of non-target objects.
[0061] If the confidence of the detection box is between Th2 and Th1, it is temporarily marked as a "candidate target" and needs to enter the subsequent motion detection and tracking module for auxiliary judgment.
[0062] In one embodiment, for the detection results determined to be candidate objects, although their confidence is lower than the first confidence threshold, they still have a certain degree of credibility and may contain potential target objects. For these candidate objects, various methods can be used for further processing to improve the accuracy and reliability of detection. For example, in an image, the objects or scene information around the candidate object may help determine whether it is a target object. If the candidate object is located in a scene where common target objects appear, or the layout and relationship with surrounding objects conform to the characteristics of the target object, then the possibility of it being a target object can be increased.
[0063] In this embodiment, the confidence threshold is an important parameter in the target detection algorithm for screening and classifying detection results. By setting different confidence thresholds, the accuracy and recall rate of detection can be controlled. When the first confidence threshold is relatively high, only detection results with relatively high confidence will be determined as target objects, which can improve the accuracy of detection and reduce false detections. However, some real target objects may be missed due to slightly lower confidence, reducing the recall rate. The opposite is true when the second confidence threshold is relatively low, which can increase the recall rate.
[0064] S105. Based on the motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result;
[0065] In one embodiment, the object detection algorithm can identify the target objects with obvious features in the target scene, but there may be problems of inaccurate identification or failure to identify small targets in the scene. For example, in an aerial monitoring scene, if a small bird is far from the video acquisition device, the number of pixels it occupies in the captured image is very small. The result is that although the bird is in the captured image, due to the small number of pixels it occupies, the object detection algorithm may not detect the bird target. At this time, a motion detection algorithm can be used to detect the dynamic targets in consecutive video frames to identify the bird target.
[0066] Specifically, through the above first confidence threshold and second confidence threshold, the detection results of the object detection algorithm are screened into three parts, namely, the clearly detected target objects (the first confidence is greater than or equal to the first confidence threshold), the candidate objects (the first confidence is between the first confidence threshold and the second confidence threshold), and the non-target objects (the first confidence is less than the second confidence threshold).
[0067] Among them, the candidate objects are the detection objects that the object detection algorithm cannot determine but may be moving targets. Therefore, it is necessary to use the motion detection algorithm to detect the motion features of the candidate objects, that is, to detect whether there will be changes in motion features in consecutive video frames, such as continuous displacement, periodic posture changes, etc.
[0068] As Figure 2 shown, when the object detection algorithm fails to identify small target objects, a motion detection algorithm, such as based on the optical flow method or the frame difference method, can be used to detect the motion features of the objects determined as candidate objects or non-target objects in the confidence threshold discrimination. For example, in traffic monitoring, the optical flow method can be used to calculate the motion feature information such as the speed and direction of the target vehicle over time.
[0069] In one embodiment, the motion detection algorithm can detect whether there are moving targets in each image according to the object coordinates in consecutive video frames. For some motion detection algorithms, it is possible to judge whether there is motion by tracking the coordinate changes of the target object in consecutive video frames. For example, the center coordinates (x, y) and width and height (w, h) of the target object are obtained in each frame through the object detection model, and then the change of these coordinate values in the time series is observed. If there are significant differences in the coordinate values in consecutive frames, it can be inferred that there is motion.
[0070] In one embodiment, the motion detection algorithm can also detect motion by comparing the differences between adjacent video frames. The pixel value differences between the front and rear frame images can be calculated, or whether motion occurs can be determined by comparing the grayscale images, feature points, etc. of the front and rear frames. For example, the frame difference method is a common comparison algorithm. By performing a difference operation on two adjacent frame images, if the difference exceeds a certain threshold, it is considered that there is a moving object.
[0071] Specifically, the motion detection algorithm can include the frame difference method, the optical flow method, and the background subtraction method, etc.
[0072] The frame difference method refers to performing a pixel-by-pixel difference operation between two adjacent frame images, regarding the pixels with differences less than the threshold as the static background, and the pixels with differences greater than the threshold as the motion area.
[0073] The optical flow method calculates the motion direction and speed of the pixels in the image by conditions such as the gray consistency constraint of the pixels between consecutive frames, forming an optical flow field, thereby detecting the moving object.
[0074] The background subtraction method first establishes a background model, usually by observing the image features of the scene without moving objects for a long time. Then, the current frame is compared with the background model, and the area with a large difference is considered to be the area with a moving object.
[0075] In one embodiment, various motion detection and tracking algorithms can be used to track the position coordinates of the target object in consecutive video frames, such as the optical flow method, the background difference method, etc.
[0076] Furthermore, based on the optical flow method, detect the motion trend of the candidate object to obtain the motion trend detection information of the candidate object; based on the background difference method, perform pixel comparison between the current image area where the candidate object is located and the preset background model to obtain the background area detection information of the candidate object; based on the motion trend detection information and the background area detection information, obtain the motion detection result corresponding to the candidate object.
[0077] Taking the optical flow method as an example, the basic algorithms include the Lucas-Kanade, Horn-Schunck, and Farneback algorithms, etc. The optical flow method calculates the motion direction of the pixels in the picture using the brightness difference at the pixel level between the front and rear frame images, plus the small motion constraint and motion continuity.
[0078] Specifically, the optical flow algorithm analyzes the brightness changes of pixel points in consecutive frame images. For example, the classic brightness constancy assumption (i.e., assuming that the brightness value of a pixel point remains unchanged during movement) is used as the basis to construct equations for solving the optical flow field. The Lucas-Kanade algorithm, etc., are typical methods based on this local optical flow calculation. When calculating, it mainly focuses on the changes within the local pixel neighborhood to estimate the motion of this local area. The small motion constraint means that within a very short time (i.e., between adjacent frames), the motion of an object will not be too large, which can simplify the difficulty of equation solving; and the motion continuity is manifested as the motion field being relatively smooth and coherent in space. Based on these assumptions, algorithms such as the Horn-Schunck algorithm can solve a relatively reasonable optical flow field and obtain the motion direction and speed estimates of each pixel point.
[0079] Similarly, the function of the background subtraction method is similar, that is, to find the moving regions in the picture for the detection and analysis of moving objects. They both achieve this function based on the differences between consecutive frame images, but only the principles and methods used are different.
[0080] The background subtraction method and the optical flow algorithm have a certain complementary effect. In some scenarios, the background subtraction method may detect motions that the optical flow algorithm cannot detect. For example, when the appearance of a moving object is significantly different from the background in terms of color, texture, etc., and the motion amplitude is relatively large, the background subtraction method can easily identify the moving region by comparing with the background model; while the optical flow algorithm may not be able to accurately detect the motion due to the overly large motion amplitude, which causes the assumption based on small motion constraints to fail. Conversely, the optical flow algorithm may perform better than the background subtraction method in scenarios where higher requirements are placed on the estimation of motion direction and speed, or where the background changes are relatively complex but the moving objects are relatively small and move smoothly. Therefore, the two can complement each other and be used in combination to improve the accuracy and robustness of motion detection.
[0081] The algorithm complexity of the background subtraction method is relatively simpler than that of the optical flow algorithm. It mainly compares each pixel of the current frame image with the pre-established background model and determines whether a pixel point belongs to the moving region according to the set threshold.
[0082] Specifically, the background subtraction method is based on the comparison between the background model and the current frame image. Usually, a background model is established in a certain way. For example, by statistically analyzing the video frames in the initial period of time, statistical quantities such as the average value and variance of each pixel are calculated to construct a simple background model, or more complex methods based on the mixture Gaussian model are used to model the background to better adapt to the situation where there are some dynamic changes in the background (such as swaying branches, fluctuating water surfaces, etc.). During actual detection, for each newly input image, it is compared with the background model pixel by pixel. If the difference in features such as color and brightness between a certain pixel point in the current frame and the corresponding pixel in the background model exceeds the set threshold, it is considered that the pixel point belongs to the moving area. Then, through post-processing operations such as connected region analysis of these moving pixel points, the complete moving target area can be obtained, and then the detection and tracking of moving objects can be realized.
[0083] In the small target scenario, since the target itself occupies fewer pixels, the target detection algorithm may fail to detect the target due to factors such as the target being not obvious and being easily occluded by the background. By reasonably combining the optical flow method and the background subtraction method and adjusting the operation weights and functions of the two according to the specific scenario and computing power conditions, the motion detection and analysis of the area where the small target is located can be better realized. Furthermore, by analyzing the motion of the pixels in the area where the small target is located, the motion situation of the target can be continuously tracked to make up for the deficiencies of the target detection algorithm. For example, it is very useful in scenarios such as monitoring the motion trajectories of some small wild animals in the distance in surveillance videos.
[0084] S106. When the motion detection result conforms to the characteristics of a moving target, determine the candidate object as the target object and obtain the target detection frame of the target object.
[0085] In this embodiment, by introducing a lower second confidence threshold under the conventional first confidence threshold adopted by the target detection algorithm, the candidate targets with the first confidence between the two are used as potential detection targets, so as to screen out the detection objects with low confidence due to the target being too small or partially occluded; by analyzing the motion trend of pixels between consecutive frames through the optical flow method, the motion direction and speed of the target are judged, so as to capture the motion characteristics when the target moves in the picture as the screening basis for the target object; by the background difference method, the foreground target is detected by comparing the difference between the current frame and the background model. Especially when the target moves in the background, the background difference method can effectively capture the motion trajectory of the target.
[0086] These three methods do not exist in isolation in the motion detection module, but are interrelated and complementary. The confidence threshold screening method can provide a preliminary candidate frame range, but whether these frames actually contain target objects requires further verification by the optical flow method and the background difference method. The optical flow method can capture the motion trend of the target, while the background difference method can detect the difference between the target and the background. When the candidate frames screened out by the confidence threshold screening method have a motion trend (detected by the optical flow method) and are significantly different from the background (detected by the background difference method), it can be more accurately determined whether there are moving target objects in these candidate frames. Through the joint detection of the three methods, the candidate targets that the target detection algorithm cannot identify are further screened. When the three methods verify that there are moving objects in the candidate targets, the candidate targets are identified as target objects, and then the candidate frame where the target object is located is used as the target detection frame that the target tracking algorithm needs to track and calculate, so as to track its motion trajectory and extract motion features.
[0087] For example, in a monitoring scene, a small animal is moving through the woods. Due to the small size of the animal, the target detection algorithm may not be able to accurately identify it. At this point, the motion detection module comes into play. The confidence threshold screening method retains those low-confidence candidate boxes that may contain small animals. The optical flow method captures the motion trend of the pixels in these boxes, indicating that an object is moving. The background difference rule finds that these boxes are significantly different from the background model, further confirming the existence of the target, and then determines the low-confidence candidate box as the target detection box of the target object. Through the collaborative work of these three methods, the presence of small animals can be detected more accurately and tracked.
[0088] In short, these three methods in the motion detection module can detect target objects more comprehensively and accurately through mutual correlation and collaboration, especially small target objects that are easily ignored by traditional target detection algorithms. The introduction of this method not only makes up for the shortcomings of traditional target detection algorithms, but also provides new ideas and solutions for target detection in complex scenes.
[0089] In one embodiment, the moving target object in the target scene can be identified and tracked by the target detection algorithm and the motion detection algorithm, and during the identification and tracking process, the size of the target object can be calculated by the size of the detection frame identified by the target detection algorithm or the motion detection algorithm. Then, the category of the target object is auxiliary identified according to the size of the target object.
[0090] Because the distance of different types of small objects, such as birds, airplanes, and drones, can be roughly estimated from the size of their detection boxes in the picture. Then, the category of the identified target object is determined by combining the motion characteristics of the target object, including speed, acceleration, motion trajectory characteristics, etc.
[0091] Specifically, for example, there are significant differences between the appearance, scale, motion trajectory, and motion characteristics of an aircraft and those of a bird. These characteristics can be collected for modeling. For example, algorithms such as the Kalman filter algorithm, support vector machines, or more complex neural network pattern recognition algorithms can be used to learn and model these characteristics, enabling the model to recognize characteristics such as the appearance, scale, behavior patterns, and motion characteristics of different types of small objects. Then, the learned model is used to perform pattern recognition on the target detection results (including parameters such as the scale of the target detection box and the motion characteristics of the target object), and then the recognition result of the model for the target object is output.
[0092] S107. Based on the target tracking algorithm, perform tracking operations on the target detection boxes of the target objects selected by the target detection algorithm and the motion detection algorithm to obtain the first motion characteristics corresponding to each of the target objects.
[0093] Exemplarily, the first motion characteristics may include characteristic information such as the motion trajectory, motion direction, speed (absolute speed or relative speed), and displacement of the target object. For example, by comparing the position changes of the target object frame by frame and combining the frame rate information of the video, the displacement vector of the object in each frame can be calculated, and then characteristics such as its motion speed and direction can be obtained. These motion characteristic information play an important role in understanding the behavior patterns and states of target objects and can be used for requirements such as trajectory tracking and behavior analysis.
[0094] In one embodiment, after obtaining the motion detection results, it is necessary to continuously track each moving target. Common methods include the Kalman Filter, SORT, DeepSORT, etc. Continuously tracking the moving target can obtain the tracking results of the target object between consecutive video frames, including relatively accurate target center positions, motion trajectories, and information such as size, speed, and acceleration that change over time.
[0095] Furthermore, based on the target tracking algorithm, track the target detection boxes corresponding to the same target object in consecutive video frames to obtain the continuous coordinate set of the target object; based on the continuous coordinate set, draw the motion trajectory of the target object to calculate the first motion characteristics corresponding to the target object.
[0096] In one embodiment, other first motion characteristics of the target object, such as position change, speed, direction, acceleration, etc., can be further calculated according to the motion trajectory of the target object.
[0097] Specifically, using an array camera or other two-dimensional tracking system, continuously obtain the positions of the target object at each time point to form a series of two-dimensional coordinate points and record the corresponding timestamps , thus, according to the time stamp order, each two-dimensional coordinate point is connected in sequence to obtain the motion trajectory of the target object, and through the motion trajectory, other first motion characteristics such as the position change, speed, and acceleration of the target object can be calculated.
[0098] Exemplarily, the position change of the target object can be deduced from the motion trajectory: in the two-dimensional space, the position coordinates of the target object first recognized in the video stream data are used as the origin, and any subsequent position of the target position can be represented by a position vector from the origin to this point represented as:
[0099]
[0100] Within the time interval the position change vector of the target object is:
[0101]
[0102] Exemplarily, the speed of the target object can be calculated from the motion trajectory:
[0103] The instantaneous velocity vector is:
[0104]
[0105] In the discrete case, the average velocity vector can be approximately calculated within the time interval as:
[0106]
[0107] The component form of the velocity vector is:
[0108]
[0109] The speed of the target object at the i-th time point can be expressed as:
[0110]
[0111] Exemplarily, the acceleration of the target object can be calculated from the motion trajectory:
[0112] The instantaneous acceleration vector can be expressed as:
[0113]
[0114] In the discrete case, the average acceleration vector can be approximately calculated within the time interval as:
[0115]
[0116] The component form of acceleration can be expressed as:
[0117]
[0118] The magnitude of the acceleration of the target object at any moment t can be expressed as:
[0119]
[0120] In this embodiment, by analyzing the two-dimensional motion trajectory of the target object, the first motion feature of the target object is extracted, so as to facilitate identifying the category of the target object according to the first motion feature, improve the accuracy of target object category recognition, and assist in judging the impact of the target object on the monitoring scene, thereby facilitating the proposal of corresponding countermeasures for the impact of the target object.
[0121] S108. Based on the feature matching algorithm, perform feature matching calculation on the first motion feature and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category;
[0122] In one embodiment, a feature database including the second motion features of at least one preset object can be pre-constructed. Each preset object has a corresponding feature vector for representing the second motion feature. These feature vectors can be obtained through experimental measurement, historical data statistics, or machine learning model training.
[0123] Exemplarily, the preset objects can be collected according to the scene application requirements. For example, in the traffic monitoring application scenario, the preset objects can include vehicles of different models (such as bicycles, electric vehicles, motorcycles, sedans, trucks, etc.), pedestrians, animals, etc.; in the airport airspace monitoring application scenario, the preset objects can include airplanes, drones, flying birds, balloons, kites, etc.
[0124] In one embodiment, based on the application scenario requirements, determine at least one of the preset objects in the target scene and the category labels of each preset object; based on the historical video stream data, collect the motion data of the preset objects to obtain the second motion features corresponding to the preset objects; based on the second motion features corresponding to the preset objects and the category labels, construct the feature database.
[0125] In one embodiment, to build a feature database, it is first necessary to clarify the application scenario requirements of the target scenario. For example, in a traffic monitoring scenario, objects such as vehicles, pedestrians, and animals need to be concerned; in an airport airspace monitoring scenario, objects such as airplanes, drones, birds, balloons, and kites are the focus. According to the scenario requirements, determine which objects are typical preset objects that need to be recognized. These objects should be representative and cover common object categories in the scenario.
[0126] In a specific embodiment, select segments containing preset objects from historical video stream data such as surveillance videos related to the scenario. For example, in a traffic monitoring scenario, select video segments containing passing vehicles, pedestrians, and animals. Track the preset objects in the video and record their motion data such as motion trajectories, speeds, and accelerations. These data can be automatically extracted using computer vision algorithms (such as object detection and tracking algorithms), or recorded through manual annotation. Clean and preprocess the collected motion data to remove noise and outliers to ensure the accuracy and consistency of the data. For example, smooth the motion trajectory data to remove mutation points caused by reasons such as unstable video frame rates. Select appropriate feature extraction methods according to the application scenario and the characteristics of object motion. For example, for vehicles, features such as average speed, acceleration, and steering angular velocity can be extracted, and features such as contours and markings can also be included; for birds, features such as contours, feather colors, feather characteristics, flight speeds, wing flapping frequencies, and curvatures of flight trajectories can be extracted. Combine the extracted motion features into feature vectors. For example, a feature vector of an airplane can include features such as speed, acceleration, steering angular velocity, and motion trajectory type (such as straight line, curve). Normalize the feature vectors to make their numerical ranges and distributions consistent for subsequent matching calculations. For example, normalize the speed and acceleration features to the [0, 1] interval. Add class labels to the feature vectors of each preset object. The class labels should clearly identify the category to which the object belongs, such as "bicycle", "drone", "pedestrian", etc., to ensure that the feature vectors of objects in the same category have the same class label and that the labels of different categories are not confused.
[0127] In one embodiment, after completing the feature extraction and class label marking of the preset objects in the target scenario, a suitable data structure and storage method can be used to build a feature database, such as a relational database or a non-relational database. For example, in a relational database, use tables to store feature vectors and class labels. Each row represents a preset object, and the columns include the various dimensions of the feature vector and the class label.
[0128] Generally, an index can be established for the feature database to improve the efficiency of feature query and matching. The index can be established based on some key dimensions of the feature vector, such as speed and type of motion trajectory. At the same time, as the application scenario changes and new data accumulates, the feature database should be updated in a timely manner to add new preset object features and optimize the existing features.
[0129] In this embodiment, a feature database based on the requirements of the application scenario is constructed. The database contains the second motion features and class labels of the preset objects, and can be used for class recognition and feature matching of the target object.
[0130] In one embodiment, the first motion features may include features such as the motion trajectory, speed, and acceleration of the target object, and the second motion features may also include features such as the motion trajectory, speed, and acceleration of the preset object.
[0131] In one embodiment, the second motion features may further include the contour features of the preset object, as well as the contour change features during the movement process. For example, a bird will flap its wings during flight, and the wing flapping states are different in different flight states such as gliding and changing direction. When collecting the first motion features of the target object, the contour features of the target object during the movement process can also be collected to assist in identifying the class of the target object.
[0132] Further, based on the preset feature database, obtain the second motion features of at least one of the preset objects; based on the feature matching algorithm, calculate the feature similarity between the first motion feature and each of the second motion features; based on the feature similarity, determine at least one candidate class corresponding to the target object.
[0133] In one embodiment, the preset feature database stores the second motion features of multiple preset objects, which are important bases for the matching analysis of the target object. Using the feature matching algorithm, by comparing the first motion feature of the target object with the second motion features of the preset objects in the feature database, the similarity between the two is calculated, so as to determine the candidate class of the target object.
[0134] In one embodiment, the feature matching algorithm may include one or more of algorithms such as Euclidean distance, cosine similarity, and Pearson correlation coefficient. The principles and applicable scenarios of each algorithm are different, and can measure the similarity between features from different perspectives.
[0135] Among them, the Euclidean distance is used to calculate the straight-line distance between two feature vectors in a multi-dimensional space. The smaller the distance, the higher the similarity. The cosine similarity is used to calculate the cosine value of the included angle between two feature vectors. The closer the value is to 1, the higher the similarity. The Pearson correlation coefficient is used to measure the linear correlation degree between two feature vectors. The value ranges from -1 to 1, and the greater the absolute value, the stronger the correlation.
[0136] In one embodiment, when the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, it is determined that the object category of the preset object is a candidate category of the target object.
[0137] In one embodiment, a feature matching algorithm is used to calculate the feature similarity between the first motion feature of the target object and the second motion feature of the preset object. According to the calculated feature similarity and the preset similarity threshold, the candidate category corresponding to the target object is determined.
[0138] If the feature similarity is greater than or equal to the similarity threshold, it is considered that the target object belongs to the category of the preset object; if the feature similarities of multiple preset objects are all greater than or equal to the similarity threshold, these preset objects are all used as candidate categories.
[0139] S109. Based on a preset weighted fusion mechanism, the first confidence level and the feature similarity are weighted and fused to determine the category matching degree between the target object and each preset object;
[0140] Generally, there are often inconsistencies or differences between the results of motion model matching and the category confidence levels output by detection models such as YOLO. For example, the YOLO model may result in inaccurate category confidence levels due to factors such as lighting changes and target occlusion; while the motion model may produce misjudgments due to short-term anomalies in the target motion path. Therefore, it is necessary to comprehensively consider these two pieces of information and, through the method of weighted fusion, make full use of their respective advantages to improve the reliability of the final category matching degree.
[0141] Specifically, a weighted fusion mechanism can be designed:
[0142]
[0143] Among them, and are weight parameters preset in advance or updated according to dynamic situations, satisfying + = 1; represents the confidence level of the motion detection model for category and comes from the output of motion detection models such as YOLO; represents the confidence level of the motion detection model for category The feature matching degree is the feature similarity calculated by the feature matching algorithm.
[0144] In one embodiment, the weight parameters and can be adjusted according to the actual application requirements.
[0145] Exemplarily, the weight parameters and can be adjusted according to the credibility of the motion detection model. If the motion detection model (such as YOLO) shows a high accuracy in a specific scenario, the value of can be appropriately increased, and vice versa; similarly, the value of is adjusted according to the performance of the motion detection model in this scenario. The value of
[0146] The weight parameters and can also be adjusted according to the real-time motion state of the target object. When the motion state of the target object is relatively stable, the result of the motion detection model may be more reliable, and the value of can be appropriately increased; while when the motion state of the target object is complex and changeable, the result of the detection model may be more valuable for reference, and the value of can be increased.
[0147] The weight parameters and can also be dynamically adjusted. A dynamic adjustment mechanism can be designed to automatically adjust the values of and according to the real-time detection confidence, motion confidence, and historical credibility records, etc. For example, when the detection confidence is continuously high, gradually increase the weight of ; when there is a significant change in the motion confidence, increase the weight of .
[0148] The weighted fusion mechanism can comprehensively consider the advantages of the motion detection model and the feature matching algorithm, improve the recognition accuracy of categories, and reduce the influence of misjudgment of a single model. For example, in a vehicle detection scenario, when the confidence of the detection model decreases due to partial occlusion of the vehicle, the motion model can supplement information through the motion features of the vehicle, and improve the overall recognition effect through weighted fusion.
[0149] S110. Determine the category recognition result of the target object based on the category matching degree between the target object and each of the preset objects.
[0150] In one embodiment, a comprehensive class matching degree is obtained by fusing the confidence degree and the feature similarity. The class matching degrees after fusing the confidence degree and the feature similarity of the target object and each preset object are compared, and according to the fused scores a ranking is performed, and the highest one or those above a certain threshold (such as 0.5) are selected as the final output class of the algorithm. In this way, both the discriminative ability of the visual model for appearance features and the judgment ability of the motion model for dynamic features are taken into account.
[0151] This embodiment provides an object recognition method based on motion detection. This method accurately identifies the moving target object in the video stream through a motion detection algorithm and obtains its first confidence degree and first motion feature, initially screening out the target objects with motion characteristics and reducing the interference of factors such as static backgrounds. By setting a first confidence degree threshold and a second confidence degree threshold, small-volume candidate objects ignored by the target detection algorithm are screened out. Through the motion detection algorithm, further motion detection is performed on the candidate objects, and screening and recognition are carried out for small target objects ignored by the target detection algorithm, making up for the deficiencies of the target detection algorithm, so as to detect the target object more comprehensively and accurately. Through the target tracking algorithm, the target object is tracked, and the first motion feature of the target object is extracted. Through the feature matching algorithm, the first motion feature is matched and calculated with the second motion feature of the preset object to determine the candidate class and the feature similarity, further limiting and screening the target object class from the perspective of motion features, increasing the dimension and accuracy of recognition. Through the preset weighted fusion mechanism, the first confidence degree and the feature similarity are weighted and fused, comprehensively considering the information of the target object in terms of both class confidence degree and motion feature similarity, determining the class matching degree, and then obtaining the class recognition result, avoiding the errors that may be brought by identifying solely based on a certain feature, thus significantly improving the accuracy of image recognition and being able to more accurately determine the class of the target object.
[0152] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the first embodiment of an object recognition device based on motion detection provided by this application. This object recognition device based on motion detection is used to execute the aforementioned object recognition method based on motion detection.
[0153] As Figure 3 shown, this object recognition device 200 based on motion detection includes: a data acquisition module 201, a target detection module 202, a first target object determination module 203, a candidate object determination module 204, a motion detection module 205, a second target object determination module 206, a target tracking module 207, a feature matching module 208, a feature fusion module 209, and a class recognition module 210.
[0154] The data acquisition module 201 is configured to acquire video stream data to be detected;
[0155] The target detection module 202 is configured to identify at least one object to be detected in the video stream data based on a target detection algorithm, and obtain a first confidence level for each category corresponding to the object to be detected;
[0156] The first target object determination module 203 is configured to determine the object to be detected as a target object when the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtain a target detection frame of the target object;
[0157] The candidate object determination module 204 is configured to determine the object to be detected as a candidate object when the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold;
[0158] The motion detection module 205 is configured to perform motion feature detection on the candidate object based on a motion detection algorithm, and obtain a motion detection result;
[0159] The second target object determination module 206 is configured to determine the candidate object as a target object when the motion detection result conforms to the motion target feature, and obtain a target detection frame of the target object;
[0160] The target tracking module 207 is configured to perform a tracking operation on the target detection frame of the target object screened out by the target detection algorithm and the motion detection algorithm based on a target tracking algorithm, and obtain a first motion feature corresponding to each target object;
[0161] The feature matching module 208 is configured to perform feature matching calculation on the first motion feature and the second motion features of at least one preset object based on a feature matching algorithm, determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset object corresponding to each candidate category;
[0162] The feature fusion module 209 is configured to perform weighted fusion on the first confidence level and the feature similarity based on a preset weighted fusion mechanism, and determine the category matching degree between the target object and each preset object;
[0163] The category recognition module 210 is configured to determine a category recognition result of the target object based on the category matching degree between the target object and each preset object.
[0164] In one embodiment, the motion detection module 205 includes:
[0165] A motion trend detection unit, configured to detect the motion trend of the candidate object based on the optical flow method, and obtain the motion trend detection information of the candidate object;
[0166] A background detection unit, configured to perform pixel comparison between the current image region where the candidate object is located and a preset background model based on the background difference method, and obtain the background region detection information of the candidate object;
[0167] A motion detection evaluation unit, configured to obtain the motion detection result corresponding to the candidate object based on the motion trend detection information and the background region detection information.
[0168] In one embodiment, the target tracking module 207 includes:
[0169] A target tracking unit, configured to track the target detection frames corresponding to the same target object in consecutive video frames based on the target tracking algorithm, and obtain the continuous coordinate set of the target object;
[0170] A first motion feature calculation unit, configured to draw the motion trajectory of the target object based on the continuous coordinate set, so as to calculate the first motion feature corresponding to the target object.
[0171] In one embodiment, the target detection module 202 further includes:
[0172] A non-target object determination unit, configured to determine that the object in the target detection frame is a non-target object when the first confidence level is less than the second confidence threshold.
[0173] In one embodiment, the feature matching module 208 includes:
[0174] A second motion feature acquisition unit, configured to acquire the second motion features of at least one of the preset objects based on a preset feature database;
[0175] A feature similarity calculation unit, configured to calculate the feature similarity between the first motion feature and each of the second motion features based on the feature matching algorithm;
[0176] A candidate category determination unit, configured to determine at least one candidate category corresponding to the target object based on the feature similarity.
[0177] In one embodiment, the candidate category determination unit includes:
[0178] A candidate category determination subunit, configured to determine that the object category of the preset object is a candidate category of the target object when the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold.
[0179] In one embodiment, the object recognition device 200 based on motion detection further includes a feature database construction module, including:
[0180] A scene analysis unit, configured to determine at least one of the preset objects in the target scene and the class labels of each of the preset objects based on the application scenario requirements;
[0181] A second motion feature acquisition unit, configured to collect motion data of the preset object based on historical video stream data, and obtain the second motion feature corresponding to the preset object;
[0182] A feature database construction unit, configured to construct the feature database based on the second motion feature corresponding to the preset object and the class label.
[0183] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing embodiments of the object recognition method based on motion detection, and will not be repeated here.
[0184] The device provided in the above embodiment can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 4 shown.
[0185] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. This computer device may be a server.
[0186] Referring to Figure 4 , this computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0187] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions, and when the program instructions are executed, the processor can be made to execute any object recognition method based on motion detection.
[0188] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0189] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When this computer program is executed by the processor, the processor can be made to execute any object recognition method based on motion detection.
[0190] This network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0191] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0192] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0193] Obtain video stream data to be detected;
[0194] Based on an object detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence level of each category corresponding to the object to be detected;
[0195] When the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, determine the object to be detected as a target object, and obtain the target detection frame of the target object;
[0196] When the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold, determine the object to be detected as a candidate object;
[0197] Based on a motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result;
[0198] When the motion detection result conforms to the motion target feature, determine the candidate object as a target object, and obtain the target detection frame of the target object;
[0199] Based on an object tracking algorithm, perform a tracking operation on the target detection frames of the target objects screened by the object detection algorithm and the motion detection algorithm, and obtain the first motion feature corresponding to each target object;
[0200] Based on the feature matching algorithm, perform feature matching calculation on the first motion feature and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category;
[0201] Based on a preset weighted fusion mechanism, perform weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object;
[0202] Based on the category matching degree between the target object and each preset object, determine the category recognition result of the target object.
[0203] In one embodiment, when the processor implements the motion detection algorithm to perform motion feature detection on the candidate object and obtain a motion detection result, it is used to implement:
[0204] Based on the optical flow method, detect the motion trend of the candidate object to obtain the motion trend detection information of the candidate object;
[0205] Based on the background difference method, perform pixel comparison on the current image area where the candidate object is located and a preset background model to obtain the background area detection information of the candidate object;
[0206] Based on the motion trend detection information and the background area detection information, obtain the motion detection result corresponding to the candidate object.
[0207] In one embodiment, when the processor implements the target tracking algorithm to perform tracking operation on the target detection frames of the target object screened by the target detection algorithm and the motion detection algorithm and obtain the first motion feature corresponding to each target object, it is used to implement:
[0208] Based on the target tracking algorithm, track the target detection frames corresponding to the same target object in consecutive video frames to obtain the continuous coordinate set of the target object;
[0209] Based on the continuous coordinate set, draw the motion trajectory of the target object to calculate the first motion feature corresponding to the target object.
[0210] In one embodiment, after the processor implements the target detection algorithm to identify at least one object to be detected in the video stream data and obtain the first confidence level corresponding to each category of the object to be detected, it is further used to implement:
[0211] When the first confidence level is less than the second confidence threshold, determine that the object in the target detection frame is a non-target object.
[0212] In one embodiment, when implementing the feature matching algorithm to perform feature matching calculation on the first motion feature and the second motion features of at least one preset object, and determining at least one candidate category corresponding to the target object and the feature similarity between the target object and the preset objects corresponding to each candidate category, the processor is configured to implement:
[0213] Based on a preset feature database, obtain the second motion features of at least one of the preset objects;
[0214] Based on the feature matching algorithm, calculate the feature similarity between the first motion feature and each of the second motion features;
[0215] Based on the feature similarity, determine at least one candidate category corresponding to the target object.
[0216] In one embodiment, when implementing determining at least one candidate category corresponding to the target object based on the feature similarity, the processor is configured to implement:
[0217] When the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, determine the object category of the preset object as a candidate category of the target object.
[0218] In one embodiment, before implementing obtaining the second motion features of at least one of the preset objects based on the preset feature database, the processor is further configured to implement:
[0219] Based on the application scenario requirements, determine at least one of the preset objects in the target scenario and the category label of each preset object;
[0220] Based on historical video stream data, collect the motion data of the preset object to obtain the second motion feature corresponding to the preset object;
[0221] Based on the second motion feature corresponding to the preset object and the category label, construct the feature database.
[0222] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, where the computer program includes program instructions, and the processor executes the program instructions to implement any one of the object recognition methods based on motion detection provided by the embodiments of the present application.
[0223] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.
[0224] As described above, the foregoing is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An object recognition method based on motion detection, characterized in that The method includes: Obtaining video stream data to be detected; Based on an object detection algorithm, identifying at least one object to be detected in the video stream data, and obtaining a first confidence level for each category corresponding to the object to be detected; When the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, determining the object to be detected as a target object, and obtaining a target detection box of the target object; When the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold, determining the object to be detected as a candidate object; Based on a motion detection algorithm, performing motion feature detection on the candidate object to obtain a motion detection result; When the motion detection result conforms to the characteristics of a moving target, determining the candidate object as a target object, and obtaining a target detection box of the target object; Based on an object tracking algorithm, performing a tracking operation on the target detection boxes of the target objects selected by the object detection algorithm and the motion detection algorithm, and obtaining a first motion feature corresponding to each target object; Based on a feature matching algorithm, performing feature matching calculation on the first motion feature and second motion features of at least one preset object, determining at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; Based on a preset weighted fusion mechanism, performing weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object; Based on the category matching degree between the target object and each preset object, determining the category recognition result of the target object.
2. The object recognition method based on motion detection according to claim 1, wherein The performing motion feature detection on the candidate object based on the motion detection algorithm to obtain a motion detection result includes: Based on the optical flow method, detecting the motion trend of the candidate object to obtain motion trend detection information of the candidate object; Based on the background difference method, performing pixel comparison between the current image region where the candidate object is located and a preset background model to obtain background region detection information of the candidate object; Based on the motion trend detection information and the background region detection information, obtaining the motion detection result corresponding to the candidate object.
3. The object recognition method based on motion detection according to claim 1, characterized in that, The performing a tracking operation on the target detection boxes of the target objects selected by the object detection algorithm and the motion detection algorithm based on the object tracking algorithm to obtain a first motion feature corresponding to each target object includes: Based on the object tracking algorithm, tracking the target detection boxes corresponding to the same target object in consecutive video frames to obtain a continuous coordinate set of the target object; Based on the continuous coordinate set, drawing a motion trajectory of the target object to calculate the first motion feature corresponding to the target object.
4. The object recognition method based on motion detection according to claim 1, wherein After identifying at least one object to be detected in the video stream data based on the object detection algorithm and obtaining a first confidence level for each category corresponding to the object to be detected, it further includes: When the first confidence level is less than the second confidence level threshold, determining the object in the target detection box as a non-target object.
5. The object recognition method based on motion detection according to claim 1, characterized in that, Based on the feature matching algorithm, perform feature matching calculations on the first motion feature and the second motion features of at least one preset object, and determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category, including: Based on a preset feature database, obtain the second motion features of at least one of the preset objects; Based on the feature matching algorithm, calculate the feature similarity between the first motion feature and each of the second motion features; Based on the feature similarity, determine at least one candidate category corresponding to the target object.
6. The object recognition method based on motion detection according to claim 5, characterized in that The determining at least one candidate category corresponding to the target object based on the feature similarity includes: When the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, determine the object category of the preset object as a candidate category of the target object.
7. The object recognition method based on motion detection according to claim 5, characterized in that, Before obtaining the second motion features of at least one of the preset objects based on the preset feature database, it further includes: Based on the application scenario requirements, determine at least one of the preset objects in the target scene and the category labels of each preset object; Based on the historical video stream data, collect the motion data of the preset object to obtain the second motion feature corresponding to the preset object; Based on the second motion feature corresponding to the preset object and the category label, construct the feature database.
8. An object recognition device based on motion detection, characterized in that, The object recognition device based on motion detection includes: A data acquisition module for acquiring video stream data to be detected; A target detection module for identifying at least one object to be detected in the video stream data based on a target detection algorithm and obtaining the first confidence level corresponding to the category of each object to be detected; A target object first determination module for determining the object to be detected as a target object when the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtaining the target detection frame of the target object; A candidate object determination module for determining the object to be detected as a candidate object when the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold; A motion detection module for performing motion feature detection on the candidate object based on a motion detection algorithm to obtain a motion detection result; A target object second determination module for determining the candidate object as a target object when the motion detection result conforms to the motion target feature, and obtaining the target detection frame of the target object; A target tracking module for performing tracking operations on the target detection frames of the target objects screened by the target detection algorithm and the motion detection algorithm based on a target tracking algorithm to obtain the first motion feature corresponding to each target object; A feature matching module for performing feature matching calculations on the first motion feature and the second motion features of at least one preset object based on a feature matching algorithm, and determining at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; A feature fusion module, configured to perform weighted fusion on the first confidence level and the feature similarity based on a preset weighted fusion mechanism, and determine the category matching degree between the target object and each of the preset objects; A category recognition module, configured to determine the category recognition result of the target object based on the category matching degree between the target object and each of the preset objects.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the object recognition method based on motion detection according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the object recognition method based on motion detection according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
A method for multi-sensor multi-vehicle tracking based on image and motion feature matching
GB202409843D0
System and method for static and moving object detection
US9454819B1