Object recognition method and device based on motion detection, equipment and storage medium
By adopting motion detection-based object recognition method in ultra-high resolution camera application scenarios, using confidence threshold screening and motion feature detection, combined with target tracking and feature matching algorithms, the problem that traditional detection algorithms are difficult to identify small objects is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510552220.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In the application scenarios of ultra-high resolution cameras, traditional detection algorithms based on single-frame visual features are difficult to accurately identify smaller birds, drones and other objects. Especially in long distances or high noise, the accuracy and confidence are difficult to meet actual needs.
The object recognition method based on motion detection is adopted, and the object to be detected in the video stream is initially identified through the object detection algorithm, the confidence threshold is set to filter out the target object and the candidate object, and the motion detection algorithm is used to detect the candidate object motion characteristics, and combined with the target tracking algorithm and the feature matching algorithm, weighted fusion is performed to improve the category matching degree.
It significantly improves the accuracy of image recognition, can more accurately determine the category of target objects, and makes up for the problem that traditional object detection algorithms lack small target recognition.
Smart Images

Figure CN120071031A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to an object recognition method, device, equipment, and storage medium based on motion detection. Background Art
[0002] With the continuous development of ultra-high-resolution camera array technology, it is possible to collect extremely wide and detailed image or video data at the same moment. However, in the actual application scenarios of ultra-high-resolution array cameras, although small target objects (such as vehicles) within 2 - 3 kilometers can be clearly seen through ultra-high-resolution technology, there are still smaller objects (especially small birds, drones, etc.) at farther distances, which often only occupy very few pixel points, resulting in difficulties for traditional single-frame vision feature-based detection algorithms to accurately identify. Taking birds as an example, even when using deep neural networks (such as YOLO series algorithms) to detect them, due to the target being too small or having too high a similarity to the background, there are still problems of coexistence of missed detections and false detections. Especially in the case of extremely far distances of birds or a large amount of noise in the picture, the accuracy and confidence of existing related recognition methods are difficult to meet the actual requirements. Therefore, how to improve the accuracy of image recognition has become an urgent technical problem to be solved currently. Summary of the Invention
[0003] This application provides an object recognition method, device, equipment, and storage medium based on motion detection, aiming to improve the accuracy of image recognition.
[0004] In a first aspect, this application provides an object recognition method based on motion detection, and the object recognition method based on motion detection includes the following steps: Obtain video stream data to be detected; Based on an object detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence level corresponding to each category of the objects to be detected; When the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, determine the object to be detected as a target object, and obtain the target detection frame of the target object; When the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold, determine the object to be detected as a candidate object; Based on a motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result; When the motion detection result conforms to the characteristics of a moving target, determine the candidate object as a target object, and obtain the target detection frame of the target object; Based on the target tracking algorithm, perform tracking operations on the target detection boxes of the target objects screened by the target detection algorithm and the motion detection algorithm to obtain the first motion features corresponding to each target object; Based on the feature matching algorithm, perform feature matching calculations on the first motion features and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; Based on a preset weighted fusion mechanism, perform weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object; Based on the category matching degree between the target object and each preset object, determine the category recognition result of the target object.
[0005] In a second aspect, the present application further provides an object recognition device based on motion detection. The object recognition device based on motion detection includes: A data acquisition module, configured to acquire video stream data to be detected; A target detection module, configured to identify at least one object to be detected in the video stream data based on a target detection algorithm, and obtain the first confidence level of each category corresponding to the object to be detected; A first target object determination module, configured to determine the object to be detected as a target object when the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtain the target detection box of the target object; A candidate object determination module, configured to determine the object to be detected as a candidate object when the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold; A motion detection module, configured to perform motion feature detection on the candidate object based on a motion detection algorithm to obtain a motion detection result; A second target object determination module, configured to determine the candidate object as a target object when the motion detection result conforms to the motion target feature, and obtain the target detection box of the target object; A target tracking module, configured to perform tracking operations on the target detection boxes of the target objects screened by the target detection algorithm and the motion detection algorithm based on a target tracking algorithm to obtain the first motion features corresponding to each target object; A feature matching module, configured to perform feature matching calculations on the first motion features and the second motion features of at least one preset object based on a feature matching algorithm to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; A feature fusion module, configured to perform weighted fusion on the first confidence level and the feature similarity based on a preset weighted fusion mechanism, and determine the category matching degree between the target object and each preset object; A category recognition module, configured to determine the category recognition result of the target object based on the category matching degree between the target object and each preset object.
[0006] In a third aspect, the present application further provides a computer device, including a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the object recognition method based on motion detection as described above are implemented.
[0007] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the object recognition method based on motion detection as described above are implemented.
[0008] The present application provides an object recognition method, device, computer device, and storage medium based on motion detection. The method of the present application accurately identifies a moving target object in a video stream through a target detection algorithm and obtains its first confidence level and first motion feature, initially screening out target objects with motion characteristics and reducing interference from factors such as static backgrounds. By setting a first confidence level threshold and a second confidence level threshold, small-volume candidate objects ignored by the target detection algorithm are screened out. Through a motion detection algorithm, further motion detection is performed on the candidate objects, and small target objects ignored by the target detection algorithm are screened and identified, making up for the deficiencies of the target detection algorithm, thereby more comprehensively and accurately detecting target objects. Through a target tracking algorithm, the target object is tracked, and the first motion feature of the target object is extracted. Through a feature matching algorithm, the first motion feature is matched and calculated with the second motion feature of a preset object to determine the candidate category and feature similarity, further limiting and screening the category of the target object from the perspective of motion features, increasing the dimension and accuracy of recognition. Through a preset weighted fusion mechanism, weighted fusion is performed on the first confidence level and the feature similarity, comprehensively considering information on both the category confidence level and the motion feature similarity of the target object, determining the category matching degree, and then obtaining the category recognition result, avoiding errors that may be caused by solely relying on a certain feature for recognition, thereby significantly improving the accuracy of image recognition and being able to more accurately determine the category of the target object. Description of the Drawings
[0009] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0010] Figure 1 It is a schematic flowchart of the first embodiment of an object recognition method based on motion detection provided by the present application; Figure 2 It is a schematic flowchart of the detection process of small target moving objects provided by the present application; Figure 3 It is a schematic structural diagram of the first embodiment of an object recognition device based on motion detection provided by the present application; Figure 4 It is a schematic block diagram of the structure of a computer device provided by the embodiments of the present application.
[0011] The realization of the purpose of the present application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0013] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.
[0014] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0015] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of an object recognition method based on motion detection provided by the present application.
[0016] As Figure 1 shown, the object recognition method based on motion detection includes steps S101 to S110.
[0017] S101. Obtain the video stream data to be detected; In one embodiment, the video stream data can be a real-time video stream captured in real time by a video capture device, or a historical video stream captured by the video capture device.
[0018] Exemplarily, the video capture device can capture the dynamic scenes of the monitoring area in real time. For example, in a traffic monitoring scenario, an array camera deployed at a road intersection or other locations can continuously transmit real-time moving images of vehicles, pedestrians, etc. to the background detection system in the form of a video stream; or, in an airport monitoring scenario, the array camera can continuously capture aerial images and transmit them to the background detection system in the form of a video stream, so as to monitor in real time possible obstacles such as birds, drones, balloons, etc. in the airspace above the airport. The real-time video stream can be used to perform real-time analysis on events in the current scene of the monitoring area.
[0019] Exemplarily, the video stream data can also be historical video stream data that has been captured and stored by the video capture device, such as video stream data periodically uploaded by the video capture device. In some scenarios, when the amount of video stream data is too large or in non-real-time analysis requirements, the data upload period of the video capture device can be set. For example, the video stream data is transmitted once every hour, that is, the video stream data captured by the video capture device in the past hour is obtained, so as to analyze the video stream data of this time period. At this time, the video stream data of this time period can be regarded as historical video stream data.
[0020] Exemplarily, the video capture device can be an array camera with ultra-high resolution.
[0021] S102. Based on the object detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence level of each category corresponding to the object to be detected; In one embodiment, a preliminary screening can be performed by running common object detection methods (such as algorithms like the YOLO series, EfficientDet or Nano Det).
[0022] In one embodiment, detection algorithms such as YOLO can be used to process and analyze the video stream data, and initially output the classes, boxes, and confidence levels of the target objects.
[0023] Generally, the confidence level reflects the credibility of the model that there is indeed such a target object in the detection box. Its value is between 0 and 1, and the closer the value is to 1, the more credible the model's judgment of the target object in the detection box is. For example, if the confidence level of the detection box is 0.95, it means that the model believes that the probability of there being a target object in this box is very high.
[0024] In one embodiment, the object detection algorithm can process each frame of the video stream data frame by frame, and identify the coordinates, contours, categories, and corresponding confidence levels of the target objects in the image. The object detection algorithm can output multiple categories that the target object may correspond to. For example, for an object flying in the air, the categories it may correspond to can include birds, drones, etc.
[0025] Exemplarily, after the object detection algorithm processes each frame of the video stream data, a detection box corresponding to each target object is obtained. For example, for a video image containing multiple pedestrians, the YOLOv7 model can generate a rectangular box (detection box) for each pedestrian, indicating the position of the pedestrian in the image.
[0026] Exemplarily, a confidence threshold, such as 0.7 or 0.8, can be preset according to the requirements for the accuracy of the detection results in the actual application scenario. For example, in a security monitoring scenario, in order to reduce false alarms, the confidence threshold can be set to 0.8 to ensure that only the detection results with a high degree of confidence in the model are retained.
[0027] Specifically, the first confidence level of each target detection box is compared with the preset confidence threshold. When the first confidence level is greater than or equal to the confidence threshold, the object within the target detection box is considered a target object; otherwise, it is considered that the object within the detection box is not a target object or the detection is not accurate enough, and it is filtered out. For example, when the confidence threshold is set to 0.7, if the confidence level of a certain detection box is 0.6, then this detection box will not be retained as a target object.
[0028] S103. When the first confidence level corresponding to the object to be detected is greater than or equal to the preset first confidence threshold, determine that the object to be detected is a target object, and obtain the target detection box of the target object; S104. When the first confidence level corresponding to the object to be detected is less than the first confidence threshold and greater than the second confidence threshold, determine that the object to be detected is a candidate object; In one embodiment, when the first confidence level is less than the second confidence threshold, determine that the object in the target detection box is a non-target object.
[0029] In one embodiment, two thresholds can be set: the first confidence threshold Th1 and the second confidence threshold Th2. Among them, Th2 < Th1.
[0030] If the first confidence level of a certain detection box is higher than Th1, it can be considered that the detection result is relatively reliable, and the object in this detection box is directly marked as "target object".
[0031] If the confidence level of the detection box is lower than Th2, it is temporarily abandoned, unless subsequent motion detection gives new clues to the same area.
[0032] In one embodiment, for the detection results determined to be non-target objects, it can be considered that they have nothing to do with the target category and can be excluded from subsequent processing and analysis, thereby reducing the waste of computing resources. However, it should be noted that sometimes non-target objects may have a certain similarity in appearance to target objects. For example, in a detection task where the target category is a car, some trucks or motorcycles may be misdetected as cars, but due to their low confidence, they are excluded from the target objects. To avoid this situation, the image can be filtered and enhanced in the preprocessing stage, such as using image segmentation technology to divide the image into different regions, or performing operations such as histogram equalization on the image, to improve the detectability of target objects and the distinguishability of non-target objects.
[0033] If the confidence of the detection box is between Th2 and Th1, it is temporarily marked as a "candidate target" and needs to enter the subsequent motion detection and tracking module to assist in the judgment.
[0034] In one embodiment, for the detection results determined to be candidate objects, although their confidence is lower than the first confidence threshold, they still have a certain degree of credibility and may contain potential target objects. For these candidate objects, various methods can be used for further processing to improve the accuracy and reliability of the detection. For example, in the image, the objects or scene information around the candidate object may help determine whether it is a target object. If the candidate object is located in a scene where common target objects appear, or the layout and relationship with surrounding objects conform to the characteristics of the target object, then the possibility of it being a target object can be increased.
[0035] In this embodiment, the confidence threshold is an important parameter in the target detection algorithm for screening and classifying detection results. By setting different confidence thresholds, the accuracy and recall rate of the detection can be controlled. When the first confidence threshold is relatively high, only detection results with relatively high confidence will be determined as target objects, which can improve the accuracy of the detection and reduce false detections. However, some real target objects may be missed due to slightly lower confidence, reducing the recall rate. The opposite is true when the second confidence threshold is relatively low, which can increase the recall rate.
[0036] S105. Based on the motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result; In one embodiment, the target detection algorithm can identify target objects with obvious features in the target scene, but there may be problems of inaccurate identification or failure to identify small targets in the scene. For example, in an aerial monitoring scene, if a small bird is far from the video acquisition device, the number of pixels it occupies in the captured image is very small. The result is that although the bird is in the captured image, due to the small number of pixels it occupies, the target detection algorithm may not detect the bird target. At this time, a motion detection algorithm can be used to detect dynamic targets in consecutive video frames to identify the bird target.
[0037] Specifically, through the above first confidence threshold and second confidence threshold, the detection results of the target detection algorithm are screened into three parts, namely, the target objects clearly detected (the first confidence is greater than or equal to the first confidence threshold), the candidate objects (the first confidence is between the first confidence threshold and the second confidence threshold), and the non-target objects (the first confidence is less than the second confidence threshold).
[0038] Among them, the candidate objects are the detection objects that the target detection algorithm cannot determine but may be moving targets. Therefore, it is necessary to use the motion detection algorithm to detect the motion characteristics of the candidate objects, that is, to detect whether there will be changes in motion characteristics in consecutive video frames, such as continuous displacement, periodic posture change, etc.
[0039] Such as Figure 2 As shown, when the target detection algorithm fails to identify small target objects, a motion detection algorithm, such as based on the optical flow method or frame difference method, can be used to detect the motion characteristics of the objects determined as candidate objects or non-target objects in the confidence threshold discrimination. For example, in traffic monitoring, the optical flow method can be used to calculate motion characteristic information such as the speed and direction of a target vehicle over time.
[0040] In one embodiment, the motion detection algorithm can detect whether there are moving targets in each image according to the object coordinates in consecutive video frames. For some motion detection algorithms, it is possible to determine whether there is motion by tracking the coordinate changes of the target object in consecutive video frames. For example, the center coordinates (x, y) and width and height (w, h) of the target object are obtained in each frame through the target detection model, and then the change of these coordinate values in the time series is observed. If there are large differences in the coordinate values in consecutive frames, it can be inferred that there is motion.
[0041] In one embodiment, the motion detection algorithm can also detect motion by comparing the differences between adjacent video frames. The pixel value differences between the front and rear frame images can be calculated, or whether motion occurs can be determined by comparing the grayscale images, feature points, etc. of the front and rear frames. For example, the frame difference method is a common comparison algorithm. By performing a differential operation on two adjacent frame images, if the difference exceeds a certain threshold, it is considered that there is a moving object.
[0042] Specifically, the motion detection algorithm can include the frame difference method, the optical flow method, the background subtraction method, etc.
[0043] The frame difference method refers to performing a pixel-by-pixel differential operation between two adjacent frame images, regarding the pixels with a difference less than the threshold as the static background, and the pixels with a difference greater than the threshold as the motion area.
[0044] The optical flow method calculates the motion direction and speed of the pixels in the image by imposing conditions such as the gray-scale consistency constraint of the pixels between consecutive frames, forming an optical flow field, thereby detecting the moving object.
[0045] The background subtraction method first establishes a background model, usually by observing the image features of the scene without moving objects for a long time. Then, the current frame is compared with the background model, and the regions with large differences are considered as the regions with moving objects.
[0046] In one embodiment, various motion detection and tracking algorithms can be used to track the position coordinates of the target object in consecutive video frames, such as the optical flow method, the background difference method, etc.
[0047] Furthermore, based on the optical flow method, detect the motion trend of the candidate object to obtain the motion trend detection information of the candidate object; based on the background difference method, perform pixel comparison between the current image region where the candidate object is located and a preset background model to obtain the background region detection information of the candidate object; based on the motion trend detection information and the background region detection information, obtain the motion detection result corresponding to the candidate object.
[0048] Taking the optical flow method as an example, the basic algorithms include the Lucas-Kanade, Horn-Schunck, and Farneback algorithms. The optical flow method calculates the motion direction of the pixels in the picture using the brightness difference at the pixel level between the front and rear frame images, plus the small motion constraint and motion continuity.
[0049] Specifically, the optical flow algorithm analyzes the brightness changes of pixel points in consecutive frame images. For example, the classic brightness constancy assumption (i.e., assuming that the brightness value of a pixel point remains unchanged during movement) is used as the basis to construct equations to solve the optical flow field. The Lucas-Kanade algorithm, etc., are typical methods based on this local optical flow calculation. When calculating, it mainly focuses on the changes within the local pixel neighborhood to estimate the motion of this local area. The small motion constraint means that within a very short time (i.e., between adjacent frames), the motion of an object will not be too large, which can simplify the difficulty of equation solving; and the motion continuity is reflected in that the motion field is relatively smooth and coherent in space. Based on these assumptions, algorithms such as the Horn-Schunck algorithm can solve a relatively reasonable optical flow field and obtain the motion direction and speed estimates of each pixel point.
[0050] Similarly, the function of the background subtraction method is similar, that is, to find the moving areas in the picture for the detection and analysis of moving objects. They both achieve this function based on the differences between consecutive frame images, but only the principles and methods used are different.
[0051] The background subtraction method and the optical flow algorithm have a certain complementary effect. In some scenarios, the background subtraction method may detect motions that the optical flow algorithm cannot detect. For example, when the appearance of a moving object is significantly different from the background in terms of color, texture, etc., and the motion amplitude is relatively large, the background subtraction method can easily identify the moving area by comparing with the background model; while the optical flow algorithm may not be able to accurately detect the motion because the large motion amplitude causes the assumption based on the small motion constraint to fail. On the contrary, the optical flow algorithm may perform better than the background subtraction method in scenarios where higher requirements are placed on the estimation of motion direction and speed, or where the background changes are complex but the moving objects are relatively small and move smoothly. Therefore, the two can complement each other and be used in combination to improve the accuracy and robustness of motion detection.
[0052] The algorithm complexity of the background subtraction method is relatively simpler than that of the optical flow algorithm. It mainly compares each pixel of the current frame image with the pre-established background model and determines whether a pixel point belongs to the moving area according to the set threshold.
[0053] Specifically, the background subtraction method is based on the comparison between the background model and the current frame image. Usually, a background model is established in a certain way. For example, a simple background model can be constructed by statistically analyzing the video frames in the initial period of time, calculating statistical quantities such as the average value and variance of each pixel, or more complex methods based on mixture Gaussian models, etc. are used to model the background to better adapt to the situation where there are some dynamic changes in the background (such as swaying branches, fluctuating water surfaces, etc.). During actual detection, for each newly input image, it is compared pixel by pixel with the background model. If the difference in features such as color and brightness between a certain pixel point in the current frame and the corresponding pixel in the background model exceeds the set threshold, it is considered that the pixel point belongs to the moving area. Then, through post-processing operations such as connected region analysis of these moving pixel points, the complete moving target area can be obtained, and then the detection and tracking of moving objects can be realized.
[0054] In the small target scenario, since the target itself occupies fewer pixels, the target detection algorithm may fail to detect the target due to factors such as the target being not obvious and easily occluded by the background. By reasonably combining the optical flow method and the background subtraction method and adjusting the operation weights and functions of the two according to the specific scenario and computing power conditions, the motion detection and analysis of the area where the small target is located can be better realized. Furthermore, by analyzing the motion of the pixels in the area where the small target is located, the motion of the target can be continuously tracked to make up for the deficiencies of the target detection algorithm. For example, it is very useful in scenarios such as monitoring the motion trajectories of some small wild animals in the distance in surveillance videos.
[0055] S106. When the motion detection result conforms to the characteristics of a moving target, determine the candidate object as the target object and obtain the target detection frame of the target object. In this embodiment, by introducing a lower second confidence threshold under the conventional first confidence threshold used in the target detection algorithm, the candidate targets with the first confidence between the two are used as potential detection targets, so as to screen out the detection objects with low confidence due to the target being too small or partially occluded; by analyzing the motion trend of pixels between consecutive frames using the optical flow method, the motion direction and speed of the target can be judged, so as to capture the motion characteristics when the target moves in the picture as the screening basis for the target object; by the background difference method, the foreground target is detected by comparing the difference between the current frame and the background model. Especially when the target moves in the background, the background difference method can effectively capture the motion trajectory of the target.
[0056] These three methods do not exist in isolation in the motion detection module, but are interrelated and complementary. The confidence threshold screening method can provide a preliminary candidate frame range, but whether these frames actually contain target objects requires further verification by the optical flow method and the background difference method. The optical flow method can capture the motion trend of the target, while the background difference method can detect the difference between the target and the background. When the candidate frames screened out by the confidence threshold screening method have a motion trend (detected by the optical flow method) and are significantly different from the background (detected by the background difference method), it can be more accurately determined whether there are moving target objects in these candidate frames. Through the joint detection of the three methods, the candidate targets that the target detection algorithm cannot identify are further screened. When the three methods verify that there are moving objects in the candidate targets, the candidate targets are identified as target objects, and then the candidate frame where the target object is located is used as the target detection frame that the target tracking algorithm needs to track and calculate, so as to track its motion trajectory and extract motion features.
[0057] For example, in a monitoring scene, a small animal is moving through the woods. Due to the small size of the animal, the target detection algorithm may not be able to accurately identify it. At this point, the motion detection module comes into play. The confidence threshold screening method retains those low-confidence candidate boxes that may contain small animals. The optical flow method captures the motion trend of the pixels in these boxes, indicating that an object is moving. The background difference rule finds that these boxes are significantly different from the background model, further confirming the existence of the target, and then determines the low-confidence candidate box as the target detection box of the target object. Through the collaborative work of these three methods, the presence of small animals can be detected more accurately and tracked.
[0058] In short, these three methods in the motion detection module can detect target objects more comprehensively and accurately through mutual correlation and collaboration, especially small target objects that are easily ignored by traditional target detection algorithms. The introduction of this method not only makes up for the shortcomings of traditional target detection algorithms, but also provides new ideas and solutions for target detection in complex scenes.
[0059] In one embodiment, the moving target object in the target scene can be identified and tracked by the target detection algorithm and the motion detection algorithm, and during the identification and tracking process, the size of the target object can be calculated by the size of the detection frame identified by the target detection algorithm or the motion detection algorithm. Then, the category of the target object is auxiliary identified according to the size of the target object.
[0060] Because the distance of different types of small objects, such as birds, airplanes, and drones, can be roughly estimated from the size of their detection boxes in the picture. Then, the category of the identified target object is determined by combining the motion characteristics of the target object, including speed, acceleration, motion trajectory characteristics, etc.
[0061] Specifically, for example, the appearance, scale, motion trajectory, and motion characteristics of an aircraft are very different from those of a bird. These characteristics can be collected for modeling. For example, the Kalman filter algorithm, support vector machine, or more complex neural network pattern recognition algorithms can be used to learn and model these characteristics, enabling the model to recognize characteristics such as the appearance, scale, behavior pattern, and motion characteristics of different types of small objects. Then, the learned model is used to perform pattern recognition on the object detection results (including parameters such as the scale of the object detection box and the motion characteristics of the target object), and then the recognition result of the model for the target object is output.
[0062] S107. Based on the object tracking algorithm, perform tracking operations on the object detection boxes of the target objects screened by the object detection algorithm and the motion detection algorithm to obtain the first motion characteristics corresponding to each of the target objects. Exemplarily, the first motion characteristics may include characteristic information such as the motion trajectory, motion direction, speed (absolute speed or relative speed), and displacement of the target object. For example, by comparing the position changes of the target object frame by frame and combining the frame rate information of the video, the displacement vector of the object in each frame can be calculated, and then characteristics such as its motion speed and direction can be obtained. These motion characteristic information play an important role in understanding the behavior pattern and state of the target object and can be used for requirements such as trajectory tracking and behavior analysis.
[0063] In one embodiment, after obtaining the motion detection results, it is necessary to continuously track each moving target. Common methods include Kalman Filter, SORT, DeepSORT, etc. Continuously tracking the moving target can obtain the tracking results of the target object between consecutive video frames, including relatively accurate target center positions, motion trajectories, and information such as size, speed, and acceleration that change over time.
[0064] Further, based on the object tracking algorithm, track the object detection boxes corresponding to the same target object in consecutive video frames to obtain a continuous coordinate set of the target object; based on the continuous coordinate set, draw the motion trajectory of the target object to calculate the first motion characteristics corresponding to the target object.
[0065] In one embodiment, other first motion characteristics of the target object, such as position change, speed, direction, acceleration, etc., can be further calculated according to the motion trajectory of the target object.
[0066] Specifically, using an array camera or other two-dimensional tracking system, continuously obtain the positions of the target object at each time point to form a series of two-dimensional coordinate points and record the corresponding timestamps , thus, in the order of time stamps, each two-dimensional coordinate point is connected in sequence to obtain the motion trajectory of the target object, and through the motion trajectory, other first motion characteristics such as the position change, speed, and acceleration of the target object can be calculated.
[0067] Exemplarily, the position change of the target object can be deduced from the motion trajectory: in a two-dimensional space, the position coordinates of the target object first recognized in the video stream data are used as the origin, and any subsequent target position can be represented by a position vector from the origin to that point as:
[0068] During the time interval , the position change vector of the target object is:
[0069] Exemplarily, the speed of the target object can be calculated from the motion trajectory: The instantaneous velocity vector is:
[0070] In the discrete case, the average velocity vector can be approximately calculated within the time interval as:
[0071] The component form of the velocity vector is:
[0072] The speed of the target object at the i-th time point can be expressed as:
[0073] Exemplarily, the acceleration of the target object can be calculated from the motion trajectory: The instantaneous acceleration vector can be expressed as:
[0074] In the discrete case, the average acceleration vector can be approximately calculated within the time interval as:
[0075] The component form of the acceleration can be expressed as:
[0076] The magnitude of the acceleration of the target object at any moment t can be expressed as:
[0077] In this embodiment, by analyzing the two-dimensional motion trajectory of the target object, the first motion feature of the target object is extracted, so as to facilitate identifying the category of the target object according to the first motion feature, improving the accuracy of target object category recognition, and assisting in judging the impact of the target object on the monitoring scene, thereby facilitating the formulation of corresponding countermeasures for the impact of the target object.
[0078] S108. Based on the feature matching algorithm, perform feature matching calculation on the first motion feature and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; In one embodiment, a feature database including the second motion features of at least one preset object can be pre-constructed. Each preset object has a corresponding feature vector for representing the second motion feature. These feature vectors can be obtained through experimental measurement, historical data statistics, or machine learning model training.
[0079] Exemplarily, the preset objects can be collected according to the application scenario requirements. For example, in the traffic monitoring application scenario, the preset objects can include vehicles of different models (such as bicycles, electric vehicles, motorcycles, cars, trucks, etc.), pedestrians, animals, etc.; in the airport airspace monitoring application scenario, the preset objects can include airplanes, drones, flying birds, balloons, kites, etc.
[0080] In one embodiment, based on the application scenario requirements, determine at least one of the preset objects in the target scenario and the category labels of each preset object; based on the historical video stream data, collect the motion data of the preset objects to obtain the second motion features corresponding to the preset objects; based on the second motion features corresponding to the preset objects and the category labels, construct the feature database.
[0081] In one embodiment, to construct a feature database, it is first necessary to clarify the application scenario requirements of the target scenario. For example, in the traffic monitoring scenario, objects such as vehicles, pedestrians, and animals need to be concerned; in the airport airspace monitoring scenario, objects such as airplanes, drones, flying birds, balloons, and kites are the focus. According to the scenario requirements, determine which objects are typical preset objects that need to be recognized. These objects should be representative and cover common object categories in the scenario.
[0082] In a specific embodiment, segments containing preset objects are screened out from historical video stream data such as scene-related surveillance videos. For example, in a traffic surveillance scenario, video segments containing passing vehicles, pedestrians, and animals are selected. The preset objects in the video are tracked, and their motion data such as motion trajectories, speeds, and accelerations are recorded. These data can be automatically extracted using computer vision algorithms (such as object detection and tracking algorithms), or can be recorded through manual annotation. The collected motion data is cleaned and preprocessed to remove noise and outliers to ensure the accuracy and consistency of the data. For example, the motion trajectory data is smoothed to remove mutation points caused by reasons such as unstable video frame rates. According to the application scenario and the characteristics of object motion, appropriate feature extraction methods are selected. For example, for vehicles, features such as average speed, acceleration, and steering angular velocity can be extracted, and features such as contours and identifiers can also be included; for flying birds, features such as contours, feather colors, feather characteristics, flying speed, wing flapping frequency, and curvature of flight trajectories can be extracted. The extracted motion features are combined into feature vectors. For example, the feature vector of an airplane can include features such as speed, acceleration, steering angular velocity, and motion trajectory type (such as straight line, curve). The feature vectors are normalized to make their numerical ranges and distributions consistent for subsequent matching calculations. For example, the speed and acceleration features are normalized to the [0, 1] interval. Category labels are added to the feature vectors of each preset object. The category labels should clearly identify the category to which the object belongs, such as "bicycle", "drone", "pedestrian", etc., ensuring that the feature vectors of objects in the same category have the same category label and that the labels of different categories are not confused.
[0083] In one embodiment, after the feature extraction and category label marking of the preset objects in the target scene are completed, a feature database can be constructed using an appropriate data structure and storage method, such as a relational database or a non-relational database. For example, in a relational database, tables are used to store feature vectors and category labels. Each row represents a preset object, and the columns include the various dimensions of the feature vector and the category label.
[0084] Generally, an index can be established for the feature database to improve the efficiency of feature query and matching. An index can be established based on certain key dimensions of the feature vector, such as speed and motion trajectory type. At the same time, as the application scenario changes and new data accumulates, the feature database is updated in a timely manner to add new preset object features and optimize existing features.
[0085] This embodiment constructs a feature database based on the requirements of the application scenario. The database contains the second motion features and category labels of the preset objects and can be used for category recognition and feature matching of target objects.
[0086] In one embodiment, the first motion feature may include features such as the motion trajectory, speed, acceleration, etc. of the target object, and the second motion feature may also include features such as the motion trajectory, speed, acceleration, etc. of a preset object.
[0087] In one embodiment, the second motion feature may further include the contour feature of the preset object, as well as the contour change feature during the movement process, etc. For example, a flying bird will have the action of flapping its wings during flight, and the wing flapping state is different in different flight states such as gliding and changing direction. When collecting the first motion feature of the target object, the contour feature of the target object during the movement process may also be collected to assist in identifying the category of the target object.
[0088] Further, based on a preset feature database, obtain the second motion features of at least one of the preset objects; based on the feature matching algorithm, calculate the feature similarity between the first motion feature and each of the second motion features; based on the feature similarity, determine at least one candidate category corresponding to the target object.
[0089] In one embodiment, the second motion features of multiple preset objects are stored in the preset feature database, and these features are important bases for the matching analysis of the target object. Using the feature matching algorithm, by comparing the first motion feature of the target object with the second motion feature of the preset object in the feature database, calculate the similarity between the two, so as to determine the candidate category of the target object.
[0090] In one embodiment, the feature matching algorithm may include one or more of algorithms such as Euclidean distance, cosine similarity, Pearson correlation coefficient, etc. The principles and applicable scenarios of each algorithm are different, and can measure the similarity degree between features from different angles.
[0091] Among them, the Euclidean distance is used to calculate the straight-line distance between two feature vectors in a multi-dimensional space. The smaller the distance, the higher the similarity. The cosine similarity is used to calculate the cosine value of the included angle between two feature vectors. The closer the value is to 1, the higher the similarity. The Pearson correlation coefficient is used to measure the linear correlation degree between two feature vectors. The value ranges from -1 to 1, and the greater the absolute value, the stronger the correlation.
[0092] In one embodiment, when the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, determine the object category of the preset object as the candidate category of the target object.
[0093] In one embodiment, use the feature matching algorithm to calculate the feature similarity between the first motion feature of the target object and the second motion feature of the preset object. According to the calculated feature similarity and the preset similarity threshold, determine the candidate category corresponding to the target object.
[0094] If the feature similarity is greater than or equal to the similarity threshold, it is considered that the target object belongs to the category of the preset object; if the feature similarities of multiple preset objects are greater than or equal to the similarity threshold, these preset objects are all used as candidate categories.
[0095] S109. Based on a preset weighted fusion mechanism, perform weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object; Generally, there are often inconsistencies or differences between the results of motion model matching and the category confidence levels output by detection models such as YOLO. For example, the YOLO model may have inaccurate category confidence levels due to factors such as lighting changes and target occlusion; while the motion model may make misjudgments due to short-term anomalies in the target motion path. Therefore, it is necessary to comprehensively consider these two pieces of information and, through weighted fusion, make full use of their respective advantages to improve the reliability of the final category matching degree.
[0096] Specifically, a weighted fusion mechanism can be designed:
[0097] Among them, and are weight parameters preset in advance or updated according to dynamic situations, satisfying + = 1; represents the confidence level of the motion detection model for category , coming from the output of motion detection models such as YOLO; represents the feature matching degree of the motion detection model for category , which is the feature similarity calculated through the feature matching algorithm.
[0098] In one embodiment, the setting of the weight parameters and can be adjusted according to actual application requirements.
[0099] Exemplarily, the weight parameters and can be adjusted according to the credibility of the motion detection model. If the motion detection model (such as YOLO) shows a high accuracy rate in a specific scenario, the value of can be appropriately increased, and vice versa, the value of is decreased; similarly, the value of is adjusted according to the performance of the motion detection model in this scenario.
[0100] The weight parameters and It can also be adjusted according to the real-time motion state of the target object. When the motion state of the target object is relatively stable, the results of the motion detection model may be more reliable, and the value can be appropriately increased; while when the motion state of the target object is complex and changeable, the results of the detection model may be more valuable for reference, and the value can be increased.
[0101] Weight parameter and can also be dynamically adjusted. A dynamic adjustment mechanism can be designed to automatically adjust and values according to the real-time detection confidence, motion confidence, and historical credibility records, etc. For example, when the detection confidence remains high, gradually increase the weight; when there is a significant change in the motion confidence, increase the weight.
[0102] The weighted fusion mechanism can comprehensively consider the advantages of the motion detection model and the feature matching algorithm, improve the recognition accuracy of categories, and reduce the impact of misjudgment of a single model. For example, in a vehicle detection scenario, when the confidence of the detection model decreases due to partial occlusion of the vehicle, the motion model can supplement information through the motion characteristics of the vehicle, and improve the overall recognition effect through weighted fusion.
[0103] S110. Determine the category recognition result of the target object based on the category matching degree between the target object and each of the preset objects.
[0104] In one embodiment, a comprehensive category matching degree is obtained by fusing the confidence and feature similarity. Compare the category matching degrees after fusing the confidence and feature similarity of the target object and each preset object, and rank according to the fused score to select the highest one or those above a certain threshold (such as 0.5) as the final output category of the algorithm. This takes into account both the discriminative ability of the visual model for appearance features and the judgment ability of the motion model for dynamic features.
[0105] This embodiment provides an object recognition method based on motion detection, which accurately identifies the moving target object in the video stream through the motion detection algorithm and obtains its first confidence and first motion feature, preliminarily screens out the target object with motion characteristics, and reduces the interference of factors such as static background. By setting the first confidence threshold and the second confidence threshold, the small volume candidate objects ignored by the target detection algorithm are screened out, and the motion detection algorithm is used to further perform motion detection on the candidate objects, and the small target objects ignored by the target detection algorithm are screened and identified, which makes up for the shortcomings of the target detection algorithm, thereby detecting the target object more comprehensively and accurately. The target object is tracked by the target tracking algorithm and the first motion feature of the target object is extracted. The first motion feature is matched and calculated with the second motion feature of the preset object by the feature matching algorithm, and the candidate category and feature similarity are determined, and the target object category is further limited and screened from the perspective of motion features, which increases the dimension and accuracy of recognition. The first confidence and feature similarity are weightedly fused through a preset weighted fusion mechanism, which comprehensively considers the information of the target object in terms of category confidence and motion feature similarity, determines the category matching degree, and then obtains the category recognition result, avoiding the errors that may be caused by identification based on a single feature, thereby significantly improving the accuracy of image recognition and being able to more accurately determine the category of the target object.
[0106] See also Figure 3 , Figure 3 It is a structural schematic diagram of a first embodiment of a motion detection-based object recognition device provided in the present application. The motion detection-based object recognition device is used to execute the aforementioned motion detection-based object recognition method.
[0107] like Figure 3 As shown, the object recognition device 200 based on motion detection includes: a data acquisition module 201, a target detection module 202, a first target object determination module 203, a candidate object determination module 204, a motion detection module 205, a second target object determination module 206, a target tracking module 207, a feature matching module 208, a feature fusion module 209 and a category recognition module 210.
[0108] The data acquisition module 201 is used to acquire the video stream data to be detected; The target detection module 202 is used to identify at least one object to be detected in the video stream data based on a target detection algorithm, and obtain a first confidence level corresponding to a category of each object to be detected; A first target object determination module 203 is used to determine that the object to be detected is a target object when a first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtain a target detection frame of the target object; A candidate object determination module 204, configured to determine the object to be detected as a candidate object when a first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold; A motion detection module 205, configured to perform motion feature detection on the candidate object based on a motion detection algorithm to obtain a motion detection result; A target object second determination module 206, configured to determine the candidate object as a target object when the motion detection result conforms to motion target characteristics, and obtain a target detection frame of the target object; A target tracking module 207, configured to perform a tracking operation on a target detection frame of the target object screened out by the target detection algorithm and the motion detection algorithm based on a target tracking algorithm to obtain a first motion feature corresponding to each target object; A feature matching module 208, configured to perform feature matching calculation on the first motion feature and a second motion feature of at least one preset object based on a feature matching algorithm, determine at least one candidate category corresponding to the target object, and a feature similarity between the target object and the preset object corresponding to each candidate category; A feature fusion module 209, configured to perform weighted fusion on the first confidence level and the feature similarity based on a preset weighted fusion mechanism to determine a category matching degree between the target object and each preset object; A category recognition module 210, configured to determine a category recognition result of the target object based on the category matching degree between the target object and each preset object.
[0109] In one embodiment, the motion detection module 205 includes: A motion trend detection unit, configured to detect a motion trend of the candidate object based on an optical flow method to obtain motion trend detection information of the candidate object; A background detection unit, configured to perform pixel comparison on a current image area where the candidate object is located and a preset background model based on a background difference method to obtain background area detection information of the candidate object; A motion detection evaluation unit, configured to obtain the motion detection result corresponding to the candidate object based on the motion trend detection information and the background area detection information.
[0110] In one embodiment, the target tracking module 207 includes: A target tracking unit, configured to track the target detection frame corresponding to the same target object in consecutive video frames based on the target tracking algorithm to obtain a continuous coordinate set of the target object; A first motion feature calculation unit, configured to draw a motion trajectory of the target object based on the continuous coordinate set, so as to calculate a first motion feature corresponding to the target object.
[0111] In one embodiment, the target detection module 202 further includes: A non-target object determination unit, configured to determine that the object in the target detection frame is a non-target object when the first confidence level is less than the second confidence level threshold.
[0112] In one embodiment, the feature matching module 208 includes: A second motion feature acquisition unit, configured to acquire a second motion feature of at least one of the preset objects based on a preset feature database; A feature similarity calculation unit, configured to calculate a feature similarity between the first motion feature and each of the second motion features based on the feature matching algorithm; A candidate category determination unit, configured to determine at least one candidate category corresponding to the target object based on the feature similarity.
[0113] In one embodiment, the candidate category determination unit includes: A candidate category determination subunit, configured to determine the object category of the preset object as a candidate category of the target object when the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold.
[0114] In one embodiment, the object recognition device 200 based on motion detection further includes a feature database construction module, including: A scene analysis unit, configured to determine at least one of the preset objects in the target scene and category labels of each of the preset objects based on application scenario requirements; A second motion feature acquisition unit, configured to collect motion data of the preset object based on historical video stream data, and obtain the second motion feature corresponding to the preset object; A feature database construction unit, configured to construct the feature database based on the second motion feature corresponding to the preset object and the category label.
[0115] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing embodiment of the object recognition method based on motion detection, and will not be described herein again.
[0116] The device provided in the above embodiment can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 shown.
[0117] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device may be a server.
[0118] Refer to Figure 4 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0119] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which when executed, can cause the processor to execute any object recognition method based on motion detection.
[0120] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0121] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any object recognition method based on motion detection.
[0122] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in
[0123] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0124] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps: Obtain video stream data to be detected; Based on the object detection algorithm, identify at least one object to be detected in the video stream data, and obtain the first confidence level of each object to be detected corresponding to its category; When the first confidence level corresponding to the object to be detected is greater than or equal to the preset first confidence level threshold, determine the object to be detected as the target object, and obtain the target detection frame of the target object; When the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than the second confidence level threshold, determine the object to be detected as a candidate object; Based on the motion detection algorithm, perform motion feature detection on the candidate object to obtain a motion detection result; When the motion detection result conforms to the motion target characteristics, determine the candidate object as the target object, and obtain the target detection frame of the target object; Based on the object tracking algorithm, perform tracking operations on the target detection frames of the target objects screened by the object detection algorithm and the motion detection algorithm to obtain the first motion features corresponding to each target object; Based on the feature matching algorithm, perform feature matching calculations on the first motion features and the second motion features of at least one preset object to determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category; Based on the preset weighted fusion mechanism, perform weighted fusion on the first confidence level and the feature similarity to determine the category matching degree between the target object and each preset object; Based on the category matching degree between the target object and each preset object, determine the category recognition result of the target object.
[0125] In one embodiment, when the processor implements performing motion feature detection on the candidate object based on the motion detection algorithm to obtain a motion detection result, it is used to implement: Based on the optical flow method, detect the motion trend of the candidate object to obtain the motion trend detection information of the candidate object; Based on the background difference method, perform pixel comparison on the current image area where the candidate object is located and the preset background model to obtain the background area detection information of the candidate object; Based on the motion trend detection information and the background area detection information, obtain the motion detection result corresponding to the candidate object.
[0126] In one embodiment, when the processor implements performing tracking operations on the target detection frames of the target objects screened by the object detection algorithm and the motion detection algorithm based on the object tracking algorithm to obtain the first motion features corresponding to each target object, it is used to implement: Based on the target tracking algorithm, track the target detection boxes corresponding to the same target object in consecutive video frames to obtain a continuous coordinate set of the target object; Based on the continuous coordinate set, draw the motion trajectory of the target object to calculate the first motion feature corresponding to the target object.
[0127] In one embodiment, after the processor implements the target detection algorithm to identify at least one object to be detected in the video stream data and obtain the first confidence level of each category corresponding to the object to be detected, it is further configured to implement: When the first confidence level is less than the second confidence threshold, determine that the object in the target detection box is a non-target object.
[0128] In one embodiment, when the processor implements the feature matching algorithm to perform feature matching calculations on the first motion feature and the second motion features of at least one preset object, determine at least one candidate category corresponding to the target object, and the feature similarity between the target object and the preset objects corresponding to each candidate category, it is configured to implement: Based on a preset feature database, obtain the second motion features of at least one of the preset objects; Based on the feature matching algorithm, calculate the feature similarity between the first motion feature and each of the second motion features; Based on the feature similarity, determine at least one candidate category corresponding to the target object.
[0129] In one embodiment, when the processor implements determining at least one candidate category corresponding to the target object based on the feature similarity, it is configured to implement: When the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, determine that the object category of the preset object is a candidate category of the target object.
[0130] In one embodiment, before the processor implements obtaining the second motion features of at least one of the preset objects based on the preset feature database, it is further configured to implement: Based on the application scenario requirements, determine at least one of the preset objects in the target scene and the category labels of each of the preset objects; Based on historical video stream data, collect the motion data of the preset object to obtain the second motion feature corresponding to the preset object; Based on the second motion feature corresponding to the preset object and the category label, construct the feature database.
[0131] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the object recognition methods based on motion detection provided by the embodiments of the present application.
[0132] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.
[0133] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for object recognition based on motion detection, characterized in that: The method comprises: Obtain the video stream data to be detected; Based on the target detection algorithm, identifying at least one object to be detected in the video stream data, and obtaining a first confidence level corresponding to a category of each object to be detected; When the first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, determining that the object to be detected is a target object, and obtaining a target detection frame of the target object; When the first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than the second confidence level threshold, determining that the object to be detected is a candidate object; Based on the motion detection algorithm, motion feature detection is performed on the candidate object to obtain a motion detection result; When the motion detection result meets the motion target feature, determining the candidate object as a target object, and obtaining a target detection frame of the target object; Based on the target tracking algorithm, tracking operation is performed on the target detection frames of the target objects selected by the target detection algorithm and the motion detection algorithm to obtain first motion features corresponding to each of the target objects; Based on a feature matching algorithm, performing feature matching calculation on the first motion feature and the second motion feature of at least one preset object to determine at least one candidate category corresponding to the target object and feature similarity between the target object and the preset objects corresponding to each candidate category; Based on a preset weighted fusion mechanism, weighted fusion is performed on the first confidence and the feature similarity to determine the category matching degree between the target object and each of the preset objects; Based on the category matching degree between the target object and each of the preset objects, a category recognition result of the target object is determined.
2. The object recognition method based on motion detection according to claim 1, characterized in that: The step of performing motion feature detection on the candidate object based on the motion detection algorithm to obtain a motion detection result includes: Based on the optical flow method, detecting the motion trend of the candidate object, and obtaining the motion trend detection information of the candidate object; Based on the background difference method, a pixel comparison is performed between the current image area where the candidate object is located and a preset background model to obtain background area detection information of the candidate object; Based on the motion trend detection information and the background area detection information, the motion detection result corresponding to the candidate object is obtained.
3. The object recognition method based on motion detection according to claim 1, characterized in that: The target tracking algorithm is used to track the target detection frames of the target objects selected by the target detection algorithm and the motion detection algorithm to obtain the first motion features corresponding to the target objects, including: Based on the target tracking algorithm, the target detection frames corresponding to the same target object in consecutive video frames are tracked to obtain a continuous coordinate set of the target object; Based on the continuous coordinate set, a motion trajectory of the target object is drawn to calculate a first motion feature corresponding to the target object.
4. The object recognition method based on motion detection according to claim 1, characterized in that: After identifying at least one object to be detected in the video stream data based on the target detection algorithm and obtaining a first confidence level of a category corresponding to each object to be detected, the method further includes: When the first confidence level is less than the second confidence level threshold, it is determined that the object in the target detection frame is a non-target object.
5. The object recognition method based on motion detection according to claim 1, characterized in that: The method of performing feature matching calculation on the first motion feature and the second motion feature of at least one preset object based on the feature matching algorithm to determine at least one candidate category corresponding to the target object and the feature similarity between the target object and the preset objects corresponding to each of the candidate categories includes: Based on a preset feature database, obtaining a second motion feature of at least one of the preset objects; Based on the feature matching algorithm, calculating the feature similarity between the first motion feature and each of the second motion features; Based on the feature similarity, at least one candidate category corresponding to the target object is determined.
6. The object recognition method based on motion detection according to claim 5, characterized in that: The determining, based on the feature similarity, at least one candidate category corresponding to the target object comprises: When the feature similarity corresponding to any preset object is greater than or equal to a preset similarity threshold, the object category of the preset object is determined as a candidate category of the target object.
7. The object recognition method based on motion detection according to claim 5, characterized in that: Before acquiring the second motion feature of at least one of the preset objects based on the preset feature database, the method further includes: Based on application scenario requirements, determining at least one of the preset objects in the target scene and a category label of each of the preset objects; Based on the historical video stream data, the motion data of the preset object is collected to obtain the second motion feature corresponding to the preset object; The feature database is constructed based on the second motion feature corresponding to the preset object and the category label.
8. An object recognition device based on motion detection, characterized in that: The object recognition device based on motion detection comprises: A data acquisition module, used to acquire video stream data to be detected; A target detection module, used to identify at least one object to be detected in the video stream data based on a target detection algorithm, and obtain a first confidence level corresponding to a category of each object to be detected; A first target object determination module, configured to determine that the object to be detected is a target object when a first confidence level corresponding to the object to be detected is greater than or equal to a preset first confidence level threshold, and obtain a target detection frame of the target object; a candidate object determination module, configured to determine that the object to be detected is a candidate object when a first confidence level corresponding to the object to be detected is less than the first confidence level threshold and greater than a second confidence level threshold; A motion detection module, used to perform motion feature detection on the candidate object based on a motion detection algorithm to obtain a motion detection result; a second target object determination module, configured to determine that the candidate object is a target object when the motion detection result meets the motion target feature, and obtain a target detection frame of the target object; A target tracking module is used to perform tracking operations on the target detection frames of the target objects selected by the target detection algorithm and the motion detection algorithm based on the target tracking algorithm to obtain first motion features corresponding to each of the target objects; a feature matching module, configured to perform feature matching calculation on the first motion feature and the second motion feature of at least one preset object based on a feature matching algorithm, to determine at least one candidate category corresponding to the target object, and a feature similarity between the target object and the preset objects corresponding to each of the candidate categories; A feature fusion module, configured to perform weighted fusion on the first confidence and the feature similarity based on a preset weighted fusion mechanism, and determine a category matching degree between the target object and each of the preset objects; The category recognition module is used to determine the category recognition result of the target object based on the category matching degree between the target object and each of the preset objects.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the object recognition method based on motion detection as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the object recognition method based on motion detection as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Basketball motion detection method and system based on deep learning
CN117935373A
A method for multi-sensor multi-vehicle tracking based on image and motion feature matching
GB202409843D0
System and method for static and moving object detection
US9454819B1