Target tracking and behavior detection method based on camera equipment
By improving the target matching algorithm and trajectory recording method, using Kalman filter and inter-frame relationship matrix processing, the problems of target loss and error matching in the prior art are solved, and the accuracy and robustness of target tracking and behavior detection are improved.
Patent Information
- Application Number
- CN202510265864.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the target tracking method is prone to the problems of target loss and mismatch, especially in complex scenarios, which leads to insufficient accuracy and robustness of target tracking and behavior detection.
By improving the target matching algorithm and trajectory recording method, Kalman filter is used to initialize and update the target state, and the occlusion situation is processed in combination with the inter-frame relationship matrix, and the motion trajectory of the target object is recorded to ensure the continuity and accuracy of the tracking.
It improves the accuracy and robustness of target tracking and behavior detection in complex scenarios, reduces the situation of target loss and mismatch, and ensures the integrity of the moving trajectory of the target object.
Smart Images

Figure CN120219772A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of camera monitoring, and particularly relates to a method for target tracking and behavior detection based on a camera device. Background Art
[0002] With the continuous improvement of people's attention to public safety and the rapid development of video monitoring devices and video processing technologies, video monitoring systems play an increasingly important role in social life. Facing the increasing number of cameras in the market, it is impossible to rely solely on manual visual inspection by monitoring personnel to achieve monitoring. Moreover, since the time of abnormal events in most monitoring scenarios is short and random, such manual monitoring not only causes huge waste of manpower, but also easily makes the monitoring personnel lax in thinking and leads to missed alarms.
[0003] If an intelligent video monitoring system is adopted, using a computer to assist or even replace humans to complete monitoring or control tasks, automatically alarm when an abnormal event occurs, and then confirmed by the monitoring personnel, the work efficiency and monitoring effect will be greatly improved, while the work intensity will be greatly reduced. However, perceiving and recognizing human behavior in complex environments is one of the hot and difficult topics in intelligent video monitoring research. Its task is to use cameras to monitor and interpret the continuous and instantaneous objects in a specific environment in real time, understand and predict the behavior and events of context-related objects, and interact with the observed objects based on the information obtained from sensors, which has important value in applications such as detection, monitoring, management, and command in public facilities, commerce, transportation, and military scenarios.
[0004] The visual behavior perception system in an intelligent environment undertakes the dual tasks of monitoring and interacting with people in the environment. Therefore, the main errors in the existing target tracking and behavior detection methods in practical applications are as follows:
[0005] (1) Abort the tracking of the lost target at the lost frame, and the target appears quickly, mistakenly regarding the target that reappears after being lost as a new target.
[0006] (2) Mistakenly match the detection value that reappears after being lost with the target before being lost.
[0007] (3) The target object is blocked for a long time, and the specific position of the pedestrian cannot be detected, resulting in the loss of the tracking target and the inability to obtain the complete movement trajectory of the pedestrian. Summary of the Invention
[0008] Aiming at the above defects, the technical problem solved by the present invention is to provide a method and system that can accurately track targets and detect behaviors. By improving the target matching algorithm and trajectory recording method, the problems of target loss and incorrect matching in the prior art are solved, and the accuracy and robustness of target tracking and behavior detection in complex scenarios are improved.
[0009] A method for target tracking and behavior detection based on a camera device is provided in the first aspect of the present invention, characterized in that the method includes:
[0010] S101: Obtain real-time monitoring video data of an object to be recognized, perform frame division processing on the video monitoring data to obtain video frames of the object to be recognized.
[0011] S102: Extract the boundary of the first video frame to obtain a regional image of the object to be recognized.
[0012] S103: Extract features from the binary matrix of the regional image of the object to be recognized, match the features with the features of the binary matrix of the target object, and determine whether the object to be recognized is the target object;
[0013] S104: Record the movement trajectory of the target object.
[0014] According to an embodiment of the present invention, in step S102, extracting the boundary of the video frame to obtain a regional image of the object to be recognized includes: performing edge detection on the object to be recognized to obtain the boundary coordinates of the object to be recognized; constructing a first coding matrix of the contour of the object to be recognized according to the boundary coordinates.
[0015] Perform binary processing on the first coding matrix of the contour of the object to be recognized to construct a second coding matrix of the contour of the object to be recognized.
[0016] Arrange the rows or columns of the second coding matrix to construct a coding chain of the object to be recognized, and obtain a regional image of the object to be recognized according to the coding chain.
[0017] According to an embodiment of the present invention, in S103, extracting features from the binary matrix of the regional image of the object to be recognized, matching the features with the features of the binary matrix of the target object, and determining whether the object to be recognized is the target object includes: setting a first matching threshold, if the feature data of the binary matrix of the object to be recognized is not greater than the feature data of the binary matrix of the target object, then return to step S101.
[0018] If the feature data of the binary matrix of the object to be recognized is greater than the feature data of the binary matrix of the target object, it is determined as the target object.
[0019] According to an embodiment of the present invention, in S104, recording the movement trajectory of the target object includes:
[0020] Confirm the target object, initialize the Kalman filter, and set the initial state vector and error covariance matrix of the target object.
[0021] Predict the target state of the current video frame of the target object; and update the state vector and the error covariance matrix; update the target object state according to the updated state vector and error covariance matrix.
[0022] Record the updated target object state.
[0023] According to an embodiment of the present invention, the S103 further includes: determining whether the target object is occluded, including:
[0024] Draw the circumscribed rectangle of the target object, record its coordinate information, set a second matching threshold, and construct the inter-frame relationship matrix P of the target object through the circumscribed matrix and coordinate information of the target object, as shown in Equation (1):
[0025] Equation (1)
[0026] Determine whether the target object is occluded, where the represents the nth target object in the mth frame; includes:
[0027] If the overlap value between the area of the circumscribed rectangle of the target object in the current video frame and the area of the rectangle of the target object in the previous video frame is greater than the second matching threshold:
[0028] Equation (2)
[0029] If the overlap value between the area of the circumscribed rectangle of the target object in the current video frame and the area of the rectangle of the target object in the previous video frame is not greater than the second matching threshold:
[0030] Equation (3)
[0031] The inter-frame relationship matrix of the target object If a non-zero value appears in the row or column of the matrix, it is determined that the target object in the previous video frame is occluded in the current video frame; if the value at the same matrix position in the target object inter-frame relationship matrix changes from 1 to 0, it is determined that the target object in the previous video frame appears again in the current video frame.
[0032] According to an embodiment of the present invention, the method further includes: recording the time when the target object changes from being occluded to appearing according to the changes in the row and column values of the inter-frame relationship matrix If the time is less than the third threshold, delete the target object; if the time is not less than the third threshold, retain the target object.
[0033] According to an embodiment of the present invention, S104 further includes: extracting the behavioral features and skeletal features of the target object, and calculating the motion state of the target object.
[0034] According to an embodiment of the present invention, the method further includes S105: judging whether the behavior of the target object is abnormal according to the motion trajectory of the target object, and if the behavior of the target object is abnormal, giving an early warning.
[0035] A second aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory, and a processor; the memory is used for storing computer instructions; the processor is used for running the computer instructions stored in the memory to implement the above method for target tracking and behavior detection based on a camera device.
[0036] A third aspect of the present invention provides a storage medium, including: a readable storage medium and computer instructions, the computer instructions are stored in the readable storage medium; the computer instructions are used for implementing the above method for target tracking and behavior detection based on a camera device.
[0037] The beneficial effects provided by the present invention: It can accurately track the target and detect its behavior. By improving the target matching algorithm and trajectory recording method, the problems of target loss and incorrect matching in the prior art are solved, and the accuracy and robustness of target tracking and behavior detection in complex scenarios are improved. Description of the Drawings
[0038] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0039] Figure 1 It is a flowchart of the method for target tracking and behavior detection based on a camera device disclosed in the embodiment of the present invention.
[0040] Through the above drawings, the clear embodiments of the present disclosure have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Detailed Embodiments
[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of the devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] In the first aspect of the present invention, a method for target tracking and behavior detection based on a camera device is provided, as Figure 1 shown. S101: Obtain real-time monitoring video data of the object to be recognized, perform frame division processing on the video monitoring data, and obtain video frames of the object to be recognized.
[0043] Based on a camera probe, obtain real-time monitoring video of the object to be recognized, and perform frame division processing on the real-time monitoring video to obtain real-time video frame information. Video capture and frame division processing. Continuously capture a video stream of the target monitoring scene through a high-definition camera probe. First, start the camera probe to ensure that it covers the predetermined monitoring range, and adjust parameters such as focal length and exposure to obtain the best image quality. Then, use a built-in or external image capture device to capture video data at a fixed number of frames per second (for example, 30 frames per second). For subsequent processing, the captured video stream needs to be subjected to frame division processing. The specific implementation is as follows: Video stream preprocessing: Divide the continuous video stream into independent frames, and each frame represents a static image containing timestamp information to record the capture time of each frame.
[0044] Frame rate adjustment: According to actual monitoring requirements, the frame rate of the video stream can be adjusted to ensure that neither important information is lost nor excessive redundant data is generated.
[0045] Image quality optimization: Perform preliminary quality optimization on each frame of the image, including denoising, contrast enhancement, etc., to improve the accuracy of subsequent image processing.
[0046] Assume that the camera probe has been calibrated and start capturing a monitoring video. First, the probe captures the video stream at a speed of 30 frames per second. The video stream enters the preprocessing stage, and the system automatically divides it into independent frames. During this process, the system checks and adjusts the frame rate. For example, if the monitoring scene changes slowly, the frame rate can be reduced to reduce the data volume.
[0047] Subsequently, each frame of the image is processed by a denoising algorithm and contrast enhancement to improve the image quality. The optimized frames are stored with timestamps attached to form a set of video frame information in the first stage.
[0048] Through the above steps, the user can obtain clear and processed video frame information, laying a foundation for subsequent target extraction and tracking.
[0049] S102: Perform boundary extraction on the first video frame to obtain the regional image of the object to be recognized.
[0050] Based on the image contour encoding algorithm, perform region extraction processing on the real-time video frame information for the object to be recognized, and obtain the tracking region image. Image contour encoding and target region extraction. In this step, we will use the image contour encoding algorithm to accurately extract the region of the object to be recognized.
[0051] This process involves the following key sub-steps, including: performing target boundary extraction processing on the object to be recognized in the real-time monitoring video frame information based on the image analysis model to obtain the boundary coordinate information of the object to be recognized; constructing a boundary encoding matrix based on the boundary coordinate information of the object to be recognized to obtain the boundary encoding matrix of the object to be recognized.
[0052] Perform binarization processing on the boundary encoding matrix of the target object based on a preset threshold to obtain a binarized boundary encoding matrix; perform permutation encoding on the binarized boundary encoding matrix to obtain an encoding chain corresponding to the boundary; obtain the target object region image based on the encoding chain corresponding to the boundary.
[0053] Target boundary extraction: Use an image analysis model, such as an edge detection operator (such as the Canny operator), to detect the edges of the object to be recognized and obtain accurate boundary coordinate information.
[0054] Construct a boundary encoding matrix: According to the extracted boundary coordinate information, construct an encoding matrix that can represent the contour of the object to be recognized. This matrix maps the points on the boundary to the elements in the matrix.
[0055] Binarization processing: To simplify subsequent processing, perform binarization processing on the boundary encoding matrix. By setting a suitable threshold, convert the elements in the matrix to 0 or 1 to form a binarized boundary encoding matrix.
[0056] Encoding chain permutation: Permute the rows or columns in the binarized boundary encoding matrix to form an encoding chain that can represent the contour of the object to be recognized.
[0057] Taking the pedestrian recognition in a monitoring scenario as an example, first, the system extracts the edges of the pedestrians in the real-time video frame through the Canny operator and obtains the edge coordinate points. Then, map these coordinate points to a two-dimensional matrix to construct a boundary encoding matrix.
[0058] Set an appropriate threshold to perform binarization processing on the boundary encoding matrix. For example, set all points greater than a certain gray value to 1 and the rest to 0. This step clearly shows the contour of the pedestrians in a binarized form.
[0059] Subsequently, convert each row or column in the binarized matrix into an encoding chain, and these encoding chains together constitute the contour features of the pedestrians.
[0060] Finally, through these encoded chains, the system can reconstruct the regional image of the object to be recognized. This image clearly depicts the outline of the pedestrian, providing precise input for subsequent object matching and tracking.
[0061] Through the above steps, the user can not only extract the target region of interest from complex surveillance scenarios but also lay a solid foundation for subsequent behavior analysis and trajectory tracking.
[0062] S103: Extract features from the binary matrix of the regional image of the object to be recognized, match them with the features of the binary matrix of the target object, and determine whether the object to be recognized is the target object. Represent the state of the target at a certain moment with an 8-dimensional vector. , where (u, v) is the center coordinate of the detection bounding box, r is the aspect ratio, h represents the height, and the other 4 variables are the speeds of the previous 4 variables. Use the Kalman filter to predict it, so as to obtain information such as the tracking position and size (u, v, r, h) of the corresponding target in the next frame. Since the moving distance of pedestrians in adjacent two frames is not very large, the IOU between the candidate box and the detection box in the image frame is used as the matching basis. The larger the IOU, the more they are considered to match. The specific formula (1) is as follows:
[0063] Equation (1)
[0064] According to the detection results of the previous frame, the information of the predicted candidate box of the pedestrian in the current frame predicted by the Kalman filter. The target detector detects the current frame to obtain the information of the detection candidate box of the current frame. First, calculate the IOU of all tracking boxes and all detection boxes at the current moment, and then use the Hungarian algorithm to select the best unique matching result according to the size of the IOU value. After the Hungarian algorithm completes the matching of the current frame, replace the predicted box of this frame with the detection box of this frame, calculate the predicted box of the next frame, and repeat the above steps until the target meets the condition to end tracking.
[0065] Due to the existence of the Kalman filter, the trajectory break caused by pedestrians being occluded for a short time can be reduced. When pedestrians are not detected for a long time, the predicted box generated by the Kalman filter will drift, resulting in prediction failure. At this time, the trajectory is considered unstable. Therefore, after reaching the time threshold, stop tracking it, store it as a normal motion trajectory, and obtain the video frame set.
[0066] Confirm the identity of the object to be recognized through the target matching algorithm, and continuously record the motion trajectory of the target object using the Kalman filter to ensure the accuracy of the normal motion trajectory.
[0067] Target matching process: Feature extraction: Extract features from the binary matrix of the object to be recognized in the tracking area image, including but not limited to shape, size, position, etc.
[0068] Matching algorithm: Use a preset target matching algorithm, such as matching based on histogram similarity or shape-based template matching, to compare the features of the object to be recognized with the features of the known target object.
[0069] Judgment of matching result: If the matching degree between the features of the object to be recognized and the features of the target object exceeds the preset matching probability threshold, it is determined that the object to be recognized is the target object.
[0070] Trajectory recording process, including: Kalman filter initialization: After confirming the target object, initialize the Kalman filter and set the initial state vector and error covariance matrix.
[0071] State prediction: Predict the target state of the current frame based on the target state of the previous frame.
[0072] Observation update: Compare the actual observation data of the current frame with the predicted state, and update the state vector and error covariance matrix.
[0073] Trajectory recording: Record the updated target state to form a normal motion trajectory.
[0074] Suppose in a monitoring scenario, the system obtains the feature information of the object to be recognized through a feature extraction algorithm. Next, use a template matching algorithm to match the shape features of the object to be recognized with the shape of the target object stored in the database.
[0075] If the matching degree exceeds 80% (the preset matching probability threshold), the system confirms the object to be recognized as the target object. Once confirmed, the system immediately starts the Kalman filter.
[0076] Taking a moving vehicle as an example, the Kalman filter predicts the position of the current frame based on the position and speed of the vehicle in the previous frame. When the actual observation data (the real position of the vehicle in the current frame) is input, the Kalman filter combines the predicted value with the observed value, updates the state of the vehicle, and reduces the prediction error.
[0077] Through continuous prediction and update, the Kalman filter can smoothly record the motion trajectory of the vehicle. These trajectory data are then stored for analyzing the motion behavior of the target object.
[0078] Through the above steps, users can not only accurately identify the target object, but also obtain the continuous motion trajectory of the target object in the monitoring scenario, providing important data support for subsequent behavior analysis and anomaly detection.
[0079] When the target object is occluded, the tracking situation is processed through the circumscribed rectangle calibration and the inter-frame relationship matrix, the occluded motion trajectory segment is recorded, and it is added to the normal motion trajectory.
[0080] When the target object is still within the image of the tracking area of the object to be recognized but is occluded for a preset number of frames and fails to be matched, the tracking time of the target object is judged at this time. When the tracking time meets the preset frame number range, the image information of this segment of the motion trajectory is recorded as the occluded motion trajectory segment and added to the normal motion trajectory to obtain the second video frame set.
[0081] Specifically, the circumscribed rectangle calibration of the target object. Among them, the binary image of the tracking area image is scanned from left to right and from top to bottom, and the marking information is recorded with a matrix of the same size as the target object. If the gray value of the currently scanned pixel is 1, it is marked as the target pixel connected to it. If it is connected to two or more targets, these targets can be considered the same and connected. If a transition from a pixel with a gray value of 1 to an isolated pixel with a gray value of 0 is found, a new target mark is assigned.
[0082] Calculate the feature information of the moving target. After obtaining the circumscribed rectangle information of the target object, it is relatively easy to calculate the centroid of the moving target. Here, the center of the circumscribed rectangle of the target object is used instead of the centroid. The specific algorithm is shown in equations (2) and (3):
[0083] Equation (2)
[0084] Equation (3)
[0085] Among them, in the formula 、 respectively represent the abscissa and ordinate of the center of the nth target object, 、 、
[0086] 、 respectively represent the four coordinates of the left, right, top, and bottom of the circumscribed rectangle.
[0087] Through the above data, an inter-frame relationship matrix is established to process the tracking situation by classifying it into the case where the target object is occluded and the case where the target object appears. The definition formula of the inter-frame relationship matrix is shown in equation (4).
[0088] Equation (4)
[0089] The number of rows and columns of the matrix respectively correspond to the size of the target linked list of the current frame and the size of the target linked list of the previous frame. The target linked list of the current frame is shown in equation (5).
[0090] Equation (5)
[0091] The previous-frame target linked list is as shown in Equation (6).
[0092] Equation (6)
[0093] Wherein, represents the feature information of the nth moving target in the (k - 1)th frame, and here the features refer to the center coordinates of the moving target and the width and height of the circumscribed rectangle.
[0094] The value of each element in the inter-frame relationship matrix P is the result of calculating the overlapping area of the circumscribed rectangles of the moving targets between adjacent frames. If the overlapping area is greater than the set threshold, it is considered that and are matched, so that
[0095] Equation (7)
[0096] Otherwise,
[0097] Equation (8)
[0098] According to the inter-frame relationship matrix, the relationship between the current frame and the moving targets in the previous frame can be determined.
[0099] When the target object is occluded, the inter-frame relationship matrix is expressed as Equation (9):
[0100] Equation (9)
[0101] Wherein, if there are multiple non-zero elements in the kth row of the inter-frame relationship matrix P, such as the hth column and the (h + 1)th column being non-zero, then the hth and (h + 1)th targets in the previous frame are occluded in the current frame.
[0102] When the target appears, the inter-frame relationship matrix is expressed as Equation (10).
[0103] Equation (10)
[0104] Wherein, if the kth row of the inter-frame relationship matrix P is all 0, then the kth target in the current frame is an emerging target.
[0105] Collect the number of frames during the tracking time from when the target object is occluded to when the target object appears. When the tracking time meets the preset frame number range, record the image information of this segment of the motion trajectory as an occluded motion trajectory segment and add it to the normal motion trajectory to obtain a video frame set.
[0106] Among them, collect the tracking time during the period when the target object is occluded and when the target object appears. When the tracking time is less than 20 frames, this track segment is discarded. In this way, targets with too short tracking time can be deleted, achieving the effect of reducing noise pollution. In areas with low pedestrian density, the detector has a better effect and can stably track the target. The obtained long track of the target object may be long enough, and the system believes that the main characteristics of the target object's current movement have been obtained and there is no need to track it anymore. This step of work is achieved by judging whether the tracking time of this target is greater than 500 frames. If the tracking time is less than 500 frames, it is considered that this track is just a segment during the movement of the target object, and it is added to the normal movement track to obtain a video frame set.
[0107] Track recording and updating under occlusion. When the target object fails to be directly matched due to occlusion in the tracking area image, this step aims to ensure that even under occlusion, the movement track of the target object can be recorded and the continuity of the track can be maintained through precise circumscribed rectangle calibration and inter-frame relationship matrix analysis.
[0108] Sub-step of circumscribed rectangle calibration. Target object positioning: Perform pixel-by-pixel scanning on the binary image of the tracking area image to locate the position of the target object in the current frame.
[0109] Drawing the circumscribed rectangle: For the located target object, draw a minimum rectangle that can completely enclose the target object, that is, the circumscribed rectangle, and record its coordinate information.
[0110] Occlusion detection: By comparing the changes in the circumscribed rectangles in consecutive frames, determine whether the target object is occluded. If the target object fails to be detected within a consecutive preset number of frames, it is considered that occlusion has occurred.
[0111] Sub-step of inter-frame relationship matrix processing. Establishing the inter-frame relationship: When the target object is occluded, establish an inter-frame relationship matrix by analyzing the target position relationship between consecutive frames.
[0112] Track prediction: Use the inter-frame relationship matrix to predict the potential movement track of the target object during occlusion.
[0113] Track fusion: When the target object reappears, fuse the predicted track with the actually observed track to ensure the continuity and accuracy of the track.
[0114] Consider a pedestrian in a surveillance scenario who fails to be directly detected due to occlusion by trees or other pedestrians in consecutive frames. First, the system draws the circumscribed rectangle of the pedestrian through pixel-by-pixel scanning. When occlusion occurs, the system starts the analysis of the inter-frame relationship matrix.
[0115] By comparing the positions of the bounding rectangles of pedestrians in consecutive frames, the system constructs an inter-frame relationship matrix to predict the possible movement trajectories of pedestrians. For example, if a pedestrian is on the left in the previous frame and on the right in the next frame, the system will predict that the pedestrian may move in a straight line during this period.
[0116] When the pedestrian reappears in the surveillance video, the system fuses the predicted trajectory with the actually observed trajectory to ensure the continuity of the trajectory data. This fusion process may involve methods such as weighted averaging to balance the predicted and actual observed data.
[0117] The recorded occluded trajectory segments are then added to the normal movement trajectory to form the complete movement trajectory of the target object.
[0118] Through this step, even when the target object is occluded, the user can obtain the approximate position and movement trend of the target object, which is crucial for maintaining the continuity and accuracy of surveillance. In addition, this method can also improve the robustness of the system to occlusion problems in complex surveillance environments.
[0119] S104: Record the movement trajectory of the target object. Extract the behavioral features and skeletal features of the target object from the set of video frames, and calculate the movement behavior and limb behavior.
[0120] Extract the behavioral features and skeletal features of the target object, calculate the movement behavior of the moving target according to the behavioral features, and calculate the limb behavior of the moving target according to the skeletal features.
[0121] Convert the behavior of the target object into behavior coordinates in a three-dimensional coordinate system; determine the bounding rectangle area according to the behavior coordinates, calculate the centroid coordinates, the width-to-height ratio of the rectangle corresponding to the behavior, and the rectangle tilt angle according to the bounding rectangle area to obtain the behavioral features; divide the behavior coordinates into skeletal point coordinates according to a preset skeletal sequence; calculate the relative displacement features of skeletal points between frames and the relative distance features of skeletal points within a frame according to the skeletal point coordinates to obtain the skeletal features.
[0122] Extract the behavioral features and skeletal features of the target object, and then calculate its movement behavior and limb behavior. The following is the detailed implementation process:
[0123] Behavioral feature extraction. Coordinate transformation: Convert the two-dimensional coordinates of the target object in the video frame into coordinates in a three-dimensional coordinate system to more accurately describe its spatial position. Determination of the bounding rectangle area: Determine the bounding rectangle area of the target object according to its three-dimensional coordinates, and record the centroid coordinates, the width-to-height ratio of the rectangle, and the rectangle tilt angle. These parameters together constitute the behavioral features. Feature recording: Record the behavioral features of the target object in each frame to form a feature sequence.
[0124] Bone feature extraction. Bone point division: According to the preset bone sequence, divide the coordinates of the object to be recognized into bone points to obtain the coordinates of key bone points. Relative displacement and distance calculation: Calculate the relative displacement features of bone points between frames and the relative distance features between bone points within a frame. These features reflect the dynamic bone structure of the target object.
[0125] Taking the motion analysis of pedestrians in a monitoring scenario as an example, first, the system converts the two-dimensional coordinates of pedestrians in each frame into three-dimensional coordinates to more accurately track their motion trajectories. Then, the system draws a circumscribed rectangle for the pedestrians and records the centroid coordinates, aspect ratio, and tilt angle. These data reflect the basic characteristics of pedestrian behavior.
[0126] Subsequently, according to the pedestrian bone model, the system divides the key parts of pedestrians (such as the head, shoulders, waist, legs, etc.) into bone points and calculates the relative displacement between these bone points between frames and the relative distance within a frame. These data constitute the bone features of pedestrians.
[0127] By analyzing these behavior features and bone features, the system can calculate the motion behavior and limb behavior of pedestrians. For example, analyze the walking speed and direction of pedestrians through the centroid movement trajectory; judge the action type of pedestrians, such as walking, running, jumping, etc., through the relative displacement of bone points.
[0128] S105: Judge whether the behavior of the target object is abnormal according to the motion trajectory of the target object. If the behavior of the target object is abnormal, give an early warning.
[0129] Finally, the system compares these behaviors and limb behaviors with the abnormal behavior library to judge whether there are abnormal behaviors in pedestrians.
[0130] Through the above steps, users can not only extract the behavior and bone features of the target object, but also analyze the motion state and limb movements of the target object through these features, providing important technical support for behavior recognition and abnormal detection in the monitoring scenario.
[0131] Judge whether the motion behavior and limb behavior are abnormal according to the abnormal behavior library. If there are abnormal behaviors, give an early warning or record them.
[0132] Judge whether the motion behavior and the limb behavior are abnormal behaviors according to the abnormal behavior library; when at least one of the motion behavior and the limb behavior belongs to the abnormal behavior library, determine that the motion target has abnormal behaviors.
[0133] Among them, the method for constructing the abnormal behavior library includes: screening key motion frames of an image to obtain target frame images, reconstructing a background image for the target frame images to obtain a video background, peeling the background from the target frame images according to the video background to obtain moving targets, and determining the corresponding abnormal behavior library according to the video background.
[0134] The steps for generating the abnormal behavior library include: extracting two adjacent images from multiple frames of images one by one as target images; performing a masking operation on the target images to obtain masked images; calculating the difference feature values of the masked images, and when the difference feature values are greater than a preset difference value, using the target images as target frame images.
[0135] Selecting images of a preset number of frames from the target frame images as sequence images; graying the sequence images to obtain grayscale images; binarizing the grayscale images to obtain binary images; separating the background regions of the target frame images according to the binary images to obtain a video background.
[0136] Performing differential operations on the target frame images according to the video background; performing morphological processing on the images after the differential operations to obtain morphological images; extracting features according to the morphological images, and generating image annotations according to the feature extraction results; generating moving targets according to the morphological images and the image annotations.
[0137] Vectorizing the video background to obtain background features; performing feature matching on the background features in a preset scene library, and taking the scene with the highest matching degree as the target scene of the video background; retrieving in the scene database according to the target scene to obtain the abnormal behavior library corresponding to the target scene.
[0138] Abnormal behavior recognition and early warning. We will use the constructed abnormal behavior library to evaluate the behavior characteristics and limb behaviors of the target object to determine whether there are abnormal behaviors. If an abnormal behavior is detected, the system will trigger an early warning mechanism or record it for subsequent analysis.
[0139] Process of abnormal behavior recognition: Analysis of behavior characteristics and limb behaviors: Combining the behavior characteristics and bone characteristics extracted from the video frame set, comprehensively analyzing the motion behaviors and limb behaviors of the target object in the monitoring scene.
[0140] Comparison with the abnormal behavior library: Comparing the motion behaviors and limb behaviors of the target object with the abnormal behavior library, which contains predefined abnormal behavior patterns.
[0141] Abnormal determination: If the motion behaviors and limb behaviors of the target object match any one of the patterns in the abnormal behavior library, it is determined that the target object has an abnormal behavior.
[0142] Early warning and record processing. Early warning trigger: Once an abnormal behavior is detected, the system immediately triggers the early warning mechanism, which can be through sound, visual cues, or directly notifying the monitoring personnel.
[0143] Behavior record: Whether the behavior is abnormal or not, the system records the behavior characteristics and body behaviors of the target object for subsequent data analysis and behavior pattern recognition.
[0144] Taking a shopping mall monitoring scenario as an example, the system extracts and analyzes the behavior characteristics and skeletal characteristics of pedestrians through the previous steps. In step six, the system compares these characteristics with the abnormal behavior library.
[0145] The abnormal behavior library may contain behavior patterns such as running fast, abnormal gathering, fighting, etc. If the behavior of the pedestrian matches any one of the patterns in the library, the system will immediately determine it as an abnormal behavior and trigger an early warning.
[0146] The processing methods of the early warning can be diverse. For example, the system can emit a high-decibel alarm sound, or remind the monitoring personnel in the form of high-brightness flashing on the display screen of the monitoring center. At the same time, the system will automatically record the moment when the early warning is triggered, the characteristic information of the target object, and the description of the abnormal behavior.
[0147] For normal behaviors that do not trigger an early warning, the system will also record them. These records help to construct normal behavior patterns and provide a more comprehensive reference for abnormal behavior detection.
[0148] Through this step, users can not only promptly discover abnormal behaviors in the monitoring scenario, but also provide rich data support for subsequent behavior analysis, thereby improving the intelligence level of the monitoring system and the ability to respond to emergencies.
[0149] The beneficial effects provided by the present invention: It can accurately track the target and detect its behavior. By improving the target matching algorithm and trajectory recording method, it solves the problems of target loss and incorrect matching in the prior art, and improves the accuracy and robustness of target tracking and behavior detection in complex scenarios.
[0150] The second aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory, and a processor; the memory is used to store computer instructions; the processor is used to run the computer instructions stored in the memory to implement the above method for target tracking and behavior detection based on a camera device.
[0151] The third aspect of the present invention provides a storage medium, including: a readable storage medium and computer instructions, the computer instructions are stored in the readable storage medium; the computer instructions are used to implement the above method for target tracking and behavior detection based on a camera device.
[0152] Obviously, the above specific implementation cases are only examples for illustrating the application of this method, rather than limitations on the implementation manners. For those of ordinary skill in the art, based on the above description, other different forms of changes and variations can be made to study other related issues. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0153] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disks, or optical disks that can store program codes.
[0154] The above-described embodiments of electronic devices and the like are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0155] Through the description of the above implementation manners, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium such as ROM / RAM, magnetic disks, optical disks, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0156] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them; although the embodiments of the present invention have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
[0157] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0158] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for target tracking and behavior detection based on a camera device, characterized in that: The method comprises: S101: Acquire real-time monitoring video data of an object to be identified, and perform frame processing on the video data to obtain video frames of the object to be identified; S102: extracting the boundary of the video frame to obtain a regional image of the object to be identified; S103: extracting features from the binary matrix of the region image of the object to be identified, matching the features with the binary matrix of the target object, and determining whether the object to be identified is the target object; S104: Recording the movement trajectory of the target object.
2. The method according to claim 1, characterized in that The step S102 extracts the boundary of the video frame to obtain the regional image of the object to be identified, including: Detect the edge of the object to be identified and obtain the boundary coordinates of the object to be identified; construct a first encoding matrix of the contour of the object to be identified according to the boundary coordinates; Binarize the first coding matrix of the contour of the object to be identified to construct a second coding matrix of the contour of the target object to be identified; Arrange the rows or columns of the second coding matrix to construct a coding chain of the object to be identified, and obtain a regional image of the object to be identified according to the coding chain.
3. The method according to claim 2, characterized in that The step S103 extracts features from the binary matrix of the region image of the object to be identified, matches the features of the binary matrix of the target object, and determines whether the object to be identified is the target object, including: A first matching threshold is set, and if the binary matrix feature data of the object to be identified is not greater than the binary matrix feature data of the target object, the process returns to step S101; If the binary matrix feature data of the object to be identified is greater than the binary matrix feature data of the target object, it is determined to be the target object.
4. The method according to claim 3, characterized in that The step of recording the target object's motion trajectory in S104 includes: Confirming the target object, initializing the Kalman filter, and setting the initial state vector and error covariance matrix of the target object; Predicting the target state of the current video frame of the target object; and updating the state vector and the error covariance matrix; updating the state of the target object according to the updated state vector and the error covariance matrix; The updated state of the target object is recorded.
5. The method according to claim 1 or 3, characterized in that: The step S103 further includes: determining whether the target object is blocked, including: Draw the bounding rectangle of the target object, record its coordinate information, set the second matching threshold, and construct the target object inter-frame relationship matrix P through the target object's bounding matrix and coordinate information to determine whether the target object is blocked, as shown in formula (1): Formula (1) Said Represents the nth target object in the mth frame; including: If the overlap value between the circumscribed rectangular area of the target object in the current frame and the rectangular area of the target object in the previous frame is greater than the second matching threshold: Formula (2) If the overlap value between the circumscribed rectangular area of the target object in the current frame and the rectangular area of the target object in the previous frame is not greater than the second matching threshold: Formula (3) The target object inter-frame relationship matrix If a non-zero value appears in the row or column value, it is determined that the target object of the previous video frame is blocked in the current video frame; if the target object inter-frame relationship matrix If the value of the same matrix position changes from 1 to 0, it is judged that the target object in the previous video frame appears again in the current video frame.
6. The method according to claim 5, characterized in that: The method further comprises: According to the inter-frame relationship matrix The changes in row and column values record the time from when the target object is blocked to when it appears. If the time is less than a third threshold, the target object is deleted; if the time is not less than the third threshold, the target object is retained.
7. The method according to claim 4 or 6, characterized in that: The S104 also includes: The behavior features and skeleton features of the target object are extracted, and the motion state of the target object is calculated.
8. The method according to claim 1, characterized in that The method further comprises S105: Whether the behavior of the target object is abnormal is determined according to the motion trajectory of the target object, and if the behavior of the target object is abnormal, an early warning is issued.
9. A smart device, characterized in that: include: transmitter, receiver, memory and processor; The memory is used to store computer instructions; The processor is used to execute the computer instructions stored in the memory to implement the method for target tracking and behavior detection based on a camera device according to any one of claims 1 to 8.
10. A storage medium, characterized in that: include: A readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; The computer instructions are used to implement the method for target tracking and behavior detection based on a camera device as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Pedestrian detection method based on HOG and D-S evidence theory multi-information fusion
CN105335701A
Pedestrian tracking method based on least square locus prediction and intelligent obstacle avoidance model
CN106023244A
Target tracking method and device based on monitoring camera
CN112991396A
Abnormal behavior analysis method and device based on intelligent linkage of multiple video cameras
CN114565882A
Video coding motion estimation method and apparatus based on multi-target tracking
WO2024244416A1