A trajectory tracking method and device
In the trajectory tracking method in the field of computer vision, the background video frame is determined using the continuous video frame sequence and the position information of the cage, and the background is deducted, which solves the problems of high cost and poor effect of the traditional method, and achieves more efficient and economical trajectory tracking.
Patent Information
- Application Number
- CN202211414553.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Traditional trajectory tracking methods based on background subtraction are costly and have poor results, especially when environment changes frequently.
By obtaining the continuous video frame sequence of the target object, the position information of the target object in each sub-video frame sequence is determined, and based on this information and the position information of the cage, the left mid-empty background video frame and the right mid-empty background video frame are determined, and the background subtraction is performed to obtain the motion trajectory.
It reduces the cost of trajectory tracking, improves the tracking effect, and maintains high accuracy in the event of environmental changes.
Smart Images

Figure CN115631217B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and particularly to a trajectory tracking method and device. Background Art
[0002] In the field of computer vision, video analysis methods for animal behavior analysis have received extensive attention. In particular, the analysis of animal movement trajectories can provide important information for various purposes. For example, the movement trajectory of macaques not only reflects the overall activity level of macaques but also records important spatial information about movement. Analyzing the movement trajectory of macaques can be used to classify different behaviors with specific phenotypes and movement characteristics such as Parkinson's disease, Huntington's disease, and Alzheimer's disease.
[0003] Traditional trajectory tracking methods based on background subtraction first obtain an empty background, then use background subtraction to obtain the animal foreground, and then track the animal foreground to obtain the animal's movement trajectory. However, traditional background subtraction requires a specific environment to provide a clean and stable background. For example, using a special cage made of transparent material and shooting under stable lighting conditions increases the cost of trajectory tracking and has poor practicability. Moreover, the background environment often changes with the shooting time, and only generating an empty background results in poor trajectory tracking effects. Summary of the Invention
[0004] In view of this, the present application provides a trajectory tracking method and device for solving the problems of high cost and poor effect of trajectory tracking in the prior art. The technical solutions are as follows:
[0005] A trajectory tracking method includes:
[0006] Obtaining a target video frame sequence composed of consecutive video frames of a target object, where the target video frame sequence includes multiple sub-video frame sequences, and each video frame included in each sub-video frame sequence includes the position information of the cage containing the target object;
[0007] Determining the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determining the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence, where the target object included in the left half-empty background video frame is in the right half area of the cage, and the target object included in the right half-empty background video frame is in the left half area of the cage;
[0008] Performing background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence;
[0009] Based on the foreground video frame sequence corresponding to each sub-video frame sequence, determine the motion trajectory of the target object in the target video frame sequence.
[0010] Optionally, determine the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determine the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence, including:
[0011] For each sub-video frame sequence:
[0012] Use the frame difference method to calculate the frame pixel difference between every two adjacent video frames in the sub-video frame sequence, and based on the calculated frame pixel differences, determine the position information of the target object in each video frame included in the sub-video frame sequence;
[0013] For each calculated frame pixel difference, if the frame pixel difference is greater than or equal to a preset high motion threshold, then both of the two adjacent video frames corresponding to the frame pixel difference are used as the first target video frames to obtain all the first target video frames included in the sub-video frame sequence;
[0014] According to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence, determine the left half-empty background video frame and the right half-empty background video frame corresponding to the sub-video frame sequence;
[0015] To obtain the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence.
[0016] Optionally, according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence, determine the left half-empty background video frame and the right half-empty background video frame corresponding to the sub-video frame sequence, including:
[0017] According to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence, determine the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence;
[0018] Select a left half-empty background video frame from the left half-empty background video frame set corresponding to the sub-video frame sequence as the left half-empty background video frame corresponding to the sub-video frame sequence, and select a right half-empty background video frame from the right half-empty background video frame set corresponding to the sub-video frame sequence as the right half-empty background video frame corresponding to the sub-video frame sequence.
[0019] Optionally, based on the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence, determining the left half-air background video frame set and the right half-air background video frame set corresponding to the sub-video frame sequence, including:
[0020] For each first target video frame included in the sub-video frame sequence, using an object detection model to determine the bounding box information of the target object in the first target video frame, and based on the bounding box information of the target object in the first target video frame and the position information of the target object in the first target video frame, determining whether the confidence of the position information of the target object in the first target video frame is greater than a preset confidence threshold. If so, taking the first target video frame as a second target video frame;
[0021] Based on the position information of the target object in each second target video frame included in the sub-video frame sequence and the position information of the cage in each second target video frame included in the sub-video frame sequence, determining the left half-air background video frame set and the right half-air background video frame set corresponding to the sub-video frame sequence.
[0022] Optionally, selecting a left half-air background video frame from the left half-air background video frame set corresponding to the sub-video frame sequence, including:
[0023] Selecting the left half-air background video frame with the earliest time from the left half-air background video frame set corresponding to the sub-video frame sequence;
[0024] Selecting a right half-air background video frame from the right half-air background video frame set corresponding to the sub-video frame sequence, including:
[0025] Selecting the right half-air background video frame with the earliest time from the right half-air background video frame set corresponding to the sub-video frame sequence.
[0026] Optionally, performing background subtraction on each sub-video frame sequence according to the left half-air background video frame and the right half-air background video frame corresponding to each sub-video frame sequence, including:
[0027] Generating an empty background video frame corresponding to each sub-video frame sequence according to the left half-air background video frame and the right half-air background video frame corresponding to each sub-video frame sequence;
[0028] Converting each video frame included in each sub-video frame sequence and the empty background video frame corresponding to each sub-video frame sequence into grayscale video frames respectively, obtaining the grayscale video frame sequence corresponding to each sub-video frame sequence and the empty background grayscale video frame;
[0029] Performing background subtraction on the grayscale video frame sequence corresponding to each sub-video frame sequence according to the empty background grayscale video frame corresponding to each sub-video frame sequence.
[0030] Optionally, based on the foreground video frame sequences corresponding to each sub-video frame sequence, determining the motion trajectory of the target object in the target video frame sequence includes:
[0031] Performing standard image processing on the foreground video frame sequences corresponding to each sub-video frame sequence to obtain the processed foreground video frame sequences corresponding to each sub-video frame sequence;
[0032] For each processed foreground video frame sequence corresponding to each sub-video frame sequence, generating a bounding box of the target object in each processed foreground video frame included in the processed foreground video frame sequence, so as to obtain the bounding boxes corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence;
[0033] Determining the motion trajectory of the target object in the target video frame sequence according to the bounding boxes corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0034] Optionally, determining the motion trajectory of the target object in the target video frame sequence according to the bounding boxes corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence includes:
[0035] Determining whether to adjust the bounding box corresponding to each processed foreground video frame according to the frame pixel difference between every two adjacent video frames in each sub-video frame sequence. If so, adjusting the bounding box corresponding to the processed foreground video frame that needs to be adjusted;
[0036] After adjustment, taking the center point of the bounding box corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence as the tracking point corresponding to each processed foreground video frame;
[0037] Determining the motion trajectory of the target object in the target video frame sequence according to the tracking points corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0038] Optionally, determining the motion trajectory of the target object in the target video frame sequence according to the tracking points corresponding to each processed foreground video frame included in each processed foreground video frame sequence corresponding to each sub-video frame sequence includes:
[0039] Combining the processed foreground video frame sequences corresponding to each sub-video frame sequence into the processed foreground video frame sequence corresponding to the target video frame sequence;
[0040] For each processed foreground video frame in the processed foreground video frame sequence corresponding to the target video frame sequence, except for the first processed foreground video frame, calculate the Euclidean distance between the tracking point corresponding to this processed foreground video frame and the tracking point corresponding to the previous processed foreground video frame of this processed foreground video frame. If the Euclidean distance is greater than a preset distance threshold, then use the tracking point corresponding to this processed foreground video frame as a target tracking point corresponding to the target video frame sequence;
[0041] Based on all the target tracking points corresponding to the target video frame sequence, determine the motion trajectory of the target object in the target video frame sequence.
[0042] A trajectory tracking device, comprising:
[0043] A target video frame sequence acquisition module, configured to acquire a target video frame sequence composed of consecutive video frames of a target object, where the target video frame sequence includes multiple sub-video frame sequences, and the position information of the cage containing the target object is included in each video frame included in each sub-video frame sequence;
[0044] A half-air background video frame determination module, configured to determine the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage included in each sub-video frame sequence, determine the left half-air background video frame and the right half-air background video frame corresponding to each sub-video frame sequence, where the right half-side area of the cage of the target object is included in the left half-air background video frame, and the left half-side area of the cage of the target object is included in the right half-air background video frame;
[0045] A foreground video frame sequence determination module, configured to perform background subtraction on each sub-video frame sequence according to the left half-air background video frame and the right half-air background video frame corresponding to each sub-video frame sequence, to obtain the foreground video frame sequence corresponding to each sub-video frame sequence;
[0046] A motion trajectory determination module, configured to determine the motion trajectory of the target object in the target video frame sequence based on the foreground video frame sequence corresponding to each sub-video frame sequence.
[0047] As can be seen from the above technical solutions, for the trajectory tracking method provided in this application, first, a target video frame sequence composed of consecutive video frames containing a target object is obtained. Considering that the target object may be in the left half area and the rear half area of the cage when moving in the cage, a right half-empty background video frame and a left half-empty background video frame without the target object can be obtained. Based on this, this application determines the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage included in each sub-video frame sequence, determines the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence. Then, this application can perform background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence. Since the target object is more prominent in the foreground video frame sequence, based on the foreground video frame sequence corresponding to each sub-video frame sequence, the motion trajectory of the target object in the target video frame sequence can be accurately determined. Since this application has no special requirements for the environment where the target object is located, the cost of trajectory tracking is reduced, and the practicability is better. Moreover, each sub-video frame sequence included in the target video frame sequence determines a left half-empty background video frame and a right half-empty background video frame, making the left half-empty background video frame and the right half-empty background video frame during background subtraction more consistent with environmental changes, and improving the trajectory tracking effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0049] Figure 1 It is a schematic flowchart of the trajectory tracking method provided by the embodiment of the present application;
[0050] Figure 2a It is a schematic diagram of the real overall recording environment and the camera position;
[0051] Figure 2b It is a schematic diagram of the camera shooting a macaque in the opposite cage;
[0052] Figure 3a It is a schematic diagram of the right half-empty background video frame set corresponding to a sub-video frame sequence;
[0053] Figure 3b It is a schematic diagram of a frame of the right half-empty background video frame corresponding to a sub-video frame sequence;
[0054] Figure 3cSchematic diagram of the left half-empty background video frame set corresponding to a sub-video frame sequence;
[0055] Figure 3d Schematic diagram of a frame of the left half-empty background video frame corresponding to a sub-video frame sequence;
[0056] Figure 3e Schematic diagram of the spliced empty background video frame;
[0057] Figure 4a Schematic diagram of the process of background subtraction for a video frame in three environments;
[0058] Figure 4b Schematic diagram of a frame of the grayscale video frame containing macaques in the third environment;
[0059] Figure 4c Schematic diagram of the empty background grayscale video frame in the first environment;
[0060] Figure 4d Schematic diagram of the foreground video frame obtained by background subtraction in the first environment;
[0061] Figure 4e Schematic diagram of the empty background grayscale video frame in the second environment;
[0062] Figure 4f Schematic diagram of the foreground video frame obtained by background subtraction in the second environment;
[0063] Figure 4g Schematic diagram of the empty background grayscale video frame in the third environment;
[0064] Figure 4h Schematic diagram of the foreground video frame obtained by background subtraction in the third environment;
[0065] Figure 5a Schematic diagram of a frame of the video frame in a sub-video frame sequence;
[0066] Figure 5b Schematic diagram of the foreground video frame obtained by background subtraction;
[0067] Figure 5c Schematic diagram of the processed foreground video frame;
[0068] Figure 6a Schematic diagram of tracking the trajectory of macaques based on 6 video frames taken during the day;
[0069] Figure 6b Schematic diagram of tracking the trajectory of macaques based on 6 video frames taken at night;
[0070] Figure 7aSchematic diagram of the change of the IoU value of MonkeyTrail over time;
[0071] Figure 7b Schematic diagram of the change of the IoU value of SSD over time;
[0072] Figure 7c Schematic diagram of the change of the IoU value of YOLOv5 over time;
[0073] Figure 7d Schematic diagram of the change of the IoU value of BSM over time;
[0074] Figure 7e Schematic diagram of the change of the IoU value of FDM over time;
[0075] Figure 7f Schematic diagram of the total activity of the macaques estimated in the experiment and the period when the animals were completely occluded;
[0076] Figure 8 Schematic diagram of the tracking success rates of MonkeyTrail, YOLOv5, SSD, BSM, and FDM varying with the overlap threshold;
[0077] Figure 9a Schematic diagram of the total activity of the first macaque in 2019 and 2020;
[0078] Figure 9b Schematic diagram of the total activity of the second macaque in 2019 and 2020;
[0079] Figure 10a Schematic diagram of the heatmap of the spatial preference of the first macaque in 2019;
[0080] Figure 10b Schematic diagram of the heatmap of the spatial preference of the first macaque in 2020;
[0081] Figure 10c Schematic diagram of the heatmap of the spatial preference of the second macaque in 2019;
[0082] Figure 10d Schematic diagram of the heatmap of the spatial preference of the second macaque in 2020;
[0083] Figure 11 Schematic diagram of the structure of the trajectory tracking device provided by the embodiment of the present application;
[0084] Figure 12 Hardware structure block diagram of the trajectory tracking device provided by the embodiment of the present application. Detailed implementation manners
[0085] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0086] The present application provides a trajectory tracking method. Next, the trajectory tracking method provided by the present application will be introduced in detail through the following embodiments.
[0087] Please refer to Figure 1 , which shows a schematic flowchart of the trajectory tracking method provided by the embodiments of the present application. The trajectory tracking method may include:
[0088] Step S101, obtaining a target video frame sequence composed of consecutive video frames of a target object.
[0089] Taking the motion trajectory tracking scenario of macaques as an example for illustration. This experiment can be carried out in accordance with the international standards for the care and use of non-human primates, and after obtaining the approval of the Animal Experiment Ethics Committee (AEEI-2019-077), adult macaques are placed in an animal room with a temperature of 18 - 26 °C and a humidity of 40% - 70%, and each macaque is separately housed in adjacent cages. During the experiment, the macaques are given a fixed amount of special monkey feed and sufficient drinking water (the macaques can drink freely), as well as a fixed amount of fresh vegetables and fruits every day. In addition, to be more consistent with the external environment, the animal room maintains a 12-hour light-dark cycle.
[0090] Refer to Figure 2a and Figure 2b as shown, Figure 2a is a schematic diagram of the real overall recording environment and the camera position, Figure 2b is a schematic diagram of the camera shooting the macaque in the opposite cage, Figure 2a The two cages at the white border in Figure 2b correspond to the two opposite cages photographed by the Camera camera shown in Figure 2a The specific positions of the camera are the positions where the black borders are located in the upper left and upper right corners, and Figure 2b the position of the camera (i.e., the camera) shown in
[0091] After setting the camera position, video shooting can be performed on the target object to obtain a target video frame sequence composed of consecutive video frames of the target object.
[0092] In this step, the target video frame sequence includes multiple sub-video frame sequences. The number of video frames included in the multiple sub-video frame sequences can be determined according to factors such as the speed of environmental change (i.e., the frequency of generating an empty background) and the movement of the macaque. For example, the sequence composed of video frames within every 40 minutes can be used as a sub-video frame sequence (if the target object stays on one side of the cage for a long time, it may be difficult to generate a virtual empty background subsequently, which may affect the tracking accuracy. However, empirical data shows that, on average, updating the background every 40 minutes shows quite good results, thus providing a sufficient time window for generating multiple virtual empty backgrounds). For example, assuming that the frames per second (fps) preset in this application is 5fps, then a sub-video frame sequence can be composed of every 40 * 60 * 5 video frames.
[0093] It should be noted that the above-mentioned macaque movement trajectory tracking scenario, as well as the number of video frames included in the sub-video frame sequence and the frames per second, are all examples and do not limit this application.
[0094] In this embodiment, the position information of the cage containing the target object in each video frame included in each video frame sequence can be pre-marked for subsequent calculation. Thus, each video frame included in each sub-video frame sequence provided in this embodiment includes the position information of the cage containing the target object.
[0095] Preferably, considering that when the camera shoots the target object, it is very likely to shoot objects other than the target object into the target video frame sequence. For the convenience of downstream processing and to avoid interference from other objects, the consecutive video frames shot by the camera can be cropped according to the position information of the cage containing the target object to obtain a target video frame sequence composed of consecutive video frames that only contain the cage and the target object in the cage.
[0096] Optionally, the position information of the cage in this step can be the coordinate information of the four vertices in the bounding box (also called the bounding box) of the cage.
[0097] Step S102: Determine the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determine the left half empty background video frame and the right half empty background video frame corresponding to each sub-video frame sequence.
[0098] Considering that in the prior art, cages made of special materials such as transparent materials are required, resulting in high costs and poor practicability. After in-depth research, the inventor of this case came up with the idea that during the movement of the target object, there are situations where it is located in the left half area and the right half area of the cage. Based on these two situations, a complete empty background can be obtained by splicing the half-empty backgrounds, so that this application has no special requirements for the environment where the target object is located (here it refers to the cage). A relatively accurate complete empty background can be obtained by using an ordinary cage used in normal life, reducing the cost of making cages made of materials such as transparent materials, and the practicability is better when using an ordinary cage.
[0099] To implement the above idea, for each sub-video frame sequence among multiple sub-video frame sequences, first determine the position information of the target object in each video frame included in the sub-video frame sequence. Optionally, the position information can be the coordinate information of the four vertices in the bounding box of the target object; then, for each video frame included in the sub-video frame sequence, based on the position information of the target object in the video frame and the position information of the cage that houses the target object in the video frame, determine whether the video frame is a left half-empty background video frame or a right half-empty background video frame, thereby obtaining the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence.
[0100] Here, the target object included in the left half-empty background video frame is in the right half area of the cage, and the target object included in the right half-empty background video frame is in the left half area of the cage.
[0101] For example, if the position coordinates of the target object in a video frame are (1,1), (1,2), (6,1), and (6,2), and the position coordinates of the cage are (0,0), (0,8), (8,0), and (8,8), then it can be considered that this video frame is the left half-empty background video frame corresponding to the corresponding sub-video sequence.
[0102] Step S103: Perform background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence.
[0103] In this step, background subtraction (BSM) can be used to perform background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence. Specifically, for each video frame included in each sub-video frame sequence, perform background subtraction on the video frame according to the left half-empty background video frame and the right half-empty background video frame corresponding to the sub-video frame sequence to obtain the foreground video frame corresponding to the video frame. Here, the influence of the empty background (mainly referring to the cage and the stationary objects other than the target object in the cage) is removed from the foreground video frame corresponding to the video frame, making the target object included therein more prominent.
[0104] Step S104: Determine the motion trajectory of the target object in the target video frame sequence based on the foreground video frame sequence corresponding to each sub-video frame sequence.
[0105] The trajectory tracking method provided in this application first obtains a target video frame sequence composed of consecutive video frames containing the target object. Considering that the target object may be in the left half area and the rear half area of the cage when moving in the cage, the right half empty background video frame and the left half empty background video frame without the target object can be obtained. Based on this, this application determines the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage included in each sub-video frame sequence, determines the left half empty background video frame and the right half empty background video frame corresponding to each sub-video frame sequence. Then, this application can perform background subtraction on each sub-video frame sequence according to the left half empty background video frame and the right half empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence. Since the target object is more prominent in the foreground video frame sequence, based on the foreground video frame sequence corresponding to each sub-video frame sequence, the motion trajectory of the target object in the target video frame sequence can be accurately determined. Since this application has no special requirements for the environment where the target object is located, the cost of trajectory tracking is reduced, and the practicability is better. Moreover, each sub-video frame sequence included in the target video frame sequence determines a left half empty background video frame and a right half empty background video frame, making the left half empty background video frame and the right half empty background video frame during background subtraction more consistent with environmental changes and improving the trajectory tracking effect.
[0106] In an embodiment of this application, the process of the foregoing "Step S102: Determine the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determine the left half empty background video frame and the right half empty background video frame corresponding to each sub-video frame sequence" is introduced.
[0107] For the convenience of description, take any one of the multiple sub-video frame sequences as an example to introduce the process of "determining the position information of the target object in this sub-video frame sequence, and based on the determined position information and the position information of the cage in this sub-video frame sequence, determine the left half empty background video frame and the right half empty background video frame corresponding to this sub-video frame sequence". It should be understood that for each sub-video frame sequence, the process of determining the left and right half empty background video frames is the same, and this application will not elaborate one by one.
[0108] Specifically, the process of "determining the position information of the target object in this sub-video frame sequence, and based on the determined position information and the position information of the cage in this sub-video frame sequence, determine the left half empty background video frame and the right half empty background video frame corresponding to this sub-video frame sequence" may include:
[0109] Step S201: Calculate the frame pixel difference between every two adjacent video frames in the sub-video frame sequence by using the frame difference method, and determine the position information of the target object in each video frame included in the sub-video frame sequence based on the calculated frame pixel differences.
[0110] In this step, the frame pixel difference between every two adjacent video frames in the sub-video frame sequence can be calculated by using the frame difference method FDM. Since the number of frame pixel differences can provide a rough estimate of the movement of the target object, and at the same time, the position of the pixel difference can provide the approximate position of the bounding box for the target object, therefore, in this step, the position information of the target object in each video frame included in the sub-video frame sequence can be determined based on the calculated frame pixel differences, and the specific determination process is the same as the prior art and will not be elaborated here.
[0111] Step S202: For each calculated frame pixel difference, if the frame pixel difference is greater than or equal to a preset high motion threshold, both of the two adjacent video frames corresponding to the frame pixel difference are used as the first target video frames to obtain all the first target video frames included in the sub-video frame sequence.
[0112] Each sub-video frame sequence includes multiple video frames. Thus, multiple frame pixel differences can be obtained in the previous step. For example, if a sub-video frame sequence includes 10 video frames, then by calculating the frame pixel difference between every two video frames, 9 frame pixel differences can be obtained.
[0113] In this step, considering that the tracking position obtained by FDM is more reliable when the target object moves actively, based on this, this step can use a preset high motion threshold (i.e., pixel difference threshold) to screen the video frames in the sub-video frame sequence.
[0114] Specifically, for each calculated frame pixel difference, in this step, the frame pixel difference can be compared with the preset high motion threshold. If the frame pixel difference is greater than or equal to the preset high motion threshold, then in this step, both of the two video frames corresponding to the frame pixel difference can be used as the first target video frames.
[0115] For example, a sub-video frame sequence includes 10 video frames, which are video frames 1 to 10 respectively. Among them, the frame pixel difference between video frames 1 and 2 is denoted as frame pixel difference 1, the frame pixel difference between video frames 2 and 3 is denoted as frame pixel difference 2, and so on. The frame pixel difference between video frames 9 and 10 is denoted as frame pixel difference 9. Assume that only frame pixel differences 1, 3, 7, 8, and 9 among frame pixel differences 1 to 9 are greater than the high motion threshold. Then in this step, it can be determined that the first target video frames included in the sub-video frame sequence include video frames 1, 2, 3, 4, and 7 to 10.
[0116] Step S203: Determine the left half-air background video frames and right half-air background video frames corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence.
[0117] Specifically, the process of this step may include:
[0118] Step S2031: Determine the left half-air background video frame set and right half-air background video frame set corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence.
[0119] In this step, all the first target video frames included in the sub-video frame sequence can be screened into a left half-air background video frame set (initial L set) and a right half-air background video frame set (initial R set). Here, the left half-air background video frame set corresponding to a sub-video frame sequence refers to the set composed of the first target video frames in which the target object is in the right half-side area of the cage in the sub-video frame sequence, and the right half-air background video frame set corresponding to a sub-video frame sequence refers to the set composed of the first target video frames in which the target object is in the left half-side area of the cage in the sub-video frame sequence.
[0120] In an optional embodiment, considering that the frame pixel difference obtained by FDM can only provide the approximate position of the target object bounding box, there may be a situation of incorrect screening when screening the initial L set and initial R set based on the position information of the target object and the cage obtained by FDM. To obtain more accurate L set and R set, this embodiment can also detect the position information of the target object in the initial L set and R set based on the target detection model, and only retain the video frames with high-confidence position information.
[0121] Specifically, the process of step S2031 may include:
[0122] Step A1: For each first target video frame included in the sub-video frame sequence, use the target detection model to determine the bounding box information of the target object in the first target video frame. According to the bounding box information of the target object in the first target video frame and the position information of the target object in the first target video frame, determine whether the confidence of the position information of the target object in the first target video frame is greater than a preset confidence threshold. If so, use this first target video frame as the second target video frame.
[0123] Here, the target detection model is a model existing in the prior art, such as the YOLOv5 model. The training process can be referred to the introduction in the prior art in detail and will not be elaborated here.
[0124] In this step, the object detection model can detect the bounding box information of the target object from the first target video frame, compare the bounding box information detected by the model with the position information of the target object determined by the FDM, and determine whether the confidence level of the position information of the target object under the FDM is greater than the preset confidence threshold. If so, the first target video frame can be used as the second target video frame. If not, the first target video frame will not be used as the second target video frame. In this way, all the second target video frames included in the sub-video frame sequence can be obtained, and the process of the object detection model screening video frames with high-confidence position information is completed.
[0125] Step A2: According to the position information of the target object in each second target video frame included in the sub-video frame sequence and the position information of the cage in each second target video frame included in the sub-video frame sequence, determine the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence.
[0126] In this step, based on the position information of the target object in each second target video frame included in the sub-video frame sequence and the position information of the cage in each second target video frame included in the sub-video frame sequence, the initial L set and R set can be screened to obtain the final L set and R set as the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence.
[0127] Step S2032: Select a left half-empty background video frame from the left half-empty background video frame set corresponding to the sub-video frame sequence as the left half-empty background video frame corresponding to the sub-video frame sequence, and select a right half-empty background video frame from the right half-empty background video frame set corresponding to the sub-video frame sequence as the right half-empty background video frame corresponding to the sub-video frame sequence.
[0128] In this step, a left half-empty background video frame can be arbitrarily selected from the left half-empty background video frame set corresponding to the sub-video frame sequence, and a right half-empty background video frame can be arbitrarily selected from the right half-empty background video frame set corresponding to the sub-video frame sequence.
[0129] To improve the quality of the empty background, when selecting the left and right half-empty background video frames, there should be a high position distinction between the tracking boxes (i.e., bounding boxes or enclosing boxes). Preferably, the nearest pair of L and R frames can be selected in chronological order. For example, select the left half-empty background video frame with the earliest time in the left half-empty background video frame set corresponding to the sub-video frame sequence, and select the right half-empty background video frame with the earliest time from the right half-empty background video frame set corresponding to the sub-video frame sequence.
[0130] In this embodiment, for each sub-video frame sequence, a pair of left semi-empty background video frames and right semi-empty background video frames are obtained according to the above process. Thus, a series of automatically generated left semi-empty background video frames and right semi-empty background video frames can be obtained for background subtraction.
[0131] After obtaining the left semi-empty background video frames and right semi-empty background video frames corresponding to each sub-video frame sequence in the foregoing embodiment, background subtraction can be performed according to step S103.
[0132] In the following embodiment, the process of step S103, "performing background subtraction on each sub-video frame sequence according to the left semi-empty background video frames and right semi-empty background video frames corresponding to each sub-video frame sequence", will be described.
[0133] Specifically, the process of step S103, "performing background subtraction on each sub-video frame sequence according to the left semi-empty background video frames and right semi-empty background video frames corresponding to each sub-video frame sequence", includes:
[0134] Step S301: Generate an empty background video frame corresponding to each sub-video frame sequence according to the left semi-empty background video frame and right semi-empty background video frame corresponding to each sub-video frame sequence.
[0135] In this step, the left semi-empty background included in the left semi-empty background video frame corresponding to each sub-video frame sequence and the right semi-empty background included in the right semi-empty background video frame can be spliced to obtain a complete empty background video frame.
[0136] See Figure 3a 、 Figure 3b 、 Figure 3c 、 Figure 3d and Figure 3e shown. Figure 3a is a schematic diagram of a set of right semi-empty background video frames corresponding to a sub-video frame sequence. Figure 3b is a schematic diagram of a frame of right semi-empty background video frame corresponding to a sub-video frame sequence. Figure 3c is a schematic diagram of a set of left semi-empty background video frames corresponding to a sub-video frame sequence. Figure 3d is a schematic diagram of a frame of left semi-empty background video frame corresponding to a sub-video frame sequence. Figure 3e is a schematic diagram of the spliced empty background video frame, in which the position of the macaque is shown by a thick black bounding box. Figure 3a 、 Figure 3b 、 Figure 3c and Figure 3d are the same sub-video frame sequences. In this embodiment, the right semi-empty background video frame shown in Figure 3a can be selected from the set of right semi-empty background video frames, and the left semi-empty background video frame shown in Figure 3b can be selected from the set of left semi-empty background video frames shown in Figure 3c and the left semi-empty background video frame shown in Figure 3dFor the left half empty background video frame shown, then Figure 3b The right half empty background area in which does not contain macaques is spliced with Figure 3d The left half empty background area in which does not contain macaques to obtain Figure 3e The empty background video frame shown, that is, Figure 3e The left half empty background area of the empty background video frame shown comes from Figure 3d The left half empty background area of Figure 3e The right half empty background area of the empty background video frame shown comes from Figure 3d The right half empty background area of
[0137] Step S302: Convert each video frame included in each sub-video frame sequence and the empty background video frame corresponding to each sub-video frame sequence into grayscale video frames respectively, to obtain the grayscale video frame sequence corresponding to each sub-video frame sequence and the empty background grayscale video frame.
[0138] In this embodiment, it is necessary to first convert the original RGB color frame into grayscale and then perform background subtraction. Based on this, first, each video frame included in each sub-video frame sequence and the empty background video frame corresponding to each sub-video frame sequence are respectively converted into grayscale video frames, to obtain the grayscale video frame sequence corresponding to each sub-video frame sequence and the empty background grayscale video frame.
[0139] Step S303: Perform background subtraction on the grayscale video frame sequence corresponding to each sub-video frame sequence according to the empty background grayscale video frame corresponding to each sub-video frame sequence.
[0140] In this step, background subtraction can be performed using background subtraction method. The specific subtraction process is the prior art and will not be elaborated here.
[0141] In order to verify the influence of environmental changes on the background subtraction effect, the inventor of this case conducted several experiments. The experimental process can refer to Figure 4a Figure 4a As the schematic diagram of the process of performing background subtraction on a video frame in three environments, the experimental process includes: performing background subtraction on Figure 4b and Figure 4c to obtain Figure 4d ; performing background subtraction on Figure 4b and Figure 4e to obtain Figure 4f ; Figure 4b and Figure 4g are subtracted from each other to obtain Figure 4h wherein, Figure 4b is the schematic diagram of a frame of grayscale video frame containing macaques in the third environment, Figure 4c is the schematic diagram of the empty background grayscale video frame in the first environment, Figure 4e is a schematic diagram of an empty background grayscale video frame in the second environment. Figure 4g It is a schematic diagram of an empty background grayscale video frame in the third environment. Figure 4d This is a schematic diagram of the foreground video frame obtained by background subtraction in the first environment. Figure 4f This is a schematic diagram of the foreground video frame obtained by background subtraction in the second environment. Figure 4h This is a schematic diagram of a foreground video frame obtained by background subtraction in the third environment.
[0142] Here, the first environment, the second environment and the third environment correspond to different shooting time periods. For example, in this embodiment, the empty background grayscale video frame in the first environment is an empty background grayscale video frame obtained based on the sub-video frame sequence of the first hour, the empty background grayscale video frame in the second environment is an empty background grayscale video frame obtained based on the sub-video frame sequence of the second hour, and the empty background grayscale video frame in the third environment is an empty background grayscale video frame obtained based on the sub-video frame sequence of the third hour.
[0143] Figure 4c , Figure 4e and Figure 4g There are changes in the brightness of the background, the appearance of details, etc. Figure 4b The shooting time and Figure 4g The corresponding shooting time is closest, making Figure 4d and Figure 4f None of the foreground video frames shown can completely eliminate the effect of the cage, but Figure 4h The foreground video frame shown successfully highlights the macaque and performs relatively well on background removal, demonstrating the necessity of frequently updating the empty background.
[0144] In summary, compared to the prior art, which usually requires removing the target object from the cage to obtain a physically created empty background, the present application can generate virtual empty background video frames based on the left and right half-empty background video frames, which is more convenient and practical. In addition, during a long shooting process, the environment changes over time (such as cage movement, changes in lighting conditions, and the introduction of new species in the frame, etc.), which makes the pre-obtained blank background mismatch with the real background, and the background subtraction effect becomes worse and worse over time. The present application generates a corresponding empty background video frame for each sub-video frame sequence, which improves the effect of background subtraction and thus improves the trajectory tracking effect of the target object.
[0145] Furthermore, after the foreground video frame sequence is obtained by background subtraction, the present embodiment can obtain the motion trajectory of the target object in the target video frame sequence according to step S104.
[0146] See also Figure 4hThe foreground video frame shown, although it can highlight the macaque to a certain extent, the macaque is still not obvious enough. To further highlight the target object in the foreground video frame, the present application can also perform standard image processing on the foreground video frame sequence obtained in the foregoing embodiment, and then perform trajectory tracking after the processing.
[0147] Based on this, the process of the above "step S104, based on the foreground video frame sequence corresponding to each sub-video frame sequence, determine the motion trajectory of the target object in the target video frame sequence" may include:
[0148] Step S401, perform standard image processing on the foreground video frame sequence corresponding to each sub-video frame sequence to obtain the processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0149] Here, the standard image processing includes but is not limited to the following processing: spatial median filtering, binarization, erosion, and dilation.
[0150] It can be understood that some image processing parameters are involved in the standard image processing. The present application can optimize these parameters according to the experimental environment used to obtain the parameters suitable for the current experimental environment. If applied to other environments, appropriate adjustments can be made to obtain the best effect.
[0151] See Figure 5a 、 Figure 5b and Figure 5c shown, where Figure 5a is a schematic diagram of a video frame in a sub-video frame sequence, Figure 5b is a schematic diagram of the foreground video frame obtained by background subtraction, Figure 5c is a schematic diagram of the processed foreground video frame. Figure 5a The video frames shown illustrate a typical situation of a daily life cage, Figure 5b is for Figure 5a the video frame shown to obtain the foreground video frame by background subtraction, Figure 5c is for Figure 5b the foreground video frame in to perform spatial median filtering, binarization, erosion, and dilation to obtain the processed foreground video frame. Compared with the foreground video frame in Figure 5b the foreground video frame in, Figure 5c can better highlight the target object.
[0152] Step S402, for the processed foreground video frame sequence corresponding to each sub-video frame sequence, generate a bounding box of the target object in each processed foreground video frame included in the processed foreground video frame sequence, so as to obtain the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0153] Specifically, in this step, for each processed foreground video frame sequence corresponding to a sub-video frame sequence, a bounding box can be formed by finding the smallest rectangular area covering the foreground (i.e., the target object) in each processed foreground video frame included in the processed foreground video frame sequence, so as to obtain the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0154] Step S403: Determine the motion trajectory of the target object in the target video frame sequence according to the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0155] There are multiple implementation manners for this step. The present application provides but is not limited to the following one implementation manner.
[0156] Optionally, the process of this step includes:
[0157] Step S4031: Determine whether it is necessary to adjust the bounding box corresponding to each processed foreground video frame according to the frame pixel difference between every two adjacent video frames in each sub-video frame sequence. If so, adjust the bounding box corresponding to the processed foreground video frame that needs to be adjusted.
[0158] Considering that standard image processing is performed in step S401, overprocessing may occur in the case of severe occlusion of the target object, resulting in over-eliminating the foreground result. To avoid the situation of over-eliminating the foreground result caused by severe occlusion, this step can determine whether it is necessary to adjust the bounding box corresponding to each processed foreground video frame according to the frame pixel difference between every two adjacent video frames in each sub-video frame sequence. If so, adjust the bounding box corresponding to the processed foreground video frame that needs to be adjusted. For example, if the bounding box of the target object in the foreground video frame 1 is too small due to standard image processing and the target object is in a high-density occlusion area, and if the position information of the target object in the video frame corresponding to this foreground video frame 1 (the video frame corresponding to this foreground video frame 1 refers to the video frame corresponding to this foreground video frame 1 in the corresponding sub-video frame sequence) determined according to the frame pixel difference calculated by FDM is also in this area, then this step can track the bounding box with the same size as the previous foreground video frame.
[0159] Step S4032: After the adjustment, use the center point of the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence as the tracking point corresponding to each processed foreground video frame.
[0160] For each processed foreground video frame sequence corresponding to a sub-video frame sequence, the center point of the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence can be used as a tracking point for each processed foreground video frame included in the processed foreground video frame sequence. For the convenience of subsequent description, in this step, the tracking points for all processed foreground video frames included in the processed foreground video frame sequence are denoted as the tracking points corresponding to the sub-video frame sequence, and thus the tracking points corresponding to each processed foreground video frame can be obtained.
[0161] For example, a processed foreground video frame sequence corresponding to a sub-video frame sequence includes 10 processed foreground video frames. A total of 10 tracking points can be obtained from the 10 processed foreground video frames (each processed foreground video frame contains 1 tracking point), and these 10 tracking points are the 10 tracking points corresponding to the sub-video frame sequence.
[0162] It should be noted that if the bounding box corresponding to a processed foreground video frame is adjusted in the previous step, the tracking point for this processed foreground video frame refers to the center point of the adjusted bounding box corresponding to this processed foreground video frame.
[0163] Step S4033: Determine the motion trajectory of the target object in the target video frame sequence according to the tracking points corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence.
[0164] In this step, the motion trajectory of the target object in the target video frame sequence can be extracted by connecting the respective tracking points, which can be used to reflect the central position of the target object's body.
[0165] See Figure 6a and Figure 6b , Figure 6a are schematic diagrams for tracking the trajectory of a macaque based on 6 video frames taken during the day. Figure 6b is a schematic diagram for tracking the trajectory of a macaque based on 6 video frames taken at night. Figure 6a and Figure 6b each contain 6 video frames, the time interval between each video frame is greater than 10 s, and Figure 6a and Figure 6b show different motions and various levels of occlusion of the macaque ( Figure 6a and Figure 6b show the macaque outlined in a black box). For Figure 6a and Figure 6b , tracking is performed from left to right and top to bottom for the video frames. One tracking point is obtained for each video frame, and the respective tracking points are connected in sequence to obtain Figure 6a and Figure 6bThe motion trajectory shown by the black broken line in [the figure], it can be seen that even in the case of severe occlusion, the present application can still obtain a relatively accurate motion trajectory, and the present application is applicable to various motion situations of the target object.
[0166] In an alternative embodiment, the process of step S4033 includes:
[0167] Step B1: Combine the processed foreground video frame sequences corresponding to each sub-video frame sequence to form the processed foreground video frame sequence corresponding to the target video frame sequence.
[0168] Step B2: For each processed foreground video frame in the processed foreground video frame sequence corresponding to the target video frame sequence except the first processed foreground video frame, calculate the Euclidean distance between the tracking point corresponding to this processed foreground video frame and the tracking point corresponding to the previous processed foreground video frame of this processed foreground video frame. If this Euclidean distance is greater than the preset distance threshold, then use the tracking point corresponding to this processed foreground video frame as a target tracking point corresponding to the target video frame sequence.
[0169] In order to reduce the interference of the subtle movements of the limbs or head on the tracking points, in this embodiment, the movements less than the preset distance threshold can be discarded. That is to say, when the Euclidean distance between the tracking point corresponding to a processed foreground video frame and the tracking point corresponding to the previous processed foreground video frame does not exceed this threshold, the change of the tracking point corresponding to this processed foreground video frame is filtered out.
[0170] Step B3: Determine the motion trajectory of the target object in the target video frame sequence according to all the target tracking points corresponding to the target video frame sequence.
[0171] Here, the distance of the trajectory represents the moving amount of the target object, that is, the total amount of motion of the target object.
[0172] To further verify the tracking effect of the trajectory tracking method provided by this application, the inventors of this case created a test dataset to quantify the performance compared with the prior art. To test its reliability in various motion states of macaques, the selected data includes continuous motion, continuous stillness, and transitions between motion and stillness. To test different occlusion levels, the data selected by this application includes cases where macaques are occluded by metal bars, grids, and their own body parts. In addition, the three monkeys used in the test dataset have different appearances, including size and fur color. No trained animals appear in the test dataset. Finally, to test the performance under different lighting conditions, the dataset includes daytime and nighttime video data. The entire test dataset contains 55 minutes of video, consisting of 13 continuous clips of different lengths to cover all the above conditions. To reduce the workload of manual labeling, the frame rate was reduced to 2 or 5 fps, and the total number of frames is 8,130. The LabelImg marking platform was used to manually select the bounding box containing the entire animal as the baseline.
[0173] The trajectory tracking method provided by this application fully utilizes the combined advantages of BSM, FDM, and YOLOv5. First, the improvements relative to these three methods and the deep learning model SSD are quantitatively analyzed through the following experiments. Before the experiment, a manually annotated dataset was prepared as the baseline, and then the results of the trajectory tracking method provided by this application were compared with those of BSM, FDM, YOLOv5, and SSD. To provide a comprehensive comparison, the test dataset contains 13 video clips from three animals, covering various motion patterns, occlusions, and lighting conditions (daytime and nighttime). For easy comparison, the FDM, BSM, YOLOv5, and SSD processes have the same corresponding functions as those in the trajectory tracking method provided by this application. For example, the image processing in BSM without background update is the same as that in the trajectory tracking method provided by this application, and the YOLOv5 model used for object detection is the same as that used in the trajectory tracking method provided by this application, etc. The tracking accuracy is determined by IoU, which measures the overlap degree between the bounding box of the target object generated by each method (i.e., the adjusted bounding box shown in step S4031) () and the bounding box generated by the baseline (), and the definition of IoU is as follows.
[0174] To visualize the tracking results of different methods, this experiment shows the IoU value of each frame in the test dataset on a continuous timeline, specifically as Figure 7a the schematic diagram of the change of the IoU value of MonkeyTrail over time shown, Figure 7b the schematic diagram of the change of the IoU value of SSD over time shown, Figure 7c the schematic diagram of the change of the IoU value of YOLOv5 over time shown, Figure 7d the schematic diagram of the change of the IoU value of BSM over time shown, and Figure 7eSchematic diagram of the change of the IoU value of FDM over time, where Figure 7a is this application, Figure 7a , Figure 7b , Figure 7c , Figure 7d and Figure 7e the thick black line in Figure 7b is the average value of IoU obtained by this application, Figure 7c the dotted line in Figure 7d is the average value of IoU obtained by SSD, Figure 7e the dotted line in
[0175] To better understand how animal activity and occlusion affect different methods, refer to Figure 7f the schematic diagram of the total activity of macaques estimated by the experiment and the period when the animal is completely occluded shown in Figure 7f the dark grayish-black part in Figure 7f represents the total activity of macaques estimated by the experiment,
[0176] Combined with Figure 7f , and by comparing Figure 7a , Figure 7b , Figure 7c , Figure 7d and Figure 7e it is found that the IoU value of MonkeyTrail is generally larger and has less variability, indicating that the MonkeyTrail method provides the most accurate and stable tracking results; in contrast, the IoU values of SSD and YOLOv5 show highly fluctuating performance, and reasonable tracking results are interrupted by the IoU zero phase, indicating that SSD and YOLOv5 simply cannot detect animals; compared with Figure 7e , it can be seen that these failures usually occur when the animal is occluded, which also indicates that SSD and YOLOv5 are highly sensitive to occlusion.
[0177] In addition, without frequent updating of the blank background, the traditional BSM shows noisy results. In this case, Figure 7c it is difficult for YOLOv5 shown in Figure 7d to obtain an accurate foreground area of macaques through standard image processing, resulting in inaccurate tracking results. In addition, compared with Figure 7e FDM shown in
[0178] In summary, combined withFigure 7a , Figure 7b , Figure 7c , Figure 7d and Figure 7e can reflect the advantages and disadvantages of the FDM, BSM, YOLOv5, and SSD methods, and the combination of them in this application can achieve accurate and stable tracking.
[0179] To provide a more comprehensive and quantitative comparison, the inventors of this case also compared the success rate with the overlap threshold that varies with the system. This threshold is commonly used to measure the performance of object tracking methods, and the results are as Figure 8 shown. Figure 8 shows a schematic diagram of the tracking success rates of MonkeyTrail, YOLOv5, SSD, BSM, and FDM varying with the overlap threshold. When the "overlap threshold" of a certain frame is greater than the set threshold, that frame is considered successful, and the percentage of the total number of successful frames in all frames is defined as the "success rate".
[0180] As Figure 8 shown, the average success rate of MonkeyTrail within the entire overlap threshold range is 0.690, the average success rate of YOLOv5 within the entire overlap threshold range is 0.662, the average success rate of SSD within the entire overlap threshold range is 0.316, the average success rate of BSM within the entire overlap threshold range is 0.396, and the average success rate of FDM within the entire overlap threshold range is 0.226.
[0181] Then referring to Figure 8 shown, compared with SSD, BSM, and FDM, the success rate of the trajectory tracking method (MonkeyTrail) provided by this application is consistently higher within the entire overlap threshold range. Although YOLOv5 seems to be more accurate in a small number of frames, experiments have shown that it fails to fully detect monkeys in approximately 15% of the frames; in contrast, the trajectory tracking method provided by this application shows more stable tracking results and produces more favorable overall performance.
[0182] To prove the practical value of the method provided by this application in analyzing the behavior of macaque monkeys, this application uses the trajectories recorded by this method to calculate the exercise volume and spatial preference of two macaque monkeys in two time periods separated by one year. Each time period contains records of five consecutive days. The exercise volume and spatial preference are useful indicators of behavioral changes caused by factors such as drug injection, surgery, and changes in external conditions. Monitoring these parameters daily can reveal their acute and chronic effects on behavior.
[0183] The total activity volume is as Figure 9a and Figure 9b shown, whereFigure 9a Schematic diagram of the total activity of the first macaque in 2019 and 2020, Figure 9b is a schematic diagram of the total activity of the second macaque in 2019 and 2020. In addition to the obvious results of the sleep and wake cycles affecting the activity patterns, both macaques showed a bimodal activity pattern during the day in the cage in the 2019 records, which may be the combined result of physiological activities (such as napping) and care (i.e., feeding, lighting); in 2020, with changes in both the lighting schedule and the care activity pattern, the total activity of both macaques became a trimodal pattern. These results illustrate the value of tracking total activity to capture behavioral changes in macaques in their daily cages. These changes not only reflect the influence of the above-mentioned non-accidental external factors, but also provide valuable information for verifying intentional treatments (such as drug administration or surgery).
[0184] Heat maps of spatial preference are as Figure 10a , Figure 10b , Figure 10c and Figure 10d shown, Figure 10a is a schematic diagram of the heat map of the spatial preference of the first macaque in 2019, Figure 10b is a schematic diagram of the heat map of the spatial preference of the first macaque in 2020, Figure 10c is a schematic diagram of the heat map of the spatial preference of the second macaque in 2019, Figure 10d is a schematic diagram of the heat map of the spatial preference of the second macaque in 2020, Figure 10a , Figure 10b , Figure 10c and Figure 10d shown. The horizontal and vertical axes of the heat maps represent the X and Y coordinates of the cage respectively. Each heat map area represents the number of times the macaque's trajectory passes through that space, normalized by the maximum number found in a region. Each heat map is obtained by averaging the trajectory data for five days.
[0185] First, the two-dimensional (2D) projection of the cage was divided into 16 regions. Then, the number of times the macaque's trajectory passed through each region (normalized by the maximum number found in a region) was calculated to determine its spatial preference. Similar to the total activity, the experimental analysis of the average spatial preference of the macaques from 11:00 am to 12:00 pm for 5 consecutive days in 2019 and 2020 was carried out. In 2019, the first macaque showed a preference for hanging at the top (specifically, see Figure 10a ), while the second macaque preferred to sit at the bottom of the cage (specifically, see Figure 10c ). One year later, the second macaque still showed a preference for sitting at the bottom of the cage (specifically, see Figure 10d) However, the first macaque changed its behavioral preference and spent more time sitting rather than hanging (for details, see Figure 10b ).
[0186] In summary, as shown in Figure 9a , Figure 9b , Figure 10a , Figure 10b , Figure 10c , Figure 10d , tracking the behavior of animals in their daily life cages can provide useful information about spatio-temporal domain movement patterns.
[0187] In summary of the above embodiments, the motion trajectories extracted by the trajectory tracking method provided in this application can be used to analyze spatial preference and exercise amount; compared with the motion recorded by pixel differences, although the motion calculated by trajectory distance is not suitable for detecting subtle motions of small body parts, it can provide more accurate results of the whole-body motion level, especially in the case of cage restrictions. The trajectory also provides spatial information missing in pixel differences. Although pose estimation can be used to analyze more detailed motion patterns, the overall motion trajectory of the animal can still provide important information.
[0188] In addition to extracting motion trajectories, the trajectory tracking method provided in this application can also provide the contour of the animal for each frame, including the bounding box or body mask, which helps future pose detection algorithms. In addition, the bounding box and its content provided by the trajectory tracking method provided in this application can be used as training samples to train or fine-tune other deep learning-based methods to achieve more complex detection and recognition.
[0189] The trajectory tracking method provided in this application uses an ordinary high-definition camera installed outside the cage to record the macaques in the daily life cage of macaques for a long time. This low-cost setup can be scaled up to automatically track many animals, thus allowing large-scale applications. In addition to behavioral tracking in future experiments, this trajectory tracking method can also be used for retrospective analysis of the stored data in the animal rooms equipped with video recording devices. In addition, the trajectory tracking method provided in this application can be easily extended to use a depth camera for three-dimensional (3D) tracking, thus providing more comprehensive information about the animal's motion patterns.
[0190] The embodiments of this application also provide a trajectory tracking device. The trajectory tracking device provided in the embodiments of this application will be described below. The trajectory tracking device described below can be correspondingly referred to the trajectory tracking method described above.
[0191] Please refer to Figure 11 , which shows the structural schematic diagram of the trajectory tracking device provided in the embodiments of this application. As shown in Figure 11As shown in the figure, the trajectory tracking device may include: a target video frame sequence acquisition module 1101, a semi-air background video frame determination module 1102, a foreground video frame sequence determination module 1103, and a motion trajectory determination module 1104.
[0192] The target video frame sequence acquisition module 1101 is configured to acquire a target video frame sequence composed of consecutive video frames of a target object, where the target video frame sequence includes a plurality of sub-video frame sequences, and each video frame included in each sub-video frame sequence includes position information of a cage containing the target object.
[0193] The semi-air background video frame determination module 1102 is configured to determine the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage included in each sub-video frame sequence, determine the left semi-air background video frame and the right semi-air background video frame corresponding to each sub-video frame sequence, where the target object included in the left semi-air background video frame is in the right half area of the cage, and the target object included in the right semi-air background video frame is in the left half area of the cage.
[0194] The foreground video frame sequence determination module 1103 is configured to perform background subtraction on each sub-video frame sequence according to the left semi-air background video frame and the right semi-air background video frame corresponding to each sub-video frame sequence, to obtain the foreground video frame sequence corresponding to each sub-video frame sequence.
[0195] The motion trajectory determination module 1104 is configured to determine the motion trajectory of the target object in the target video frame sequence based on the foreground video frame sequence corresponding to each sub-video frame sequence.
[0196] In summary, the working principle of the trajectory tracking device disclosed in this embodiment is the same as that of the trajectory tracking method disclosed in the above embodiment, and will not be elaborated here.
[0197] An embodiment of the present application further provides a trajectory tracking device. Optionally, Figure 12 The hardware structure block diagram of the trajectory tracking device is shown. Referring to Figure 12 , the hardware structure of the trajectory tracking device may include: at least one processor 1201, at least one communication interface 1202, at least one memory 1203, and at least one communication bus 1204;
[0198] In the embodiment of the present application, the number of the processor 1201, the communication interface 1202, the memory 1203, and the communication bus 1204 is at least one, and the processor 1201, the communication interface 1202, and the memory 1203 complete mutual communication through the communication bus 1204;
[0199] The processor 1201 may be a central processing unit (CPU), or a specific application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0200] The memory 1203 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory;
[0201] Among them, the memory 1203 stores a program, and the processor 1201 can call the program stored in the memory 1203. The program is used for:
[0202] Obtaining a target video frame sequence composed of consecutive video frames of a target object. Among them, the target video frame sequence includes multiple sub-video frame sequences, and each video frame included in each sub-video frame sequence includes the position information of the cage containing the target object;
[0203] Determining the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determining the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence. Among them, the target object included in the left half-empty background video frame is in the right half area of the cage, and the target object included in the right half-empty background video frame is in the left half area of the cage;
[0204] Performing background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence;
[0205] Based on the foreground video frame sequence corresponding to each sub-video frame sequence, determining the motion trajectory of the target object in the target video frame sequence.
[0206] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0207] The embodiment of the present application also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned trajectory tracking method is implemented.
[0208] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0209] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0210] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other.
[0211] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A trajectory tracking method, characterized in that, Including: Obtaining a target video frame sequence composed of consecutive video frames of a target object, where the target video frame sequence includes a plurality of sub-video frame sequences, and the position information of the cage containing the target object is included in each video frame included in each sub-video frame sequence; Determining the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determining the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence, where the target object included in the left half-empty background video frame is in the right half-side area of the cage, and the target object included in the right half-empty background video frame is in the left half-side area of the cage; Performing background subtraction on each sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence to obtain the foreground video frame sequence corresponding to each sub-video frame sequence; Based on the foreground video frame sequence corresponding to each sub-video frame sequence, determining the motion trajectory of the target object in the target video frame sequence, the determining the position information of the target object in each sub-video frame sequence, and based on the determined position information and the position information of the cage in each sub-video frame sequence, determining the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence, including: For each sub-video frame sequence: Calculating the frame pixel difference between every two adjacent video frames in the sub-video frame sequence by using the frame difference method, and determining the position information of the target object in each video frame included in the sub-video frame sequence based on the calculated frame pixel differences; For each calculated frame pixel difference, if the frame pixel difference is greater than or equal to a preset high motion threshold, both of the two adjacent video frames corresponding to the frame pixel difference are used as the first target video frames to obtain all the first target video frames included in the sub-video frame sequence; Determining the left half-empty background video frame and the right half-empty background video frame corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence; To obtain the left half-empty background video frame and the right half-empty background video frame corresponding to each sub-video frame sequence; The determining the motion trajectory of the target object in the target video frame sequence based on the foreground video frame sequence corresponding to each sub-video frame sequence includes: Performing standard image processing on the foreground video frame sequence corresponding to each sub-video frame sequence to obtain the processed foreground video frame sequence corresponding to each sub-video frame sequence; For the processed foreground video frame sequence corresponding to each sub-video frame sequence, generating a bounding box of the target object in each processed foreground video frame included in the processed foreground video frame sequence to obtain the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub-video frame sequence; Determine the motion trajectory of the target object in the target video frame sequence according to the bounding boxes corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each said sub-video frame sequence.
2. The trajectory tracking method according to claim 1, characterized in that, The determining of the left half-empty background video frame and the right half-empty background video frame corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence includes: Determine the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence; Select a left half-empty background video frame from the left half-empty background video frame set corresponding to the sub-video frame sequence as the left half-empty background video frame corresponding to the sub-video frame sequence, and select a right half-empty background video frame from the right half-empty background video frame set corresponding to the sub-video frame sequence as the right half-empty background video frame corresponding to the sub-video frame sequence.
3. The trajectory tracking method according to claim 2, characterized in that, The determining of the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence according to the position information of the target object in each first target video frame included in the sub-video frame sequence and the position information of the cage in each first target video frame included in the sub-video frame sequence includes: For each first target video frame included in the sub-video frame sequence, use an object detection model to determine the bounding box information of the target object in the first target video frame, and according to the bounding box information of the target object in the first target video frame and the position information of the target object in the first target video frame, determine whether the confidence of the position information of the target object in the first target video frame is greater than a preset confidence threshold. If so, regard the first target video frame as a second target video frame; Determine the left half-empty background video frame set and the right half-empty background video frame set corresponding to the sub-video frame sequence according to the position information of the target object in each second target video frame included in the sub-video frame sequence and the position information of the cage in each second target video frame included in the sub-video frame sequence.
4. The trajectory tracking method according to claim 2, characterized in that, The selecting of a left half-empty background video frame from the left half-empty background video frame set corresponding to the sub-video frame sequence includes: Select the left half-empty background video frame with the earliest time from the left half-empty background video frame set corresponding to the sub-video frame sequence; The selecting of a right half-empty background video frame from the right half-empty background video frame set corresponding to the sub-video frame sequence includes: Select the right half-empty background video frame with the earliest time from the right half-empty background video frame set corresponding to the sub-video frame sequence.
5. The trajectory tracking method according to claim 1, characterized in that, The background subtraction of each said sub-video frame sequence according to the left half-empty background video frame and the right half-empty background video frame corresponding to each said sub-video frame sequence includes: Generate an empty background video frame corresponding to each sub - video frame sequence according to the left half - empty background video frame and the right half - empty background video frame corresponding to each sub - video frame sequence. Convert each video frame included in each sub - video frame sequence and the empty background video frame corresponding to each sub - video frame sequence into grayscale video frames respectively, to obtain a grayscale video frame sequence and an empty background grayscale video frame corresponding to each sub - video frame sequence. Perform background subtraction on the grayscale video frame sequence corresponding to each sub - video frame sequence according to the empty background grayscale video frame corresponding to each sub - video frame sequence.
6. The trajectory tracking method according to claim 1, characterized in that, The determining the motion trajectory of the target object in the target video frame sequence according to the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub - video frame sequence includes: Determine whether it is necessary to adjust the bounding box corresponding to each processed foreground video frame according to the frame pixel difference between every two adjacent video frames in each sub - video frame sequence. If so, adjust the bounding box corresponding to the processed foreground video frame that needs to be adjusted. After adjustment, use the center point of the bounding box corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub - video frame sequence as the tracking point corresponding to each processed foreground video frame. Determine the motion trajectory of the target object in the target video frame sequence according to the tracking point corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub - video frame sequence.
7. The trajectory tracking method according to claim 6, characterized in that,The determining the motion trajectory of the target object in the target video frame sequence according to the tracking point corresponding to each processed foreground video frame included in the processed foreground video frame sequence corresponding to each sub - video frame sequence includes: Form the processed foreground video frame sequence corresponding to the target video frame sequence by the processed foreground video frame sequences corresponding to each sub - video frame sequence. For each processed foreground video frame except the first processed foreground video frame included in the processed foreground video frame sequence corresponding to the target video frame sequence, calculate the Euclidean distance between the tracking point corresponding to this processed foreground video frame and the tracking point corresponding to the previous processed foreground video frame of this processed foreground video frame. If this Euclidean distance is greater than the preset distance threshold, then use the tracking point corresponding to this processed foreground video frame as a target tracking point corresponding to the target video frame sequence. Determine the motion trajectory of the target object in the target video frame sequence according to all the target tracking points corresponding to the target video frame sequence.
8. An apparatus for the trajectory tracking method according to claim 1, characterized in that, Include: A target video frame sequence acquisition module, configured to acquire a target video frame sequence composed of consecutive video frames of a target object, where the target video frame sequence includes multiple sub - video frame sequences, and the position information of the cage containing the target object is included in each video frame included in each sub - video frame sequence. The half-air background video frame determination module is configured to determine the position information of the target object in each of the sub-video frame sequences, and based on the determined position information and the position information of the cage included in each of the sub-video frame sequences, determine the left half-air background video frame and the right half-air background video frame corresponding to each of the sub-video frame sequences, wherein the target object included in the left half-air background video frame is in the right half-side area of the cage, and the target object included in the right half-air background video frame is in the left half-side area of the cage; The foreground video frame sequence determination module is configured to perform background subtraction on each of the sub-video frame sequences according to the left half-air background video frame and the right half-air background video frame corresponding to each of the sub-video frame sequences, to obtain the foreground video frame sequence corresponding to each of the sub-video frame sequences; The motion trajectory determination module is configured to determine the motion trajectory of the target object in the target video frame sequence based on the foreground video frame sequence corresponding to each of the sub-video frame sequences.
Citation Information
Patent Citations
Target tracking method and device, electronic equipment and storage medium
CN112825193A
Method for improving tracking using dynamic background compensation with centroid compensation
US20150178568A1