A Multi-Object Detection and Tracking Method, Device, Equipment and Medium Based on Video Frames
By processing video frames and comparing databases, dynamically adjusting the number of video frame inputs and tracking frame IDs, the problem of repeated reporting in gas station video frame detection is solved, and accurate compliance detection of gas refueling behavior is achieved.
Patent Information
- Application Number
- CN202311117580.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The existing gas station video frame detection methods are prone to repeated reporting of the same target, resulting in the inability to match personnel registration information and refueling information, and the inability to normally detect whether the refueling behavior is compliant.
By acquiring the video stream and performing continuous video frame processing, reorganizing the video frames into a picture sequence according to the bearing threshold, inputting the target detection network to obtain the target frame picture, and obtaining the tracking frame ID through the target tracking network. Combining with database comparison, dynamically adjusting the number of video frame inputs and tracking frame IDs, and using factors such as the target movement speed and time difference to determine the ID relationship to avoid misjudgment.
The accuracy and efficiency of refueling behavior detection is improved, detection inaccuracy caused by excessive video frames is avoided, and the time-consuming similarity comparison is reduced when the data is large, and the operation efficiency of the method is improved.
Smart Images

Figure CN117173213B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a multi-object detection and tracking method, device, equipment and medium based on video frames. Background Art
[0002] The refined oil in gas stations can be directly filled and sold to cars and motorcycles, or sold to purchasers with compliant oil drums. For example, in scenarios where some engineering vehicles need gasoline filling and some scooters also need gasoline filling, etc., a refined oil retail registration system has been basically deployed in various cities across the country. Users purchasing bulk oil must register, and only after the registration is completed can the gas station conduct sales. However, due to the limited control strength of gas stations and the deliberate illegal operations of some people, etc., it is very likely that barreled oil will be directly taken out of the gas station without registration. And the danger of barreled gasoline is very high. If it is misused by people, it will bring huge potential safety hazards to society.
[0003] During the detection process of refueling behavior, due to the frequent occlusion of cars or pedestrians in gas stations, when using traditional tracking algorithms such as SORT, DeepSort, StrongSort or BOT-Sort, the detection effect is very unsatisfactory. The problem of limited model input caused by unlimited increase in the number of video frames and then inaccurate detection is likely to cause duplicate reporting of the same target, which results in the inability to match the personnel registration information and refueling information, and leads to the inability to normally detect whether the refueling behavior is compliant. Summary of the Invention
[0004] Aiming at the problem that related tracking algorithms are likely to cause duplicate reporting of the same target, which results in the inability to match the personnel registration information and refueling information, and leads to the inability to normally detect whether the refueling behavior is compliant, the present invention provides a multi-object detection and tracking method, device, equipment and medium based on video frames.
[0005] In a first aspect, the technical solution of the present invention provides a multi-object detection and tracking method based on video frames, including the following steps:
[0006] Obtain a video stream, and perform continuous video frame processing on the obtained video stream to obtain video frames;
[0007] When the number of obtained video frames is greater than the bearing threshold, reorganize the video frames according to the bearing threshold to form a picture sequence, and input the regenerated picture sequence into a target detection network to obtain a target box picture; wherein, the bearing threshold is the maximum number of video frames that the system can parse in real time; specifically, it refers to the maximum number of frames of the video that the system can process under the current system temperature through the current detection network and target tracking network.
[0008] Copy the target box image, input it into the target tracking network to obtain the tracking box ID and the tracking box image, compare the tracking box image corresponding to the obtained tracking box ID with the tracking box image in the database, determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID;
[0009] Obtain the tracking box score, the tracking box image, and the current time corresponding to the re-determined tracking box ID and store them in the database;
[0010] Display the tracking effect on the image.
[0011] As a further limitation of the technical solution of the present invention, before the step of re-organizing the video frames according to the carrying threshold to form an image sequence when the number of obtained video frames is greater than the carrying threshold and inputting the image sequence into the target detection network to obtain the target box image, it includes:
[0012] Judge whether the carrying threshold is determined;
[0013] If it is determined, judge whether the number of obtained video frames is less than the carrying threshold;
[0014] If so, generate an image sequence from the obtained video frames according to the time sequence;
[0015] Input the image sequence into the target detection network to obtain the target box moving speed, the target box position, the target box category and the corresponding score, and obtain the target box image according to the target box position;
[0016] Execute the steps: copy the target box image, input it into the target tracking network to obtain the tracking box ID and the tracking box image, compare the tracking box image corresponding to the obtained tracking box ID with the tracking box image in the database, determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID;
[0017] If it is not determined, generate an image sequence from the obtained video frames according to the time sequence, input the generated image sequence into the target detection network, and test the maximum number of video frames that the system can parse in real time;
[0018] Execute the steps: judge whether the number of obtained video frames is less than the carrying threshold.
[0019] Before the system performs the current detection, first obtain the temperature of the current system, and then obtain the carrying threshold of the system at the current temperature from the database. When there is no such carrying threshold in the database, the carrying threshold is first determined, and the carrying threshold at this temperature is written into the database for direct call next time.
[0020] As a further limitation of the technical solution of the present invention, the steps of copying the target box picture, inputting it into the target tracking network to obtain the tracking box ID and the tracking box picture, comparing the tracking box picture corresponding to the obtained tracking box ID with the tracking box pictures in the database, judging the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re - determining the tracking box ID include:
[0021] Judge whether there is a tracking box picture corresponding to the target box picture in the database;
[0022] If so, copy the tracking box picture and the corresponding tracking box ID in the database;
[0023] Execute the step: display the tracking effect on the picture;
[0024] If not, copy the target box picture and input it into the target tracking network. Obtain the tracking box position, the tracking box score, and the tracking box ID through the target tracking network, and obtain the tracking box picture according to the tracking box position;
[0025] Compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re - determine the tracking box ID.
[0026] As a further limitation of the technical solution of the present invention, the steps of comparing the tracking box picture corresponding to the obtained tracking box ID with the tracking box pictures in the database, judging the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re - determining the tracking box ID include:
[0027] Perform a database comparison on the obtained tracking box ID to judge whether it is a newly added tracking box ID;
[0028] If it is not a newly added tracking box ID, judge whether the tracking box score corresponding to the obtained tracking box ID is greater than the tracking box score threshold;
[0029] If so, obtain the tracking box ID, the corresponding tracking box position, the tracking box score, the tracking box picture, and the current time and store them in the database;
[0030] If not, confirm that the tracking box ID is a newly added tracking box ID, compare the tracking picture corresponding to the newly added tracking box ID with the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re - determine the tracking box ID;
[0031] If it is a newly added tracking box ID, execute the step: compare the tracking picture corresponding to the newly added tracking box ID with the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re - determine the tracking box ID.
[0032] As a further limitation of the technical solution of the present invention, the tracking picture corresponding to the newly added tracking box ID is compared with the tracking box pictures in the database, and the relationship between the tracking box ID and the original tracking box ID is judged according to the comparison result. Redetermining the tracking box ID includes:
[0033] Compare the aspect ratios of the tracking pictures corresponding to the newly added tracking box ID with the tracking box pictures in the database one by one to see if they meet the ratio threshold;
[0034] If so, calculate the similarity between the tracking picture corresponding to the newly added tracking box ID and the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and redetermine the tracking box ID;
[0035] If not, determine that the newly added tracking box ID is the new tracking box ID, and organize the tracking box position, tracking box score, tracking box picture corresponding to the new tracking box ID and the current time to be written into the database.
[0036] As a further limitation of the technical solution of the present invention, the steps of calculating the similarity between the tracking picture corresponding to the newly added tracking box ID and the tracking box pictures in the database, and judging the relationship between the tracking box ID and the original tracking box ID according to the comparison result to redetermine the tracking box ID include:
[0037] Calculate the similarity between the tracking box pictures corresponding to the newly added tracking box ID and the tracking box pictures in the database one by one;
[0038] Judge whether the similarity meets the similarity threshold;
[0039] If the similarity threshold is met, judge whether the time difference and speed relationship between the tracking box pictures corresponding to the two tracking box IDs that meet the similarity threshold meet the relationship threshold;
[0040] If so, change the newly added tracking box ID to the original tracking box ID in the database and update the information of the original tracking box ID in the database to the database;
[0041] If not, execute the steps: determine that the newly added tracking box ID is the new tracking box ID, and organize the tracking box position, tracking box score, tracking box picture corresponding to the new tracking box ID and the current time to be written into the database;
[0042] If the similarity threshold is not met, execute the steps: determine that the newly added tracking box ID is the new tracking box ID, and organize the tracking box position, tracking box score, tracking box picture corresponding to the new tracking box ID and the current time to be written into the database.
[0043] As a further limitation of the technical solution of the present invention, the number of video frames is determined by the target movement speed, and the specific relationship is as follows:
[0044]
[0045] Among them, F is the number of video frames selected per second, max_frequent is the maximum number of video frames that the system can parse per second, obj_v is the maximum speed at which the detection target moves, obj_s is the minimum size of the detection target, and obj_num is the number of detection targets.
[0046] As a further limitation of the technical solution of the present invention, the step of inputting the picture sequence into the target detection network to obtain the target box picture includes:
[0047] Input the picture sequence into the target detection network to obtain the target box movement speed, target box position, target box category and corresponding scores;
[0048] According to the coordinates of the target box position, calculate the target box size and aspect ratio;
[0049] Select the last picture of the picture sequence and crop it according to the target box size to obtain the target box picture.
[0050] In a second aspect, the technical solution of the present invention also provides a multi-target detection and tracking device based on video frames, including a target detection module and a target tracking module;
[0051] The target detection module is used to obtain a video stream, perform continuous video frame processing on the obtained video stream to obtain video frames; when the number of obtained video frames is greater than the carrying threshold, reorganize the video frames according to the carrying threshold to form a picture sequence, and input the picture sequence into the target detection network to obtain the target box picture; among them, the carrying threshold is the maximum number of video frames that the system can parse in real time;
[0052] The target tracking module is used to obtain the tracking box ID and the tracking box picture according to the target box picture, compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box picture in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; obtain the tracking box score, the tracking box picture corresponding to the re-determined tracking box ID, and the current time and store them in the database; display the tracking effect on the picture. [[ID=*28]]
[0053] As a further limitation of the technical solution of the present invention, the target detection module is further configured to determine whether the bearing threshold is determined; if it is determined, determine whether the number of video frames obtained is less than the bearing threshold; if so, generate a picture sequence from the obtained video frames according to the time sequence; input the picture sequence into the target detection network to obtain the target box moving speed, the target box position, the target box category and the corresponding score, and obtain the target box picture according to the target box position; if not, update the bearing threshold to the target detection network; reorganize the video frames according to the bearing threshold to generate a picture sequence; if it is not determined, generate a picture sequence from the obtained video frames according to the time sequence, input the generated picture sequence into the target detection network, and test the maximum number of video frames that the system can parse in real time.
[0054] As a further limitation of the technical solution of the present invention, the tracking module is specifically configured to determine whether there is a tracking box picture corresponding to the target box picture in the database; if so, copy the tracking box picture and the corresponding tracking box ID in the database; display the tracking effect on the picture; if not, obtain the tracking box position, the tracking box score and the tracking box ID according to the target box picture, and obtain the tracking box picture according to the tracking box position; compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box picture in the database, and determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID.
[0055] As a further limitation of the technical solution of the present invention, the tracking module is further configured to compare the obtained tracking box ID with the database to determine whether it is a newly added tracking box ID; if it is not a newly added tracking box ID, determine whether the tracking box score corresponding to the obtained tracking box ID is greater than the tracking box score threshold; if so, obtain the tracking box ID, the corresponding tracking box position, the tracking box score, the tracking box picture and the current time and store them in the database; if not, confirm that the tracking box ID is a newly added tracking box ID, compare the tracking picture corresponding to the newly added tracking box ID with the tracking box picture in the database, and determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; if it is a newly added tracking box ID, execute the steps: compare the tracking picture corresponding to the newly added tracking box ID with the tracking box picture in the database, and determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID.
[0056] As a further limitation of the technical solution of the present invention, the tracking module is further configured to compare the aspect ratio of the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database one by one to determine whether it meets the ratio threshold; if so, calculate the similarity between the tracking image corresponding to the newly added tracking box ID and the tracking box images in the database, determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; if not, determine the newly added tracking box ID as the new tracking box ID, and organize the tracking box position, tracking box score, tracking box image corresponding to the new tracking box ID, and the current time to be written into the database.
[0057] As a further limitation of the technical solution of the present invention, the tracking module is further configured to calculate the similarity between the tracking box images corresponding to the newly added tracking box ID and the tracking box images in the database one by one; determine whether the similarity meets the similarity threshold; if the similarity threshold is met, determine whether the time difference and speed relationship between the tracking box images corresponding to the two tracking box IDs that meet the similarity threshold meet the relationship threshold; if so, change the newly added tracking box ID to the original tracking box ID in the database and update the information of the original tracking box ID in the database to the database.
[0058] In a third aspect, the technical solution of the present invention further provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the multi-object detection and tracking method based on video frames as described in the first aspect.
[0059] In a fourth aspect, the technical solution of the present invention further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the multi-object detection and tracking method based on video frames as described in the first aspect.
[0060] From the above technical solutions, it can be seen that the present invention has the following advantages:
[0061] 1. First, measure the number of video frames that the system hardware can parse in real time to avoid the inaccurate speed detection of the target box caused by the unlimited increase in the number of video frames.
[0062] 2. Analyze consecutive video frames, and obtain the moving speed of the target box through the time interval between input video frames and the target detection position in the video frames.
[0063] 3. Use the maximum number of video frames that can be parsed by the target detection speed and hardware as a two-way limit to dynamically change the number of video frames input to the detection network. For example, if the target moves slowly, the minimum number of frames can be used; if the target moves relatively fast, it can be set to the maximum number of video frames that can be parsed by the hardware.
[0064] 4. The target tracking network adopted first performs feature extraction on the picture data through a CNN network to limit the differences in the dimensions of different input parameters, and then fuses it with other data to avoid the large differences in the data of different input data sources, resulting in a reduced impact of small input positions on the output.
[0065] 5. The database adopts a dynamic update method, setting the maximum number of stored IDs, the data information stored for each ID, and the latest data information for each ID, avoiding the serious time-consuming of similarity comparison due to too much data volume and improving the efficiency of this method.
[0066] 6. When comparing different IDs, first compare the aspect ratios of the tracking boxes to avoid directly entering the similarity detection network and improve the operation efficiency of this method.
[0067] 7. Use a comprehensive determination of the time difference and speed to determine the relationship between this ID and the original ID, avoiding misjudgment between IDs.
[0068] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect.
[0069] Therefore, compared with the prior art, the present invention has prominent substantive features and significant progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0071] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0072] Figure 2 It is a schematic flowchart of the method according to another embodiment of the present invention.
[0073] Figure 3 It is a schematic diagram of the target detection process in the embodiment of the present invention.
[0074] Figure 4 It is a schematic diagram of the target tracking process in the embodiment of the present invention.
[0075] Figure 5 It is a schematic diagram of the similarity calculation process in the embodiment of the present invention.
[0076] Figure 6 It is a schematic block diagram of a device according to an embodiment of the present invention. Detailed implementation manners
[0077] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0078] As Figure 1 shown, the embodiment of the present invention provides a multi-object detection and tracking method based on video frames. The method is implemented through an object detection network and an object tracking network in the system, and includes the following steps:
[0079] Step 1: Obtain a video stream, perform continuous video frame processing on the obtained video stream, and obtain video frames;
[0080] Perform image preprocessing, scale the image size to a specific size; the specific size is set to 640*640. During the process of scaling the image size, scale it without changing the aspect ratio of the image, scale the maximum side to 640, and fill the smaller side with grayscale; perform continuous video frame processing, and form a sequence of images from the parsed images according to the time sequence;
[0081] Step 2: When the number of obtained video frames is greater than the bearing threshold, reorganize the video frames into a sequence of images according to the bearing threshold, and input the sequence of images into the object detection network to obtain object box images; wherein, the bearing threshold is the maximum number of video frames that the system can parse in real time;
[0082] Step 3: Copy the object box images and input them into the object tracking network to obtain tracking box IDs and tracking box images, compare the tracking box images corresponding to the obtained tracking box IDs with the tracking box images in the database, judge the relationship between the tracking box IDs and the original tracking box IDs according to the comparison results, and re-determine the tracking box IDs;
[0083] Step 4: Obtain the tracking box scores, tracking box images corresponding to the re-determined tracking box IDs, and the current time and store them in the database;
[0084] Step 5: Display the tracking effect on the image.
[0085] It should be noted that in the embodiments of the present invention, the minimum number of video frames should be greater than 2 frames, and the maximum should be less than the maximum number of video frames that the system can parse; that is, the bearing threshold.
[0086] The number of video frames is further determined by the target moving speed, and the specific relationship is as follows:
[0087]
[0088] Where F is the number of video frames selected per second, max_frequent: the maximum number of video frames that the system can parse per second, obj_v: the maximum speed of the detected target moving, obj_s: the minimum size of the detected target, obj_num: the number of detected targets. It should be noted that in the embodiments of the present invention, the data structure stored in the database includes time, target box pictures (that is, detection pictures), tracking box positions, tracking box IDs, tracking box speeds, and tracking box scores; that is, F needs to meet the above 5 conditions, F is greater than or equal to 3, F is less than or equal to max_frequent, and F is proportional to obj_v, obj_s, and obj_num.
[0089] As Figure 2 shown, the embodiments of the present invention also provide a multi-target detection and tracking method based on video frames. The method is implemented through a target detection network and a target tracking network in the system, and includes the following steps:
[0090] S1: Obtain a video stream, and perform continuous video frame processing on the obtained video stream to obtain video frames;
[0091] S2: Determine whether the bearing threshold is determined;
[0092] If so, execute step S3; if not, execute step S8;
[0093] S3: Determine whether the number of obtained video frames is less than the bearing threshold;
[0094] If so, execute step S4; if not, execute step S6; [[ID=:30]]
[0095] S4: Generate a picture sequence from the obtained video frames according to the time sequence;
[0096] S5: Input the picture sequence into the target detection network to obtain the target box moving speed, target box position, target box category, and corresponding scores, and obtain the target box picture according to the target box position; execute step S9;
[0097] Based on the coordinates of the target box position, calculate the size and aspect ratio of the target box; based on the detected target box position, i.e., the input image sequence, obtain the target box image of the detected object; use the last image in the image sequence as the template image, and crop the target box image from this image according to the size of the target box.
[0098] S6: Update the bearing threshold to the target detection network.
[0099] S7: Reorganize the video frames according to the bearing threshold to generate an image sequence; execute step S5.
[0100] S8: Generate an image sequence from the acquired video frames according to the time sequence, input the generated image sequence into the target detection network, and test the maximum number of video frames that the system can parse in real time; execute step S3.
[0101] S9: Determine whether there is a tracking box image corresponding to the target box image in the database.
[0102] If yes, execute step S10; if no, execute step S11.
[0103] S10: Copy the tracking box image and the corresponding tracking box ID in the database; execute step S21.
[0104] In fact, in this step: extract the target tracking image, target tracking box position, and target tracking box ID data from the database.
[0105] S11: Copy the target box image and input it into the target tracking network. Obtain the tracking box position, tracking box score, and tracking box ID through the target tracking network. Obtain the tracking box image according to the tracking box position; copy the target box image as the tracking box image, and assign values to the tracking box ID sequentially from 0 according to the number of targets; input the template image, target detection box image, score, target box movement speed, target box category, target box area, target box aspect ratio, and target tracking box image and target tracking box ID into the target tracking network; through the target tracking network, obtain the tracking box position, tracking box ID, and tracking box score, and obtain the tracking box image according to the target box position and the template image.
[0106] S12: Compare the obtained tracking box ID with the database to determine whether it is a newly added tracking box ID.
[0107] If no, execute step S13; if yes, execute step S16.
[0108] S13: Determine whether the tracking box score corresponding to the obtained tracking box ID is greater than the tracking box score threshold.
[0109] If yes, execute step S14; if no, execute step S15.
[0110] S14: Obtain the tracking box ID and the corresponding tracking box position, tracking box score, tracking box image, and the current time, and store them in the database; execute step S22;
[0111] S15: Confirm that the tracking box ID is a newly added tracking box ID;
[0112] S16: Compare the aspect ratio of the tracking image corresponding to the newly added tracking box ID with the aspect ratios of the tracking box images in the database one by one to check if it meets the ratio threshold;
[0113] If yes, execute step S18; if no, execute step 17;
[0114] S17: Determine that the newly added tracking box ID is a new tracking box ID, organize the tracking box position, tracking box score, tracking box image, and the current time corresponding to the new tracking box ID, and write them into the database; execute step S22;
[0115] S18: Calculate the similarity between the tracking box image corresponding to the newly added tracking box ID and the tracking box images in the database one by one;
[0116] S19: Determine whether the similarity meets the similarity threshold;
[0117] If yes, execute step S20; if no, execute step S17;
[0118] S20: Determine whether the time difference and speed relationship between the tracking box images corresponding to the two tracking box IDs that meet the similarity threshold meet the relationship threshold;
[0119] It should be noted that the time difference and speed relationship are comprehensively determined based on the position, time, and speed of the tracking box in the tracking image. The specific determination relationship is as follows:
[0120] <0>
[0121] track_v = track_v ID0 +track_v ID1
[0122] track_time = time ID0 +time ID1
[0123] Among them, α and β are specifically defined parameters;
[0124] track_position ID0 : The position of ID0;
[0125] track_position ID1: Position of ID1;
[0126] dis(track_position ID0 -track_position ID1 ): Distance between ID0 and ID1;
[0127] track_v ID0 : Speed of ID0;
[0128] track_v ID1 : Speed of ID1;
[0129] time ID0 : Time of ID0;
[0130] time ID1 : Time of ID1;
[0131] track_v: Sum of the speeds of ID1 and ID2;
[0132] track_time: Time difference between ID1 and ID2.
[0133] If so, execute step S21; if not, execute step S17;
[0134] S21: Change the ID of the newly added tracking box to the original tracking box ID in the database and update the information of the original tracking box ID in the database to the database; the database only stores a certain amount of IDs, and each ID only saves a certain amount of picture information and the ID information of the latest time corresponding to the ID;
[0135] S22: Display the tracking effect on the picture.
[0136] The process of the target detection network obtaining the target box picture is as Figure 3 shown:
[0137] STEP1: Input a set number of consecutive video frames;
[0138] STEP2: Arrange the video frames in time sequence, starting from the first frame and successively taking a certain number of subsequent video frames to form a picture sequence;
[0139] The video frames are set in groups of 3, that is, 0 / 1 / 2 as a group, 1 / 2 / 3 as a group, and so on;
[0140] STEP3: Enter the first layer of LSTM and Dropout networks;
[0141] STEP4: Enter the second layer of LSTM and Dropout networks;
[0142] STEP5: Enter the feature extraction CNN network;
[0143] STEP6: Feed the extracted feature map into the RPN network to obtain the target recommendation regions of the picture;
[0144] STEP7: Feed the obtained feature map and the target recommendation regions into the ROI Align network simultaneously to obtain the feature map of the required size;
[0145] STEP7.1: Divide the bbox region equally according to the size required for output. It is very likely that the vertices after equal division do not fall on real pixel points;
[0146] STEP7.2: Take another fixed 4 points in each block, and weight (bilinear interpolation) the values of the 4 real pixel points closest to it to obtain the value of this blue point;
[0147] STEP7.3: 4 new values will be calculated within one block. Take the max among these new values as the output value of this block, and finally a 2x2 output can be obtained;
[0148] STEP8: The feature map of the required size passes through the head layer to obtain the speed detection value and the target detection region; the head layer branches the network into two branches, one of which is used for speed detection operations, and the other is used for target detection operations;
[0149] STEP9: The speed detection passes through two 1*1 CNN layers to obtain the segmentation result; that is, the moving speed of the target box;
[0150] STEP10: The target detection network passes through the CNN layer to obtain the target box position, target box category, and score respectively;
[0151] STEP11: The program ends;
[0152] The tracking process of the target tracking network is as Figure 4 shown:
[0153] S-1: According to the target detection network, obtain the moving speed of the target box, target box position, target box category, and score, and further calculate to obtain the target box area and target box aspect ratio;
[0154] S-2: Obtain the previous target box picture and the previous target detection box ID from the database;
[0155] S-3: The pictures corresponding to the target box, the target box picture, and the previous tracking box picture all pass through the CNN network for feature extraction;
[0156] S-4: After CNN feature extraction, it is input into the CNN network together with the moving speed of the target box, the target box category, the target box area, the aspect ratio of the target box, the score, the ID of the previous tracking box, the area of the previous tracking box, and the aspect ratio of the previous tracking box;
[0157] S-4.1: Perform weight limitation on the input parameters;
[0158] S-5: After passing through the CNN network, the position of the tracking box, the ID of the tracking box, the speed of the tracking box, and the score of the tracking box are obtained;
[0159] S-6: The program ends;
[0160] The similarity calculation process is as Figure 5 shown:
[0161] SS1: Receive images corresponding to two different comparison IDs, the image corresponding to the tracking box ID1 is image I, and the image corresponding to the tracking box ID2 is image II;
[0162] SS2: Scale image I and image II to the same size;
[0163] SS2.1: For the model with a smaller width, boundary pixel filling is used to make the effective area located in the middle, and the widths of the two pictures are the same;
[0164] SS3: Image I enters the RCNN neural network I to associate the information of each picture;
[0165] Preferably, the RCNN neural network I uses an LSTM neural network;
[0166] SS4: Image I enters the Resnet50 neural network I for feature extraction;
[0167] SS5: Image I enters the 1*1 CNN neural network I for picture granularity classification;
[0168] SS6: The neural network processing flow of image II is the same as that of image I, and the neural network parameters of image II reuse the neural network parameters of image I;
[0169] SS7: Obtain the similarity between image I and image II through MSE;
[0170] SS8: The program ends.
[0171] As Figure 6 shown, the embodiment of the present invention also provides a multi-object detection and tracking device based on video frames, including a target detection module and a target tracking module;
[0172] A target detection module, which is used to obtain a video stream, perform continuous video frame processing on the obtained video stream to obtain video frames; when the number of obtained video frames is greater than the carrying threshold, reorganize the video frames according to the carrying threshold to form a picture sequence, and input the picture sequence into a target detection network to obtain a target box picture; wherein, the carrying threshold is the maximum number of video frames that the system can parse in real time.
[0173] A target tracking module, which is used to obtain a tracking box ID and a tracking box picture according to the target box picture, compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box picture in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; obtain the tracking box score, the tracking box picture and the current time corresponding to the re-determined tracking box ID and store them in the database; display the tracking effect on the picture.
[0174] It should be noted that the target detection module is also used to judge whether the carrying threshold is determined; if it is determined, judge whether the number of obtained video frames is less than the carrying threshold; if so, generate a picture sequence from the obtained video frames according to the time sequence; input the picture sequence into the target detection network to obtain the target box moving speed, the target box position, the target box category and the corresponding score, and obtain the target box picture according to the target box position; if not, update the carrying threshold to the target detection network; reorganize the video frames according to the carrying threshold to generate a picture sequence; if it is not determined, generate a picture sequence from the obtained video frames according to the time sequence, input the generated picture sequence into the target detection network, and test the maximum number of video frames that the system can parse in real time.
[0175] The tracking module is specifically used to determine whether there is a tracking box image corresponding to the target box image in the database; if so, copy the tracking box image and the corresponding tracking box ID in the database; display the tracking effect on the image; if not, obtain the tracking box position, tracking box score, and tracking box ID according to the target box image, and obtain the tracking box image according to the tracking box position; compare the tracking box image corresponding to the obtained tracking box ID with the tracking box images in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; specifically used to compare the obtained tracking box ID with the database to determine whether it is a newly added tracking box ID; if it is not a newly added tracking box ID, judge whether the tracking box score corresponding to the obtained tracking box ID is greater than the tracking box score threshold; if so, obtain the tracking box ID, the corresponding tracking box position, tracking box score, tracking box image, and the current time and store them in the database; if not, confirm that the tracking box ID is a newly added tracking box ID, compare the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; if it is a newly added tracking box ID, execute the steps: compare the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID. Specifically used to compare the aspect ratios of the tracking images corresponding to the newly added tracking box ID with the tracking box images in the database one by one to see if they meet the ratio threshold; if so, calculate the similarity between the tracking image corresponding to the newly added tracking box ID and the tracking box images in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; if not, determine that the newly added tracking box ID is a new tracking box ID, and organize the tracking box position, tracking box score, tracking box image, and the current time corresponding to the new tracking box ID and write them into the database. Specifically used to calculate the similarity between the tracking box images corresponding to the newly added tracking box ID and the tracking box images in the database one by one; judge whether the similarity meets the similarity threshold; if the similarity threshold is met, judge whether the time difference and speed relationship between the tracking box images corresponding to the two tracking box IDs that meet the similarity threshold meet the relationship threshold; if so, change the newly added tracking box ID to the original tracking box ID in the database and update the information of the original tracking box ID in the database to the database.
[0176] The number of video frames is determined by the target movement speed, and the specific relationship is as follows:
[0177]
[0178] Among them, F is the number of video frames selected per second, max_frequent is the maximum number of video frames that the system can parse per second, obj_v is the maximum speed at which the detection target moves, obj_s is the minimum size of the detection target, and obj_num is the number of detection targets.
[0179] It should be noted that the picture sequence is input into the object detection network to obtain the moving speed of the object box, the position of the object box, the category of the object box, and the corresponding score; according to the coordinates of the position of the object box, the size and aspect ratio of the object box are calculated; the last picture of the picture sequence is selected and cropped according to the size of the object box to obtain the object box picture.
[0180] The embodiment of the present invention also provides an electronic device, which includes: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following method: Step 1: Obtain a video stream, and perform continuous video frame processing on the obtained video stream to obtain video frames; Step 2: When the number of obtained video frames is greater than the carrying threshold, reorganize the video frames according to the carrying threshold to form a picture sequence, and input the picture sequence into the object detection network to obtain an object box picture; where the carrying threshold is the maximum number of video frames that the system can parse in real time; Step 3: Copy the object box picture and input it into the object tracking network to obtain the tracking box ID and the tracking box picture, and compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box picture in the database, and judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; Step 4: Obtain the tracking box score, the tracking box picture corresponding to the re-determined tracking box ID, and the current time and store them in the database; Step 5: Display the tracking effect on the picture.
[0181] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0182] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method provided in the above method embodiment. For example, it includes: Step 1: Obtain a video stream, and perform continuous video frame processing on the obtained video stream to obtain video frames; Step 2: When the number of obtained video frames is greater than the carrying threshold, reorganize the video frames according to the carrying threshold to form a picture sequence, and input the picture sequence into the target detection network to obtain a target box picture; wherein, the carrying threshold is the maximum number of video frames that the system can parse in real time; Step 3: Copy the target box picture and input it into the target tracking network to obtain a tracking box ID and a tracking box picture, and compare the tracking box picture corresponding to the obtained tracking box ID with the tracking box picture in the database, and determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; Step 4: Obtain the tracking box score, the tracking box picture corresponding to the re-determined tracking box ID, and the current time and store them in the database; Step 5: Display the tracking effect on the picture.
[0183] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and all should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi-object detection and tracking method based on video frames, characterized in that, It includes the following steps: Obtain a video stream, perform continuous video frame processing on the obtained video stream, and obtain video frames; When the number of obtained video frames is greater than the carrying threshold, reorganize the video frames according to the carrying threshold to form a picture sequence, and input the newly generated picture sequence into the target detection network to obtain target box pictures; wherein, the carrying threshold is the maximum number of video frames that the system can parse in real time; Copy the target box pictures and input them into the target tracking network to obtain the tracking box ID and tracking box pictures, compare the tracking box pictures corresponding to the obtained tracking box ID with the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; Obtain the tracking box score, tracking box pictures corresponding to the re-determined tracking box ID, and the current time and store them in the database; Display the tracking effect on the picture.
2. The multi-object detection and tracking method based on video frames according to claim 1, wherein Before the step of reorganizing the video frames according to the carrying threshold to form a picture sequence when the number of obtained video frames is greater than the carrying threshold and inputting the picture sequence into the target detection network to obtain target box pictures, it includes: Judge whether the carrying threshold is determined; If it is determined, judge whether the number of obtained video frames is less than the carrying threshold; If so, generate a picture sequence from the obtained video frames according to the time sequence; Input the picture sequence into the target detection network to obtain the target box moving speed, target box position, target box category and corresponding scores, and obtain the target box pictures according to the target box position; Execute the steps: copy the target box pictures and input them into the target tracking network to obtain the tracking box ID and tracking box pictures, compare the tracking box pictures corresponding to the obtained tracking box ID with the tracking box pictures in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; If it is not determined, generate a picture sequence from the obtained video frames according to the time sequence, input the generated picture sequence into the target detection network, and test the maximum number of video frames that the system can parse in real time; Execute the steps: judge whether the number of obtained video frames is less than the carrying threshold; Among them, the number of video frames is determined by the target moving speed, and the specific relationship is as follows: F is the number of video frames selected per second, max_frequent is the maximum number of video frames that the system can parse per second, obj_v is the maximum speed of the detected target moving, obj_s is the minimum size of the detected target, and obj_num is the number of detected targets.
3. The multi-object detection and tracking method based on video frames according to claim 1 or 2, characterized in that, The steps of copying the target box pictures and inputting them into the target tracking network to obtain the tracking box ID and tracking box pictures, comparing the tracking box pictures corresponding to the obtained tracking box ID with the tracking box pictures in the database, judging the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determining the tracking box ID include: Judge whether there are tracking box pictures corresponding to the target box pictures in the database; If so, copy the tracking box pictures and the corresponding tracking box ID in the database; Execute the steps: display the tracking effect on the picture; If not, copy the target box image and input it into the target tracking network. Obtain the tracking box position, tracking box score, and tracking box ID through the target tracking network, and obtain the tracking box image according to the tracking box position. Compare the tracking box image corresponding to the obtained tracking box ID with the tracking box images in the database. Determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID.
4. The multi-object detection and tracking method based on video frames according to claim 3, characterized in that, The steps of comparing the tracking box image corresponding to the obtained tracking box ID with the tracking box images in the database, determining the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determining the tracking box ID include: Perform a database comparison on the obtained tracking box ID to determine whether it is a newly added tracking box ID. If it is not a newly added tracking box ID, determine whether the tracking box score corresponding to the obtained tracking box ID is greater than the tracking box score threshold. If so, obtain the tracking box ID, the corresponding tracking box position, tracking box score, tracking box image, and the current time, and store them in the database. If not, confirm that the tracking box ID is a newly added tracking box ID. Compare the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database. Determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID. If it is a newly added tracking box ID, execute the steps: Compare the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database. Determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID.
5. The multi-object detection and tracking method based on video frames according to claim 4, characterized in that, Comparing the tracking image corresponding to the newly added tracking box ID with the tracking box images in the database, determining the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determining the tracking box ID include: Compare the aspect ratios of the tracking images corresponding to the newly added tracking box ID and the tracking box images in the database one by one to determine whether they meet the ratio threshold. If so, calculate the similarity between the tracking image corresponding to the newly added tracking box ID and the tracking box images in the database. Determine the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID. If not, determine that the newly added tracking box ID is a new tracking box ID, and organize the tracking box position, tracking box score, tracking box image, and the current time corresponding to the new tracking box ID and write them into the database.
6. The multi-object detection and tracking method based on video frames according to claim 5, characterized in that The steps of calculating the similarity between the tracking image corresponding to the newly added tracking box ID and the tracking box images in the database, determining the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determining the tracking box ID include: Calculate the similarity between the tracking box image corresponding to the newly added tracking box ID and the tracking box images in the database one by one. Determine whether the similarity meets the similarity threshold. If the similarity threshold is met, determine whether the time difference and speed relationship between the tracking box images corresponding to the two tracking box IDs that meet the similarity threshold meet the relationship threshold. If so, change the newly added tracking box ID to the original tracking box ID in the database and update the information of the original tracking box ID in the database to the database. Otherwise, execute the steps: Determine that the ID of the newly added tracking box is the new tracking box ID, and organize the position, score, image of the tracking box corresponding to the new tracking box ID, and the current time to be written into the database; If the similarity threshold is not met, execute the steps: Determine that the ID of the newly added tracking box is the new tracking box ID, and organize the position, score, image of the tracking box corresponding to the new tracking box ID, and the current time to be written into the database.
7. The multi-object detection and tracking method based on video frames according to claim 1, wherein The steps of inputting the image sequence into the object detection network to obtain the object box image include: Inputting the image sequence into the object detection network to obtain the moving speed, position, category, and corresponding score of the object box; Calculating the size and aspect ratio of the object box according to the coordinates of the object box position; Selecting the last image of the image sequence and cropping it according to the size of the object box to obtain the object box image.
8. A multi-object detection and tracking device based on video frames, characterized in that, It includes an object detection module and an object tracking module; The object detection module is used to obtain the video stream, perform continuous video frame processing on the obtained video stream to obtain video frames; when the number of obtained video frames is greater than the bearing threshold, reorganize the video frames according to the bearing threshold to form an image sequence, and input the image sequence into the object detection network to obtain the object box image; where the bearing threshold is the maximum number of video frames that the system can parse in real time; The object tracking module is used to obtain the tracking box ID and the tracking box image according to the object box image, compare the tracking box image corresponding to the obtained tracking box ID with the tracking box image in the database, judge the relationship between the tracking box ID and the original tracking box ID according to the comparison result, and re-determine the tracking box ID; obtain the tracking box score, tracking box image corresponding to the re-determined tracking box ID, and the current time and store them in the database; display the tracking effect on the image.
9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the video frame-based multi-object detection and tracking method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the video frame-based multi-object detection and tracking method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Tracking method and device based on target detection
CN110111363A
Human body behavior recognition method and system based on multi-target tracking
CN110399808A