Target object status detection method
By analyzing the surveillance video frames through the target detection model and tracking model, the real-time and accuracy issues of personnel detention detection are solved, and automated detention detection is realized.
Patent Information
- Application Number
- CN202210229660.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-03-09
AI Technical Summary
The real-time and accuracy of the existing technology for detecting stranded personnel cannot be guaranteed, resulting in the possibility that stranded personnel may not be discovered in time.
The surveillance video frames are analyzed through the target detection model and tracking model to obtain the identity information of the target object. The presence of the target object in the previous and next video frames is determined based on the identity information, and the retention time is updated to determine whether it is in a retention state.
It realizes the automatic detection of stranded personnel, improves the real-time and accuracy of detection, and eliminates the need for human intervention.
Smart Images

Figure CN114821457B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular relates to a method for detecting the state of a target object. Background Art
[0002] As society pays more and more attention to humanitarian care and people's life safety, the demand for detecting whether people are stranded is also increasing. For example, in places such as stations, airports, and shopping malls, if the elderly, children, and disabled people are stranded in areas with less traffic and are not discovered and cannot get help, they are likely to be in danger.
[0003] Currently, video surveillance systems can be used to determine whether people are stranded. Existing surveillance equipment is mainly used to record and store video images of corresponding scenes. The actual determination of whether people are stranded is completed by staff.
[0004] To ensure the real-time detection of detained personnel, staff must watch the surveillance video images at all times. However, long-term monitoring work will cause staff to lose concentration or produce visual fatigue, which may easily lead to detained personnel not being discovered in time, and the real-time and accuracy of detained personnel detection cannot be guaranteed. Summary of the Invention
[0005] The embodiment of the present application provides a method for detecting the status of a target object, which can at least solve the problem in the prior art that the real-time and accuracy of personnel detention detection cannot be guaranteed.
[0006] The present invention provides a method for detecting the state of a target object, the method comprising:
[0007] Get the video frame to be detected;
[0008] Inputting the video frame to be detected into the target detection model, detecting the target object in the video frame to be detected by the target detection model, and obtaining a first video frame marked with a detection frame;
[0009] Matching the target object in the first video frame with at least one first object to determine identity information of the target object in the first video frame;
[0010] determining whether the target object exists in a second video frame according to the identity information, where the second video frame is a frame preceding the video frame to be detected;
[0011] When the target object exists in the second video frame, updating the residence time of the target object;
[0012] When the stay time is greater than the first threshold, it is determined that the target object is in a stay state.
[0013] The state detection method of the target object of the embodiment of the present application can obtain a video frame to be detected, input the video frame to be detected into a target detection model, detect the target object in the video frame to be detected by the target detection model, obtain a first video frame marked with a detection frame, and then input the first video frame into a tracking model, determine the identity information of the target object in the first video frame by the tracking model, and then determine whether the target object appears in a second video frame based on the identity information. The second video frame is the previous frame of the video frame to be detected. If it appears in the second video frame, the residence time of the target object is updated, and when the residence time is greater than the first threshold, it is determined that the target object is in a detained state. In this way, automatic detection of personnel detention can be achieved without human intervention, thereby improving the real-time and accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 This is a flow chart of a method for detecting the state of a target object provided by one embodiment of the present application;
[0016] Figure 2a This is a schematic diagram of a process application scenario of a target object status detection method provided by an embodiment of the present application;
[0017] Figure 2b This is a schematic diagram of a process application scenario of another target object status detection method provided by an embodiment of the present application;
[0018] Figure 2c This is a schematic diagram of a process application scenario of another target object status detection method provided by an embodiment of the present application;
[0019] Figure 3a This is a schematic diagram of a process application scenario of another target object status detection method provided by an embodiment of the present application;
[0020] Figure 3b This is a schematic diagram of a process application scenario of another target object status detection method provided by an embodiment of the present application;
[0021] Figure 3c This is a schematic diagram of a process application scenario of another target object status detection method provided by an embodiment of the present application;
[0022] Figure 4 This is a schematic structural diagram of a target object status detection device provided by one embodiment of the present application;
[0023] Figure 5 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0024] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0025] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0026] Figure 1 FIG. 1 is a flow chart showing a method for detecting the state of a target object provided by an embodiment of the present application. It should be noted that the method for detecting the state of a target object can be applied to a device for detecting the state of a target object, such as Figure 1 As shown, the target object status detection method may include the following steps:
[0027] S110, obtaining a video frame to be detected;
[0028] S120, inputting the video frame to be detected into the target detection model, detecting the target object in the video frame to be detected by the target detection model, and obtaining a first video frame marked with a detection frame;
[0029] S130, matching the target object in the first video frame with at least one first object to determine identity information of the target object in the first video frame;
[0030] S140, determining whether the target object exists in a second video frame according to the identity information, where the second video frame is a frame preceding the video frame to be detected;
[0031] S150, when the target object exists in the second video frame, updating the retention time of the target object;
[0032] S160: When the retention time is greater than a first threshold, determine that the target object is in a retention state.
[0033] Thus, it is possible to obtain a video frame to be detected, input the video frame to be detected into a target detection model, detect the target object in the video frame to be detected by the target detection model, obtain a first video frame marked with a detection frame, and then input the first video frame into a tracking model, determine the identity information of the target object in the first video frame by the tracking model, and then determine whether the target object appears in a second video frame based on the identity information. The second video frame is the previous frame of the video frame to be detected. If it appears in the second video frame, the retention time of the target object is updated, and when the retention time is greater than a first threshold, it is determined that the target object is in a retained state. In this way, automatic detection of personnel retention can be achieved without manual intervention, thereby improving the real-time and accuracy of detection.
[0034] Regarding S110 , the video frame to be detected may be the current frame of the target video. The target video may be a surveillance video of the target area, which may be obtained by a surveillance device. Of course, the target video may also be other videos, and this is not limited here. The target area may be an area where personnel are to be detected, such as a train station, a shopping mall, or an airport. The target video may include multiple video frames to be detected.
[0035] In S120 , the target video can be input into the target detection model in real time, so that each input is the current video frame, i.e., the video frame to be detected. The video frame to be detected is input into the target detection model, and the target detection model can detect the target object in the video frame to be detected, thereby obtaining a first video frame marked with a detection frame. The first video frame includes at least one target object, each of which is marked with a detection frame. The target object is the object for which determination of whether it is in a retained state is required.
[0036] Here, the target detection model can be Yolov5. Before using this target detection model for detection, the target detection model needs to be trained. Specifically, surveillance video of the target area can be collected first. Multiple video frames can be periodically extracted from the surveillance video. People with an occlusion rate of no more than 30% are manually framed and labeled using a labeling tool. The labeled multiple video frames are divided into a test set and a training set in a ratio of 1:5. The target detection model is then trained with the training set, and the training effect is tested with the test set.
[0037] Regarding S130, the first object is an object whose identity information has been determined. The target object is matched with at least one first object. If the target object matches one of the at least one first object, the identity information of the target object can be determined to be the identity information of the object, that is, the target object and the object are the same object. If the target object does not match any of the at least one first object, new identity information that is different from any of the first objects can be set for the target object. The identity information can be an identifier, such as a number.
[0038] In some implementations, to determine the identity information of each target object in the first video frame, S130 may specifically include:
[0039] For each detection box in the first video frame, perform the following steps:
[0040] Obtain the first feature information of the detection frame;
[0041] matching the first feature information with at least one second feature information;
[0042] When it is determined that the first feature information matches the target feature information in the at least one second feature information, the identity information of the target object corresponding to the first feature information is determined to be the identity information of the first object corresponding to the target feature information.
[0043] Here, the identity information of each target object can be determined separately. The target object has been marked by a detection frame. Therefore, each detection frame can be processed separately. The first feature information corresponding to the target object can be the first feature information of the detection frame. The first feature information can include the first position information of the detection frame and the first contour feature of the target object in the detection frame. The first contour feature can be obtained by inputting the first video frame into a re-identification model, a wide ResNet or a pedestrian re-identification Fast-Reid model, wherein obtaining the first contour feature of the target object through Fast-Reid will be more accurate. At least one second feature information can be predicted based on the third feature information corresponding to at least one first object. The second feature information can include the second position information of the predicted box corresponding to the first object and the second contour feature of the first object when it appears in at least one video frame.
[0044] In some embodiments, the identity information of the target object can be determined using a multi-target tracking model (deepsort). Specifically, the first feature information corresponding to the target object can be matched with at least one second feature information corresponding to at least one first object. If the first feature information matches the target feature information in the at least one second feature information, it can be determined that the target object and the first object corresponding to the target feature information are the same object. Therefore, the identity information of the target object can be determined to be the identity information of the first object corresponding to the target feature information.
[0045] In this way, by matching the first feature information corresponding to each target object with the second feature information corresponding to at least one first object, the identity information of the target object can be determined when the first feature information matches the target feature information in at least one second feature information.
[0046] In some embodiments, the contour features of the target object are usually saved each time it appears. However, since the target object may overlap with other objects when obtaining the contour features, what is actually obtained is the fusion features of the target object and other objects. The larger the overlapping area, the less accurate the contour features of the target object. Therefore, in order to make the saved contour features of the target object more accurate, the IOU value of every two target objects in the first video frame can be calculated. If the IOU value is greater than the IOU threshold, it indicates that the overlapping area of the two target objects is large. Therefore, the contour features of the two target objects in the first video frame are not saved.
[0047] In some embodiments, in order to obtain the second feature information, before matching the first feature information with the at least one second feature information, the method may further include:
[0048] Obtaining third feature information corresponding to at least one first object;
[0049] Kalman filter prediction is performed based on the at least one third feature information to obtain at least one second feature information.
[0050] Here, the third feature information corresponding to at least one stored first object can be obtained, and then based on the third feature information, the feature information of at least one first object at the moment corresponding to the video frame to be detected is predicted through Kalman filtering to obtain at least one second feature information.
[0051] In this way, the second feature information can be obtained, which is convenient for matching with the first feature information.
[0052] In some embodiments, in order to determine target characteristic information that matches the first characteristic information, matching the first characteristic information with at least one second characteristic information may specifically include:
[0053] performing cascade matching on the first feature information and at least one fourth feature information;
[0054] If the cascade matching is successful, determining that the first feature information matches target feature information in at least one fourth feature information;
[0055] If the cascade matching fails, performing an intersection-over-union (IOU) matching on the first feature information and at least one second feature information;
[0056] When the IOU matching is successful, it is determined that the first feature information matches the target feature information in the at least one second feature information.
[0057] Here, the second feature information may include fourth feature information and fifth feature information. The fourth feature information may be predicted based on feature information that has been successfully matched a preset number of times, and the fifth feature information may be predicted based on feature information that has not been successfully matched a preset number of times. Cascade matching may be performed on the first feature information and at least one fourth feature information. If the cascade matching is successful, it may be determined that the first feature information matches the target feature information in the at least one fourth feature information. If the cascade matching fails, an IOU matching may be performed on the first feature information and at least one second feature information. If the IOU matching is successful, it may be determined that the first feature information matches the target feature information in the at least one second feature information.
[0058] In this way, target feature information that matches the first feature information can be determined, thereby facilitating the determination of the identity information of the target object.
[0059] In some embodiments, the first feature information may not match any of the second feature information. Therefore, in order to determine the identity information of the target object, after performing IOU matching on the first feature information and at least one of the second feature information, the method may further include:
[0060] When the IOU matching fails, the identity information of the target object corresponding to the first feature information is determined to be the first identity information.
[0061] Here, the first identity information is different from the identity information of any first object. If the IOU match fails, it can be indicated that the first feature information does not match any second feature information, and there is no object in the first object that matches the target object. Therefore, the target object is an object that has newly appeared in the target area, and new identity information can be assigned to the target object.
[0062] In this way, the identity information of the target object can be determined when the first feature information does not match any feature information in the second feature information.
[0063] In addition, after the IOU matching, the stored third feature information can also be cleaned up. Specifically, if the eighth feature information in at least one fourth feature information does not match the first feature information corresponding to any target object in the video frame to be detected, and the number of consecutive mismatches reaches a preset threshold, it can be indicated that the first object corresponding to the eighth feature information does not appear in multiple consecutive video frames, indicating that the first object corresponding to the eighth feature information has left the target area. Therefore, the eighth feature information can be deleted. Since the eighth feature information is predicted based on the third feature information, the third feature information corresponding to the eighth feature information can also be deleted. In addition, if the ninth feature information in at least one fifth feature information does not match the first feature information corresponding to any target object in the video frame to be detected, the ninth feature information can be deleted. Since the ninth feature information is predicted based on the third feature information, the third feature information corresponding to the ninth feature information can be deleted.
[0064] In some embodiments, when performing cascade matching, the existing DeepSort matches the first feature information in the order of the last appearance time of the first object corresponding to each fourth feature information from latest to earliest. However, Figure 2a As shown, in the T-2 (T is a positive integer) video frame, the identity information of 201 is determined to be 001, and the identity information of 202 is 002. Figure 2b In the T-1th video frame shown, 201 is blocked by 202, resulting in 201 not being detected. Figure 2c In the Tth video frame shown, when determining the identity information of 203, if the order of the existing DeepSort in cascade matching is followed, 203 will be matched with 202 first, and only when the matching degree of 203 and 202 is lower than the threshold will 203 be matched with 201. Figure 2b In the T-1th video frame shown, 201 is blocked by 202 and both appear in the detection frame of 002. Therefore, when the contour feature of 002 is obtained in the T-1th video frame, the fusion feature of 201 and 202 is obtained. This will result in a high matching degree when 203 is matched with 202 in the T-1th video frame, which may exceed the threshold, resulting in a successful match, thereby mistakenly believing that 203 is 202 in the T-1th video frame and determining the identity information of 203 as 002. However, in fact, 203 is 201 in the T-2th video frame, and the identity information should be 001. In order to make the cascade matching more accurate, the above-mentioned cascade matching of the first feature information and the at least one fourth feature information may specifically include:
[0065] Determine the first time period before the target time according to the preset duration;
[0066] matching the first feature information with fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree;
[0067] When the maximum value of at least one matching degree is greater than the third threshold, determining the sixth feature information corresponding to the maximum value as the target feature information;
[0068] When the maximum value of at least one matching degree is less than or equal to the third threshold, a second time period before the first time period is determined according to a preset duration, and the second time period is updated to the first time period. Based on the updated first time period, the first feature information is matched with the fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree, until the target feature information is determined or the number of matches reaches the fourth threshold.
[0069] Here, the target time may be the time corresponding to the video frame to be detected, and the first time period may be the time period immediately preceding the target time. For example, if the preset time duration is 5 seconds and the target time is 09:25:30, the first time period may be 09:25:25-09:25:29. The first feature information is matched with the fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree. If the maximum value of the at least one matching degree is greater than a third threshold, the maximum value is the matching degree between the first feature information and the sixth feature information, and the sixth feature information is therefore determined to be the target feature information. If the maximum value of the at least one matching degree is less than or equal to the third threshold, the time period immediately preceding the first time period is determined to be the second time period according to the preset time duration, and the second time period is updated to the first time period. The process then returns to matching the first feature information with the fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree. In other words, the first feature information is matched with the fourth feature information corresponding to the first object that last appeared within the second time period to obtain at least one matching degree.
[0070] In some examples, the target time is 09:25:30, then the first time period may be 09:25:25-09:25:29, if the time at which the first object corresponding to the fourth feature information a last appeared is 09:25:26, and the time at which the first object corresponding to the fourth feature information b last appeared is 09:25:27, then the first feature information is matched with the fourth feature information a and the fourth feature information b respectively, if the matching degree between the first feature information and the fourth feature information a is greater than the matching degree between the first feature information and the fourth feature information b, then the matching degree between the first feature information and the fourth feature information a is compared with the third threshold, if it is greater than the third threshold, then the fourth feature information a is determined to be the sixth feature information, i.e., the target feature information, if it is less than or equal to the third threshold, then The second time period is determined to be 09:25:20-09:25:24. If the last appearance time of the first object corresponding to the fourth feature information c is 09:25:21, and the last appearance time of the first object corresponding to the fourth feature information d is 09:25:22, then the first feature information can be matched with the fourth feature information c and the fourth feature information d respectively. If the matching degree between the first feature information and the fourth feature information c is greater than the matching degree between the first feature information and the fourth feature information d, then the matching degree between the first feature information and the fourth feature information c is compared with the third threshold. If it is greater than the third threshold, then the fourth feature information c is determined to be the sixth feature information, that is, the target feature information. The above process is repeated until the target feature information is determined or the number of matches reaches the fourth threshold.
[0071] In this way, the problem of inaccurate matching results caused by the existing cascade matching order can be avoided to a certain extent.
[0072] In some embodiments, the cascade matching of the first feature information and the at least one fourth feature information may specifically include:
[0073] Calculating the cosine distance between the first feature information and each fourth feature information based on the first feature information and the third feature included in the at least one fourth feature information to obtain a cost matrix;
[0074] Calculating the Mahalanobis distance between the first feature information and each fourth feature information respectively according to the first position information and the third position information included in the at least one fourth feature information;
[0075] Constrain the cost matrix through Mahalanobis distance;
[0076] The constrained cost matrix is input into the Hungarian algorithm to obtain the matching cascade result.
[0077] Here, the cost matrix includes multiple cosine distances. If the Mahalanobis distance between the first feature information and a certain fourth feature information is greater than a preset Mahalanobis distance threshold, the cosine distance corresponding to the first feature information and the fourth feature information in the cost matrix can be set to infinitesimal or to a preset constant less than a fifth threshold. Here, setting the cosine distance less than the fifth threshold indicates a mismatch between the two.
[0078] In this way, through the above process, the first feature information and at least one fourth feature information can be cascade matched to obtain a cascade matching result, thereby determining the target feature information in at least one fourth feature information that matches the first feature information, or determining that the first feature information does not match any fourth feature information.
[0079] In some embodiments, in order to more accurately determine the cosine distance between the first feature information and each fourth feature information, the cosine distance between the first feature information and each fourth feature information is calculated based on the first profile feature and the third profile feature included in the at least one fourth feature information. Specifically, the cosine distance may include:
[0080] Each third contour feature includes a fourth contour feature of the second object corresponding to the fourth feature information when it appears in at least one video frame. For each third contour feature, the following steps are performed:
[0081] calculating a cosine distance between the first contour feature and each fourth contour feature to obtain at least one first cosine distance;
[0082] determining a second cosine distance of the at least one first cosine distance that is greater than a fifth threshold;
[0083] When the ratio of the number of second cosine distances to the number of first cosine distances is greater than a sixth threshold, a maximum value of at least one first cosine distance is determined as the cosine distance between the first feature information and each fourth feature information.
[0084] Here, since each third outline feature includes a fourth outline feature when the second object appears in at least one video frame, each third outline feature includes at least one fourth outline feature. By calculating the cosine distances of the first outline feature and each fourth outline feature, at least one first cosine distance can be obtained. Among the at least one first cosine distance, the distance greater than a fifth threshold is determined as a second cosine distance. If the proportion of the second cosine distances is greater than a sixth threshold, the maximum value among the at least one first cosine distance can be determined as the final cosine distance. If the proportion of the second cosine distances is less than or equal to the sixth threshold, the final cosine distance can be determined to be a preset constant less than the fifth threshold.
[0085] In this way, by taking the maximum value of at least one first cosine distance as the final cosine distance when the proportion of the second cosine distance is greater than the sixth threshold, and setting the final cosine distance to a preset constant less than the fifth threshold when the proportion of the second cosine distance is less than or equal to the sixth threshold, the influence of extreme values on the determination of the final cosine distance can be avoided, thereby facilitating more accurate determination of the cosine distance of the first feature information and each fourth feature information.
[0086] In some embodiments, in order to more accurately determine the Mahalanobis distance between the first feature information and each fourth feature information, the Mahalanobis distance between the first feature information and each fourth feature information is calculated based on the first position information and the third position information included in the at least one fourth feature information, and specifically may include:
[0087] The first position information includes the first vertex coordinates, first width, and first height of the detection box. Each third position information includes the second vertex coordinates, second width, and second height of the prediction box. For each third position information, perform the following steps:
[0088] Calculate the Mahalanobis distance between the coordinates of the first vertex and the coordinates of the second vertex to determine the first Mahalanobis distance;
[0089] Calculate the Mahalanobis distance between the first width and the second width to determine a first Mahalanobis distance;
[0090] Calculate the Mahalanobis distance between the first height and the second height, and determine the third Mahalanobis distance;
[0091] calculating the Mahalanobis distance between the first aspect ratio and the second aspect ratio to determine a fourth Mahalanobis distance;
[0092] A weighted calculation is performed on the first Mahalanobis distance, the second Mahalanobis distance, the third Mahalanobis distance, and the fourth Mahalanobis distance to determine the Mahalanobis distance between the first feature information and each fourth feature information.
[0093] Here, for each fourth feature information, the Mahalanobis distances between the first feature information and each fourth feature information in dimensions such as vertex coordinates, height, width, and aspect ratio can be calculated respectively to obtain the first Mahalanobis distance, the second Mahalanobis distance, the third Mahalanobis distance, and the fourth Mahalanobis distance. Then, a weighted calculation is performed to ultimately determine the Mahalanobis distance between the first feature information and the fourth feature information. The first aspect ratio is the ratio of the first width to the first height, and the second aspect ratio is the ratio of the second width to the second height.
[0094] In this way, through the above process, the Mahalanobis distance between the first feature information and each fourth feature information can be determined more accurately.
[0095] In some embodiments, as Figures 3a-3cAs shown, during the process of 301 and 302 crossing each other, the detection frame of 301 suddenly becomes smaller and then suddenly becomes larger, resulting in inaccurate position information of the detection frame of 301, thereby adversely affecting the matching of the first feature information and the fourth feature information. In order to improve the accuracy of the matching of the first feature information and the fourth feature information, before performing the weighted calculation on the first Mahalanobis distance, the second Mahalanobis distance, the third Mahalanobis distance, and the fourth Mahalanobis distance to determine the Mahalanobis distance between the first feature information and each fourth feature information, the method may further include:
[0096] Calculate the difference between the first height and the second height;
[0097] When the difference is greater than the seventh threshold, the weight of the third Mahalanobis distance is reset to 0.
[0098] Here, the difference between the first height corresponding to the first feature information and the second height corresponding to the fourth feature information can be calculated. If the difference is greater than the seventh threshold, it can be determined that the target object corresponding to the first feature information may have crossed another object, causing the detection frame size to suddenly change. Because the change in detection frame size is more obvious in height, the weight of the third Mahalanobis distance corresponding to height can be reset to 0 when calculating the Mahalanobis distance.
[0099] In this way, inaccurate matching caused by the target object crossing other objects can be avoided.
[0100] In S140 , it may be checked whether there is an object with the same identity information as the target object in the second video frame, that is, it is determined whether the target object exists in the second video frame.
[0101] In some embodiments, if the target object does not exist in the second video frame, it may be because the target object has left the target area and then re-entered it, or it may be because the target object is blocked by other objects in the second video frame. In order to determine the state of the target object, after S140, the method may further include:
[0102] When the target object does not exist in the second video frame, updating the residence time and disappearance time of the target object;
[0103] If the disappearance time is less than or equal to the second threshold, determining whether the retention time is greater than the first threshold;
[0104] When the retention time is greater than a first threshold, determining that the target object is in a retention state;
[0105] When the disappearance time is greater than the second threshold, the feature information of the target object is initialized.
[0106] Here, if the target object does not exist in the second video frame, the target object's residence time and disappearance time can be updated simultaneously. By judging whether the disappearance time is greater than the second threshold value, it is judged whether the target object re-enters after leaving the target area or is blocked by other objects in the second video frame. If the disappearance time is less than or equal to the second threshold value, it is indicated that the target object is blocked by other objects in the second video frame and does not leave the target area. Therefore, it is necessary to set the disappearance time of the target object to 0 and judge whether the target object is in a detained state. If the residence time of the target object is greater than the first threshold value, it can be determined that the target object is in a detained state. If the residence time of the target object is less than or equal to the first threshold value, it can be determined that the target object is not in a detained state. If the disappearance time is greater than the second threshold value, it is indicated that the target object does not exist in the second video frame because the target object has left the target area and re-entered the target area in the video frame to be detected. Therefore, it is necessary to initialize the feature information of the target object, delete the feature information of the target object stored before, and save the feature information of the target object obtained from the video frame to be detected. The feature information can include the position information and contour features of the target object.
[0107] In this way, through the above process, it is possible to determine whether the target object that appears in the video frame to be detected but does not appear in the second video frame has re-entered the target area after leaving, or is blocked by other objects in the second video frame, and after determining that the target object is blocked in the second video frame, determine whether the target object is in a stagnant state; after determining that the target object has re-entered after leaving the target area, initialize the feature information of the target object so as to restart the detection of the target object's current state.
[0108] Regarding S150 , if the target object exists in the second video frame, the residence time of the target object may be updated, for example, by adding 1 to the residence time of the target object.
[0109] Regarding S160, after the retention time is updated, it is determined whether the retention time is greater than a first threshold. If it is greater than the first threshold, it can be determined that the target object is in a retention state.
[0110] In addition, the identity information determined each time can be stored in a database. If the first identity information stored in the database does not match the identity information of all target objects in the video frame to be detected, the residence time and disappearance time of the object corresponding to the first identity information can be updated at the same time, and then it is determined whether the disappearance time is greater than a second threshold. If it is greater than the second threshold, it means that the object corresponding to the first identity information has left, so the first identity information can be deleted from the database.
[0111] Based on the same inventive concept, the embodiment of the present application also provides a state detection device for a target object. Figure 4The target object status detection device provided in the embodiment of the present application is described in detail.
[0112] Figure 4 A schematic structural diagram of a target object status detection device provided by an embodiment of the present application is shown.
[0113] like Figure 4 As shown, the target object state detection device may include:
[0114] The acquisition module 401 is used to acquire the video frame to be detected;
[0115] A detection module 402 is configured to input a video frame to be detected into a target detection model, detect a target object in the video frame to be detected by the target detection model, and obtain a first video frame marked with a detection frame;
[0116] A matching module 403 is configured to match the target object in the first video frame with at least one first object to determine identity information of the target object in the first video frame, where the first object is an object whose identity information has been determined;
[0117] A first determination module 404 is configured to determine whether the target object exists in a second video frame according to the identity information, where the second video frame is a frame preceding the video frame to be detected;
[0118] A first updating module 405 is configured to update the residence time of the target object when the target object exists in the second video frame;
[0119] The second determining module 406 is configured to determine that the target object is in a retention state when the retention time is greater than a first threshold.
[0120] Thus, it is possible to obtain a video frame to be detected, input the video frame to be detected into a target detection model, detect the target object in the video frame to be detected by the target detection model, obtain a first video frame marked with a detection frame, and then input the first video frame into a tracking model, determine the identity information of the target object in the first video frame by the tracking model, and then determine whether the target object appears in a second video frame based on the identity information. The second video frame is the previous frame of the video frame to be detected. If it appears in the second video frame, the retention time of the target object is updated, and when the retention time is greater than a first threshold, it is determined that the target object is in a retained state. In this way, automatic detection of personnel retention can be achieved without manual intervention, thereby improving the real-time and accuracy of detection.
[0121] In some embodiments, if the target object does not exist in the second video frame, it may be because the target object has left the target area and then re-entered it, or it may be because the target object is blocked by other objects in the second video frame. In order to determine the state of the target object, the apparatus may further include:
[0122] a second updating module, configured to, after determining whether the target object exists in the second video frame according to the identity information, update the residence time and disappearance time of the target object if the target object does not exist in the second video frame;
[0123] a third determining module, configured to determine whether the retention time is greater than the first threshold when the disappearance time is less than or equal to the second threshold;
[0124] a fourth determining module, configured to determine that the target object is in a detention state when the detention time is greater than a first threshold;
[0125] The initialization module is used to initialize the feature information of the target object after updating the residence time and disappearance time of the target object and when the disappearance time is greater than a second threshold, the feature information including the position information and contour feature of the target object.
[0126] In some implementations, to determine the identity information of each target object in the first video frame, the matching module 403 may specifically include:
[0127] an acquisition submodule, configured to, for each detection frame in the first video frame, respectively perform the following steps: acquiring first feature information of the detection frame, where the first feature information includes first position information of the detection frame and a first contour feature of a target object in the detection frame, where the first contour feature is obtained by inputting the first video frame into a person re-identification Fast-Reid model;
[0128] a matching submodule, configured to, for each detection frame in the first video frame, respectively perform: matching the first feature information with at least one second feature information, where the at least one second feature information is predicted based on third feature information corresponding to at least one first object, the second feature information including second position information of the predicted frame corresponding to the first object and second contour features of the first object when it appears in the at least one video frame;
[0129] A determination submodule is used to perform, for each detection frame in the first video frame, respectively: when it is determined that the first feature information matches the target feature information in at least one second feature information, determine that the identity information of the target object corresponding to the first feature information is the identity information of the first object corresponding to the target feature information.
[0130] In some implementations, to determine target feature information that matches the first feature information, the matching submodule may specifically include:
[0131] a cascade matching unit, configured to perform cascade matching on the first feature information and at least one fourth feature information, wherein the second feature information includes fourth feature information and fifth feature information, the fourth feature information being predicted based on feature information that has been successfully matched continuously for a preset number of times, and the fifth feature information being predicted based on feature information that has not been successfully matched continuously for a preset number of times;
[0132] a first determining unit, configured to determine, when the cascade matching succeeds, that the first feature information matches target feature information in at least one fourth feature information;
[0133] An IOU matching unit is configured to perform an intersection-over-union (IOU) matching on the first feature information and at least one second feature information when the cascade matching fails;
[0134] The second determining unit is configured to determine, when the IOU matching is successful, whether the first feature information matches the target feature information in the at least one second feature information.
[0135] In some embodiments, in order to determine the identity information of the target object, the matching submodule may further include:
[0136] The third determination unit is used to determine that the identity information of the target object corresponding to the first feature information is the first identity information after IOU matching is performed on the first feature information and at least one second feature information, if the IOU matching fails, and the first identity information is different from the identity information of any first object.
[0137] In some implementations, to make the cascade matching more accurate, the cascade matching unit may specifically include:
[0138] A first determining subunit is configured to determine a first time period before a target moment according to a preset duration, where the target moment is a moment corresponding to the video frame to be detected;
[0139] a matching subunit, configured to match the first feature information with fourth feature information corresponding to the first object that last appeared within the first time period, to obtain at least one matching degree;
[0140] a second determining subunit, configured to determine, when a maximum value of at least one matching degree is greater than a third threshold, the sixth feature information corresponding to the maximum value as the target feature information;
[0141] An updating subunit is used to determine a second time period before the first time period according to a preset duration when the maximum value of at least one matching degree is less than or equal to a third threshold, and to update the second time period to the first time period, and to return, based on the updated first time period, to match the first feature information with the fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree, until the target feature information is determined or the number of matches reaches the fourth threshold.
[0142] In some implementations, the cascade matching unit may specifically include:
[0143] A first calculation subunit is configured to calculate, based on the first contour feature and the third contour feature included in at least one fourth characteristic information, a cosine distance between the first characteristic information and each fourth characteristic information to obtain a cost matrix;
[0144] a second calculation subunit, configured to calculate the Mahalanobis distance between the first feature information and each fourth feature information respectively based on the first position information and the third position information included in the at least one fourth feature information;
[0145] The constraint subunit is used to constrain the cost matrix through the Mahalanobis distance;
[0146] The algorithm subunit is used to input the constrained cost matrix into the Hungarian algorithm to obtain the matching cascade result.
[0147] In some embodiments, in order to more accurately determine the cosine distance between the first feature information and each fourth feature information, the first calculation subunit is specifically configured to:
[0148] Each third contour feature includes a fourth contour feature of the second object corresponding to the fourth feature information when it appears in at least one video frame. For each third contour feature, the following steps are performed:
[0149] calculating a cosine distance between the first contour feature and each fourth contour feature to obtain at least one first cosine distance;
[0150] determining a second cosine distance of the at least one first cosine distance that is greater than a fifth threshold;
[0151] When the ratio of the number of second cosine distances to the number of first cosine distances is greater than a sixth threshold, a maximum value of at least one first cosine distance is determined as the cosine distance between the first feature information and each fourth feature information.
[0152] In some embodiments, in order to more accurately determine the Mahalanobis distance between the first feature information and each fourth feature information, the second calculation subunit is specifically configured to:
[0153] The first position information includes the first vertex coordinates, first width, and first height of the detection box. Each third position information includes the second vertex coordinates, second width, and second height of the prediction box. For each third position information, perform the following steps:
[0154] Calculate the Mahalanobis distance between the coordinates of the first vertex and the coordinates of the second vertex to determine the first Mahalanobis distance;
[0155] Calculate the Mahalanobis distance between the first width and the second width to determine a first Mahalanobis distance;
[0156] Calculate the Mahalanobis distance between the first height and the second height, and determine the third Mahalanobis distance;
[0157] calculating a Mahalanobis distance between a first aspect ratio and a second aspect ratio to determine a fourth Mahalanobis distance, wherein the first aspect ratio is a ratio of the first width to the first height, and the second aspect ratio is a ratio of the second width to the second height;
[0158] A weighted calculation is performed on the first Mahalanobis distance, the second Mahalanobis distance, the third Mahalanobis distance, and the fourth Mahalanobis distance to determine the Mahalanobis distance between the first feature information and each fourth feature information.
[0159] In some embodiments, in order to improve the accuracy of matching the first feature information and the fourth feature information, the second calculation subunit is further specifically configured to:
[0160] Before performing weighted calculation on the first Mahalanobis distance, the second Mahalanobis distance, the third Mahalanobis distance, and the fourth Mahalanobis distance to determine the Mahalanobis distance between the first feature information and each fourth feature information, calculating the difference between the first height and the second height;
[0161] When the difference is greater than the seventh threshold, the weight of the third Mahalanobis distance is reset to 0.
[0162] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present application is shown.
[0163] like Figure 5 As shown, the electronic device 5 is a structural diagram of an exemplary hardware architecture of an electronic device that can implement the target object state detection method and target object state detection device according to the embodiments of the present application. The electronic device can refer to the electronic device in the embodiments of the present application.
[0164] The electronic device 5 may include a processor 501 and a memory 502 storing computer program instructions.
[0165] Specifically, the processor 501 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0166] The memory 502 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 502 may include removable or non-removable (or fixed) media. Where appropriate, the memory 502 may be internal or external to the integrated gateway disaster recovery device. In certain embodiments, the memory 502 is a non-volatile solid-state memory. In certain embodiments, the memory 502 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Therefore, generally, the memory 502 includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of the present application.
[0167] The processor 501 reads and executes computer program instructions stored in the memory 502 to implement any one of the target object status detection methods in the above embodiments.
[0168] In one example, the electronic device may further include a communication interface 503 and a bus 504. Figure 5 As shown, the processor 501 , the memory 502 , and the communication interface 503 are connected via a bus 504 and communicate with each other.
[0169] The communication interface 503 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0170] Bus 504 comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 504 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0171] The electronic device can execute the target object state detection method in the embodiment of the present application, thereby realizing the combination Figures 1 to 4 A method and device for detecting the state of a target object are described.
[0172] In addition, in conjunction with the target object status detection method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the target object status detection methods in the above embodiments is implemented.
[0173] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0174] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0175] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0176] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.
[0177] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. A method for detecting the state of a target object, characterized in that: include: Get the video frame to be detected; Inputting the video frame to be detected into a target detection model, detecting the target object in the video frame to be detected by the target detection model, and obtaining a first video frame marked with a detection frame; Match the target object in the first video frame with at least one first object to determine the identity information of the target object in the first video frame, where the first object is an object with determined identity information, wherein the first feature information of the target object is cascade matched with the fourth feature information in the second feature information of at least one of the first objects, and the fourth feature information is predicted based on the feature information that has been successfully matched for a preset number of times in a row, and the cascade matching includes: determining the Mahalanobis distance between the first feature information and each fourth feature information based on the first position information in the first feature information and the third position information included in at least one of the fourth feature information, the first position information including the first vertex coordinates, the first width and the first height of the detection box, and each third position information including the second vertex coordinates, the second width and the second height of the prediction box, and determining the matching cascade result based on the Mahalanobis distance. determining, according to the matching cascade result, the identity information of the target object, wherein the Mahalanobis distance is obtained by weighted calculation of a first Mahalanobis distance, a second Mahalanobis distance, a third Mahalanobis distance, and a fourth Mahalanobis distance, wherein the first Mahalanobis distance is determined by calculating the Mahalanobis distance between the first vertex coordinate and the second vertex coordinate, the second Mahalanobis distance is determined by calculating the Mahalanobis distance between the first width and the second width, the third Mahalanobis distance is determined by calculating the Mahalanobis distance between the first height and the second height, and the fourth Mahalanobis distance is determined by calculating the Mahalanobis distance between a first aspect ratio and a second aspect ratio, wherein the first aspect ratio is a ratio of the first width to the first height, and the second aspect ratio is a ratio of the second width to the second height, and when it is determined that the difference between the first height of the first position information and the second height of the third position information is greater than a seventh threshold, resetting the weight of the third Mahalanobis distance to 0; determining whether the target object exists in a second video frame according to the identity information, where the second video frame is a frame preceding the video frame to be detected; In a case where the target object exists in the second video frame, updating the residence time of the target object; When the retention time is greater than a first threshold, it is determined that the target object is in a retention state.
2. The method according to claim 1, characterized in that After determining whether the target object exists in the second video frame according to the identity information, the method further includes: When the target object does not exist in the second video frame, updating the residence time and disappearance time of the target object; If the disappearance time is less than or equal to the second threshold, determining whether the retention time is greater than the first threshold; When the retention time is greater than a first threshold, determining that the target object is in a retention state; When the disappearance time is greater than a second threshold, feature information of the target object is initialized, where the feature information includes position information and contour features of the target object.
3. The method according to claim 1, characterized in that Matching the target object in the first video frame with at least one first object to determine identity information of the target object in the first video frame includes: For each detection frame in the first video frame, perform the following steps respectively: Obtaining first feature information of the detection frame, where the first feature information includes first position information of the detection frame and a first contour feature of a target object in the detection frame, where the first contour feature is obtained by inputting the first video frame into a person re-identification Fast-Reid model; matching the first feature information with at least one second feature information, where the at least one second feature information is predicted based on third feature information corresponding to at least one first object, the second feature information including second position information of a prediction box corresponding to the first object and a second outline feature of the first object when it appears in at least one video frame; When it is determined that the first feature information matches the target feature information in at least one second feature information, the identity information of the target object corresponding to the first feature information is determined to be the identity information of the first object corresponding to the target feature information.
4. The method according to claim 3, characterized in that The matching of the first feature information with at least one second feature information includes: performing cascade matching on the first feature information and at least one fourth feature information, where the second feature information includes fourth feature information and fifth feature information, the fourth feature information being predicted based on feature information that has been successfully matched continuously a preset number of times, and the fifth feature information being predicted based on feature information that has not been successfully matched continuously a preset number of times; If the cascade matching is successful, determining that the first feature information matches target feature information in at least one fourth feature information; If the cascade matching fails, performing an intersection-over-union (IOU) matching on the first feature information and at least one second feature information; When the IOU matching is successful, it is determined that the first feature information matches the target feature information in at least one second feature information.
5. The method according to claim 4, characterized in that After performing IOU matching on the first feature information and at least one second feature information, the method further includes: In the event that the IOU matching fails, it is determined that the identity information of the target object corresponding to the first feature information is the first identity information, and the first identity information is different from the identity information of any first object.
6. The method according to claim 4, characterized in that The cascade matching of the first feature information and at least one fourth feature information includes: Determine a first time period before a target moment according to a preset duration, wherein the target moment is a moment corresponding to the video frame to be detected; matching the first feature information with fourth feature information corresponding to a first object that last appeared within the first time period to obtain at least one matching degree; When a maximum value among the at least one matching degree is greater than a third threshold, determining the sixth feature information corresponding to the maximum value as the target feature information; When the maximum value of the at least one matching degree is less than or equal to the third threshold, a second time period before the first time period is determined according to a preset duration, and the second time period is updated to the first time period. Based on the updated first time period, the first feature information is matched with the fourth feature information corresponding to the first object that last appeared within the first time period to obtain at least one matching degree, until the target feature information is determined or the number of matches reaches the fourth threshold.
7. The method according to claim 4, characterized in that The cascade matching of the first feature information and at least one fourth feature information includes: Calculating the cosine distance between the first feature information and each fourth feature information according to the first feature information and the third feature included in the at least one fourth feature information to obtain a cost matrix; Calculating the Mahalanobis distance between the first feature information and each fourth feature information respectively according to the first position information and the third position information included in the at least one fourth feature information; Constraining the cost matrix by the Mahalanobis distance; The constrained cost matrix is input into the Hungarian algorithm to obtain the matching cascade result.
8. The method according to claim 7, characterized in that The calculating, based on the first contour feature and the third contour feature included in the at least one fourth characteristic information, the cosine distance between the first characteristic information and each fourth characteristic information includes: Each third contour feature includes a fourth contour feature of the second object corresponding to the fourth feature information when it appears in at least one video frame. For each third contour feature, the following steps are performed respectively: calculating a cosine distance between the first contour feature and each fourth contour feature to obtain at least one first cosine distance; determining a second cosine distance of the at least one first cosine distance that is greater than a fifth threshold; When the ratio of the number of the second cosine distances to the number of the first cosine distances is greater than a sixth threshold, a maximum value of the at least one first cosine distance is determined as the cosine distance between the first feature information and each fourth feature information.
Citation Information
Patent Citations
Suspicious object detection method and device
CN106846357A
Behavior analysis method, equipment and device
CN111401296A
Pedestrian reloading and re-identification method for open space
CN112541421A
Video-based one-way passenger flow information detection method in two-way passenger flow channel
CN112560641A