Target behavior early warning method based on artificial intelligence
By applying the target behavior warning method based on artificial intelligence in places with large traffic, the problem of difficulty in timely warning abnormal situations in the existing technology is solved, and timely warning of abnormal aggregation or discrete behaviors is achieved, which improves the ability to deal with risks and reduces costs.
Patent Information
- Application Number
- CN202410423936.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-04-10
AI Technical Summary
In places with large flow of people, it is difficult to detect and warn of abnormal situations in a timely manner, resulting in the risk of riots and spread.
Using the target behavior warning method based on artificial intelligence, each video image is preprocessed by acquiring a set of video images, and inputting it into the target behavior warning model. The model includes a target detection unit, a trajectory detection unit, an index calculation unit and a behavior warning unit, which is used to detect personnel positions, motion trajectories, abnormal gathering or discrete behaviors, and to perform early warnings.
It has achieved timely early warning of abnormal situations in places with large flow of people, improved the timeliness of responding to abnormal situations, reduced risks and costs, and reduced the demand for fixed-point personnel and the need for mobile inspections.
Smart Images

Figure CN119131671B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based target behavior early warning method. Background Art
[0002] With the development of artificial intelligence in the field of image recognition and target detection, target tracking and behavior recognition based on machine vision have been increasingly applied to video surveillance systems. The currently popular target behavior recognition models are mainly divided into three categories: single-person behavior recognition, multi-person behavior recognition, and crowd social behavior recognition. For places with large flow of people (such as indoor offline game entertainment venues, outdoor commercial pedestrian streets, etc.), due to the large flow of people in the scene, if an abnormal situation occurs, the impact is greater, it is easy to cause riots and is easier to spread, so appropriate personnel are needed to maintain order.
[0003] However, the existing order maintenance plan usually arranges fixed personnel (such as fixed-location dispatch or mobile patrol), but this plan is difficult to detect situations in a timely manner for places with larger areas. (Special personnel are required to monitor, but with a large number of surveillance deployments, it is difficult for a small number of personnel to take care of them all the time). Mobile patrols require a larger number of personnel, which can improve timeliness to a certain extent, but the increased cost is higher.
[0004] Therefore, how to achieve timely warning when abnormal situations occur in such scenarios is a technical problem that needs to be solved in this field. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a target behavior warning method based on artificial intelligence to achieve timely warning when abnormal situations occur in scenes with large flow of people.
[0006] In order to achieve the above purpose, the embodiments of the present application are implemented in the following manner:
[0007] In a first aspect, an embodiment of the present application provides an artificial intelligence-based target behavior warning method, comprising: obtaining a video image set obtained based on monitoring a target area, wherein the video image set includes multiple frames of video images; preprocessing each frame of the video image to obtain a corresponding input image; inputting the input image frame by frame into a preset target behavior warning model, determining the position of a person based on the input image through the target behavior warning model, further detecting the movement trajectory of the person, and issuing a warning when abnormal aggregation behavior or abnormal discrete behavior is determined in the target area based on the movement trajectory.
[0008] In combination with the first aspect, in a first possible implementation method of the first aspect, the target behavior warning model includes a target detection unit, a trajectory detection unit, an index calculation unit and a behavior warning unit, wherein the target detection unit is used to perform target detection on each frame of input image to determine the position of each person in the input image; the trajectory detection unit is used to track each person based on the person's position to determine the movement trajectory of each person; the index calculation unit is used to determine the person aggregation index and the person dispersion index of each sub-area in the target area based on the movement trajectory of each person, wherein the target area includes multiple sub-areas; the behavior warning unit is used to determine whether there is abnormal aggregation behavior or abnormal dispersion behavior based on the person aggregation index and the person dispersion index of each sub-area, and to issue a warning when abnormal aggregation behavior or abnormal dispersion behavior exists.
[0009] In combination with the first possible implementation manner of the first aspect, in a second possible implementation manner of the first aspect, the target detection unit adopts a YOLOv5 model, and improves a positioning regression function of the YOLOv5 model in target detection, and the improved positioning regression function is:
[0010]
[0011]
[0012] Among them, L detect is the positioning regression function of target detection, N is the number of feature maps for target detection, i represents the i-th feature map of target detection, L box represents the bounding box loss function of the feature map, L obj represents the target confidence loss function of the feature map, L cls represents the target category confidence loss function of the feature map, λ1, λ2, and λ3 are the loss functions L box , L obj , L cls The weight parameter, B i is the number of target boxes that match the prior anchor points in the i-th feature map, s i ×s i Indicates the number of grids divided into the i-th feature map.
[0013] In combination with the second possible implementation manner of the first aspect, in a third possible implementation manner of the first aspect, the loss function satisfy:
[0014]
[0015]
[0016]
[0017] v j =σ
[0018]
[0019] Among them, IOU j ∈(0, 1) represents B in the YOLOv5 model i is the prediction box evaluation index corresponding to the jth target box matched to the prior anchor point in the i-th feature map, which is used to measure the overlap between the prediction box and the jth target box matched to the prior anchor point. j and They represent the jth predicted box in the i-th feature map and the jth target box matched to the prior anchor point, respectively. Represents the predicted box b j The center point and target box The Euclidean distance between the center points, c j is the predicted box b j With target box The diagonal length of the minimum closed module, is the penalty coefficient, is the predicted box b j The center point and target box The Manhattan distance between the center points of j is the intermediate variable, v j Represents the predicted box b j With target box The aspect ratio similarity is measured using KL divergence. Represents the predicted box b j The probability density of the aspect ratio, w j Represents the predicted box b j The width, h j Represents the predicted box b j Height, Represents the target box The probability density of the aspect ratio is, Represents the target box The width of Represents the target box The average height, KL represents the divergence, and σ is the sigmoid activation function.
[0020] In combination with the third possible implementation of the first aspect, in a fourth possible implementation of the first aspect, the loss function satisfy:
[0021]
[0022]
[0023]
[0024]
[0025] Among them, y k is the true label of the kth grid in the ith feature map, x k is the predicted label of the kth grid in the ith feature map, b k is the kth grid in the ith feature map, For grid b k The target box that matches the prior anchor point corresponding to the i-th feature map, For grid b k With target box The distance between the center points.
[0026] In combination with the third possible implementation of the first aspect, in a fifth possible implementation of the first aspect, the loss function satisfy:
[0027]
[0028]
[0029] Among them, y l is the true label of the lth target box matched to the prior anchor in the i-th feature map, x l is the predicted label of the lth predicted box in the ith feature map.
[0030] In combination with the first possible implementation manner of the first aspect, in a sixth possible implementation manner of the first aspect, the trajectory detection unit uses DeepSORT to implement multi-target tracking.
[0031] In combination with the first possible implementation method of the first aspect, in the seventh possible implementation method of the first aspect, the index calculation unit is specifically used to: determine the current destination area and the current departure area of each person based on the movement trajectory of each person output by the trajectory detection unit; determine the personnel aggregation index and the personnel dispersion index of each sub-area based on the current destination area and the current departure area of each person.
[0032] In combination with the seventh possible implementation manner of the first aspect, in an eighth possible implementation manner of the first aspect, the index calculation unit is specifically configured to: calculate the personnel gathering index of each sub-area by using the following formula:
[0033]
[0034] in, represents the updated personnel gathering index of sub-region s, represents the personnel gathering index of sub-region s before updating, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 Represents the number of people in sub-area s.
[0035] In combination with the seventh possible implementation manner of the first aspect, in a ninth possible implementation manner of the first aspect, the index calculation unit is specifically configured to:
[0036] The personnel dispersion index of each sub-region is calculated using the following formula:
[0037]
[0038] in, represents the population dispersion index of sub-area s, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 Represents the number of people in sub-area s.
[0039] Beneficial effects:
[0040] 1. By acquiring a set of video images obtained based on monitoring the target area, preprocessing each frame of the video image to obtain a corresponding input image, and inputting the input image frame by frame into a preset target behavior warning model, and the target behavior warning model includes a target detection unit, a trajectory detection unit, an index calculation unit and a behavior warning unit. The target detection unit performs target detection on each frame of the input image to determine the position of each person in the input image; the trajectory detection unit tracks each person based on the position of the person to determine the movement trajectory of each person; the index calculation unit determines the personnel aggregation index and personnel dispersion index of each sub-area in the target area based on the movement trajectory of each person; the behavior warning unit determines whether there is abnormal aggregation behavior or abnormal dispersion behavior based on the personnel aggregation index and personnel dispersion index of each sub-area, and issues a warning when abnormal aggregation behavior or abnormal dispersion behavior exists. In this way, artificial intelligence technology can be used to monitor abnormalities in the target area, and abnormal aggregation behavior or abnormal discrete behavior can be discovered in time or even predicted in advance (mainly abnormal aggregation situations can be predicted in advance to a certain extent), so as to issue timely warnings. Personnel can quickly arrive at the scene from fixed points to maintain order, greatly improving the timeliness of responding to abnormal situations and effectively reducing risks and costs (a small number of personnel are stationed at fixed points for order, and there is no need for a large number of mobile personnel to patrol).
[0041] 2. Based on the YOLOv5 model, the positioning regression function (understood as the loss function) of the YOLOv5 model in target detection is improved, and the bounding box loss function L of the feature map is calculated. box When calculating the target confidence loss function L of the feature map, the measurement method of the loss function is improved, and a penalty term is introduced to better handle the situation where the prediction box and the target box overlap but the center point is far away, so that the YOLOv5 model can be more sensitive to this situation and more accurately identify the person in the image. obj When using the classic binary cross entropy classification loss function (corresponding to the first term in the formula), the weight coefficient (corresponding to the second term in the formula) is introduced into the loss term of the prediction box using the IoU and center point distance of the sample. It can be adjusted according to the IoU and center point distance of the sample to adapt to the situation where the IoU is low or the center point distance is small. By introducing the classification probability into Focal Loss, the loss weight of easy-to-classify samples can be reduced, so that the YOLOv5 model pays more attention to difficult-to-classify samples and improves the classification effect. When calculating the target category confidence loss function L of the feature map cls When , the cross entropy loss function is used and the sigmoid function is introduced for adjustment, which is conducive to accelerating the convergence of the model.
[0042] 3. Designing a calculation method for the personnel gathering index of each sub-area can effectively measure the speed and degree of personnel gathering in the sub-area, so as to timely discover and even predict abnormal gathering behavior to a certain extent. Designing a calculation method for the personnel dispersion index of each sub-area can discover the rapid dispersion characterized by abnormal dispersion behavior, and can take into account the number of people in the sub-area, effectively reducing the false alarm rate caused by the rapid departure of a small number of people, thereby improving the accuracy of early warning.
[0043] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 A flowchart of an artificial intelligence-based target behavior warning method provided in an embodiment of the present application.
[0046] Figure 2This is the structural diagram of the target behavior early warning model.
[0047] Icons: 10-target behavior warning model; 11-target detection unit; 12-trajectory detection unit; 13-index calculation unit; 14-behavior warning unit. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0049] See also Figure 1 , Figure 1 A flowchart of a target behavior early warning method based on artificial intelligence provided in an embodiment of the present application. In this embodiment, the target behavior early warning method based on artificial intelligence is run by a server and may include step S10, step S20 and step S30.
[0050] In order to implement monitoring and early warning of the target area, the server may execute step S10.
[0051] Step S10: Acquire a video image set obtained by monitoring the target area, wherein the video image set includes multiple frames of video images.
[0052] In this embodiment, the server may obtain a monitoring video stream sent by a monitoring device that monitors the target area, and then extract frames (eg, extract 1 every 5 frames) to obtain a video image set including multiple frames of video images.
[0053] After obtaining the video image set, the server may execute step S20.
[0054] Step S20: pre-process each frame of video image to obtain a corresponding input image.
[0055] In this embodiment, the server may preprocess each frame of the video image, such as scaling, cropping (for example, cropping to 416×416, 608×608, 640×640, and this embodiment takes 640×640 as an example), etc., to obtain a corresponding input image.
[0056] After obtaining the input image, the server may execute step S30.
[0057] Step S30: Input the input image frame by frame into the preset target behavior warning model 10, determine the position of the person based on the input image through the target behavior warning model, further detect the movement trajectory of the person, and issue a warning when abnormal aggregation behavior or abnormal discrete behavior is determined in the target area based on the movement trajectory.
[0058] This step mainly relies on the target behavior warning model 10 set up in the server to run. Figure 2, Figure 2 It is a structural block diagram of the target behavior warning model 10.
[0059] In this embodiment, the target behavior warning model 10 may include a target detection unit 11 , a trajectory detection unit 12 , an index calculation unit 13 and a behavior warning unit 14 .
[0060] Exemplarily, the target detection unit 11 mainly uses the input image to perform target detection (detecting people in the input image), and adopts the architecture of the YOLOv5 model (YOLOv5s used in this embodiment). In order to improve the detection effect, this solution further improves the positioning regression function of the YOLOv5 model in target detection to adapt it to human target detection in scenes with large traffic. Since the architecture of the YOLOv5 model has not been adjusted, it will not be described here. You can refer to the architecture of the existing YOLOv5 model (for example, "Explanation of the structural principle of the YOLOv5 network model").
[0061] Since the YOLOv5 model uses a single-stage target detection algorithm, it can detect and locate multiple targets at the same time, with relatively high speed and accuracy. In the YOLOv5 model, the positioning regression function L detect By L box (bounding box loss function), L obj (target confidence loss function) and L cls (target category confidence loss function), this scheme is for the bounding box loss function L box , target confidence loss function L obj And the target category confidence loss function L cls All of them have been improved to stably achieve target detection in scenes with large flow of people in this scheme.
[0062] The improved positioning regression function of the YOLOv5 model is:
[0063]
[0064]
[0065] Among them, L detect is the positioning regression function of target detection (that is, it can be understood as a loss function), N is the number of feature maps for target detection, i represents the i-th feature map of target detection, and L box represents the bounding box loss function of the feature map, L obj represents the target confidence loss function of the feature map, L cls represents the target category confidence loss function of the feature map, λ1, λ2, and λ3 are the loss functions L box , L obj , L clsThe weight parameter, B i is the number of target boxes that match the prior anchor points in the i-th feature map, s i ×s i Indicates the number of grids divided into the i-th feature map.
[0066] Here, the prior anchor is a technique used to assist target detection. It is a set of predefined boxes used to capture targets of different scales and aspect ratios. The YOLOv5 model divides the image into grids (s i ×s i ), each grid cell is responsible for predicting a set of bounding boxes and corresponding category probabilities. In order to adapt to targets of different sizes and shapes, each grid cell is usually associated with a set of prior anchor points.
[0067] Specifically, the loss function satisfy:
[0068]
[0069]
[0070]
[0071] v j =σ
[0072]
[0073] Among them, IOU j ∈(0, 1) represents B in the YOLOv5 model i is the prediction box evaluation index corresponding to the jth target box matched to the prior anchor point in the i-th feature map, which is used to measure the overlap between the prediction box and the jth target box matched to the prior anchor point. j and They represent the jth predicted box in the i-th feature map and the jth target box matched to the prior anchor point, respectively. Represents the predicted box b j The center point and target box The Euclidean distance between the center points, c j is the predicted box b j With target box The diagonal length of the minimum closed module, is the penalty coefficient, is the predicted box b j The center point and target box The Manhattan distance between the center points of j is the intermediate variable, v j Represents the predicted box b j With target box The aspect ratio similarity is measured by KL divergence (the model's learning tendency for aspect ratio can be adjusted during training to better adapt to the shape characteristics of different targets). Represents the predicted box b j The probability density of the aspect ratio, w j Represents the predicted box b j The width, h j Represents the predicted box b j Height, Represents the target box The probability density of the aspect ratio is, Represents the target box The width of Represents the target box The height of , KL represents the divergence, and σ is the sigmoid activation function.
[0074] Specifically, the loss function satisfy:
[0075]
[0076]
[0077]
[0078]
[0079] Among them, y k is the true label of the kth grid in the ith feature map, x k is the predicted label of the kth grid in the ith feature map, b k is the kth grid in the ith feature map, For grid b k The target box that matches the prior anchor point corresponding to the i-th feature map, For grid b k With target box The distance between the center points.
[0080] Specifically, the loss function satisfy:
[0081]
[0082]
[0083] Among them, y l is the true label of the lth target box matched to the prior anchor in the i-th feature map, x l is the predicted label of the lth predicted box in the ith feature map.
[0084] Based on the YOLOv5 model, the positioning regression function of the YOLOv5 model in target detection is improved, and the bounding box loss function L is calculated in the feature map. box When the loss function is improved, the KL divergence is used for measurement. The model's learning tendency for aspect ratio can be adjusted during training to better adapt to the shape characteristics of different targets. At the same time, a penalty term is introduced to better handle the situation where the prediction box and the target box overlap but the center point is far away, so that the YOLOv5 model can be more sensitive to this situation and more accurately identify people in the image. obj When using the classic binary cross entropy classification loss function (corresponding to the first term in the formula), the weight coefficient (corresponding to the second term in the formula) is introduced into the loss term of the prediction box using the IoU and center point distance of the sample. It can be adjusted according to the IoU and center point distance of the sample to adapt to the situation where the IoU is low or the center point distance is small. By introducing the classification probability into Focal Loss, the loss weight of easy-to-classify samples can be reduced, so that the YOLOv5 model pays more attention to difficult-to-classify samples and improves the classification effect. When calculating the target category confidence loss function L of the feature map cls When , the cross entropy loss function is used and the sigmoid function is introduced for adjustment, which is conducive to accelerating the convergence of the model.
[0085] It should be noted that the YOLOv5 model can be trained separately before use, using images in such scenes with large flow of people as training samples (the people targets in the images are labeled) to train the YOLOv5 model, and after the training is completed, it can be used as the target detection unit 11.
[0086] Exemplarily, the trajectory detection unit 12 mainly tracks each person through the personnel position, and determines the motion trajectory of each person (including trajectory and speed information). Therefore, this embodiment adopts the DeepSORT model as the trajectory detection unit 12 to realize multi-target tracking. Since the existing DeepSORT model is very suitable for the scene with large flow of people in this embodiment, because this part is adjusted in an improved manner, by obtaining the prediction box determined and output by the YOLOv5 model based on the input image, real-time multi-target tracking is performed, and the motion trajectory of the target can be output, which can include elements such as position information, target ID, and confidence. At the same time, since the Kalman filter can be used in the DeepSORT model to determine the speed of the tracking target, the position and speed of the target can be predicted. In the optimization scheme based on this embodiment, in order to realize the prediction of the current destination area of the personnel, the motion trajectory output by the DeepSORT model (trajectory detection unit 12) can also include speed information and predicted position in addition to the elements of the motion trajectory (position information, target ID, confidence, etc.), so as to predict the current destination area, so as to monitor the abnormal aggregation behavior that may occur in advance, which will not be repeated here.
[0087] Exemplarily, the index calculation unit 13 is mainly used to determine the personnel aggregation index and the personnel dispersion index of each sub-area in the target area (including multiple sub-areas) based on the movement trajectory of each person.
[0088] Specifically, the index calculation unit 13 may determine the current destination area and the current departure area of each person based on the movement trajectory of each person output by the trajectory detection unit 12 .
[0089] Taking the motion trajectory including position information, target ID, confidence, etc. as an example, the index calculation unit 13 can determine the sub-area corresponding to the position based on the position information of the target ID (each target ID corresponds to a person), thereby determining the current destination area of this target ID. Thus, the current destination area of each person can be determined. And, the index calculation unit 13 can determine the current departure area of the target ID based on the last destination area of each person (marked with the target ID) (here, the motion trajectory without prediction is taken as an example, so the last destination area is the last location area) and the current destination area. If the destination areas of the two times are different, the last destination area is determined to be the current departure area of this target ID.
[0090] Of course, since the time accuracy of the monitored motion trajectory is relatively high (i.e., the time interval between two adjacent motion trajectories output is relatively short), the index calculation unit 13 can obtain the motion trajectory output by the trajectory detection unit 12 at intervals (for example, obtaining a motion trajectory once every 3 times, or obtaining a motion trajectory once every 5 times, which is not limited here) to avoid a reduction in detection accuracy due to higher time accuracy (i.e., subdividing a large number of gatherings or departures in a short period of time into multiple time points, resulting in statistical fragmentation).
[0091] After determining the current destination area and the current departure area of each person, the index calculation unit 13 may determine the personnel gathering index and the personnel dispersion index of each sub-area based on the current destination area and the current departure area of each person.
[0092] Specifically, the index calculation unit 13 may use the following formula to calculate the personnel gathering index of each sub-region:
[0093]
[0094] in, represents the updated personnel gathering index of sub-region s, represents the personnel gathering index of sub-region s before updating, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 represents the number of people in sub-area s (here we take the method without considering prediction as an example, c s0 That is, it indicates the number of people who used sub-area s as the destination area last time).
[0095] And, the index calculation unit 13 can calculate the personnel dispersion index of each sub-area using the following formula:
[0096]
[0097] in, represents the population dispersion index of sub-area s, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 represents the number of people in sub-area s (here we take the method without considering prediction as an example, c s0 That is, it indicates the number of people who used sub-area s as the destination area last time).
[0098] Based on this embodiment, a further optimized technical solution can predict the current destination area of each person. Then, the elements of the motion trajectory need location information, target ID, confidence, speed information and predicted location to predict the current destination area. For the predicted current destination area, when calculating the personnel aggregation index and personnel dispersion index using formulas (11) and (12), c s0 It is the number of people in sub-area s (that is, the number of people actually in sub-area s), which can no longer be equal to the number of people who took sub-area s as the destination area last time. This needs to be noted.
[0099] After calculating the personnel gathering index and personnel dispersion index of each sub-area, the behavior warning unit 14 is mainly used to determine whether there is abnormal gathering behavior or abnormal dispersion behavior based on the personnel gathering index and personnel dispersion index of each sub-area, and to issue a warning when abnormal gathering behavior or abnormal dispersion behavior exists.
[0100] Exemplarily, the behavior warning unit 14 can determine whether the personnel gathering index and the personnel dispersion index reach the corresponding thresholds (the thresholds of the personnel gathering index in different sub-areas may be different, and the personnel gathering index is usually between 10-50, but not limited to this; and the threshold of the personnel dispersion index is only between 0 and 1, and does not include endpoint values). If the personnel gathering index reaches the corresponding threshold, it is considered that abnormal gathering behavior exists; if the personnel dispersion index reaches the corresponding threshold, it is considered that abnormal dispersion behavior exists.
[0101] When there is abnormal aggregate behavior or abnormal discrete behavior, the behavior warning unit 14 can issue a warning.
[0102] In summary, the embodiment of the present application provides a target behavior warning method based on artificial intelligence, which obtains a set of video images obtained by monitoring the target area, pre-processes each frame of the video image, obtains a corresponding input image, and inputs the input image frame by frame into a preset target behavior warning model 10, and the target behavior warning model 10 includes a target detection unit 11, a trajectory detection unit 12, an index calculation unit 13 and a behavior warning unit 14. The target detection unit 11 performs target detection on each frame of the input image to determine the position of each person in the input image; the trajectory detection unit 12 tracks each person based on the position of the person to determine the movement trajectory of each person; the index calculation unit 13 determines the personnel aggregation index and personnel dispersion index of each sub-area in the target area based on the movement trajectory of each person; the behavior warning unit 14 determines whether there is abnormal aggregation behavior or abnormal dispersion behavior based on the personnel aggregation index and personnel dispersion index of each sub-area, and issues a warning when abnormal aggregation behavior or abnormal dispersion behavior exists. In this way, artificial intelligence technology can be used to monitor abnormalities in the target area, and abnormal aggregation behavior or abnormal discrete behavior can be discovered in time or even predicted in advance (mainly abnormal aggregation situations can be predicted in advance to a certain extent), so as to issue timely warnings. Personnel can quickly arrive at the scene from fixed points to maintain order, greatly improving the timeliness of responding to abnormal situations and effectively reducing risks and costs (a small number of personnel are stationed at fixed points for order, and there is no need for a large number of mobile personnel to patrol).
[0103] Based on the YOLOv5 model, the positioning regression function (understood as the loss function) of the YOLOv5 model in target detection is improved, and the bounding box loss function L of the feature map is calculated. box When calculating the target confidence loss function L of the feature map, the measurement method of the loss function is improved, and a penalty term is introduced to better handle the situation where the prediction box and the target box overlap but the center point is far away, so that the YOLOv5 model can be more sensitive to this situation and more accurately identify the person in the image. obj When using the classic binary cross entropy classification loss function (corresponding to the first term in the formula), the weight coefficient (corresponding to the second term in the formula) is introduced into the loss term of the prediction box using the IoU and center point distance of the sample. It can be adjusted according to the IoU and center point distance of the sample to adapt to the situation where the IoU is low or the center point distance is small. By introducing the classification probability into Focal Loss, the loss weight of easy-to-classify samples can be reduced, so that the YOLOv5 model pays more attention to difficult-to-classify samples and improves the classification effect. When calculating the target category confidence loss function L of the feature map cls When , the cross entropy loss function is used and the sigmoid function is introduced for adjustment, which is conducive to accelerating the convergence of the model.
[0104] By designing a calculation method for the gathering index of people in each sub-area, we can effectively measure the speed and degree of gathering of people in the sub-area, so as to timely discover and even predict abnormal gathering behavior to a certain extent. By designing a calculation method for the dispersion index of people in each sub-area, we can discover the rapid dispersion characterized by abnormal dispersion behavior, and can take into account the number of people in the sub-area, effectively reducing the false alarm rate caused by the rapid departure of a small number of people, thereby improving the accuracy of early warning.
[0105] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0106] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A target behavior early warning method based on artificial intelligence, characterized in that: include: Acquire a video image set obtained by monitoring the target area, wherein the video image set includes multiple frames of video images; Preprocess each frame of video image to obtain the corresponding input image; Input the input image frame by frame into the preset target behavior warning model, determine the position of the person based on the input image through the target behavior warning model, further detect the movement trajectory of the person, and issue a warning when abnormal aggregation behavior or abnormal discrete behavior is determined in the target area based on the movement trajectory; The target behavior warning model includes a target detection unit, a trajectory detection unit, an index calculation unit and a behavior warning unit. The target detection unit is used to perform target detection on each frame of input image to determine the position of each person in the input image; The trajectory detection unit is used to track each person based on their position and determine the movement trajectory of each person; The index calculation unit is used to determine the personnel gathering index and personnel dispersion index of each sub-area in the target area based on the movement trajectory of each person, wherein the target area includes multiple sub-areas; The behavior warning unit is used to determine whether there is abnormal gathering behavior or abnormal dispersion behavior based on the personnel gathering index and personnel dispersion index of each sub-area, and to issue a warning when abnormal gathering behavior or abnormal dispersion behavior exists; The target detection unit adopts the YOLOv5 model and improves the positioning regression function of the YOLOv5 model in target detection. The improved positioning regression function is: Among them, L detect is the positioning regression function of target detection, N is the number of feature maps for target detection, i represents the i-th feature map of target detection, L box represents the bounding box loss function of the feature map, L obj represents the target confidence loss function of the feature map, L cls represents the target category confidence loss function of the feature map, λ1, λ2, and λ3 are the loss functions L box , L obj , L cls The weight parameter, B i is the number of target boxes that match the prior anchor points in the i-th feature map, s i ×s i Indicates the number of grids divided by the i-th feature map; Loss Function satisfy: Among them, IOU j ∈(0,1) represents B in the YOLOv5 model i is the prediction box evaluation index corresponding to the jth target box matched to the prior anchor point in the i-th feature map, which is used to measure the overlap between the prediction box and the jth target box matched to the prior anchor point. j and They represent the jth predicted box in the i-th feature map and the jth target box matched to the prior anchor point, respectively. Represents the predicted box b j The center point and target box The Euclidean distance between the center points, c j is the predicted box b j With target box The diagonal length of the minimum closed module, is the penalty coefficient, is the predicted box b j The center point and target box The Manhattan distance between the center points of j is the intermediate variable, v j Represents the predicted box b j With target box The aspect ratio similarity is measured using KL divergence. Represents the predicted box b j The aspect ratio probability density, w j Represents the predicted box b j The width, h j Represents the predicted box b j Height, Represents the target box The probability density of the aspect ratio is, Represents the target box The width of Represents the target box The height of , KL represents the divergence, and σ is the sigmoid activation function; Loss Function satisfy: Among them, y k is the true label of the kth grid in the ith feature map, x k is the predicted label of the kth grid in the ith feature map, b k is the kth grid in the ith feature map, For grid b k The target box that matches the prior anchor point corresponding to the i-th feature map, For grid b k With target box The distance between the center points; Loss Function satisfy: Among them, y l is the true label of the lth target box matched to the prior anchor in the i-th feature map, x l is the predicted label of the lth predicted box in the ith feature map.
2. The target behavior early warning method based on artificial intelligence according to claim 1 is characterized in that: The trajectory detection unit uses DeepSORT to achieve multi-target tracking.
3. The target behavior early warning method based on artificial intelligence according to claim 1 is characterized in that: The index calculation unit is specifically used for: Based on the movement trajectory of each person output by the trajectory detection unit, determine the current destination area and the current departure area of each person; Based on the current destination area and the current departure area of each person, the personnel aggregation index and the personnel dispersion index of each sub-area are determined.
4. The target behavior early warning method based on artificial intelligence according to claim 3 is characterized in that: The index calculation unit is specifically used for: The following formula is used to calculate the population concentration index of each sub-region: in, represents the updated personnel gathering index of sub-region s, represents the personnel gathering index of sub-region s before updating, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 Represents the number of people in sub-area s.
5. The target behavior early warning method based on artificial intelligence according to claim 3 is characterized in that: The index calculation unit is specifically used for: The personnel dispersion index of each sub-region is calculated using the following formula: in, represents the population dispersion index of sub-area s, c s Indicates the number of people who use sub-area s as the current destination area, d s Indicates the number of people who have left the sub-area s at this time, c s0 Represents the number of people in sub-area s.
Citation Information
Patent Citations
Method and system for monitoring abnormal behaviors of personnel in lightweight government affair hall
CN117037269A