Monitoring video data abnormal behavior detection method based on deep learning
Through a deep learning-based monitoring video data abnormal behavior detection method, combined with the difference in smoke concentration data and pixel value, the problem of insufficient accuracy in detecting abnormal behavior in the smoking and fire-free places in the prior art is solved, and higher detection accuracy and anti-interference ability are achieved.
Patent Information
- Application Number
- CN202510349471.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art has the accuracy problem in detecting abnormal behavior of personnel in smoking and fire-free places, especially in the occlusion state or in a specific angle, and interference factors such as high temperature steam affect the detection accuracy.
The abnormal behavior detection method of monitoring video data based on deep learning is adopted. By obtaining the video images of the character in the monitoring video in real time, detecting the position information of the person's front, obtaining the movement status of the head and hands, identifying suspected abnormal behaviors, and determining the detection results based on the smoke concentration data and the difference in pixel values.
It improves the accuracy of detecting abnormal behaviors in occlusion states or specific angles, reduces the impact of interference such as high-temperature steam, and enhances the accuracy of monitoring abnormal behaviors of people in smoking and fire-free places.
Smart Images

Figure CN120014711A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data anomaly detection, and in particular to a method for detecting abnormal behavior in surveillance video data based on deep learning. Background Art
[0002] As key areas for safety management, no-smoking and no-fire places often involve the production, storage and operation of flammable and explosive substances. For example, chemical plants, oil and gas stations, storage centers and other areas strictly prohibit fireworks. These areas not only require the use of a large number of high-temperature and high-pressure equipment during the production process, but also contain a variety of flammable and explosive chemicals, which may cause fires or explosions if not handled with care. Therefore, strengthening production safety management and strictly monitoring unsafe behaviors (such as smoking and open flame operations) are key links to ensure safety.
[0003] When detecting unsafe behaviors of people in no-smoking and no-fire places, the existing technology usually uses a combination of surveillance cameras and regularized detection to detect abnormal behaviors, and uses a target detection algorithm (YOLO) to identify abnormal behaviors. The algorithm identifies abnormal actions under surveillance cameras after annotating abnormal images and normal images. However, the algorithm has certain limitations. For example, when an employee smokes in an obstructed state or at a specific angle, the detection system may find it difficult to capture the complete smoking action or trajectory, resulting in missed reports. At the same time, there may be interference factors such as high-temperature steam, smoke or dust generated by equipment operation in no-smoking and no-fire places, which may lead to inaccurate detection of abnormal behaviors of people under surveillance cameras, thereby affecting the accuracy of the detection system.
[0004] Therefore, how to accurately identify the abnormal behavior of people in no-smoking and no-fire places has become an urgent problem to be solved. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides a method for detecting abnormal behavior in surveillance video data based on deep learning, so as to solve the problem of how to accurately identify the abnormal behavior of people in no-smoking and no-fire places.
[0006] An embodiment of the present invention provides a method for detecting abnormal behavior in surveillance video data based on deep learning, the method comprising the following steps:
[0007] For any person appearing in the surveillance video in the target place, a person video image of the person is obtained in real time, and the person front position information in the person video image is detected. When the person front position information cannot be detected, the person video image is marked as a target image, and multiple consecutive frames of the target image are obtained until the number of target images reaches a preset number threshold, thereby obtaining a target image set;
[0008] Obtaining head position information and hand position information of any person in each frame of the target image set, obtaining a motion coefficient for characterizing the overall motion state of the head and hand of any person according to the head position information and hand position information of any person in each frame of the target image set, and identifying whether any person in the target image set has suspected abnormal behavior according to the motion coefficient;
[0009] If any of the persons in the target image set has suspected abnormal behavior, the smoke concentration data in the target place and the pixel values of the pixel points in each frame of the target image in the target image set are obtained, and the abnormal behavior detection result of any of the persons is determined based on the smoke concentration data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set.
[0010] Preferably, after acquiring the character video image of any character in real time and detecting the character front position information in the character video image, the method further includes:
[0011] When the front position information of a person is detected, the target detection model is used to identify abnormal behavior of any person in the person video image.
[0012] Preferably, the step of obtaining the motion coefficient for characterizing the overall motion state of the head and hand of any person according to the head position information and hand position information of any person in each frame of the target image in the target image set includes:
[0013] For any target image in the target image set, according to the head position information and hand position information of any person in the target image, the relative angle and relative distance between the head and the hand of any person in the target image are obtained, and the relative angle and relative distance between the head and the hand of any person in each frame of the target image in the target image set are respectively obtained, and a relative angle data sequence and a relative distance data sequence are correspondingly obtained;
[0014] Performing second-order difference processing on the relative angle data sequence to obtain a corresponding second-order difference data sequence, obtaining the mean of the second-order difference data sequence, recorded as a first mean, respectively calculating the perfect square difference between each second-order difference data in the second-order difference data sequence and the first mean, correspondingly obtaining the mean of the perfect square difference, recorded as a second mean, obtaining a first ratio of the second mean to a preset stationarity parameter, substituting the opposite of the first ratio into an exponential function with a natural constant as the base, and obtaining a first exponential function result;
[0015] For any relative distance except the first relative distance in the relative distance data sequence, obtain a first difference between the any relative distance and its previous relative distance, obtain a first product of the first difference and a preset distance scaling parameter, perform cosine processing on the first product to obtain a cosine value, obtain a Gaussian function value of the any relative distance, perform weighted summation processing on the cosine value and the Gaussian function value to obtain a weighted summation result of the any relative distance, respectively obtain a weighted summation result of each relative distance in the relative distance data sequence, and obtain a corresponding mean of the weighted summation results, which is recorded as a third mean;
[0016] A weighted sum is performed on the first exponential function result and the third mean value to obtain a motion coefficient for characterizing the overall motion state of the head and hands of any one of the characters.
[0017] Preferably, the step of identifying, based on the motion coefficient, whether any of the characters in the target image set has suspected abnormal behavior includes:
[0018] Setting a first motion coefficient threshold and a second motion coefficient threshold, wherein the first motion coefficient threshold is less than the second motion coefficient threshold;
[0019] If the motion coefficient is greater than or equal to the first motion coefficient threshold and less than or equal to the second motion coefficient threshold, it is confirmed that any of the characters has suspected abnormal behavior in the target image set.
[0020] Preferably, the step of obtaining the smoke concentration data in the target location and the pixel values of the pixels in each frame of the target image in the target image set includes:
[0021] Obtaining a sampling period corresponding to the target image set, recorded as a target sampling period, and obtaining a smoke concentration data sequence in the target place during the target sampling period;
[0022] By using a preset smoke detection algorithm, the smoke region of the head of any person in each frame of the target image in the target image set is obtained, and the pixel value of the pixel point in each head smoke region is correspondingly obtained.
[0023] Preferably, the abnormal behavior detection result of any person is determined according to the smoke concentration data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set, including:
[0024] Obtaining a smoke significance score within the target sampling period according to a difference in smoke concentration data in the smoke concentration data sequence;
[0025] According to the pixel value difference of each pixel point in the head smoke area in the target image set, counting the number of target images with smoke diffusion features in the target image set;
[0026] If the smoke prominence score is greater than or equal to a preset smoke prominence score threshold, and the number is greater than or equal to a first preset number, it is determined that any of the characters has abnormal behavior.
[0027] Preferably, obtaining the smoke significance score within the target sampling period according to the difference of smoke concentration data in the smoke concentration data sequence includes:
[0028] Perform first-order difference processing on the smoke concentration data sequence to obtain a first-order difference data sequence, obtain the mean of the first-order difference data sequence, and record it as the fourth mean. For any first-order difference data in the first-order difference data sequence, obtain the perfect square difference between the arbitrary first-order difference data and the fourth mean, obtain the addition result of the perfect square difference between a constant 1 and a first preset multiple, and obtain the inverse of the addition result, which is recorded as the smoke data score corresponding to the arbitrary first-order difference data. Obtain the smoke data score corresponding to each first-order difference data in the first-order difference data sequence, and obtain the mean of the smoke data scores. Use the mean of the smoke data scores as the smoke significance score within the target sampling period.
[0029] Preferably, counting the number of target images with smoke diffusion features in the target image set according to the pixel value difference of each pixel point in the head smoke area in the target image set includes:
[0030] For any target image in the target image set except the first frame target image, obtain the overlapping area between the head smoke area in the any target image and the head smoke area in the previous target image of the any target image, respectively calculate the difference between the corresponding pixel values of each pixel point in the overlapping area in the any target image and in the previous target image, and obtain the mean of the difference in the overlapping area, which is recorded as the fifth mean, and perform normalization processing on the fifth mean to obtain the smoke diffusion degree of the any target image;
[0031] Obtain the number of pixel points in the head smoke area in any of the target images, recorded as a first number, obtain the number of pixel points in the head smoke area in a target image previous to the any of the target images, recorded as a second number, calculate a second difference between the first number and the second number, substitute the opposite of the second difference of the second preset multiple into an exponential function with a natural constant as a base to obtain a second exponential function result, calculate the difference between a constant 1 and the result of the second exponential function, and obtain the smoke diffusion area of any of the target images;
[0032] Performing weighted summation on the smoke diffusion degree and the smoke diffusion area of any target image to obtain the smoke diffusion coefficient of any target image;
[0033] The smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image in the target image set except the first target image are obtained, and the number of target images with smoke diffusion characteristics in the target image set is counted according to the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image in the target image set except the first target image.
[0034] Preferably, counting the number of target images with smoke diffusion features in the target image set according to the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image frame except the first target image frame in the target image set includes:
[0035] If the smoke diffusion degree and the smoke diffusion area of any target image are greater than the constant 0, and the smoke diffusion coefficient of any target image is greater than a preset smoke diffusion coefficient threshold, it is determined that any target image has smoke diffusion features, and the target images with smoke diffusion features in the target image set are screened to obtain the number of target images with smoke diffusion features in the target image set.
[0036] Preferably, after determining the abnormal behavior detection result of any person according to the smoke density data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set, the method further includes:
[0037] Acquire in real time a second preset number of frames of new person video images consecutively following the target image set; if the second preset number of frames of new person video images are all target images, add the second preset number of frames of new person video images to the target image set, remove the target images before the second preset number of frames in the target image set, and obtain a new target image set.
[0038] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0039] The present invention targets any person appearing in a surveillance video in a target place, obtains a person video image of the any person in real time, and detects the person's front position information in the person video image; when the person's front position information cannot be detected, marks the person video image as a target image, obtains multiple consecutive frames of the target image, until the number of target images reaches a preset number threshold, and obtains a target image set; obtains the head position information and hand position information of the any person in each frame of the target image in the target image set, obtains a motion coefficient for characterizing the overall motion state of the head and hand of the any person according to the head position information and hand position information of the any person in each frame of the target image in the target image set, and identifies whether the any person has suspected abnormal behavior in the target image set according to the motion coefficient; if the any person has suspected abnormal behavior in the target image set, obtains the smoke concentration data in the target place and the pixel value of the pixel point in each frame of the target image in the target image set, and determines the abnormal behavior detection result of the any person according to the smoke concentration data in the target place and the pixel value difference of the pixel point in each frame of the target image in the target image set. Among them, when the front position information of a person cannot be detected, the motion coefficient used to characterize the overall motion state of the head and hands of any person is obtained through the head position information and hand position information of any person in each frame of the target image, and preliminary identification is performed, so that the person can be accurately identified even when performing abnormal behaviors in an occluded state or at a specific angle; when any person is identified to have suspected abnormal behavior, the abnormal behavior of any person is detected based on the smoke concentration data in the target place and the pixel value difference of the pixel points in each frame of the target image, thereby reducing the interference of high-temperature steam and the like on the identification of abnormal behaviors, making the monitoring of abnormal behaviors of people in the target place, including but not limited to smoking, more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0041] Figure 1 This is a flowchart of a method for detecting abnormal behavior in surveillance video data based on deep learning provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0042] Embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, but should not be understood as limiting the present disclosure.
[0043] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0044] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.
[0045] See also Figure 1 , is a flow chart of a method for detecting abnormal behavior in surveillance video data based on deep learning provided in Embodiment 1 of the present invention, such as Figure 1 As shown, the method may include:
[0046] Step S101, for any person appearing in the surveillance video in the target place, a person video image of the person is obtained in real time, and the person's front position information in the person video image is detected. When the person's front position information cannot be detected, the person video image is marked as a target image, and multiple consecutive frames of the target image are obtained until the number of target images reaches a preset threshold, thereby obtaining a target image set.
[0047] No smoking and no fire places include chemical plants, oil and gas stations, storage centers, etc. Abnormal behaviors of people in no smoking and no fire places include smoking, open flame operation, etc. This embodiment takes the chemical plant as the target place and the smoking behavior of people in the target place as an example to detect the smoking behavior of people in the chemical plant.
[0048] Deploy surveillance cameras in different areas of the chemical plant to ensure that the collected data can cover various scenarios. For example, focus on production workshops, equipment maintenance areas, boiler rooms, and rest areas where people may gather. The surveillance camera monitors at a frame rate of 30FPS, which is not limited here. According to the specific implementation scenario settings, any person appearing in the surveillance video in the chemical plant is recorded as the target person, and the target person's video image is obtained in real time. The spatiotemporal transformer model is used to detect the target person's video image obtained in real time, and detect the person's front position information in the target person's video image obtained in real time.
[0049] Among them, the spatiotemporal transformer model is a model that is learned and trained in advance. Historical video data in the chemical plant is collected, and representative sample data is selected from the historical video data. The labels are refined through manual labeling and video analysis tools to obtain a trained character position information data set, where the labels include the front position and the back position. Then, the spatiotemporal transformer model is used for learning and training to obtain a trained spatiotemporal transformer model. Since the spatiotemporal transformer model is a prior art, it will not be repeated here.
[0050] When the front position information of the person is detected, the target detection model is used to identify the abnormal smoking behavior of any person in the video image of the person. Specifically, the YOLO model is trained through the trained smoking behavior data set to obtain a trained YOLO model (that is, a trained target detection model). Since the training of the YOLO model is a prior art, it will not be described in detail here. The video image of the person with the front position information of the person is detected is sent to the trained target detection model. The trained target detection model analyzes the key parts (head, hand) in the input video image of the person, outputs the target frame with the key parts, matches the target frame with the trained smoking behavior data set, detects whether there is smoking behavior in the target frame (detects the hand close to the mouth or the bright spot of the cigarette butt), if the match is successful, it is determined that the target person has abnormal smoking behavior, if the match is unsuccessful, it is determined that the target person does not have abnormal smoking behavior.
[0051] When the front position information of a person cannot be detected, it is difficult to capture the complete smoking action or trajectory using the trained target detection model (YOLO model), which may result in missed detections and affect the detection results of the target person's smoking behavior. Therefore, when the front position information of a person cannot be detected, it is necessary to further analyze the key parts of the target person so as to accurately detect the target person's abnormal smoking behavior.
[0052] Since smoking behavior is usually a periodic and repetitive action, that is, the hands move back and forth close to and away from the head, the target person's abnormal smoking behavior can be further identified based on the movement trajectory of the key parts of the target person.
[0053] In this embodiment, the period of the target person's smoking action is set to 5 seconds, which is not limited here and can be set according to the specific implementation scenario. When the front position information of the person cannot be detected, the real-time acquired video image of the target person is marked as the target image, and multiple consecutive frames of target images are acquired until the number of target images reaches a preset threshold (i.e., 150 frames of target images in the period of one smoking action). The 150 consecutive frames of target images are combined into a target image set for analyzing whether the target person has abnormal smoking behavior.
[0054] At this point, the target image set is obtained.
[0055] Step S102, obtaining the head position information and hand position information of any person in each frame of the target image in the target image set, obtaining a motion coefficient for characterizing the overall motion state of the head and hand of any person based on the head position information and hand position information of any person in each frame of the target image in the target image set, and identifying whether any person in the target image set has suspected abnormal behavior based on the motion coefficient.
[0056] After obtaining the target image set, it is necessary to further identify the target person's abnormal smoking behavior based on the motion trajectory of the key parts of the target person. Since the key parts directly related to the target person's smoking behavior are the target person's head and hands, the skeleton key point detection technology can be used to obtain the target person's head position information and hand position information in each frame of the target image set. The skeleton key point detection technology belongs to the existing technology and will not be repeated here. Then, based on the target person's head position information and hand position information in each frame of the target image set, the motion coefficient used to characterize the overall motion state of the target person's head and hands is obtained, and then the motion coefficient is used to identify the target person's abnormal smoking behavior in the target image set.
[0057] Skeleton key point detection technology can identify the key point positions of the target person, such as the head, hands, etc., from each frame of the target image, and generate a key point coordinate set, which includes the head position coordinates and hand position coordinates of the target person. Therefore, the target person's motion trajectory can be analyzed by analyzing the coordinate changes of the key parts of the target person in combination with the dynamic characteristics in the time dimension. Since the target person's head and hand motion state changes during the smoking process are: the relative distance between the hand and the head is periodically enlarged and reduced, and the relative angle between the hand and the head is also periodically enlarged and reduced. Therefore, the hand position coordinates and head position coordinates of the target person in each frame of the target image can be obtained, and then the motion coefficient used to characterize the overall motion state of the head and hand of the target person is obtained according to the change characteristics of the relative distance and relative angle between the hand position coordinates and the head position coordinates of the target person in each frame of the target image, and then the motion coefficient of the target person is used to judge whether the motion trajectory of the target person is consistent with the motion trajectory of the smoking behavior.
[0058] Among them, according to the change characteristics of the relative distance and relative angle between the hand position coordinates and the head position coordinates of the target person in each frame of the target image, the method for obtaining the motion coefficient used to characterize the overall motion state of the head and hands of the target person is as follows:
[0059] (1) For any target image in the target image set, based on the head position information and hand position information of any person in the target image, the relative angle and relative distance between the head and hand of any person in the target image are obtained, and the relative angle and relative distance between the head and hand of any person in each frame of the target image in the target image set are respectively obtained, and a relative angle data sequence and a relative distance data sequence are correspondingly obtained.
[0060] In one embodiment, taking the i-th frame target image in the target image set as an example, a coordinate system is constructed in the i-th frame target image using the skeleton key point detection technology to obtain the key point coordinates of the head position information and the hand position information of the target person in the i-th frame target image, which are respectively recorded as (x head,i ,y head,i ) and (x hand,i ,y hand,i ), the formula for calculating the relative angle and relative distance between the head and hand of the target person in the i-th frame target image is:
[0061]
[0062] Among them, θ i is the relative angle between the head and hand of the target person in the i-th frame target image; x hand,i is the horizontal coordinate of the target person's hand position information in the i-th frame target image; head,i y is the horizontal coordinate of the head position information of the target person in the i-th frame target image; hand,i y is the ordinate of the target person's hand position information in the i-th frame target image; head,i is the ordinate of the head position information of the target person in the target image of the i-th frame; d i is the relative distance between the head and hand of the target person in the i-th frame target image; arctan() is the inverse tangent function; i is the sequence number of the target image in the target image set.
[0063] It should be noted that the relative angle and relative distance belong to the prior art and will not be described in detail here.
[0064] According to the method for obtaining the relative angle and relative distance between the head and hand of the target person in the i-th frame target image, the relative angle and relative distance between the head and hand of the target person in each frame target image of the target image set are respectively obtained, and the relative angle between the head and hand of the target person in each frame target image of the target image set is composed into a relative angle data sequence, and the relative distance between the head and hand of the target person in each frame target image of the target image set is composed into a relative distance data sequence.
[0065] (2) Obtaining motion coefficients for characterizing the overall motion state of the head and hands of any of the characters based on the relative angle data sequence and the relative distance data sequence.
[0066] Specifically, a second-order difference processing is performed on the relative angle data sequence to obtain a corresponding second-order difference data sequence, a mean of the second-order difference data sequence is obtained, recorded as a first mean, and the perfect square difference between each second-order difference data in the second-order difference data sequence and the first mean is respectively calculated to obtain a corresponding mean of the perfect square difference, recorded as a second mean, and a first ratio of the second mean to a preset stationarity parameter is obtained, and the opposite of the first ratio is substituted into an exponential function with a natural constant as the base to obtain a first exponential function result;
[0067] For any relative distance except the first relative distance in the relative distance data sequence, obtain a first difference between the any relative distance and its previous relative distance, obtain a first product of the first difference and a preset distance scaling parameter, perform cosine processing on the first product to obtain a cosine value, obtain a Gaussian function value of the any relative distance, perform weighted summation processing on the cosine value and the Gaussian function value to obtain a weighted summation result of the any relative distance, respectively obtain a weighted summation result of each relative distance in the relative distance data sequence, and obtain a corresponding mean of the weighted summation results, which is recorded as a third mean;
[0068] A weighted sum is performed on the first exponential function result and the third mean value to obtain a motion coefficient for characterizing the overall motion state of the head and hands of any one of the characters.
[0069] In one embodiment, the calculation expression of the motion coefficient used to characterize the overall motion state of the target person's head and hands is:
[0070]
[0071] Where C is the motion coefficient of the target person in the target image set; θ ′ j is the jth second-order difference data in the second-order difference data sequence; j is the sequence number of the second-order difference data in the second-order difference data sequence; V is the number of second-order difference data in the second-order difference data sequence; F is the mean of the second-order difference data sequence; τ is the preset stability parameter; i is the sequence number of the target image in the target image set; N is the number of target images in the target image set; e is a natural constant; Δd i is the difference between the relative distance in the target image of the i-th frame and the relative distance in the target image of the i-1-th frame in the target image set; k is the preset distance scaling parameter; cos() is the cosine function; d iis the corresponding relative distance in the target image of the i-th frame in the target image set; μ d is the mean of the relative distance data sequence; d is the preset parameter; α is the weight coefficient of relative angle change; β is the weight coefficient of relative distance change; γ is the weight coefficient of periodic distance change; δ is the weight coefficient of relative distance; the reference values of α and β are 0.4 and 0.6 respectively, the reference values of γ and δ are 0.5, the reference value of τ is 4, and the reference value of k is π is the ratio of the circumference of a circle, λ is the ratio of the circumference of a circle d The reference value of is twice the standard deviation of the relative distance function sequence, which is not limited here and can be set according to the specific implementation scenario.
[0072] It should be noted that is the Gaussian function value corresponding to the relative distance in the target image of the i-th frame. The Gaussian function belongs to the prior art and will not be described here; Indicates the fluctuation range of the relative angle between the target person's head and hands. The larger it is, the greater the fluctuation range of the relative angle between the target person's head and hands is, and the greater the target person's motion coefficient is. The smaller it is, the smaller the fluctuation range of the relative angle between the target person's head and hand is, and the smaller the target person's motion coefficient is. Since smoking behavior has a fixed motion amplitude, that is, the relative angle between the head and hand of the target person with smoking behavior has a fixed fluctuation range, if the fluctuation range of the relative angle between the head and hand of the target person is too large or too small, it does not conform to the motion trajectory of smoking behavior, that is, if the motion coefficient is too large or too small, it does not conform to the motion trajectory of smoking behavior; Indicates the periodicity of the movement of the relative distance between the target person's head and hands. The larger it is, the greater the cycle of the change in the relative distance between the target person's head and hands, and the greater the target person's motion coefficient. The smaller it is, the smaller the cycle of the change of the relative distance between the target person's head and hand is, and the smaller the motion coefficient of the target person is. Since smoking behavior has a fixed cycle, that is, the change of the relative distance between the head and hand of the target person with smoking behavior has a fixed cycle, therefore, if the cycle of the change of the relative distance between the head and hand of the target person is too large or too small, it does not conform to the motion trajectory of the smoking behavior, that is, if the motion coefficient is too large or too small, it does not conform to the motion trajectory of the smoking behavior.
[0073] Furthermore, a motion coefficient threshold is set, and whether the target person has suspected abnormal behavior is identified based on the motion coefficient threshold.
[0074] Specifically, a first motion coefficient threshold and a second motion coefficient threshold are set, and the first motion coefficient threshold is smaller than the second motion coefficient threshold. If the motion coefficient is greater than or equal to the first motion coefficient threshold and less than or equal to the second motion coefficient threshold, it is confirmed that any of the characters in the target image set has suspected abnormal behavior.
[0075] In one embodiment, the first motion coefficient threshold is set to 0.5, and the second motion coefficient threshold is set to 0.9. There is no limitation here and it can be set according to the specific implementation scenario. When the motion coefficient is less than 0.5, it means that the fluctuation range of the relative angle between the head and hand of the target person is too small, or the cycle of the relative distance change between the head and hand of the target person is too small, that is, the relative change between the head and hand of the target person is relatively drastic, which does not conform to the motion trajectory of smoking behavior, and it is confirmed that the target person does not have any smoking behavior; when the motion coefficient is greater than 0.9, it means that the fluctuation range of the relative angle between the head and hand of the target person is too large, or the cycle of the relative distance change between the head and hand of the target person is too large, that is, the relative change between the head and hand of the target person is relatively weak, which does not conform to the motion trajectory of smoking behavior, and it is confirmed that the target person does not have any smoking behavior; when the motion coefficient is greater than or equal to 0.5 and less than or equal to 0.9, it means that the relative change between the head and hand of the target person conforms to the motion trajectory of smoking behavior, and it is confirmed that the target person has suspected smoking behavior.
[0076] At this point, it is confirmed whether the target person has suspected smoking behavior in the target image set.
[0077] Step S103, if any of the persons in the target image set has suspected abnormal behavior, then the smoke concentration data in the target place and the pixel values of the pixel points in each frame of the target image in the target image set are obtained, and the abnormal behavior detection result of any of the persons is determined based on the smoke concentration data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set.
[0078] When the motion coefficient is greater than or equal to 0.5 and less than or equal to 0.9, it means that the relative change between the target person's head and hands is similar to the motion trajectory of smoking behavior, confirming that the target person has suspected smoking behavior. However, since the motion trajectories of drinking water, wiping sweat, etc. are also similar to the motion trajectory of smoking behavior, it cannot be directly determined that the target person has smoking behavior. Therefore, further analysis is needed to determine whether the target person has smoking behavior.
[0079] Since smoke will be generated in the head area of the smoker during the smoking process, and infrared sensors are deployed in the monitoring area of the chemical plant, which can collect smoke concentration data in the chemical plant, it is possible to judge the smoke generation in the head area of the target person based on the smoke concentration data in the chemical plant, and then judge whether the target person has smoking behavior. However, in a chemical plant, not only smoking will generate smoke, but there may also be high-temperature steam in the chemical plant, which will affect the smoke concentration data collected by the smoke sensor. Therefore, the smoking behavior of the target person cannot be accurately detected based only on the smoke concentration data in the chemical plant. In the process of smoking, not only smoke will be generated in the head area of the target person, but the generated smoke also has diffusion characteristics. Therefore, the smoke concentration data of the target person's head area and the diffusion characteristics of the smoke can be combined to accurately judge whether the target person has smoking behavior. In the target image, the pixel values of different pixels will change with the change of smoke concentration. After the smoke diffuses, the smoke concentration becomes smaller, and the pixel values of the pixels in the target image will become larger. Therefore, the sampling period corresponding to the target image set is obtained, recorded as the target sampling period, and the smoke concentration data sequence in the chemical plant during the target sampling period is obtained; and the preset smoke detection algorithm is used to obtain the head smoke area of any person in each frame of the target image in the target image set, and the corresponding pixel value of each pixel in the head smoke area is obtained for subsequent analysis of whether the target person has smoking behavior.
[0080] In one embodiment, assuming that the sampling period corresponding to the target image set is from 2:00:00 pm to 2:00:05 pm, then during the sampling period in the target image set, the smoke concentration data of the chemical plant is obtained at a collection frequency of once per second, and the smoke concentration data in the sampling period corresponding to the target image set are combined into a smoke concentration data sequence, which is not limited here and can be set according to the specific implementation scenario; a smoke detection algorithm is used to obtain the smoke area of the target person's head in each frame of the target image in the target image set, and the pixel values of the pixel points of the smoke area of the target person's head in each frame of the target image are obtained, wherein the smoke detection algorithm belongs to the prior art and will not be repeated here.
[0081] Furthermore, according to the smoke concentration data in the chemical plant and the pixel value difference of each pixel in each frame of the target image in the target image set, the method for determining whether the target person has a smoking behavior is as follows:
[0082] (1) Obtaining a smoke significance score within the target sampling period based on the difference in smoke concentration data in the smoke concentration data sequence.
[0083] Specifically, a first-order difference processing is performed on the smoke concentration data sequence to obtain a first-order difference data sequence, and the mean of the first-order difference data sequence is obtained, which is recorded as a fourth mean. For any first-order difference data in the first-order difference data sequence, the perfect square difference between the arbitrary first-order difference data and the fourth mean is obtained, and the addition result of the perfect square difference between a constant 1 and a first preset multiple is obtained, and the inverse of the addition result is obtained, which is recorded as a smoke data score corresponding to the arbitrary first-order difference data. The smoke data score corresponding to each first-order difference data in the first-order difference data sequence is obtained, and the mean of the smoke data scores is obtained, and the mean of the smoke data scores is used as the smoke significance score within the target sampling period.
[0084] In one embodiment, the formula for calculating the smoke significance score within the target sampling period is:
[0085]
[0086] Where S is the smoke significance score during the target sampling period; H t is the tth first-order difference data in the first-order difference data sequence; t is the sequence number of the first-order difference data in the first-order difference data sequence; μ H is the mean of the first-order difference data sequence; ρ is the first preset multiple, the reference value is 2, there is no restriction here, and it can be set according to the specific implementation scenario; T is the number of first-order difference data in the first-order difference data sequence; 1 is a constant.
[0087] It should be noted that due to the diffusion characteristics of smoke, the smoke concentration will vary at each sampling time. Therefore, H t ―μ H The larger it is, the greater the difference between adjacent smoke concentration data in the smoke concentration data sequence is, the more it conforms to the characteristics of smoke, and the greater the smoke significance score is.
[0088] (2) According to the pixel value difference of each pixel point in the head smoke area in the target image set, the number of target images with smoke diffusion features in the target image set is counted.
[0089] aFor any target image in the target image set except the first frame target image, obtain the overlapping area of the head smoke area in the any target image and the head smoke area in the previous target image of the any target image, calculate the difference between the corresponding pixel values of each pixel point in the overlapping area in the any target image and in the previous target image, and obtain the corresponding mean of the difference in the overlapping area, which is recorded as the fifth mean, and perform normalization processing on the fifth mean to obtain the smoke diffusion degree of the any target image.
[0090] In one embodiment, taking the i-th frame of the target image in the target image set as an example, the formula for calculating the smoke diffusion degree of the i-th frame of the target image is:
[0091]
[0092] Where D is the smoke diffusion degree of the target image in the i-th frame; M is the number of pixels in the overlapping area between the target image in the i-th frame and the head smoke area in the target image in the i-1-th frame; I i―1,m is the pixel value of the mth pixel point in the overlapping area in the i-1th frame target image; I i,m is the pixel value of the mth pixel in the overlapping area in the i-th frame of the target image; m is the serial number of the pixel in the overlapping area; i is the serial number of the target image in the target image set; 255 is the maximum grayscale value.
[0093] It should be noted that the greater the difference in pixel values of the pixel points in the overlapping area between the i-th frame target image and the head smoke area in the i-1-th frame target image, the greater the degree of diffusion of the smoke in the i-th frame target image, the greater the degree of smoke diffusion in the i-th frame target image, and the greater the smoke diffusion coefficient of the i-th frame target image.
[0094] b. Obtain the number of pixel points in the head smoke area in any of the target images, recorded as the first number; obtain the number of pixel points in the head smoke area in the previous target image of any of the target images, recorded as the second number; calculate the second difference between the first number and the second number; substitute the opposite of the second difference of the second preset multiple into an exponential function with a natural constant as the base to obtain a second exponential function result; calculate the difference between the constant 1 and the result of the second exponential function to obtain the smoke diffusion area of any of the target images.
[0095] In one embodiment, taking the i-th frame of the target image in the target image set as an example, the formula for calculating the smoke diffusion area of the i-th frame of the target image is:
[0096]
[0097] Wherein, E is the smoke diffusion area of the target image in the i-th frame; e is a natural constant; p is the second preset multiple, and the reference value is 0.01, which is not limited here and can be set according to the specific implementation scenario; R i is the first quantity; R i―1 is the second quantity; i is the sequence number of the target image in the target image set; 1 is a constant.
[0098] It should be noted that, the larger the difference between the first number and the second number, the larger the area of the head smoke region in the i-th frame target image, the larger the smoke diffusion area of the i-th frame target image, and the larger the smoke diffusion coefficient of the i-th frame target image.
[0099] c. Performing weighted summation on the smoke diffusion degree and the smoke diffusion area of any target image to obtain the smoke diffusion coefficient of any target image.
[0100] In one implementation, taking the i-th target image in the target image set as an example, the smoke diffusion coefficient of the i-th target image is calculated:
[0101]
[0102] Among them, A i is the smoke diffusion coefficient of the target image in the i-th frame; D is the smoke diffusion degree of the target image in the i-th frame; E is the smoke diffusion area of the target image in the i-th frame; is the weight coefficient of the smoke diffusion degree of the target image in the i-th frame; ω is the weight coefficient of the smoke diffusion area of the target image in the i-th frame; The reference value of and ω is 0.5, which is not limited here and can be set according to the specific implementation scenario.
[0103] It should be noted that, the greater the smoke diffusion degree of the i-th target image, the more obvious the smoke diffusion characteristics in the i-th target image, and the greater the smoke diffusion coefficient of the i-th target image; the larger the smoke diffusion area of the i-th target image, the larger the area of the head smoke region in the i-th target image, and the greater the smoke diffusion coefficient of the i-th target image.
[0104] d. Obtain the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image frame in the target image set except the first target image frame, and count the number of target images with smoke diffusion characteristics in the target image set based on the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image frame in the target image set except the first target image frame.
[0105] Specifically, if the smoke diffusion degree and the smoke diffusion area of any target image are greater than the constant 0, and the smoke diffusion coefficient of any target image is greater than a preset smoke diffusion coefficient threshold, it is determined that any target image has smoke diffusion features, and the target images with smoke diffusion features in the target image set are screened to obtain the number of target images with smoke diffusion features in the target image set.
[0106] In one embodiment, taking the i-th frame target image in the target image set as an example, the smoke diffusion coefficient threshold is set to 0.2, which is not limited here and can be set according to the specific implementation scenario. If the smoke diffusion coefficient of the i-th frame target image in the target image set is greater than 0.2, and the smoke diffusion degree of the i-th frame target image and the smoke diffusion area of the i-th frame target image are both greater than 0, then it is determined that there is a smoke diffusion feature in the i-th frame target image. According to the method for judging whether there is a smoke diffusion feature in the i-th frame target image, it is judged whether there is a smoke diffusion feature in each frame target image in the target image set, and the target images with smoke diffusion features in the target image set are screened to obtain the number of target images with smoke diffusion features in the target image set.
[0107] (3) If the smoke prominence score is greater than or equal to a preset smoke prominence score threshold, and the number is greater than or equal to a first preset number, it is determined that any of the characters has abnormal smoking behavior.
[0108] In one embodiment, the smoke significance score threshold is set to 0.8, which is not limited here and can be set according to the specific implementation scenario. If the smoke significance score within the target sampling period is greater than or equal to 0.8, and there are 90 frames or more of the target images with smoke diffusion characteristics within the target sampling period, it is determined that the target person has abnormal smoking behavior, which is not limited here and can be set according to the specific implementation scenario.
[0109] After determining abnormal smoking behavior within the target sampling period, obtain in real time new person video images of a second preset number of frames consecutively following the target image set; if the new person video images of the second preset number of frames are all target images, add the new person video images of the second preset number of frames to the target image set, remove the target images of the first second preset number of frames in the target image set, obtain a new target image set, and determine whether there is smoking behavior in the new target image set.
[0110] For example: 30 consecutive frames (i.e. 1 second) of new person video images after the target image set (i.e. 5 seconds) are acquired in real time. If the 30 consecutive frames (i.e. 1 second) of new person video images after the target image set (i.e. 5 seconds) are all target images, the target images of the first 30 frames (i.e. 1 second) in the target image set are removed, and the 30 consecutive frames (i.e. 1 second) of new person video images after the target image set (i.e. 5 seconds) are added to the target image set to obtain a new target image set. According to the above-mentioned method for judging smoking behavior in the target image set, it is judged whether there is smoking behavior in the new target image set; if there is at least one frame of new person video image in the 30 consecutive frames (i.e. 1 second) of new person video images after the target image set (i.e. 5 seconds) in which the front position information of the person is detected (i.e. the new person video image is not the target image), the trained target detection model is directly used to identify the abnormal smoking behavior in the new person video image in which the front position information of the person is detected.
[0111] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.
Claims
1. A method for detecting abnormal behavior in surveillance video data based on deep learning, characterized in that: The method comprises: For any person appearing in the surveillance video in the target place, a person video image of the person is obtained in real time, and the person front position information in the person video image is detected. When the person front position information cannot be detected, the person video image is marked as a target image, and multiple consecutive frames of the target image are obtained until the number of target images reaches a preset number threshold, thereby obtaining a target image set; Obtaining head position information and hand position information of any person in each frame of the target image set, obtaining a motion coefficient for characterizing the overall motion state of the head and hand of any person according to the head position information and hand position information of any person in each frame of the target image set, and identifying whether any person in the target image set has suspected abnormal behavior according to the motion coefficient; If any of the persons in the target image set has suspected abnormal behavior, the smoke concentration data in the target place and the pixel values of the pixel points in each frame of the target image in the target image set are obtained, and the abnormal behavior detection result of any of the persons is determined based on the smoke concentration data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set.
2. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 1 is characterized in that: After acquiring the video image of any person in real time and detecting the front position information of the person in the video image of the person, the method further includes: When the front position information of a person is detected, the target detection model is used to identify abnormal behavior of any person in the person video image.
3. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 1 is characterized in that: The step of obtaining a motion coefficient for characterizing the overall motion state of the head and hand of any person according to the head position information and the hand position information of any person in each frame of the target image set includes: For any target image in the target image set, according to the head position information and hand position information of any person in the target image, the relative angle and relative distance between the head and the hand of any person in the target image are obtained, and the relative angle and relative distance between the head and the hand of any person in each frame of the target image in the target image set are respectively obtained, and a relative angle data sequence and a relative distance data sequence are correspondingly obtained; Performing second-order difference processing on the relative angle data sequence to obtain a corresponding second-order difference data sequence, obtaining the mean of the second-order difference data sequence, recorded as a first mean, respectively calculating the perfect square difference between each second-order difference data in the second-order difference data sequence and the first mean, correspondingly obtaining the mean of the perfect square difference, recorded as a second mean, obtaining a first ratio of the second mean to a preset stationarity parameter, substituting the opposite of the first ratio into an exponential function with a natural constant as the base, and obtaining a first exponential function result; For any relative distance except the first relative distance in the relative distance data sequence, obtain a first difference between the any relative distance and its previous relative distance, obtain a first product of the first difference and a preset distance scaling parameter, perform cosine processing on the first product to obtain a cosine value, obtain a Gaussian function value of the any relative distance, perform weighted summation processing on the cosine value and the Gaussian function value to obtain a weighted summation result of the any relative distance, respectively obtain a weighted summation result of each relative distance in the relative distance data sequence, and obtain a corresponding mean of the weighted summation results, which is recorded as a third mean; A weighted sum is performed on the first exponential function result and the third mean value to obtain a motion coefficient for characterizing the overall motion state of the head and hands of any one of the characters.
4. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 1 is characterized in that: The step of identifying, based on the motion coefficient, whether any of the characters in the target image set has suspected abnormal behavior includes: Setting a first motion coefficient threshold and a second motion coefficient threshold, wherein the first motion coefficient threshold is less than the second motion coefficient threshold; If the motion coefficient is greater than or equal to the first motion coefficient threshold and less than or equal to the second motion coefficient threshold, it is confirmed that any of the characters has suspected abnormal behavior in the target image set.
5. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 1, characterized in that: The step of obtaining the smoke concentration data in the target location and the pixel values of the pixels in each frame of the target image in the target image set includes: Obtain a sampling period corresponding to the target image set, recorded as a target sampling period, and obtain a smoke concentration data sequence in the target place during the target sampling period; By using a preset smoke detection algorithm, the smoke region of the head of any person in each frame of the target image in the target image set is obtained, and the pixel value of the pixel point in each head smoke region is correspondingly obtained.
6. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 5 is characterized in that: The step of determining the abnormal behavior detection result of any person according to the smoke density data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set includes: Obtaining a smoke significance score within the target sampling period according to a difference in smoke concentration data in the smoke concentration data sequence; According to the pixel value difference of each pixel point in the head smoke area in the target image set, counting the number of target images with smoke diffusion features in the target image set; If the smoke prominence score is greater than or equal to a preset smoke prominence score threshold, and the number is greater than or equal to a first preset number, it is determined that any of the characters has abnormal behavior.
7. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 6 is characterized in that: The step of obtaining the smoke significance score within the target sampling period according to the difference of the smoke concentration data in the smoke concentration data sequence includes: Perform first-order difference processing on the smoke concentration data sequence to obtain a first-order difference data sequence, obtain the mean of the first-order difference data sequence, and record it as the fourth mean. For any first-order difference data in the first-order difference data sequence, obtain the perfect square difference between the arbitrary first-order difference data and the fourth mean, obtain the addition result of the perfect square difference between a constant 1 and a first preset multiple, and obtain the inverse of the addition result, which is recorded as the smoke data score corresponding to the arbitrary first-order difference data. Obtain the smoke data score corresponding to each first-order difference data in the first-order difference data sequence, and obtain the mean of the smoke data scores. Use the mean of the smoke data scores as the smoke significance score within the target sampling period.
8. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 6 is characterized in that: The counting of the number of target images with smoke diffusion features in the target image set according to the pixel value difference of each pixel point in the head smoke area in the target image set comprises: For any target image in the target image set except the first frame target image, obtain the overlapping area between the head smoke area in the any target image and the head smoke area in the previous target image of the any target image, respectively calculate the difference between the corresponding pixel values of each pixel point in the overlapping area in the any target image and in the previous target image, and obtain the mean of the difference in the overlapping area, which is recorded as the fifth mean, and perform normalization processing on the fifth mean to obtain the smoke diffusion degree of the any target image; Obtain the number of pixel points in the head smoke area in any of the target images, recorded as a first number, obtain the number of pixel points in the head smoke area in a target image previous to the any of the target images, recorded as a second number, calculate a second difference between the first number and the second number, substitute the opposite of the second difference of the second preset multiple into an exponential function with a natural constant as a base to obtain a second exponential function result, calculate the difference between a constant 1 and the result of the second exponential function, and obtain the smoke diffusion area of any of the target images; Performing weighted summation on the smoke diffusion degree and the smoke diffusion area of any target image to obtain the smoke diffusion coefficient of any target image; The smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image in the target image set except the first target image are obtained, and the number of target images with smoke diffusion characteristics in the target image set is counted according to the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image in the target image set except the first target image.
9. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 8, characterized in that: The counting of the number of target images with smoke diffusion features in the target image set according to the smoke diffusion degree, smoke diffusion area and smoke diffusion coefficient of each target image frame except the first target image frame in the target image set includes: If the smoke diffusion degree and the smoke diffusion area of any target image are greater than the constant 0, and the smoke diffusion coefficient of any target image is greater than a preset smoke diffusion coefficient threshold, it is determined that any target image has smoke diffusion features, and the target images with smoke diffusion features in the target image set are screened to obtain the number of target images with smoke diffusion features in the target image set.
10. The method for detecting abnormal behavior of surveillance video data based on deep learning according to claim 1, characterized in that: After determining the abnormal behavior detection result of any person according to the smoke density data in the target place and the pixel value difference of the pixel points in each frame of the target image in the target image set, the method further includes: Acquire in real time a second preset number of frames of new person video images consecutively following the target image set; if the second preset number of frames of new person video images are all target images, add the second preset number of frames of new person video images to the target image set, remove the target images before the second preset number of frames in the target image set, and obtain a new target image set.
Citation Information
Patent Citations
Abnormal behavior monitoring method and device, equipment and storage medium
CN113112528A
Abnormal behavior detection method and device, computer equipment and storage medium
CN113435362A
Outdoor place smoking behavior identification method with low false alarm rate
CN118570621A
Method and equipment for determining position of hand on the basis of depth image
JP2014235743A
Behavior detection method and apparatus, computer device, storage medium, and program
WO2023273132A1
Cited By
Personnel behavior supervision method based on AI electronic chest card lightweight model training
CN121121864A