Image processing abnormal behavior recognition method and system for security system

By acquiring the local and overall attention of key points on the human body and adjusting the step size of the ST-GCN algorithm, the limitations of the spatiotemporal graph convolution algorithm in long-term action sequence recognition are overcome, and the accuracy of abnormal behavior recognition is improved.

CN122116475APending Publication Date: 2026-05-29DAO YUNXING (XIAN) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DAO YUNXING (XIAN) TECH CO LTD
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Spatiotemporal graph convolution algorithms have limitations when dealing with complex long-term action sequences, leading to a decrease in the accuracy of abnormal behavior recognition.

Method used

By acquiring the local and overall attention levels of key human body points in community surveillance videos, the step size of the ST-GCN algorithm is adjusted to capture complex action features at different time scales, thereby improving the ability to recognize complex actions.

Benefits of technology

It improves the accuracy of community safety monitoring results, effectively identifies complex actions, and reduces misjudgments and omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116475A_ABST
    Figure CN122116475A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and more particularly to an image processing abnormal behavior recognition method and system for a security system. The method comprises the following steps: obtaining the local attention degree of each human body key point in each frame of community monitoring personnel grayscale image according to the comprehensive speed and acceleration of the human body key point in the community monitoring personnel grayscale image; obtaining the overall attention degree of each human body key point according to the local attention degree and the difference between the surrounding area of each human body key point in multiple frames of community monitoring personnel grayscale images; obtaining the adaptive step length of each human body key point according to the overall attention degree of the human body key point; obtaining the community monitoring personnel feature map according to the adaptive step length of the human body key point; and performing community safety monitoring based on the community monitoring personnel feature map. The present application improves the accuracy of the community safety monitoring result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for identifying abnormal behavior in image processing for security systems. Background Technology

[0002] In a community, security systems serve as the first line of defense for community safety. Through real-time monitoring and data analysis, they can promptly detect and warn of potential security risks. Security systems can analyze people's movement trajectories and behavioral patterns to identify abnormal behaviors. For example, if someone violates community safety regulations, such as fighting, tailgating, or illegal climbing, the system will automatically issue an alert to remind security personnel to be aware of potential anomalies. The surveillance cameras in the security system can collect video footage of the monitored area.

[0003] Among them, the Spatiotemporal Graph Convolutional (ST-GCN) algorithm is often used to identify human posture in videos of monitored areas: First, extract key points of human joints from each frame of the video to construct a spatial graph describing human posture; then, connect the spatial graphs of multiple consecutive frames in the time dimension to form a spatiotemporal graph that simultaneously contains spatial connections between joints and the temporal motion of the joints themselves; finally, by analyzing this spatiotemporal graph, identify whether there is any abnormal behavior of the personnel.

[0004] However, spatiotemporal graph convolution algorithms mainly focus on short temporal dependencies between adjacent frames, which limits their ability to handle complex long-term action sequences. For example, for abnormal behaviors that require long-term observation to identify, spatiotemporal graph convolution algorithms cannot effectively identify the coherence and overall pattern of the entire action. This limitation may lead to some complex abnormal behaviors being misjudged or missed, thereby reducing the accuracy of abnormal behavior identification. Summary of the Invention

[0005] To address the technical problem that spatiotemporal graph convolution algorithms primarily focus on short-term temporal dependencies between adjacent frames, which limits their ability to handle complex long-term action sequences and leads to reduced accuracy in identifying abnormal behavior, this invention provides an image processing method and system for identifying abnormal behavior in security systems.

[0006] In a first aspect, the present invention provides an image processing method for abnormal behavior recognition in security systems, employing the following technical solution: An image processing method for abnormal behavior recognition in security systems includes the following steps: The community surveillance video is divided into several frames of grayscale images of community surveillance personnel; several key human body points are obtained in each frame of the grayscale image of community surveillance personnel. Obtain the local attention level of each human body key point in each frame of the grayscale image of community surveillance personnel. , This indicates the degree of local attention given to the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; as well as a1 and a2 represent the maximum acceleration and the maximum combined velocity of the i-th human key point in all frames of grayscale images of community monitoring personnel; a1 and a2 both represent weight parameters. Represents a linear normalization function; Based on the local attention level and the differences in the surrounding area of ​​each human body key point among the grayscale images of community monitoring personnel in multiple frames, the overall attention level of each human body key point is obtained; based on the overall attention level of each human body key point, the adaptive step size of each human body key point is obtained; based on the adaptive step size of the human body key points, the feature map of community monitoring personnel is obtained, and community security monitoring is carried out based on the feature map of community monitoring personnel.

[0007] The innovation of this invention lies in obtaining the local attention level of each human keypoint in each frame of the grayscale image of community monitoring personnel based on the comprehensive velocity and acceleration of human keypoints in the grayscale image; obtaining the overall attention level of each human keypoint based on the local attention level and the differences in the surrounding area of ​​each human keypoint across multiple frames of grayscale images of community monitoring personnel; obtaining the adaptive step size of each human keypoint based on the overall attention level; obtaining the feature map of community monitoring personnel based on the adaptive step size; performing community security monitoring based on the feature map of community monitoring personnel; and improving the accuracy of community security monitoring results by adjusting the step size, enabling the ST-GCN algorithm to more effectively capture complex action features at different time scales.

[0008] Preferably, the acquisition of the acceleration of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes: Obtain the combined velocity of the i-th human body key points in the grayscale image of the community monitoring personnel in the t-th frame; , This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t-1)-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This indicates the time interval between two adjacent grayscale images of community monitoring personnel; Indicates taking the absolute value; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t.

[0009] By comprehensively considering the overall movement trend of key points in the human body, rather than just unidirectional changes, the accuracy of calculations can be improved.

[0010] Preferably, the acquisition of the comprehensive velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes: Obtain the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; obtain the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; , This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; Let represent the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t.

[0011] Preferably, the step of obtaining the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes:

[0012] In the formula, This represents the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; Let x represent the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-2; norm() represents the linear normalization function.

[0013] Preferably, the step of obtaining the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes...

[0014] In the formula, This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the ordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; The value of the ordinate of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-2 represents the value of the ordinate; norm() represents the linear normalization function.

[0015] Preferably, obtaining the overall attention level of each key point on the human body includes: Obtain the HOG features of the i-th human body key points in each frame of the grayscale image of community monitoring personnel;

[0016] In the formula, This represents the overall level of attention paid to key points on the i-th person's body. This represents grayscale images of all community monitoring personnel in all frames; This represents the HOG feature of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; This represents the mean HOG feature of the i-th human keypoint across all frames of grayscale images of community surveillance personnel. This represents the mean local attention level of the i-th human key point across all frames of grayscale images of community surveillance personnel. Indicates taking the absolute value; Represents a linear normalization function; This represents an exponential function with the natural constant as its base.

[0017] By analyzing the differences in HOG features of the area surrounding key points of the human body across multiple frames of grayscale images of community surveillance personnel, the changes in texture features of key points of the human body can be effectively captured.

[0018] Preferably, obtaining the HOG features of the i-th human body key points in each frame of the grayscale image of the community monitoring personnel includes: A neighborhood parameter n is preset. In the grayscale image of the community monitoring personnel in the t-th frame, the i-th human key point is used as the center of the window to obtain a window with a size of n×n, and this window is recorded as the neighborhood range of the i-th human key point. The HOG algorithm is used to obtain the HOG features of the neighborhood range of the i-th human key point, and these features are denoted as the HOG features of the i-th human key point in the grayscale image of the community monitoring personnel in frame t.

[0019] Preferably, the adaptive step size for obtaining each human body key point includes: A hyperparameter B is preset; the initial custom step size for each human body key point is preset. ;

[0020] In the formula, This represents the adaptive step size of the i-th human body key point; This represents the overall level of attention paid to key points on the i-th person's body. This represents an exponential function with the natural constant as its base.

[0021] By adjusting the step size, the ST-GCN algorithm can more effectively capture complex action features at different time scales, thereby improving its ability to recognize complex actions.

[0022] Preferably, the step of obtaining a feature map of community monitoring personnel and conducting community security monitoring based on the feature map includes: The adaptive step size of the surveillance video and all human key points is input into the ST-GCN network to obtain the feature map of community monitoring personnel; the feature map of community monitoring personnel is then input into the trained neural network to obtain the community security monitoring results.

[0023] This improved the accuracy of community safety monitoring results.

[0024] Secondly, this invention provides an image processing abnormal behavior recognition system for security systems, employing the following technical solution: An image processing abnormal behavior recognition system for security systems includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the aforementioned image processing abnormal behavior recognition method for security systems.

[0025] By adopting the above technical solution, the image processing abnormal behavior recognition method for security systems is generated into a computer program and stored in a memory for loading and execution by a processor. This allows for the creation of a terminal device based on the memory and processor, facilitating its use.

[0026] This invention has the following technical effects: It obtains the local attention level of each human keypoint in each frame of a grayscale image of community monitoring personnel based on the comprehensive velocity and acceleration of human keypoints in the grayscale image; it obtains the overall attention level of each human keypoint based on the local attention level and the differences in the surrounding area of ​​each human keypoint across multiple frames of grayscale images; it obtains the adaptive step size of each human keypoint based on the overall attention level; it obtains a feature map of community monitoring personnel based on the adaptive step size; it performs community security monitoring based on the feature map; and by adjusting the step size, the ST-GCN algorithm can more effectively capture complex action features at different time scales, thereby improving the ability to recognize complex actions and thus improving the accuracy of community security monitoring results. Attached Figure Description

[0027] Figure 1 This is a flowchart of the image processing abnormal behavior recognition method for security systems according to an embodiment of the present invention; Figure 2 This is a diagram showing the abnormal behavior recognition results of the present invention. Detailed Implementation

[0028] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0029] This invention discloses an image processing method for identifying abnormal behavior in security systems, referring to... Figure 1 This includes steps S1-S4: S1: Divide the community surveillance video into several frames of grayscale images of community surveillance personnel.

[0030] It should be noted that the surveillance camera will be installed in the community according to the actual situation, and it can capture the posture video of people to detect abnormal pedestrian behavior in the community. By processing the posture video of all people in frames, it is convenient to carry out subsequent implementations. Since the surveillance camera works for a long time and there is a lot of noise in the environment, it may cause noise in the framed images, thus affecting the image quality. Therefore, it is necessary to preprocess each frame of the image to improve the quality of each frame.

[0031] In this embodiment of the invention, the specific method for obtaining several frames of grayscale images of community monitoring personnel is as follows: The community surveillance video collected by the surveillance cameras installed in the community is input into the OpenCV library, and the functions in the library are used to perform frame segmentation on the community surveillance video to obtain several frames of images of community surveillance personnel. For any frame of community monitoring personnel image, the grayscale image of the community monitoring personnel is obtained by using the multi-scale Retinex algorithm and the illumination normalization algorithm to perform grayscale processing and illumination normalization processing. In this process, after the above preprocessing operations, each frame of grayscale image of community monitoring personnel not only removes noise from the image, but also improves the contrast and detail of the image and reduces the impact of illumination changes on subsequent analysis. The multi-scale Retinex algorithm and illumination normalization algorithm are existing technologies, and will not be described in detail here.

[0032] S2: Obtain several human key points in each frame of the grayscale image of community monitoring personnel; based on the positional changes of each human key point in adjacent frames of grayscale images of community monitoring personnel, obtain the comprehensive velocity and acceleration of each human key point in each frame of grayscale images of community monitoring personnel; based on the comprehensive velocity and acceleration of the human key points in the grayscale images of community monitoring personnel, obtain the local attention level of each human key point in each frame of grayscale images of community monitoring personnel.

[0033] It should be noted that human keypoints typically refer to the joint positions of the human body, such as the head, shoulders, elbows, wrists, hips, knees, and ankles. These keypoints can be located in the grayscale image of the community monitoring personnel using the OpenPose algorithm. Since human keypoints in surveillance videos captured by a surveillance camera usually exhibit a certain degree of continuity between adjacent grayscale images of the community monitoring personnel, the local attention level of each keypoint can be obtained by analyzing the changes between keypoints in adjacent grayscale images. Local attention level reflects the degree of motion change of each keypoint within a short period of time. Keypoints that change significantly within a short period of time are more likely to reflect the key dynamic parts of the action, and therefore, these keypoints usually have high reference value when analyzing human motion. Local attention level can also take into account changes in acceleration, which can be obtained based on the positional changes of the keypoints in each grayscale image of the community monitoring personnel.

[0034] In this embodiment of the invention, the specific method for obtaining several key human body points in each frame of a grayscale image of community monitoring personnel is as follows: For any given frame of a grayscale image of community monitoring personnel, the OpenPose algorithm is used to detect all human key points in the grayscale image of community monitoring personnel.

[0035] Among them, a total of 17 human body key points were obtained through the OpenPose algorithm; and the OpenPose algorithm is existing technology, so it will not be described in detail here.

[0036] In this embodiment of the invention, since the central difference method calculates the velocity of the human key points in the current frame of the grayscale image of community monitoring personnel by considering the positional changes of human key points in the grayscale images of community monitoring personnel in two consecutive frames, it can comprehensively consider the overall movement trend of human key points, rather than just unidirectional changes, thereby improving the accuracy of the calculation. Therefore, the specific method for obtaining the comprehensive velocity of each human key point in each frame of the grayscale image of community monitoring personnel based on the positional changes of each human key point in adjacent frames of the grayscale image of community monitoring personnel is as follows: For any frame of grayscale image of community monitoring personnel, construct a rectangular coordinate system with the pixel at the lower left corner of the grayscale image of community monitoring personnel as the origin, input the grayscale image of community monitoring personnel into the rectangular coordinate system, and obtain the position coordinates of each human body key point in the grayscale image of community monitoring personnel. The formula for calculating the horizontal velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t is as follows:

[0037] In the formula, This represents the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t-2)-th frame; This indicates the time interval between two adjacent grayscale images of community monitoring personnel; The expression represents taking the absolute value; norm() represents the linear normalization function.

[0038] The formula for calculating the vertical velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t is as follows:

[0039] In the formula, This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the ordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; This represents the ordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-2; This indicates the time interval between two adjacent grayscale images of community monitoring personnel; The expression represents taking the absolute value; norm() represents the linear normalization function.

[0040] The method for calculating the combined velocity of the i-th human keypoint in the t-th frame of the community monitoring personnel's grayscale image, based on its vertical and horizontal velocities, is as follows:

[0041] In the formula, This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; Let represent the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t.

[0042] It should be noted that the greater the difference in the horizontal coordinates of the i-th human key point in the grayscale images of community monitoring personnel in frames t-2, t-1, t+1, and t+2, and the smaller the time interval, the greater the horizontal velocity of the i-th human key point in frame t. Similarly, the greater the difference in the vertical coordinates of the i-th human key point in the grayscale images of community monitoring personnel in frames t-2, t-1, t+1, and t+2, and the smaller the time interval, the greater the vertical velocity of the i-th human key point in frame t.

[0043] In this embodiment of the invention, since the central difference method is used to approximate the acceleration, in order to more accurately reflect the change in velocity, the comprehensive velocity of the human key points in each frame of the grayscale image of the community monitoring personnel is considered. This method can more comprehensively capture the change in the comprehensive velocity of the human key points in each frame of the image, avoiding errors caused by sudden changes in velocity; and since only the magnitude of the acceleration is considered, and not the direction of the acceleration, the absolute value of the acceleration is sufficient; therefore, the method for calculating the acceleration of the i-th human key point in the t-th frame of the grayscale image of the community monitoring personnel is as follows: ; In the formula, This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t-1)-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This indicates the time interval between two adjacent grayscale images of community monitoring personnel; Indicates taking the absolute value; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t.

[0044] In this embodiment of the invention, local attention can reflect the degree of motion change of each human key point in a short period of time; human key points that change significantly in a short period of time can reflect the key dynamic parts of the action; that is, when the overall velocity and acceleration of a human key point in a certain frame of grayscale image are large, it indicates that the key point has changed significantly in a short period of time, which may be a turning point or a key part of the action; therefore, acceleration should be given a large weight. The method for calculating the local attention of the i-th human key point in the t-th frame of the grayscale image of community monitoring personnel is as follows: Two weight parameters are preset: a1=0.6 and a2=0.4. ; In the formula, This indicates the degree of local attention given to the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the maximum acceleration of the i-th human keypoint across all frames of grayscale images of community surveillance personnel. This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the maximum combined velocity of the i-th human keypoint across all frames of grayscale images of community surveillance personnel; norm() represents the linear normalization function.

[0045] Thus, the local attention level of each human body key point in each frame of the grayscale image of community monitoring personnel is obtained.

[0046] S3: Based on the local attention level and the differences in the surrounding area of ​​each human body key point among the grayscale images of community monitoring personnel in multiple frames, obtain the overall attention level of each human body key point.

[0047] It should be noted that the HOG algorithm is a feature descriptor used for object detection and recognition, which is particularly suitable for capturing local texture and edge information in images. The HOG feature is a high-dimensional vector that describes the gradient direction distribution of a local region. By analyzing the changes in HOG features around human keypoints, the changes in their texture features can be effectively captured, thereby obtaining the overall attention level of each human keypoint.

[0048] In this embodiment of the invention, if there exists a human body keypoint that changes significantly in a short period of time and is repetitive over a long period of time, it may correspond to an important action feature, and such a human body keypoint reflects relatively stable texture features in the image; then the method for calculating the overall attention of the i-th human body keypoint is as follows: With a preset neighborhood parameter n=10, in the grayscale image of the community monitoring personnel in the t-th frame, take the i-th human key point as the window center, obtain a window with a size of n×n, and record this window as the neighborhood range of the i-th human key point; The HOG algorithm is used to obtain the HOG features of the neighborhood range of the i-th human key point, and these features are denoted as the HOG features of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame. The formula for calculating the overall attention level of the i-th key point on the human body is: ; In the formula, This represents the overall level of attention paid to key points on the i-th person's body. This represents grayscale images of all community monitoring personnel in all frames; This represents the HOG feature of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; This represents the mean HOG feature of the i-th human keypoint across all frames of grayscale images of community surveillance personnel. This represents the mean local attention level of the i-th human key point across all frames of grayscale images of community surveillance personnel. Indicates taking the absolute value; Represents a linear normalization function; This represents an exponential function with the natural constant as its base.

[0049] It should be noted that the smaller the difference between the HOG features of the i-th human keypoint in each frame of the grayscale image of the community monitoring personnel and its mean, the more similar the HOG features of the i-th human keypoint are in multiple frames of images, that is, the smaller the texture difference and the more repetitive it is over a long period of time, and the greater the overall attention of the i-th human keypoint. The greater the local attention of the i-th human keypoint in multiple frames of grayscale images, the more significant the changes of the i-th human keypoint are in a short period of time, and the greater the overall attention of the i-th human keypoint. The HOG algorithm is existing technology, and will not be described in detail here.

[0050] S4: Based on the overall attention level of key human body points, obtain the adaptive step size of each key human body point; based on the adaptive step size of key human body points, obtain the feature map of community monitoring personnel; and conduct community security monitoring based on the feature map of community monitoring personnel.

[0051] In this embodiment of the invention, the method for calculating the adaptive step size of each human body key point based on the overall attention level of the human body key points is as follows: A hyperparameter B=0.5 is preset; the initial custom step size for each human body key point is preset. =1; ; In the formula, This represents the adaptive step size of the i-th human body key point; This represents the overall level of attention paid to key points on the i-th person's body. This represents an exponential function with the natural constant as its base.

[0052] It should be noted that the hyperparameter B is used to control the growth rate of the function. When the overall attention to human key points is higher, the value of the exponential function will decrease significantly. This means that for human key points with high overall attention, the stride will be smaller, which may give these human key points a shorter time span in the convolution operation, so as to better recognize the details of human movements.

[0053] In this embodiment of the invention, a feature map of community monitoring personnel is obtained based on the adaptive step size of key human body points; the specific method for community security monitoring based on the feature map of community monitoring personnel is as follows: The adaptive step size of the surveillance video and all human key points is input into the ST-GCN network to obtain the feature map of community monitoring personnel; The feature maps of community monitoring personnel are input into a trained neural network to obtain community security monitoring results; in this embodiment, the neural network used is ResNet50, and the method for obtaining the dataset used to train this neural network is as follows: Feature images of community monitoring personnel are collected, and each feature image is manually labeled. Specifically, if abnormal behavior is observed in a feature image, the community security monitoring result for that feature image is marked as abnormal; if normal behavior is observed, the result is marked as normal. This labeling is then recorded as the tag for each feature image. A large number of feature images of community monitoring personnel and their corresponding tags are collected to form a dataset. This dataset is then used to train the neural network, using the cross-loss function as the loss function. The ST-GCN network and its specific training process are well-known aspects of neural networks, and this embodiment will not elaborate on the specific training process.

[0054] Figure 2This is the abnormal behavior recognition result image of the present invention. The result image clearly shows that the abnormal behavior category is labeled as "Climb", corresponding to climbing behavior and with an 85% recognition confidence level. It clearly points to high-risk abnormal behavior in community security scenarios. This shows that the present invention, through the whole process design of image preprocessing, key point dynamic feature extraction, overall attention calculation, and deep learning classification, can accurately identify target abnormal behavior and quantify the reliability. It effectively avoids the influencing factors such as lighting interference and key point motion error. It not only achieves accurate positioning of abnormal behavior categories, but also provides a reliable basis for security response through high confidence output. It fully verifies the accuracy and engineering practicality of the method in community security abnormal behavior recognition.

[0055] This invention also discloses an image processing abnormal behavior recognition system for security systems, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the image processing abnormal behavior recognition method for security systems provided by this invention. The system also includes other components well-known to those skilled in the art, such as a communication bus and a communication interface; their configurations and functions are known in the art and will not be described further here.

[0056] In this invention, the aforementioned memory can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0057] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. An image processing method for abnormal behavior recognition in security systems, characterized in that, include: The community surveillance video was divided into several frames of grayscale images of community surveillance personnel. Acquire several key human body points in each frame of grayscale image of community monitoring personnel; Obtain the local attention level of each human body key point in each frame of the grayscale image of community surveillance personnel. , This indicates the degree of local attention given to the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; as well as a1 and a2 represent the maximum acceleration and the maximum combined velocity of the i-th human key point in all frames of grayscale images of community monitoring personnel; a1 and a2 both represent weight parameters. Represents a linear normalization function; Based on the local attention level and the differences in the surrounding area of ​​each human body key point among the grayscale images of community monitoring personnel in multiple frames, the overall attention level of each human body key point is obtained. Based on the overall attention level of each human body key point, obtain the adaptive step size of each human body key point; Based on the adaptive step size of key human body points, feature maps of community monitoring personnel are obtained, and community security monitoring is carried out based on these feature maps.

2. The image processing abnormal behavior recognition method for security systems according to claim 1, characterized in that, The acquisition of the acceleration of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes: Obtain the combined velocity of the i-th human body key points in the grayscale image of the community monitoring personnel in the t-th frame; , This represents the acceleration of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t-1)-th frame; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This indicates the time interval between two adjacent grayscale images of community monitoring personnel; Indicates taking the absolute value; This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t.

3. The image processing abnormal behavior recognition method for security systems according to claim 1 or 2, characterized in that, The acquisition of the comprehensive velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes: Obtain the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; obtain the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; , This represents the overall velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; Let represent the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in frame t.

4. The image processing abnormal behavior recognition method for security systems according to claim 3, characterized in that, The acquisition of the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame includes: In the formula, This represents the horizontal velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; Let x represent the x-coordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-2; norm() represents the linear normalization function.

5. The image processing abnormal behavior recognition method for security systems according to claim 3, characterized in that, The process of obtaining the vertical velocity of the i-th human body key point in the grayscale image of the community monitoring personnel in the t-th frame includes: In the formula, This represents the vertical velocity of the i-th human key point in the grayscale image of the community monitoring personnel in the t-th frame; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+1)-th frame; This represents the ordinate value of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-1; This represents the ordinate value of the i-th human key point in the grayscale image of the community monitoring personnel in the (t+2)-th frame; The value of the ordinate of the i-th human body key point in the grayscale image of the community monitoring personnel in frame t-2 represents the value of the ordinate; norm() represents the linear normalization function.

6. The image processing abnormal behavior recognition method for security systems according to claim 1, characterized in that, The acquisition of the overall attention level for each key point on the human body includes: Obtain the HOG features of the i-th human body key points in each frame of the grayscale image of community monitoring personnel; In the formula, This represents the overall level of attention paid to key points on the i-th person's body. This represents grayscale images of all community monitoring personnel in all frames; This represents the HOG feature of the i-th human key point in the grayscale image of the community monitoring personnel in frame t; This represents the mean HOG feature of the i-th human keypoint across all frames of grayscale images of community surveillance personnel. This represents the mean local attention level of the i-th human key point across all frames of grayscale images of community surveillance personnel. Indicates taking the absolute value; Represents a linear normalization function; This represents an exponential function with the natural constant as its base.

7. The image processing abnormal behavior recognition method for security systems according to claim 6, characterized in that, The acquisition of the HOG features of the i-th human body key points in each frame of the grayscale image of the community monitoring personnel includes: A neighborhood parameter n is preset. In the grayscale image of the community monitoring personnel in the t-th frame, the i-th human key point is used as the center of the window to obtain a window with a size of n×n, and this window is recorded as the neighborhood range of the i-th human key point. The HOG algorithm is used to obtain the HOG features of the neighborhood range of the i-th human key point, and these features are denoted as the HOG features of the i-th human key point in the grayscale image of the community monitoring personnel in frame t.

8. The image processing abnormal behavior recognition method for security systems according to claim 1, characterized in that, The adaptive step size for obtaining each human body key point includes: A hyperparameter B is preset; the initial custom step size for each human body key point is preset. ; In the formula, This represents the adaptive step size of the i-th human body key point; This represents the overall level of attention paid to key points on the i-th person's body. This represents an exponential function with the natural constant as its base.

9. The image processing abnormal behavior recognition method for security systems according to claim 1, characterized in that, The process of obtaining a feature map of community monitoring personnel and conducting community security monitoring based on that feature map includes: The adaptive step size of the surveillance video and all human key points is input into the ST-GCN network to obtain the feature map of community monitoring personnel; the feature map of community monitoring personnel is then input into the trained neural network to obtain the community security monitoring results.

10. An image processing-based abnormal behavior recognition system for security systems, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement the image processing abnormal behavior recognition method for security systems according to any one of claims 1-9.